For long-running agents, the useful limit is the environment, not the number of steps. Jump Trading shows how
Jump Trading lets GPT-6 Astra agents run long quantitative research loops on their own, inside a monitored environment with human-defined criteria and review before any signal is used.
Listen to this article · 4 min
AI-generated narration of the full article.

One of the most demanding and regulated users of AI agents does not control them by keeping runs short. It lets them run long, and controls where they run.
That is the lesson in a case study OpenAI published on October 6 about Jump Trading, a quantitative trading firm that builds predictive models from market data, news, events and alternative data.
The problem: research does not fit on a short leash
The usual way to keep an agent safe is a short leash: few steps, frequent approvals, narrow tasks. That works for a support bot. It fails for research, where the value comes from following a thread for a long time and changing direction when the evidence says so. An agent that stops for approval every few steps is just a slower analyst.
Lucas Baker, Jump’s head of LLM research and development, describes agents running GPT-6 Astra that now take on quantitative studies, not only coding tasks.
The architecture: people design the box, agents work inside it
As Jump describes it:
- people define the problem, the environment and the evaluation criteria before the agent starts;
- agents analyze their own findings, judge them against those criteria and redirect the work without a person reviewing every round;
- the longest runs still have regular human check-ins on data, runtime, what counts as important, and intermediate results.
The environment is the boundary. An agent inside a monitored research sandbox, with defined data and compute, can run for a long time because its actions cannot reach anything that matters without passing a gate. Autonomy inside the box is cheap; crossing its edge is expensive.
Evaluation criteria do the work that step-by-step approval used to do. Writing down what a good result looks like before the run lets the agent judge its own progress and lets people review outcomes instead of moves. It is the discipline behind our argument that in regulated AI, the model never grades its own work: the criteria come from outside the agent.
- People define problem, data and criteria
- Agents research, evaluate and redirect
- Periodic human check-ins
- Acceptance gate: same review as any signal
- Production trading workflows
- Monitored research environment: People define problem, data and criteria · Agents research, evaluate and redirect · Periodic human check-ins
Human directionAgent autonomy
The controls: an old gate, not a new one
Jump lists its safety measures as system design: boundaries, constraints, steerability, observability and human review of changes. The decisive one is the acceptance gate. Signals produced by agents are scoped and reviewed like any other signal, treated as informative but possibly wrong, before they reach trading workflows. The firm did not invent a new control for agents; it routed agent output into one it already trusted.
Nothing in the case study describes agents executing trades. The agents do research; people decide what reaches production. The productivity effects described are Jump’s and OpenAI’s own account.
The transferable lesson
Any team that wants agents doing analysis rather than answering questions faces the same choice between frequent approvals and designed boundaries. The pieces carry over directly: a research environment with its own data copies, compute budget and logging, and no path to production; criteria written before the run; check-ins about direction rather than every step; and agent output routed through the acceptance process that already exists, whether that is model validation, a change advisory board or a credit committee.
From there, autonomy can be earned rather than configured: once an agent’s results consistently pass the gate, the environment can be widened, one boundary at a time. Short leashes keep agents safe and mostly useless for hard problems. The boundary that matters is the edge of the environment.