← All insights

News analysisLens: Global4 min read

OpenAI's Agents API turns the agent loop into a managed service. The rest is still yours.

Managed harness, subagents, context compaction and now computer use. What an enterprise can hand to the model provider, and what it cannot.

Listen to this article · 5 min

AI-generated narration of the full article.

Close-up of control knobs and faders, in La Madre duotone, beside the words The agent loop
Photo: rawpixel (CC0)

For most of the last two years, building an agent meant building a loop: call the model, parse the tool request, run the tool, feed the result back, manage the context window, recover from failures, repeat. Every serious team wrote its own version, and most of them were fragile.

OpenAI’s Agents API, in public beta since September 10, offers to run that loop for you. At DevDay on September 29, OpenAI added computer use. Taken together, these announcements mark a clear line between what a model provider can now operate and what an enterprise still has to own.

What OpenAI now runs

According to OpenAI, the Agents API exposes the same harness and infrastructure that powers Codex:

  • A managed harness, operated and maintained by OpenAI, that handles session orchestration, context compaction and recovery.
  • Durable sessions that continue work across turns.
  • Tool search and programmatic tool calling, with support for MCP, custom functions and built-in tools such as web search.
  • Subagents that split a task into parallel pieces, each with its own context.
  • A choice of compute: an OpenAI-hosted sandbox, your own infrastructure, or one of several sandbox partners.

The harness is open source, so teams can inspect the logic that coordinates model calls, tools and context. There is no additional fee beyond tokens and tools. The API is in public beta, and OpenAI says it is iterating toward general availability.

The September 29 changelog adds the detail that matters most for enterprises: computer use in the Agents API runs in an OpenAI-hosted browser, “with website access approvals and sign-in handled by your application.”

That sentence describes the boundary well. OpenAI operates the browser. Your application decides which sites the agent may use and how it authenticates.

The responsibility line

What the Agents API takes on, and what stays with the enterprise. Tools and actions are shared: the API runs them, you decide which exist.CORE LAYERSCROSS-CUTTING CONTROLSBusiness application & experienceAgent runtime: orchestration & stateTools, actions & integrationsEnterprise context: data & groundingModelsIdentity & authorizationPolicy & securityEvaluationHuman approvalObservability & auditLifecycle & deploymentOperating ownership: who is accountable after go-liveNow operated by the provider (beta)Still owned by the enterprise

Core layers

  • Business application & experience
  • Agent runtime: orchestration & state
  • Tools, actions & integrations
  • Enterprise context: data & grounding
  • Models

Cross-cutting controls

  • Identity & authorization
  • Policy & security
  • Evaluation
  • Human approval
  • Observability & audit
  • Lifecycle & deployment
  • Operating ownership: who is accountable after go-live

Now operated by the provider (beta)Still owned by the enterprise

What the Agents API takes on, and what stays with the enterprise. Tools and actions are shared: the API runs them, you decide which exist.

Identity. When an agent calls your systems, whose credentials does it use? A managed harness does not answer that. Sign-in for computer use is explicitly left to your application, and the same logic applies to every MCP server and custom function you connect.

Authorization. The harness will call the tools it has. Which tools exist, with which scopes, for which users, is your design. Tool search makes it cheap to expose many tools, which is precisely why the list needs review.

Enterprise data. Choosing an OpenAI-hosted sandbox, your own infrastructure or a partner is a data decision as much as a compute decision. Files, intermediate results and session state live wherever you put them.

Audit. Your auditors will want your records of what the agent did in your systems, not a provider’s. Plan to log tool calls and actions on your side, with the business context that explains them.

Evaluation. OpenAI says it improves the harness alongside each model launch. That is useful, and it also means behavior can change when you upgrade. Keep a regression suite and rerun it before switching versions.

Application-level controls. Approval steps, spending limits, rate limits per user and kill switches are application logic. The API gives you the hooks; it does not choose the thresholds.

Computer use: power and a new attack surface

Computer use lets agents work with software that has no API: legacy portals, supplier websites, internal tools nobody will ever integrate. For many enterprises, that is where the manual work actually is.

It also means an agent is clicking through interfaces with real credentials. Three practical rules:

  1. Prefer an API when one exists. Computer use is a fallback, not a default.
  2. Keep an allowlist of sites and require approval for new ones, which is what OpenAI’s design already expects from your application.
  3. Use dedicated, least-privilege accounts for agent sign-in, never a person’s own credentials.

What to reconsider now

If your team maintains its own agent loop, the question is no longer whether you can build one. It is whether maintaining it is still a good use of engineering time. For many teams it will not be.

What should not be outsourced is the part that makes the system yours: who it acts as, what it may touch, what evidence it leaves, how you know it works, and who is accountable for it. We describe these layers in our 2026 stack guide.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case