OpenAI's Agents API puts the Codex harness behind one endpoint: what developers get, and what they give up

In public beta since September 10, with computer use added September 29. How sessions, sandboxes and approvals work, what it costs, and the catches.

ByShajanthanFounder & Editor
Published
Reading5 MIN
Diagram of OpenAI's Agents API: app, managed Codex harness, and three sandbox options
What’s new, in 20 seconds
  1. OpenAI's Agents API, in public beta since September 10, 2026, runs the open-source Codex agent harness as a managed service: OpenAI handles sessions, context compaction and recovery.
  2. There's no extra fee: you pay model, tool and container rates. Data residency is US-only and Zero Data Retention isn't supported.
  3. On September 29 OpenAI added computer use in a hosted browser. Your app must approve every new website, but not individual clicks.
Contents

The most significant developer change OpenAI shipped between early September and early October 2026 is a new product, not a model. On September 10, 2026, OpenAI released the Agents API in public beta. It is a managed service that runs the same agent harness behind Codex, so developers don't have to build their own agent loop. On September 29, at DevDay, OpenAI added computer use, which lets those agents operate a browser that OpenAI hosts.

In the changelog's words, you "build agents with a managed Codex harness while OpenAI handles session orchestration, context compaction, and recovery."

What it is

Until now, OpenAI gave developers two ways to build agents. You could call the Responses API and write the loop yourself, or you could use the open-source Agents SDK, which runs the loop inside your own application. The Agents API adds a third option, where OpenAI runs the loop. OpenAI's documentation compares the three:

OptionWho runs the agent loopWhere state livesOpenAI's "integration effort"
Agents APIOpenAI (managed Codex harness)Saved sessions, turns and items on OpenAI's sideLow
Agents SDKYour applicationYour storage / SDK sessionsMedium
Responses APIYou build itYou manage history or use ConversationsHigh

The API is built on four concepts:

  • Agent: the model, instructions, tools and MCP servers.
  • Environment: an optional sandbox where the agent edits files and runs commands.
  • Session: a durable agent instance that persists across turns.
  • Events and items: what goes in and comes out.

The harness itself is the Apache-2.0 openai/codex project, so you can read the code that OpenAI is running for you.

According to OpenAI's launch post and docs, the managed harness handles:

  • automatic context compaction for long sessions
  • tool search, which loads tool definitions only when they're needed
  • programmatic tool calling
  • MCP servers, custom functions and built-in tools such as web search
  • subagents, configured with a multi_agent block (the docs' example sets max_concurrent_subagents to 4)
  • steering an agent mid-task and resuming sessions

Where the agent runs

You choose the compute:

  • OpenAI-hosted sandbox. OpenAI says this uses the same infrastructure as Codex and ChatGPT. You can preload files, packages, skills and plugins.
  • Self-hosted sandbox. The agent works in your own environment. MarkTechPost reports that this works by running codex exec-server, which connects out to OpenAI over WebSocket with a restricted key.
  • Partner sandboxes. At launch OpenAI named nine providers: Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop and Vercel.
  • No sandbox, for agents that only call tools.

What the code looks like

Requests go to POST https://api.openai.com/v1/agents/sessions with the header OpenAI-Beta: agents=v1. The official SDKs add this header automatically. In Python, the beta lives under client.beta.agents. OpenAI's quickstart installs the SDK with pip install --upgrade openai. This section reflects openai 3.26.1, still the latest on PyPI on October 9, 2026. This is the quickstart's first example, as published in OpenAI's docs on October 9, 2026. It wraps the streamed session in .with_result_collection() so you can read the final result after the event stream ends:

PYTHON
from openai import OpenAI

with OpenAI() as client:
    with client.beta.agents.sessions.create(
        agent={
            "model": "gpt-6-astra",
            "instructions": "Write clean code, run it, and report the actual output.",
        },
        environment={"type": "openai_hosted"},
        input="Create tree.py, a Python script that prints a readable tree of the files in the current directory. Run it and show me the output.",
        stream=True,
    ).with_result_collection() as stream:
        for event in stream:
            print(event.to_json(indent=None), flush=True)
        result = stream.get_final_result()
    print(result.output_text)
    session_id = result.session_id

The API is in beta, so check the live quickstart before copying this code.

Some details are easy to get wrong:

  • API key scopes. The key needs api.agents.read and api.agents.write for session operations, plus api.responses.write for model inference. OpenAI says to keep the key outside the agent's sandbox.
  • Checking whether a turn succeeded. Watch for agent.session.turn.completed, then read the agent's reported result. A completed turn doesn't mean every tool call succeeded. agent.session.idle on its own doesn't mean success. turn.failed, turn.cancelled and session.failed signal failure or cancellation.
  • Follow-up input. Open the event stream before you send follow-up input, or you can miss early events.
  • Cleanup. Call client.beta.agents.sessions.delete(...) when you're done, after saving any files you need.

The docs' examples use gpt-6-astra. The docs don't list which other models are supported. See our GPT-6 launch story for models and prices.

Computer use: approvals by website, not by action

The September 29 update adds a computer_use tool. You enable it with an openai_hosted environment and desktop: { enabled: true }. Setting include_screenshots: true lets your app show progress. Screenshots are left out of API output by default.

The safety model is worth understanding before you ship anything:

  • Every new website origin needs approval from your app, even public sites. Turning on network access doesn't grant it. The session reports agent.session.requires_action. Your app then answers a browser_origin_access request with approve, deny or cancel.
  • Sign-ins arrive as browser_authentication requests that your application handles.
  • Approving a site does not approve individual actions. OpenAI's documentation says outright that origin approval doesn't require confirmation before consequential actions such as purchases or deletions. If you need that guarantee, OpenAI says to restrict the browser or use a runtime you control. A confirmation step implemented as a function tool only works if the agent chooses to call it.

That gap is the kind of thing our case against unattended agents is about.

Pricing and limits

Last verified: October 8, 2026

  • Fees: OpenAI says "there are no additional fees for using the Agents API." You pay the model's token rates, standard tool rates, and standard container rates for OpenAI-hosted sandboxes.
  • Data residency: United States only for now.
  • Zero Data Retention: not supported, and using a self-hosted sandbox doesn't make the API ZDR-eligible.
  • Quotas: session limits and rate limits are not documented.

Who it's for, and the trade-off

OpenAI's launch post cites customer results: Ciridae says an evaluation score rose from 0.71 to 0.85, SafetyKit reports 60% lower cost per case, and Hypha says failed agent responses fell 86%. These are vendor-supplied testimonials, not independent measurements.

The real trade is control versus convenience. The Agents API removes the hardest parts of running long agent tasks: state, compaction, retries and sandbox wiring. In exchange, your agent's loop, state and data live on OpenAI's side, under a beta header, with no ZDR and US-only residency. Teams with strict data rules, or that want to switch models easily, may still prefer running the loop themselves, whether with OpenAI's Agents SDK or Anthropic's equivalent (see build your first agent with the Claude Agent SDK).

Other API changes in the same window

  • September 29: gpt-6.1-sol launched, with multi-agent delegation to subagents in beta through the Responses API. An "Ultrafast" mode for gpt-6-astra also launched in the Responses API, without regional residency.
  • October 6: the Decisions API entered beta with gpt-6-luna. It turns text and images into typed answers, and OpenAI claims it is "10x faster than the Responses API." Separately, API usage tiers were cut from five to three: Build, Launch and Grow.
  • Agents SDK releases: openai-agents for Python reached 0.23.1 on October 2. @openai/agents for JavaScript reached 0.19.0 on October 5. Check the release notes before upgrading, since several recent minor versions carried migration notes.
  • Reminder: OpenAI's changelog says the Assistants API shut down on August 26, 2026. The replacement is the Responses and Conversations APIs.

What to do next

If you already run agents on the Responses API or the Agents SDK, there's no forced migration. The Agents API is an additional option, not a replacement. To try it, create a key with the scopes above, run the quickstart in an OpenAI-hosted sandbox, and design approval handling first if you plan to use computer use.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on OpenAI & developer tools

The Week in AI

Get the cluster, not just the headline.

0