AgentsNews

The Week in AI #2: Sonnet 5.5, OpenAI DevDay and a gated Gemini 4

September 28 – October 4, 2026: two $2/$10 workhorse models, OpenAI's always-on dots, Google's cyber-first Gemini 4 Argon, and Copilot learns to use your desktop.

ByShajanthanFounder & Editor
Published
Reading6 MIN
The Week in AI #2 graphic showing a seven-day strip from September 28 to October 4, 2026 with releases marked
What’s new, in 20 seconds
  1. Claude Sonnet 5.5 (Sep 28) and GPT-6.1 Sol (Sep 29) both cost $2/$10 per million tokens, and both companies pitch them as close to their flagship models.
  2. Google's Gemini 4 Argon (Sep 30) went only to vetted cyber defenders at first. OpenAI's DevDay brought always-on 'dots' agents and a $500-a-month Pro 500 plan.
  3. GitHub Copilot gained computer use on Oct 1. NVIDIA launched an open platform to contain misbehaving agents, and Anthropic committed $100M to training 10,000 engineers.
Contents

Covering Monday, September 28 to Sunday, October 4, 2026.

Two patterns stood out this week for anyone building with coding agents. The mid-priced models got much closer to the flagships: Anthropic and OpenAI each shipped a $2/$10-per-million-token model and pitched it as near their top tier. Meanwhile, the most capable new model of the week, Google's Gemini 4 Argon, went only to vetted cybersecurity teams. Agents also got more autonomy (OpenAI's always-on dots, Copilot's computer use), and the industry shipped more tools to contain them. Here are the seven things that mattered.

1. Claude Sonnet 5.5 launches, and Sonnet 4.5 gets a retirement date (September 28 and 30)

On September 28, 2026, Anthropic released Claude Sonnet 5.5 (claude-sonnet-5-5) at $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Anthropic claims it generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task. In Anthropic's own benchmarks it beats Opus 5.5 on Terminal-Bench 4.0 (70.6% vs. 66.4%), although Anthropic says Opus 5.5 "remains clearly stronger at complex, open-ended work." Migrating from Sonnet 5 isn't drop-in: to turn off up-front thinking you now use the new between_tools setting, and forced tool use returns an error. GitHub Copilot added Sonnet 5.5 the same day for Pro plans and up.

On September 30, Anthropic deprecated Claude Sonnet 4.5, which retires from the Claude API on November 30, 2026.

Why it matters: for most agent workloads, Sonnet is the model teams actually pay for. If you're still on Sonnet 4.5, you have until November 30 to move; Anthropic recommends Sonnet 5.5 as the replacement. See which Claude model to use.

2. OpenAI DevDay: GPT-6.1 Sol, Ultrafast and a $500 plan (September 29)

OpenAI held DevDay 2026 in San Francisco on September 29 and published a recap the same day bundling dozens of launches. The headline is GPT-6.1 Sol (gpt-6.1-sol in the API), at $2/$10 per million tokens, one-fifth of GPT-6 Astra's $10/$50. OpenAI says it "nearly matches GPT-6 Astra's intelligence" on agentic coding, computer use and professional work, a vendor claim. GPT-6 Astra Ultrafast generates up to 8x faster in Codex (about 300 tokens per second, per OpenAI) and up to 6x in the API. It's limited to the new Pro 500 plan, which OpenAI's release notes price at $500 a month (the recap describes it as 25 times the ChatGPT Plus allowance), and to Enterprise.

For developers, OpenAI also announced:

  • an Agents API with computer use and multi-agent features;
  • a Decisions API in limited preview;
  • Bedrock Managed Agents, built with Amazon so OpenAI agents run entirely in AWS;
  • a refreshed Codex: cloud environments, a voice-steerable CLI with an /agents view, a new Code Review experience and Codex Security Cloud.

Why it matters: Sonnet 5.5 and GPT-6.1 Sol launched a day apart at identical list prices, so mid-tier pricing is converging. For the developer platform changes, see our story on OpenAI's Agents API. For Codex specifically, see our Codex explainer, and for the price picture behind it, our GPT-6 launch coverage.

3. OpenAI's "dots": agents that never switch off (September 29)

Alongside DevDay, OpenAI introduced dots, "always-on" agents powered by GPT-6 Astra. Each one runs on its own cloud computer and browser and connects to more than 4,000 apps through plugins. They're rolling out to Pro and Business Premium users in eligible markets. Pro access excludes the EEA, Switzerland and the UK at launch, according to OpenAI's release notes. OpenAI says background research uses read-only tools, so dots can't send messages or change app content unprompted. Custom Rules let users allow, require approval for, or block specific actions. Sensitive tasks such as password changes always stay with the user.

Why it matters: this is a consumer-scale bet on persistent, proactive agents, the model our explainer on what "agentic" really means describes. OpenAI's own caveat is that dots "can still make mistakes."

4. Gemini 4 Argon arrives, but only for cyber defenders (September 30)

On September 30, Google DeepMind announced Gemini 4 Argon, its new frontier model for long-horizon software engineering, enterprise knowledge work and cyber defense. It's rolling out first to "trusted cyber defenders" through Google's Fairwind Program. Google says it is working with the US government's voluntary pre-release access process. A wider release will start "with paid API customers and Google AI Ultra subscribers," with no date given. Google is raising the output limit to 1M tokens, up from 64K, and lists introductory pricing of $2/$10 per million tokens, rising to $4/$20 afterward. Google reports 77.9% on DeepSWE v1.1. Google's launch post leads with that single score; the head-to-head table against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 sits on Google DeepMind's Gemini model page. Google published those numbers alongside its own evaluation methodology, so treat them as vendor-reported. Our Argon story goes through it.

Why it matters: Google is now gating its strongest model behind a vetted-access program, much as Anthropic restricts its Mythos models to vetted security users: Project Glasswing participants at launch, and since October 6 the tiers of its expanded Cyber Verification Program. Cyber capability, not just safety in general, is shaping who gets frontier models first. More in our Gemini coverage.

5. GitHub Copilot learns to use your desktop (September 30 – October 1)

On October 1, GitHub put computer use into public preview. Copilot CLI and the Copilot app can now click, type and scroll inside other desktop apps on macOS and Windows. It's off by default and asks before controlling each app. The same day brought dynamic workflows, code-defined multi-agent processes available on all Copilot plans. A day earlier, the HydraFusion research preview, which coordinates several models in one turn, expanded to VS Code and the Copilot app. GPT-6.1 Sol reached Copilot on September 29 for Pro+ plans and up.

Why it matters: Copilot now covers the "no API, no CLI" software that other agents couldn't reach, and GitHub's own docs warn about misclicks and on-screen manipulation. Full details in our story on Copilot's computer use.

6. NVIDIA wants a hardware watchdog for agents (September 28)

On September 28, NVIDIA launched the Open Agent Safety Platform. It pairs OpenShell, open-source runtime software that sets boundaries for agents and traces their actions, with Sentry, a reference design that runs an out-of-band watchdog on BlueField-4 DPUs. NVIDIA says Sentry can quarantine an agent that steps outside its boundaries "in milliseconds." Launch partners include Anthropic, Microsoft, SpaceXAI, CrowdStrike and Palo Alto Networks, and a Linux Foundation-governed Open Secure AI Alliance backs the wider effort. The same day, NVIDIA announced a $150 billion increase to its share repurchase authorization.

Why it matters: agent containment is moving from prompts and permission dialogs down into infrastructure. Our case against unattended agents explains why that layer matters.

7. Anthropic puts $100 million into training customers' engineers (October 2)

On October 2, Anthropic launched Claude Frontier Academy, backed by a $100 million commitment to train 10,000 "Frontier Deployed Engineers" by the end of 2027. The first program pairs a multi-day in-person course and graded practical with a 12-week residency, in which each engineer leads a real Claude project at their own organization. Cohorts are already running in San Francisco, New York and London, with partners including Accenture, Deloitte, McKinsey and Morgan Stanley.

Why it matters: enterprise AI adoption is bottlenecked on people, not just models, and Anthropic is spending to close that gap. See our Anthropic explainer for how the company makes money.

Also this week

  • Claude Code v2.1.287 (October 1) added Claude Mods, which let plugins change deeper behavior, plus a built-in "You should know" mod that flags things you or Claude might miss. More in our Claude Code mods explainer.
  • Anthropic's SDKs (September 30) moved the Admin API out of beta across Python, TypeScript and other languages.
  • The European Commission (September 29) opened a consultation on how technology, including AI, affects copyright.

Since then (October 5–7)

On October 6, Cursor added remote control for local agents from its iOS app. On October 7, Anthropic released Claude Haiku 5.5 (claude-haiku-5-5, 1M-token context) and cut Sonnet 5.5's cache-read price to $0.10 per million tokens. OpenAI began rolling out GPT-6 with Intelligent UI to all ChatGPT users. GitHub made Copilot's local sandboxing generally available, and OpenAI's Codex CLI 0.161.0 made GPT-6.1 Sol its default model. We'll cover these in #3.

Last verified: October 9, 2026. Missed last week? Read The Week in AI #1.

SourcesIntroducing Claude Sonnet 5.5 — Anthropic, Sep 28, 2026 · Claude Platform release notes — Anthropic · DevDay 2026 — OpenAI · DevDay 2026 Recap — OpenAI, Sep 29, 2026 · Introducing GPT-6.1 Sol — OpenAI, Sep 29, 2026 · Introducing dots — OpenAI, Sep 29, 2026 · ChatGPT release notes — OpenAI Help Center · Gemini 4 Argon — Google, Sep 30, 2026 · Gemini model page and benchmark table — Google DeepMind · The latest AI news we announced in September 2026 — Google, Oct 2, 2026 · Fairwind Program — Google · GitHub Copilot can now interact with desktop apps with computer use — GitHub Changelog, Oct 1, 2026 · Dynamic workflows in Copilot CLI and the Copilot app — GitHub Changelog, Oct 1, 2026 · HydraFusion in VS Code and the GitHub Copilot app — GitHub Changelog, Sep 30, 2026 · GitHub Copilot weekly releases — September 28 — GitHub Changelog, Oct 2, 2026 · NVIDIA Launches Open Agent Safety Platform — NVIDIA, Sep 28, 2026 · NVIDIA Announces a $150 Billion Share Repurchase Authorization Increase — NVIDIA, Sep 28, 2026 · Claude Frontier Academy — Anthropic, Oct 2, 2026 · Expanding the Cyber Verification Program — Anthropic, Oct 6, 2026 · Project Glasswing — Anthropic · Model deprecations — Claude Platform docs · Claude Code releases — GitHub · Commission seeks feedback on the effect of technology on copyright — European Commission, Sep 29, 2026 · GPT-6 and Intelligent UI for everyone — OpenAI, Oct 7, 2026 · Local sandboxing for GitHub Copilot now generally available — GitHub Changelog, Oct 7, 2026 · Remote control for local agents — Cursor Changelog, Oct 6, 2026 · openai/codex releases — GitHub

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
TopicCoding agentsCoding agents are AI systems that take a software task, such as fixing a bug or adding a feature, and carry it out… 15 stories, 2 guides, 1 comparisons.Open the hub
Comments
0

More on Coding agents & agents

The Week in AI

Get the cluster, not just the headline.

0