- Anthropic's Claude Opus 5.5 (Sep 22) cut Opus prices to $4/$20 per million tokens. OpenAI's GPT-6 Sol and Luna reached ChatGPT Work and Codex the same day, and SpaceXAI's Grok 4.7 arrived a day earlier.
- Coding tools added all three within about a day: GitHub Copilot listed Opus 5.5, GPT-6 Sol, GPT-6 Luna and Grok 4.7; Cursor shipped Grok 4.7 at launch.
- Cursor launched Rollouts and Security Review, and GitHub previewed local sandboxing in the Copilot app.
Contents
Covering Monday, September 21 to Sunday, September 27, 2026.
This was the week the coding agent market got three new frontier models in roughly 48 hours. Anthropic, OpenAI and SpaceXAI all shipped models pitched at long-running coding and agent work, and the tools developers actually use (GitHub Copilot, Cursor, Claude Code, Codex) switched them on within about a day. The model layer is moving fast enough that the real differences are shifting to harnesses, safety controls and price. Here are the six things worth knowing.
1. Anthropic releases Claude Opus 5.5 and cuts the Opus price (September 22)
On September 22, 2026, Anthropic released Claude Opus 5.5 (claude-opus-5-5), the first model in its Claude 5.5 family. It costs $4 per million input tokens and $20 per million output tokens, down from $5/$25 for Opus 5. It has a 1M-token context window by default and 128k maximum output. A fast mode research preview runs up to 2.5x faster at $8/$40. Anthropic claims Opus 5.5 performs at about the level of Claude Fable 5.1 on most work and is roughly 40% cheaper than Opus 5 on typical workloads. It is available in Claude Code and on AWS, Google Cloud and Microsoft Azure.
Developers should note the breaking changes. Thinking can no longer be switched off (you control it with the effort parameter instead), and forced tool use (tool_choice of any or tool) now returns an error. Anthropic also said, unusually candidly, that it sees signs the model "often suspects it is being evaluated," which complicates its own safety testing.
Why it matters: a 20% cut in Opus's per-token price changes the cost math for long agent sessions, where Opus-class models do much of the work. See our guide to which Claude model to use.
2. GPT-6 Sol and GPT-6 Luna reach ChatGPT Work and Codex (September 22)
The same day, OpenAI's ChatGPT release notes recorded GPT-6 Sol and GPT-6 Luna arriving in ChatGPT Work and Codex. They are the smaller siblings of GPT-6 Astra. OpenAI's September 22 launch post said they were "not yet available in Chat" and made both available in the API as gpt-6-sol and gpt-6-luna. (On October 7, OpenAI made GPT-6 Sol the Chat model for paid ChatGPT tiers and Luna the one for Free and Go.) GitHub added both to Copilot that day: Sol on Pro+, Max, Business and Enterprise, and Luna on those plans plus Pro. GitHub describes Sol as a balanced model for interactive and agentic coding and Luna as the cheapest option in the GPT-6 family.
Why it matters: GPT-6 is now available below the expensive Astra tier, which is where most coding-agent usage actually runs. Our GPT-6 launch coverage has the details. OpenAI's launch post priced GPT-6 Sol at $2 per million input tokens and $10 per million output tokens in the API, and Luna at $0.10/$0.50.
3. SpaceXAI's Grok 4.7 debuts inside Cursor (September 21)
On September 21, SpaceXAI released Grok 4.7, which it calls its "most capable model for coding and knowledge work." It said the model was "available today in Cursor and Grok Build." It costs from $2 per million input tokens and $6 per million output tokens; SpaceXAI says it is "served at the same price and speed as Grok 4.6." SpaceXAI reports 46.3% on CursorBench 4.0, up from 40.4% for Grok 4.6, along with gains on Terminal-Bench 4.0. All of these are vendor benchmarks. GitHub made Grok 4.7 available in Copilot the same day.
Why it matters: Cursor announced on August 14 that SpaceX had completed its acquisition of the company, and Grok is becoming Cursor's house model. Cursor's paid plans promise more included usage for Grok models. That matters if you're weighing Claude Code vs Codex vs Cursor on cost.
4. Cursor ships bots for the "last mile" (September 23)
On September 23, Cursor launched Rollouts and Security Review for Teams and Enterprise plans. Rollouts writes a monitoring plan on each pull request, then checks logs, metrics and traces as the change deploys, flagging regressions per environment. Cursor says it "does not merge or roll back on its own today." Security Review posts one comment per PR about exploitable bugs, such as injection, authorization bypass and exposed secrets. The same day, Cursor engineers said harness changes had cut users' token costs by 7% without lowering agent quality, by Cursor's own measurement.
Why it matters: coding tools are moving past writing code to watching it in production, and Cursor is drawing a clear line at not letting the bot roll back on its own. More in our story on Cursor's remote-control release.
5. GitHub previews a sandbox for Copilot's desktop app (September 22–24)
On September 23, GitHub put local sandboxing for the Copilot desktop app into public preview. It limits an agent session's access to files, network resources and credentials, configured per project or with /sandbox on. The day before, GitHub added OpenTelemetry support so admins can track agent activity in existing monitoring tools. On September 24, it announced a "default policy for new features" for Business and Enterprise accounts, which takes effect October 22. Preview features stay opt-in.
Why it matters: this was groundwork. A week later Copilot gained computer use, and on October 7 the sandbox became generally available. See our story on Copilot's computer use.
6. Also this week
- Anthropic API (September 23–24): cache diagnostics came out of beta, and Anthropic resumed billing for some refusals that happen before any output (categories
bio,frontier_llmandreasoning_extraction). - Google (September 22): Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS became generally available in the Gemini API, with a new Voices endpoint and voice design tools.
- ChatGPT (September 23): Live voice now supports plugins on web, iOS and Android.
What we didn't include
We found no policy or funding story from this week that we could confirm against a primary source and that rose to the level of the items above. If one surfaces, we'll add it here with a correction note.
Last verified: October 9, 2026. Next week: Claude Sonnet 5.5, OpenAI DevDay and Gemini 4 Argon, in The Week in AI #2.
About this storyBased on the sources linked below. Editorial standards




