Explainer

What "agentic" really means in 2026

Every AI product now calls itself agentic. Here's the definition the labs themselves use, how the tool loop works, and how to tell a real agent from a relabeled chatbot.

By ShajanthanUpdated 6 min read
ByShajanthanFounder & Editor
Published
Reading6 MIN
Diagram of an AI agent loop with a human approval step before tool calls
In 20 seconds
  1. Anthropic, OpenAI and Google broadly agree: an agent is a system where the model decides its own next steps and tool calls to reach a goal; a workflow follows steps a developer wrote in advance.
  2. Most of what matters in practice is autonomy: who approves each action, what the agent can reach, and how long it runs unsupervised.
  3. Coding agents (Claude Code, Codex, GitHub's Copilot cloud agent) are the clearest real examples. Gartner calls relabeled chatbots and automation tools 'agent washing.'
Contents

"Agentic" has become the default adjective for AI products, from coding tools to customer-service bots to spreadsheet add-ons. The term does have a useful technical meaning, and the major labs have written it down. This explainer uses their definitions to describe what an agent is, how it differs from a scripted workflow, and why the real question is how much autonomy you hand over. It then applies that to the products where agents are most concrete today, coding agents.

The definitions the labs actually use

Anthropic. In "Building effective agents" (December 19, 2024), Anthropic uses agentic systems as an umbrella term and splits it in two:

  • Workflows: "systems where LLMs and tools are orchestrated through predefined code paths."
  • Agents: "systems where LLMs dynamically direct their own processes and tool usage."

The same post's main advice is restraint: start with the simplest solution, and remember that agentic systems "trade latency and cost for better task performance." Their autonomy also raises the risk of compounding errors, so they should be tested in sandboxes.

OpenAI. Its guide "A practical guide to building agents" says: "Agents are systems that independently accomplish tasks on your behalf." It is explicit about what doesn't count: simple chatbots, single-turn LLM calls and sentiment classifiers are not agents, because the model doesn't control how the workflow runs. It names three components (a model, tools and instructions) and recommends getting the most out of a single agent before splitting work across several.

Google. Google Cloud's explainer defines AI agents as "software systems that use AI to pursue goals and complete tasks on behalf of users," with "reasoning, planning, and memory" and "a level of autonomy." It separates agents (proactive, goal-oriented) from assistants (reactive, where the user makes the decisions) and from bots (rule-following). It traces the core idea to ReAct, a 2022 research paper that interleaved model "reasoning" with "acting" through tools.

The three definitions share a common core: a model, in a loop, choosing its own next action toward a goal. Each one draws the line at whether the model or a developer decides what happens next. Model size and output length don't come into it.

How the loop works

An agent's basic cycle is short:

  1. The model receives a goal, instructions, and a list of tools it may call, such as "run a shell command," "read a file," "search the web" or "open a pull request."
  2. It responds with either a final answer or a tool call, a structured request to run one of those tools with specific arguments.
  3. The surrounding software (the harness) runs the tool, or asks a human first, and returns the result as an observation.
  4. The model reads the observation and decides the next step. The loop repeats until the model stops, a limit is reached, or a human intervenes.

Everything outside step 2 is ordinary software, and that is where safety lives: which tools exist, which calls need approval, and what the tool can reach once it runs. Standards such as the Model Context Protocol define how tools are described and connected, so the same agent can use many of them. If you want to see the loop in code, our guide to building a first agent with the Claude Agent SDK walks through it.

Workflow or agent? A quick test

QuestionWorkflowAgent
Who decides the sequence of steps?The developer, in codeThe model, at run time
Can it call a tool you didn't anticipate in that order?NoYes, within the tools it's given
Does it know when it's done?The code decidesThe model decides (within limits)
Typical failureBreaks on inputs the script didn't foreseeWanders, loops, or takes a wrong action confidently

Workflows aren't a lesser category. Anthropic and OpenAI both recommend them, or even a single well-prompted call, whenever the task is predictable.

Autonomy is the real variable

Calling something "an agent" doesn't tell you how much it does without you. A 2025 paper by K. J. Kevin Feng, David W. McDonald and Amy X. Zhang proposes five levels based on the user's role:

  1. Operator. You drive and the AI helps on request.
  2. Collaborator. You work side by side.
  3. Consultant. The AI leads and asks you for input.
  4. Approver. The AI acts and you sign off on key steps.
  5. Observer. The AI acts on its own and you watch.

The same product can sit at different levels depending on settings. A 2025 paper by Margaret Mitchell and co-authors argues that "risks to people increase with the autonomy of a system" and that fully autonomous agents shouldn't be built. That is a position in a live debate, not a consensus. On the other side, METR has measured the length of software tasks frontier models can complete at a 50% success rate, and found it has doubled roughly every seven months since 2019. That trend is the main argument made for letting agents run longer without supervision.

Real agents you can use today

Last verified: October 8, 2026

  • Claude Code (Anthropic) describes itself as "an agentic coding tool that reads your codebase, edits files, runs commands, and integrates with your development tools." It runs in the terminal, in VS Code and JetBrains, in a desktop app and on the web. It can also be triggered from Slack, GitHub Actions or scheduled "routines." From version 2.1.283, auto mode is the built-in starting permission mode for interactive terminal and VS Code sessions: a second model (a classifier) reviews actions instead of asking you about each one. Non-interactive runs (claude -p, the Agent SDK) generally start in the ask-first Manual mode, and organizations can turn auto mode off. That is roughly the "approver" level, with a machine doing much of the approving.
  • Codex (OpenAI) is available in the ChatGPT apps, as a CLI, as an IDE extension and as Codex Cloud. Locally, a version-controlled folder starts in a workspace-write sandbox with on-request approvals, and network access is off by default.
  • GitHub Copilot cloud agent (GitHub's docs now use this name for what was the Copilot coding agent) takes a task from an issue, a chat or the agents panel and works "in an ephemeral cloud development environment." It can open one pull request per task, and sessions are capped at 59 minutes. It can't approve or merge its own pull request.
  • ChatGPT's general-purpose agent. OpenAI's help center now says the original ChatGPT agent "is no longer available" and points users to ChatGPT Work and a cloud browser. OpenAI says GPT-6 Astra in ChatGPT Work and Codex can operate "the same applications people use every day—even when those applications don't have an API," with confirmation policies before consequential actions.
  • Computer and browser use for developers. Anthropic's computer use toolset left beta on August 19, 2026, and the same day it launched a browser use tool for driving a browser your application hosts.

For a feature-by-feature comparison of the coding tools, see Claude Code vs Codex vs Cursor and our Codex explainer.

Where "agentic" is marketing

In a June 25, 2025 forecast, Gartner used the term "agent washing" for vendors that relabel assistants, robotic process automation (RPA) tools and chatbots as agentic AI without substantial agentic capability. It estimated that only about 130 of the thousands of vendors making the claim are genuine. In the same forecast it predicted that over 40% of agentic AI projects will be canceled by the end of 2027 because of cost, unclear value or weak risk controls.

Questions to ask of any product sold as an agent:

  • Does the model choose the next step, or is it a fixed script with a language model filling in text?
  • What tools can it call, and against which systems: read-only or write?
  • Who approves what? Is there a per-action approval, a classifier, or nothing?
  • What does it run inside? A sandbox, a container, a cloud VM, or your laptop with your credentials?
  • How do you see what it did? Look for logs, diffs, commits and session replays.
  • What happens when it's wrong? Can you roll back, and does it stop, retry or keep going?

If a vendor can't answer the last three, the autonomy is a liability. Our analysis of real agent incidents and how sandboxing works covers what goes wrong when those answers are weak.

The bottom line

"Agentic" isn't a capability tier. It describes who is in control of the next step. The useful questions are narrower and more practical: what can this loop reach, who signs off, and how would you know if it went wrong?

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on Coding agents & agents

The Week in AI

New guides and explainers, every Friday.

0