Explainer

How to read a system card: a builder's guide to AI safety documentation

System cards run to hundreds of pages. What the sections mean, which safety frameworks they report against, how to read eval tables critically, and a checklist.

By ShajanthanUpdated 7 min read
ByShajanthanFounder & Editor
Published
Reading7 MIN
Annotated evaluation table showing the questions to ask when reading a system card
In 20 seconds
  1. A system card is a vendor's own report on what a model can do, how it was tested and what safeguards ship with it. It's the most detailed public evidence you'll get, but it's still the vendor's account.
  2. Anthropic, OpenAI and Google report against different frameworks (the RSP, the Preparedness Framework and the Frontier Safety Framework), so their risk labels don't map one-to-one.
  3. Read eval tables for setup, not just scores: who ran them, which snapshot, which effort level, whether safeguards were on, and whether the benchmark could be contaminated.
Contents

Every major model launch now comes with a system card: a long document in which the developer reports what the model can do, how it was tested for dangerous capabilities and misbehavior, and what safeguards ship with it. Anthropic's card for Claude Opus 5.5 has a table of contents running past page 220 (Anthropic). Few people read them end to end, and most coverage pulls out a benchmark or two. This guide explains what these documents are, how they're structured, which frameworks they report against, and how to read their evaluation tables without being misled. For a worked example, see what the Claude Opus 5.5 system card actually says.

From model cards to system cards

The idea comes from a 2018 paper by Margaret Mitchell, Timnit Gebru and colleagues. "Model Cards for Model Reporting" recommended "that released models be accompanied by documentation detailing their performance characteristics": evaluations broken down across groups, intended uses and evaluation procedures (arXiv).

Today's frontier labs use two overlapping formats:

  • Model cards, the term Google DeepMind uses, are short, structured overviews. Google's Gemini cards have consistent sections: Model Information, Model Data, Implementation and Sustainability, Distribution, Evaluation, Intended Usage and Limitations, and Ethics and Content Safety (Gemini 3.8 Flash model card).
  • System cards, used by Anthropic and OpenAI, describe the deployed system: the model plus its classifiers, fallbacks and product safeguards. They are much longer and center on catastrophic-risk testing.

The difference matters. A model card tells you about the model. A system card should also tell you what stands between the model and a misuser, and what happens when those safeguards are switched off for testing.

What's inside: the common sections

Section names differ, but the big three cover similar ground.

Last verified: October 8, 2026

TopicAnthropic (Opus 5.5 card)OpenAI (GPT‑6 Sol/Luna October card)Google (Gemini model cards)
Training data and processIntroduction: training data, crowd workersModel Data and TrainingModel Data
Catastrophic-risk evalsRSP evaluations (chemical and biological, AI R&D, alignment risk); separate cyber chapterPreparedness: Capabilities, SafeguardsFrontier Safety section referencing the FSF
Misuse and content safetySafeguards and harmlessnessModel Safety (safe completions, under-18, vision)Ethics and Content Safety
Jailbreaks and prompt injectionAgentic safety; cyber safeguards robustnessRobustness: jailbreaks, multi-turn, prompt injectionVaries
Alignment and honestyAlignment assessmentAlignment (obeying restrictions, avoiding deception)Not a standard section
OtherModel welfareHealth, hallucinationsSustainability, distribution
CapabilitiesCapabilities benchmarksMostly in launch postsEvaluation

Sources: Anthropic, OpenAI, Google DeepMind.

The frameworks behind the risk labels

Every card reports against its company's own safety framework, and the labels aren't interchangeable.

Anthropic: Responsible Scaling Policy (RSP). The current version is 3.4, effective July 8, 2026. Version 3.0 (February 24, 2026) was "a comprehensive rewrite" and introduced periodic Risk Reports that "quantify risk across all our deployed models". Reports were published in February and August 2026 (Anthropic). The RSP page still refers to ASL-3 deployment and security standards. Recent system cards, however, express determinations in capability terms. The Opus 5.5 card, for example, says the model has "CB-1" (non-novel chemical/biological weapons) but not "CB-2" (novel weapons) capability, and it points to the Risk Report for much of the detail.

OpenAI: Preparedness Framework. The version 2 update, published April 15, 2025, tracks three categories: biological and chemical, cybersecurity, and AI self-improvement. It defines two thresholds. "High" capability "could amplify existing pathways to severe harm". "Critical" capability "could introduce unprecedented new pathways to severe harm". Capabilities Reports and Safeguards Reports go to an internal Safety Advisory Group, and leadership makes the final call (OpenAI). Version 2 is still the one OpenAI's cards cite: the October 7, 2026 GPT‑6 Sol/Luna card links to it (checked October 9, 2026). That card treats both models as High in biology/chemistry and cybersecurity, and below High in AI self-improvement (OpenAI).

Google DeepMind: Frontier Safety Framework (FSF). The third version was published September 22, 2025, and updated as FSF 3.1 on April 17, 2026. It defines Critical Capability Levels (CCLs), including harmful manipulation, misalignment (instrumental reasoning) and machine-learning R&D. When relevant CCLs are reached, Google conducts "safety case reviews prior to external launches" (Google DeepMind). Current Gemini model cards report whether a model reached "Tracked or Critical Capability Levels (T/CCLs)". The Gemini 3.8 Flash card, for example, carries over the finding that Gemini 3.7 Flash "did not reach any" of them, evaluated under the April 2026 framework (Gemini 3.8 Flash card).

The key point: each company sets its own thresholds, tests against them, and decides whether they've been crossed. "High" at OpenAI is not "CB-1" at Anthropic, and neither is a Google CCL.

Where regulation comes in

In the EU, general-purpose AI (GPAI) providers have had documentation duties since August 2, 2025. The Commission's enforcement powers, including fines, began on August 2, 2026, and models already on the market before August 2025 must comply by August 2, 2027 (European Commission). Under Article 53, providers must keep technical documentation for authorities, give downstream integrators enough information to understand capabilities and limitations, maintain a copyright policy, and publish a training-content summary (AI Act, Art. 53). Models presumed to carry systemic risk, including those trained with more than 10^25 FLOPs (Art. 51), must also run evaluations including adversarial testing, assess and mitigate systemic risks, report serious incidents and maintain cybersecurity (Art. 55).

Much of that documentation goes to the EU AI Office rather than the public. Public system cards are voluntary and are not the same thing as the regulatory filings. Our EU AI Act explainer covers what applies now.

How to read an eval table critically

Most misreadings come from treating a single number as a fact about the world. Ask these questions:

Who ran it? Most evaluations are vendor-run. Anthropic says "the majority of evaluations of Claude Opus 5.5 were run in-house." Named external testers, such as METR or the U.S. Center for AI Standards and Innovation (CAISI) in Anthropic's case, usually cover specific areas, not the whole card.

Which model, exactly? Cards often test pre-release snapshots. Anthropic notes that "some sections instead use an earlier or alternate snapshot". A number for snapshot X may not describe the model you call.

Were safeguards on? Dangerous-capability tests usually run without production safeguards to measure the underlying model. Misuse tests run with them. OpenAI notes that many of its evaluations run without system-level safeguards and so reflect only baseline model behavior (OpenAI). Both kinds of result are legitimate, but don't mix them up.

How hard did they try (elicitation)? Results depend on reasoning effort, tools, scaffolding and time limits. OpenAI's GPT‑6.1 Sol launch post reports some results at maximum reasoning effort and others at medium (OpenAI). A dangerous-capability test run at low effort can understate risk. See our reasoning-models explainer for how effort changes behavior.

Could the benchmark be contaminated? Public benchmarks leak into training data. OpenAI's October 7 card flags that its ExploitBench results "may be inflated by exposure to historical vulnerabilities in the public benchmark." Anthropic's Opus 5.5 card lists a Humanity's Last Exam blocklist in its appendix. Look for signs like these that the vendor checked.

Is the difference real? Look for confidence intervals. Anthropic reports harmless-response rates with ± ranges and calls some AI R&D scores "not statistically distinguishable". OpenAI describes some multi-turn jailbreak changes as having overlapping confidence intervals. A one-point gain with overlapping intervals is noise.

Is the test set representative? Safety benchmarks are built from hard cases on purpose. OpenAI says its error rates on such sets do not reflect typical production traffic. Good for finding weaknesses; bad for estimating how often you'll hit one.

Does the model know it's being tested? OpenAI's October card notes that models may recognize evaluation settings, which can affect results. That makes alignment numbers harder to interpret, and more capable models make it harder still.

What's missing? Absences tell you something too. Check for missing sections, models released without cards, and conclusions that rely on documents you can't see, such as unpublished internal reports or Risk Reports summarized rather than reproduced.

A reading checklist

  1. Note the date and the model ID. Confirm the card covers the exact model and version you use.
  2. Read the executive summary, then the risk determination. Write down the label (e.g. "High", "CB-1", "no CCL reached") and which framework version it refers to.
  3. For each headline number, record who ran it, which snapshot, safeguards on or off, effort level and the confidence interval.
  4. Find the regressions. Good cards list them. They're often more informative than the improvements.
  5. Read the agentic-safety and prompt-injection sections if you deploy agents. These map most directly to your risk. See the case against unattended agents.
  6. Check external testing. Who tested, what they tested, and whether their findings are quoted or only summarized.
  7. List what's absent. Training compute, parameter counts, unpublished reports, and areas tested only by the vendor.
  8. Separate capability claims from safety claims. Benchmarks in system cards are still vendor-reported. Treat them like any launch claim until someone independent reproduces them.
  9. Re-check after updates. Vendors publish addenda and model updates with their own cards, as OpenAI did on October 7, 2026.

The same discipline applies to vendor research announcements. Our analysis of OpenAI's Navier–Stokes claim shows what it looks like when a strong verifier exists, and how much still depends on the specification.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on System cards & research

The Week in AI

New guides and explainers, every Friday.

0