Explainer

What an NPU actually does, and why the TOPS number on your laptop matters less than you think

Copilot+ PCs need a 40-TOPS 'AI engine'. Here's what that chip accelerates, which apps really use it, and why tools like Ollama mostly ignore it.

By ShajanthanUpdated 7 min read
ByShajanthanFounder & Editor
Published
Reading7 MIN
Diagram of a laptop chip with CPU, GPU and NPU sharing memory, showing that Windows AI features use the NPU while local chatbots like Ollama use the GPU
In 20 seconds
  1. An NPU is a low-power accelerator for the low-precision matrix math inside neural networks. Microsoft requires 40+ TOPS of NPU performance for Copilot+ PC features like Recall and Windows Studio Effects.
  2. TOPS is a peak, best-case figure. On Intel's newest laptop chips, the GPU is rated higher than the NPU (122 vs 50 TOPS on the top model).
  3. Popular local-LLM tools such as Ollama run on the GPU and CPU. The NPU is used mainly by Windows' built-in AI features and apps written for Windows ML, Core ML or vendor SDKs.
Contents

Almost every new laptop now advertises a neural processing unit (NPU) and a TOPS figure next to it. Microsoft has made one a requirement: its Copilot+ PC label demands an NPU that can do more than 40 trillion operations per second. But what does that chip do? Will it make the AI models you run faster? And should it affect what you buy? This explainer covers what NPUs accelerate, why TOPS is a weak way to compare them, which software uses them, and what that means if you want to run small models locally.

What an NPU is for

Running a neural network is mostly multiplying large grids of numbers (matrices) together and adding up the results. A CPU can do this, but it's built for flexible, branching code. A GPU is much better: thousands of simple cores doing the same operation in parallel. An NPU narrows the job further. It's a block of silicon dedicated to these multiply-and-add operations, usually at low precision, meaning 8-bit integers (int8) or smaller instead of the 16- or 32-bit floating-point numbers used in training. Lower precision means smaller numbers to move and simpler circuits to compute them, so more work per watt.

Microsoft's developer documentation defines the NPU as "a specialized computer chip for AI-intensive processes like real-time translations and image generation." The selling point is efficiency, not raw speed. An NPU can run a model continuously, such as blurring your background on every video frame or transcribing speech, without spinning up the GPU and draining the battery.

That's why NPU workloads tend to be small, always-on models: camera effects, noise suppression, live captions, on-device search indexing, and small language models of a few billion parameters. Apple, for example, says its on-device foundation model has about 3 billion parameters, stored at 2 bits per weight. Shrinking a model like that is what makes it a candidate for this kind of hardware. Our analysis of small "mini" models explains why labs build them.

The Copilot+ PC bar: 40+ TOPS

Microsoft's documentation says "many of the new Windows AI features require an NPU with the ability to run at 40+ TOPS." Its business page for Copilot+ PCs adds a storage requirement of at least 256 GB, with 50 GB free for Recall. It lists these features as exclusive to Copilot+ PCs:

  • Recall (preview): a searchable timeline of what you've seen on screen. It also requires Windows Hello Enhanced Sign-in Security.
  • Click to Do: contextual actions on whatever is on screen.
  • Live Captions with real-time translation. Basic Live Captions works on all Windows 11 PCs.
  • Windows Studio Effects, which Microsoft says requires "a qualified Neural Processing Unit."

Microsoft notes that features "vary by device and market." On Windows, Task Manager shows NPU usage alongside CPU and GPU on devices that have one, which is the easiest way to see whether something is actually using it.

Today's NPUs, by the vendors' numbers

Last verified: October 8, 2026. All figures are vendor-published peak numbers, usually for int8.

ChipNPU peak TOPSAlso on the chipSource
Qualcomm Snapdragon X2 Elite80 (two SKUs list 85)Qualcomm says "up to 80" in its marketing copyQualcomm product brief
Intel Core Ultra X9 388H (Series 3, "Panther Lake")50Arc B390 GPU: 122 TOPSIntel ARK
Intel Core Ultra 5 322 (Series 3)46GPU: 18 TOPSIntel ARK
Intel Core Ultra 7 258V (Series 2, Lunar Lake)47GPU: 64; platform total 115Intel ARK
AMD Ryzen AI 9 HX 470 (Ryzen AI 400)55"Overall" 86 TOPSAMD
AMD Ryzen AI 9 HX 370 (Ryzen AI 300)50"Overall" 80 TOPSAMD
Apple M5 / M5 Pro / M5 MaxNot published16-core Neural Engine, plus a "Neural Accelerator" in every GPU coreApple

Microsoft's NPU page lists first-generation Snapdragon X, AMD Ryzen AI 300 and Intel Core Ultra 200V machines among Copilot+ PCs. Apple has stopped quoting a TOPS figure for its Neural Engine. With the M5, its headline AI claim is about the GPU: Apple says "over 4x the peak GPU compute performance for AI compared to M4," based on its own testing.

Why TOPS is a weak metric

TOPS (trillions of operations per second) is easy to print on a sticker and hard to interpret:

  1. It's a peak, not what you'll actually get. It assumes every multiply unit is busy every cycle. Real models spend time waiting on memory, converting data formats, or running operations the NPU doesn't support, which fall back to the CPU.
  2. Precision and fine print vary. Vendors usually quote int8. A figure at int4, or one that counts "sparse" operations (skipping zeros), can look much bigger for the same hardware. Qualcomm's own X2 Elite brief says "up to 80" in one place and lists 85 for two models in its spec table.
  3. Memory bandwidth limits language models. When a model generates text, each new token requires reading essentially all of its weights from memory. Token speed is therefore often set by memory bandwidth, not compute. That's why Apple advertises bandwidth (153 GB/s on the M5, up to 614 GB/s on the M5 Max) and why a high NPU TOPS figure doesn't guarantee fast chat.
  4. The NPU isn't necessarily the fastest AI engine on the chip. Intel lists its top Series 3 chip's GPU at 122 int8 TOPS against 50 for its NPU. AMD and Intel also quote "overall" or "platform" figures that add CPU, GPU and NPU together, which no single workload gets.
  5. Software decides. An NPU only helps if the app was built for it, which leads to the most important point.

Which software actually uses the NPU

Windows' own features. Studio Effects, Recall, Click to Do and live translation are designed for the NPU. Microsoft says these features "ship in the latest releases of Windows and are available via APIs in Microsoft Foundry on Windows."

Phi Silica, and its successor. Microsoft's built-in small language model, Phi Silica, runs on the NPU on Copilot+ PCs. Microsoft says it uses speculative decoding, where a small draft model proposes tokens that the main model checks in parallel. It can also run on recent NVIDIA and AMD GPUs via an experimental path. Developer access goes through the Windows App SDK and is a "Limited Access Feature." It isn't available in China. Microsoft's documentation, updated October 2, 2026, says Phi Silica is being replaced by a new on-device model, Aion Instruct: it reaches Windows Insider devices in November 2026 and retail devices in January 2027, when Phi Silica will be removed.

Apps built on Windows ML. Microsoft now recommends Windows ML, which is built on ONNX Runtime, as the way to run models on the NPU, replacing its earlier DirectML recommendation. Windows ML downloads a vendor "execution provider" for each chip: QNN for Qualcomm's Hexagon NPU, VitisAI for AMD's NPU, OpenVINO for Intel. NPU support requires Windows 11 24H2 or later.

Vendor toolkits. Intel's OpenVINO GenAI has documentation for running language models on its NPU. AMD's Ryzen AI Software (version 1.8) runs language models through an ONNX Runtime GenAI flow on its Strix and Krackan Point chips. Its release notes list support for Ryzen AI 300 and Max 300 chips but don't yet mention Ryzen AI 400.

Apple's Core ML. Apple says Core ML "optimizes on-device performance by leveraging the CPU, GPU, and Neural Engine." Developers can let the system choose (.all) or restrict a model to CPU and Neural Engine (.cpuAndNeuralEngine). The operating system decides where each part of a model runs.

What usually doesn't: popular local-LLM tools. If you run open models with Ollama, your NPU most likely sits idle. Ollama's hardware documentation covers NVIDIA GPUs, AMD GPUs via ROCm, Apple's Metal (and, in v0.40.0, Apple's MLX framework), and Vulkan, but no NPUs. LM Studio's system requirements don't mention NPUs either. llama.cpp, the engine underneath many of these tools, lists a Hexagon backend for Snapdragon chips, an OpenVINO backend for Intel CPUs, GPUs and NPUs marked "in progress," and a CANN backend for Huawei's Ascend NPUs. AMD's own docs say llama.cpp on Ryzen AI runs on the integrated GPU only. Running local language models on laptop NPUs is possible through vendor toolkits, but it's not yet the default path in the tools most enthusiasts use.

What this means if you're buying a laptop

  • If you want Windows' AI features (Studio Effects in video calls, Recall, live translated captions), you need a Copilot+ PC, so you need an NPU of 40+ TOPS. Every current Snapdragon X2, Intel Core Ultra Series 3 and AMD Ryzen AI 400 chip listed above clears that bar. Beyond 40, small TOPS differences are unlikely to matter for these features.
  • If you want to run local language models with Ollama or LM Studio, memory matters more than the NPU. Prioritize RAM (16 GB minimum, 32 GB if you can), memory bandwidth, and a capable GPU or Apple silicon. Our guide to choosing a small model for your laptop explains the memory math.
  • If battery life during AI tasks matters, an NPU is the efficient way to run continuous, small workloads, but only in apps written to use it. Check whether the specific app you care about supports your chip's NPU.
  • Don't compare TOPS across brands. The numbers aren't measured the same way, and Apple doesn't publish one. Look for independent tests of the actual apps you use.

Confirmed vs uncertain: the Copilot+ requirements, chip figures and software support above come from Microsoft, Qualcomm, Intel, AMD, Apple, Ollama and llama.cpp documentation. How much a given NPU speeds up a given app, and how much battery it saves, depends on the app. Vendors don't publish comparable figures, so look for independent measurements of the specific app you care about.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on Small models & hardware

The Week in AI

New guides and explainers, every Friday.

0