Explainer

How to choose a small AI model for your laptop: the memory math and the current shortlist

Parameters, quantization and context length decide whether a model fits. Here's how to do the arithmetic, and which open models to try at 8, 16 and 32 GB.

By ShajanthanUpdated 7 min read
ByShajanthanFounder & Editor
Published
Reading7 MIN
Bar chart showing how quantization shrinks a 4-billion-parameter model's memory use and how longer context adds memory on top
In 20 seconds
  1. A model's weights need roughly parameters × bits per weight ÷ 8 bytes: a 4B model is about 8 GB at 16-bit and about 2–2.5 GB at 4-bit.
  2. Context costs memory too: at 128K tokens, Llama 3.2 3B's attention cache alone needs about 14 GiB, far more than its 2 GB of weights.
  3. Good starting points in October 2026: Qwen 3.5 4B, Phi-4-mini and Llama 3.2 3B on 8 GB; Gemma 4 E4B, Qwen 3.5 9B and Ministral 3 8B on 16 GB; gpt-oss-20b on 24–32 GB.
Contents

Running a small model on your own laptop is now easy. Ollama and LM Studio download and run open models with a click or one command. The hard part is choosing one that fits. Pick a model that's too big and it spills out of GPU memory, slows to a crawl, or doesn't load. This explainer shows how to estimate memory before you download, lists current small models with verified sizes and licenses, and explains why vendor benchmark tables can't make the choice for you.

Sizes, context lengths and licenses below come from the model cards, vendor pages and the Ollama and LM Studio catalogs, checked October 8, 2026; licenses and benchmark figures were re-checked on the model cards on October 9, 2026.

Step 1: Work out the memory for the weights

A model's weights are its parameters, stored at some precision. The basic formula is:

memory for weights ≈ number of parameters × bits per weight ÷ 8

  • 16-bit (FP16/BF16), the precision most models are released in: 2 bytes per parameter. A 4B model needs about 8 GB.
  • 8-bit (GGUF Q8_0: 8-bit values plus a scale for every 32 weights, so about 8.5 bits): a 4B model needs about 4.2 GB.
  • 4-bit (GGUF Q4_K, which Hugging Face's docs list at 4.5 bits per weight): a 4B model needs about 2.2 GB.

Ollama's default downloads are usually 4-bit builds. The estimate holds up well. Meta's Llama 3.2 3B has 3.21 billion parameters, which at 4.5 bits works out to 1.8 GB. Ollama's llama3.2:3b download is 2.0 GB, because common 4-bit formats keep some tensors at higher precision. Downloads for models of about 4B parameters vary by build: Ollama lists phi4-mini:3.8b at 2.5 GB and qwen3.5:4b, which also accepts image input, at 3.3–4.0 GB.

Lower precision saves memory but costs some quality. Below 4 bits (the 3-bit and 2-bit GGUF types), quality usually drops more noticeably for small models. Some vendors train models to tolerate low precision, a technique called quantization-aware training. Google released such versions of Gemma 4 in June 2026, and Apple's on-device model is stored at 2 bits.

Two traps:

  • "Effective" parameters. Google's Gemma 4 E4B has 4.5 billion "effective" parameters but 8 billion counting its per-layer embedding tables, according to its model card. Ollama's gemma4:e4b download is 6.6–9.5 GB depending on the build, much larger than "4B" suggests.
  • Mixture-of-experts (MoE). Gemma 4 26B uses 3.8B active parameters per token out of 25.2B total. It computes like a small model, but all the weights must sit in memory: Ollama lists it at 16–19 GB.

Step 2: Add the context (the part people forget)

While a model reads your prompt and writes its answer, it keeps an attention cache (the "KV cache") for every token in the context. The size is:

KV cache ≈ 2 × layers × KV heads × head dimension × bytes per value × tokens

Take Llama 3.2 3B. Its configuration file lists 28 layers, 8 key-value heads and a head dimension of 128. At 16-bit, that's about 112 KiB per token:

ContextKV cache (Llama 3.2 3B, 16-bit)vs. the 2.0 GB 4-bit weights
4,096 tokens~0.44 GiBsmall
32,768 tokens~3.5 GiBlarger than the weights
131,072 tokens (its maximum)~14 GiB7× the weights

That's why "supports 128K context" doesn't mean your laptop can use 128K. Ollama handles this by setting a conservative default: according to its docs, 4K tokens of context on machines with under 24 GiB of GPU memory, 32K for 24–48 GiB, and 256K above that. You can raise it (see our Ollama guide). Ollama can also store the cache at 8-bit or 4-bit (OLLAMA_KV_CACHE_TYPE=q8_0), roughly halving or quartering it at some cost in quality. Some newer models change this math. Qwen 3.5, for example, mixes Gated DeltaNet linear-attention blocks with full-attention blocks, according to its model card, so check the card rather than assuming the formula applies unchanged.

Step 3: Leave headroom

Your operating system, browser and the inference runtime all need memory too. As a rough rule, don't plan to use more than about two-thirds of total RAM for model plus context on a laptop you're also working on. The hardware matters too:

  • Apple silicon Macs share one pool of memory between CPU and GPU, so a 16 GB Mac can run a model that a PC with a 4 GB graphics card can't. Speed is mostly set by memory bandwidth: Apple lists 153 GB/s for the M5, 307 GB/s for the M5 Pro and up to 614 GB/s for the M5 Max. Ollama v0.40.0 now runs supported models on Apple's MLX framework by default on these machines.
  • Windows and Linux laptops with an NVIDIA or AMD GPU are fast only while the model fits in the GPU's own memory (VRAM). If it doesn't, Ollama splits it between GPU and CPU. ollama ps will show something like 48%/52% CPU/GPU, and generation slows sharply.
  • Laptops without a usable GPU run on the CPU and system RAM. That works for 1–4B models, but slowly.
  • The NPU (the "AI accelerator" in Copilot+ PCs) generally isn't used by these tools: Ollama's hardware documentation covers CPUs and GPUs but not NPUs, and LM Studio's system requirements don't mention them. Here's what an NPU actually does.

LM Studio's own guidance: 16 GB of RAM recommended on Macs (8 GB "may work" with smaller models and modest context), 16 GB of RAM and at least 4 GB of dedicated VRAM recommended on Windows. Intel Macs aren't supported. Windows on Snapdragon X Elite is.

The shortlist: small open models worth trying (October 2026)

Sizes are Ollama library download sizes. Licenses come from the model cards or vendor pages. Ranges reflect different builds of the same tag.

Last verified: October 8, 2026.

8 GB RAM (or a 4–6 GB GPU)

Model (Ollama tag)DownloadContextLicenseNotes
Llama 3.2 1B (llama3.2:1b)1.3 GB128KLlama 3.2 Community LicenseTiny; Meta built it by pruning and distilling larger Llama 3.1 models
Llama 3.2 3B (llama3.2:3b)2.0 GB128KLlama 3.2 Community LicenseTwo years old (September 2024) but still a common baseline
Qwen 3.5 4B (qwen3.5:4b)3.3–4.0 GB256KApache 2.0Text and image input; Qwen 3.5 thinks before answering by default
Phi-4-mini 3.8B (phi4-mini:3.8b)2.5 GB128KMITMicrosoft; trained heavily on synthetic data; function calling
Granite 4 3B (granite4:3b)2.1 GB128KApache 2.0IBM; aimed at enterprise tasks like RAG and tool calling

16 GB RAM (or an 8–12 GB GPU)

Model (Ollama tag)DownloadContextLicenseNotes
Gemma 4 E4B (gemma4:e4b)6.6–9.5 GB128KApache 2.0Text, image and audio input per Google's model card
Qwen 3.5 9B (qwen3.5:9b)6.6–7.6 GB256KApache 2.0Strongest vendor-reported scores in this list
Ministral 3 8B (ministral-3:8b)6.0 GB256KApache 2.0Mistral; built for edge use, with image input
Gemma 4 12B (gemma4:12b)7.7–8.0 GB256KApache 2.0Released June 3, 2026

24–32 GB RAM (or a 16 GB+ GPU)

Model (Ollama tag)DownloadContextLicenseNotes
gpt-oss-20b (gpt-oss:20b)14 GB128KApache 2.0OpenAI says it runs "with just 16 GB of memory"; reasoning model
Ministral 3 14B (ministral-3:14b)9.1 GB256KApache 2.0
Gemma 4 26B MoE (gemma4:26b)16–19 GB256KApache 2.03.8B active parameters, so faster than its size suggests

Most of these are also in LM Studio's catalog: Gemma 4, Qwen 3.5, Ministral 3, Granite 4, Phi-4 (the 14B model; Phi-4-mini has no separate entry) and gpt-oss all appeared there on October 8, 2026 (re-checked October 9). LM Studio usually offers several quantizations per model, so you can trade quality against size directly.

"Apache 2.0" and "MIT" are standard permissive licenses. The Llama 3.2 Community License is not: it carries an acceptable-use policy, requires "Built with Llama" attribution when you redistribute, and needs a separate license from Meta above 700 million monthly users. Its rights grant for the multimodal Llama 3.2 models excludes people and companies based in the EU. See our open-weight licenses comparison before you build on any of these.

What the benchmarks say, and don't

Every vendor publishes scores, but they aren't comparable across vendors. Each lab runs its own tests with its own prompts and settings. A few vendor-reported examples:

ModelVendor-reported scoresSource
Gemma 4 E2BMMLU Pro 60.0%; LiveCodeBench v6 44.0%Google model card
Gemma 4 E4BMMLU Pro 69.4%; LiveCodeBench v6 52.0%; GPQA Diamond 58.6%Google model card
Qwen 3.5 9BMMLU-Pro 82.5; LiveCodeBench v6 65.6; GPQA Diamond 81.7Qwen model card
Phi-4-miniMMLU (5-shot) 67.3; GSM8K (8-shot, chain-of-thought) 88.6Microsoft model card

What these numbers can tell you: within one vendor's family, bigger usually scores higher (Gemma 4 E4B beats E2B on every benchmark in Google's table). What they can't tell you: whether Qwen 3.5 9B is better than Gemma 4 E4B for your task, on your laptop, at the 4-bit precision you'll actually run. Vendors usually report full-precision results, and thinking mode can inflate both scores and response time.

Speed numbers are even less portable. Tokens per second depends on memory bandwidth, GPU, quantization, context length and runtime version. Measure on your own machine. In Ollama, ollama run <model> --verbose prints an "eval rate" in tokens per second after each answer. The API returns eval_count and eval_duration, so tokens per second is eval_count ÷ eval_duration × 10⁹.

A practical way to choose

  1. Start from your memory. With 8 GB, try a 3–4B model at 4-bit. With 16 GB, an 8–9B model at 4-bit, or a 4B at 8-bit. With 32 GB, try 14–20B models or a MoE.
  2. Keep the default context at first, then raise it only as far as ollama ps still shows 100% GPU, or no swapping on a Mac.
  3. Try two or three models from different families on ten prompts from your real work (summaries, emails, code snippets, data extraction). Judge the answers, not the leaderboard.
  4. Check the license for anything you'll ship.
  5. Re-check every few months. Small models improve quickly, and labs often release them alongside or soon after their flagships. Our analysis of mini-model releases has the record. If you're weighing local models against a cheap API model, see what each really costs.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on Small models & models

The Week in AI

New guides and explainers, every Friday.

0