- A model's weights need roughly parameters × bits per weight ÷ 8 bytes: a 4B model is about 8 GB at 16-bit and about 2–2.5 GB at 4-bit.
- Context costs memory too: at 128K tokens, Llama 3.2 3B's attention cache alone needs about 14 GiB, far more than its 2 GB of weights.
- Good starting points in October 2026: Qwen 3.5 4B, Phi-4-mini and Llama 3.2 3B on 8 GB; Gemma 4 E4B, Qwen 3.5 9B and Ministral 3 8B on 16 GB; gpt-oss-20b on 24–32 GB.
Contents
Running a small model on your own laptop is now easy. Ollama and LM Studio download and run open models with a click or one command. The hard part is choosing one that fits. Pick a model that's too big and it spills out of GPU memory, slows to a crawl, or doesn't load. This explainer shows how to estimate memory before you download, lists current small models with verified sizes and licenses, and explains why vendor benchmark tables can't make the choice for you.
Sizes, context lengths and licenses below come from the model cards, vendor pages and the Ollama and LM Studio catalogs, checked October 8, 2026; licenses and benchmark figures were re-checked on the model cards on October 9, 2026.
Step 1: Work out the memory for the weights
A model's weights are its parameters, stored at some precision. The basic formula is:
memory for weights ≈ number of parameters × bits per weight ÷ 8
- 16-bit (FP16/BF16), the precision most models are released in: 2 bytes per parameter. A 4B model needs about 8 GB.
- 8-bit (GGUF
Q8_0: 8-bit values plus a scale for every 32 weights, so about 8.5 bits): a 4B model needs about 4.2 GB. - 4-bit (GGUF
Q4_K, which Hugging Face's docs list at 4.5 bits per weight): a 4B model needs about 2.2 GB.
Ollama's default downloads are usually 4-bit builds. The estimate holds up well. Meta's Llama 3.2 3B has 3.21 billion parameters, which at 4.5 bits works out to 1.8 GB. Ollama's llama3.2:3b download is 2.0 GB, because common 4-bit formats keep some tensors at higher precision. Downloads for models of about 4B parameters vary by build: Ollama lists phi4-mini:3.8b at 2.5 GB and qwen3.5:4b, which also accepts image input, at 3.3–4.0 GB.
Lower precision saves memory but costs some quality. Below 4 bits (the 3-bit and 2-bit GGUF types), quality usually drops more noticeably for small models. Some vendors train models to tolerate low precision, a technique called quantization-aware training. Google released such versions of Gemma 4 in June 2026, and Apple's on-device model is stored at 2 bits.
Two traps:
- "Effective" parameters. Google's Gemma 4 E4B has 4.5 billion "effective" parameters but 8 billion counting its per-layer embedding tables, according to its model card. Ollama's
gemma4:e4bdownload is 6.6–9.5 GB depending on the build, much larger than "4B" suggests. - Mixture-of-experts (MoE). Gemma 4 26B uses 3.8B active parameters per token out of 25.2B total. It computes like a small model, but all the weights must sit in memory: Ollama lists it at 16–19 GB.
Step 2: Add the context (the part people forget)
While a model reads your prompt and writes its answer, it keeps an attention cache (the "KV cache") for every token in the context. The size is:
KV cache ≈ 2 × layers × KV heads × head dimension × bytes per value × tokens
Take Llama 3.2 3B. Its configuration file lists 28 layers, 8 key-value heads and a head dimension of 128. At 16-bit, that's about 112 KiB per token:
| Context | KV cache (Llama 3.2 3B, 16-bit) | vs. the 2.0 GB 4-bit weights |
|---|---|---|
| 4,096 tokens | ~0.44 GiB | small |
| 32,768 tokens | ~3.5 GiB | larger than the weights |
| 131,072 tokens (its maximum) | ~14 GiB | 7× the weights |
That's why "supports 128K context" doesn't mean your laptop can use 128K. Ollama handles this by setting a conservative default: according to its docs, 4K tokens of context on machines with under 24 GiB of GPU memory, 32K for 24–48 GiB, and 256K above that. You can raise it (see our Ollama guide). Ollama can also store the cache at 8-bit or 4-bit (OLLAMA_KV_CACHE_TYPE=q8_0), roughly halving or quartering it at some cost in quality. Some newer models change this math. Qwen 3.5, for example, mixes Gated DeltaNet linear-attention blocks with full-attention blocks, according to its model card, so check the card rather than assuming the formula applies unchanged.
Step 3: Leave headroom
Your operating system, browser and the inference runtime all need memory too. As a rough rule, don't plan to use more than about two-thirds of total RAM for model plus context on a laptop you're also working on. The hardware matters too:
- Apple silicon Macs share one pool of memory between CPU and GPU, so a 16 GB Mac can run a model that a PC with a 4 GB graphics card can't. Speed is mostly set by memory bandwidth: Apple lists 153 GB/s for the M5, 307 GB/s for the M5 Pro and up to 614 GB/s for the M5 Max. Ollama v0.40.0 now runs supported models on Apple's MLX framework by default on these machines.
- Windows and Linux laptops with an NVIDIA or AMD GPU are fast only while the model fits in the GPU's own memory (VRAM). If it doesn't, Ollama splits it between GPU and CPU.
ollama pswill show something like48%/52% CPU/GPU, and generation slows sharply. - Laptops without a usable GPU run on the CPU and system RAM. That works for 1–4B models, but slowly.
- The NPU (the "AI accelerator" in Copilot+ PCs) generally isn't used by these tools: Ollama's hardware documentation covers CPUs and GPUs but not NPUs, and LM Studio's system requirements don't mention them. Here's what an NPU actually does.
LM Studio's own guidance: 16 GB of RAM recommended on Macs (8 GB "may work" with smaller models and modest context), 16 GB of RAM and at least 4 GB of dedicated VRAM recommended on Windows. Intel Macs aren't supported. Windows on Snapdragon X Elite is.
The shortlist: small open models worth trying (October 2026)
Sizes are Ollama library download sizes. Licenses come from the model cards or vendor pages. Ranges reflect different builds of the same tag.
Last verified: October 8, 2026.
8 GB RAM (or a 4–6 GB GPU)
| Model (Ollama tag) | Download | Context | License | Notes |
|---|---|---|---|---|
Llama 3.2 1B (llama3.2:1b) | 1.3 GB | 128K | Llama 3.2 Community License | Tiny; Meta built it by pruning and distilling larger Llama 3.1 models |
Llama 3.2 3B (llama3.2:3b) | 2.0 GB | 128K | Llama 3.2 Community License | Two years old (September 2024) but still a common baseline |
Qwen 3.5 4B (qwen3.5:4b) | 3.3–4.0 GB | 256K | Apache 2.0 | Text and image input; Qwen 3.5 thinks before answering by default |
Phi-4-mini 3.8B (phi4-mini:3.8b) | 2.5 GB | 128K | MIT | Microsoft; trained heavily on synthetic data; function calling |
Granite 4 3B (granite4:3b) | 2.1 GB | 128K | Apache 2.0 | IBM; aimed at enterprise tasks like RAG and tool calling |
16 GB RAM (or an 8–12 GB GPU)
| Model (Ollama tag) | Download | Context | License | Notes |
|---|---|---|---|---|
Gemma 4 E4B (gemma4:e4b) | 6.6–9.5 GB | 128K | Apache 2.0 | Text, image and audio input per Google's model card |
Qwen 3.5 9B (qwen3.5:9b) | 6.6–7.6 GB | 256K | Apache 2.0 | Strongest vendor-reported scores in this list |
Ministral 3 8B (ministral-3:8b) | 6.0 GB | 256K | Apache 2.0 | Mistral; built for edge use, with image input |
Gemma 4 12B (gemma4:12b) | 7.7–8.0 GB | 256K | Apache 2.0 | Released June 3, 2026 |
24–32 GB RAM (or a 16 GB+ GPU)
| Model (Ollama tag) | Download | Context | License | Notes |
|---|---|---|---|---|
gpt-oss-20b (gpt-oss:20b) | 14 GB | 128K | Apache 2.0 | OpenAI says it runs "with just 16 GB of memory"; reasoning model |
Ministral 3 14B (ministral-3:14b) | 9.1 GB | 256K | Apache 2.0 | |
Gemma 4 26B MoE (gemma4:26b) | 16–19 GB | 256K | Apache 2.0 | 3.8B active parameters, so faster than its size suggests |
Most of these are also in LM Studio's catalog: Gemma 4, Qwen 3.5, Ministral 3, Granite 4, Phi-4 (the 14B model; Phi-4-mini has no separate entry) and gpt-oss all appeared there on October 8, 2026 (re-checked October 9). LM Studio usually offers several quantizations per model, so you can trade quality against size directly.
"Apache 2.0" and "MIT" are standard permissive licenses. The Llama 3.2 Community License is not: it carries an acceptable-use policy, requires "Built with Llama" attribution when you redistribute, and needs a separate license from Meta above 700 million monthly users. Its rights grant for the multimodal Llama 3.2 models excludes people and companies based in the EU. See our open-weight licenses comparison before you build on any of these.
What the benchmarks say, and don't
Every vendor publishes scores, but they aren't comparable across vendors. Each lab runs its own tests with its own prompts and settings. A few vendor-reported examples:
| Model | Vendor-reported scores | Source |
|---|---|---|
| Gemma 4 E2B | MMLU Pro 60.0%; LiveCodeBench v6 44.0% | Google model card |
| Gemma 4 E4B | MMLU Pro 69.4%; LiveCodeBench v6 52.0%; GPQA Diamond 58.6% | Google model card |
| Qwen 3.5 9B | MMLU-Pro 82.5; LiveCodeBench v6 65.6; GPQA Diamond 81.7 | Qwen model card |
| Phi-4-mini | MMLU (5-shot) 67.3; GSM8K (8-shot, chain-of-thought) 88.6 | Microsoft model card |
What these numbers can tell you: within one vendor's family, bigger usually scores higher (Gemma 4 E4B beats E2B on every benchmark in Google's table). What they can't tell you: whether Qwen 3.5 9B is better than Gemma 4 E4B for your task, on your laptop, at the 4-bit precision you'll actually run. Vendors usually report full-precision results, and thinking mode can inflate both scores and response time.
Speed numbers are even less portable. Tokens per second depends on memory bandwidth, GPU, quantization, context length and runtime version. Measure on your own machine. In Ollama, ollama run <model> --verbose prints an "eval rate" in tokens per second after each answer. The API returns eval_count and eval_duration, so tokens per second is eval_count ÷ eval_duration × 10⁹.
A practical way to choose
- Start from your memory. With 8 GB, try a 3–4B model at 4-bit. With 16 GB, an 8–9B model at 4-bit, or a 4B at 8-bit. With 32 GB, try 14–20B models or a MoE.
- Keep the default context at first, then raise it only as far as
ollama psstill shows100% GPU, or no swapping on a Mac. - Try two or three models from different families on ten prompts from your real work (summaries, emails, code snippets, data extraction). Judge the answers, not the leaderboard.
- Check the license for anything you'll ship.
- Re-check every few months. Small models improve quickly, and labs often release them alongside or soon after their flagships. Our analysis of mini-model releases has the record. If you're weighing local models against a cheap API model, see what each really costs.
About this storyBased on the sources linked below. Editorial standards




