- Ollama v0.40.0 installs with one command on macOS, Linux or Windows and serves models on localhost:11434 by default.
- You get a CLI, a native REST API and OpenAI- and Anthropic-compatible endpoints, so existing SDK code can point at a local model.
- The local server has no built-in authentication. Researchers have counted more than 175,000 Ollama hosts reachable from the internet, so don't bind it to 0.0.0.0 without a firewall or VPN.
Contents
Ollama is the simplest way to download an open-weight model and run it on your own computer. It wraps model downloads, a local inference server, a command-line chat and an API into one install. This guide covers installation on macOS, Windows and Linux, your first model, calling it from code, customizing it with a Modelfile, and keeping it secure.
How this guide was verified: every command below comes from Ollama's official documentation (docs.ollama.com), its download pages and its GitHub repository, checked on October 8, 2026. The commands were checked against the documentation, not executed, and example output is described from the docs.
Documented against: Ollama v0.40.0, the release marked "Latest" on GitHub on October 8, 2026. Ollama ships updates often, so check ollama -v against the releases page.
Before you start
- Memory. This matters more than anything else. A model has to fit in your GPU memory, or your system RAM if you're on a CPU or an Apple silicon Mac, with room left for the context. A model of about 4 billion parameters at 4-bit quantization is roughly a 2.5–4 GB download in Ollama's library (2.5 GB for
phi4-mini:3.8b, 3.3–4.0 GB forqwen3.5:4b). Our laptop model guide explains the math in detail. - Disk. On Windows the program itself needs at least 4 GB, and Ollama's docs warn that models can take "tens to hundreds of GB."
- Operating system. macOS Sonoma (14) or newer; Windows 10 22H2 or newer; or a mainstream 64-bit Linux (x86-64 or ARM64).
- GPU (optional). Apple silicon uses the GPU automatically. Intel Macs run on the CPU only. On Windows, NVIDIA needs driver 551.61 or newer. Ollama's GPU page lists NVIDIA cards from compute capability 5.0 upward, AMD Radeon cards via ROCm v7, and Vulkan support (on by default on Windows and Linux) as a fallback for other GPUs.
Step 1: Install Ollama
macOS or Linux. Run the official install script:
curl -fsSL https://ollama.com/install.sh | shOn a Mac you can instead download Ollama.dmg from ollama.com/download/mac and drag the app into Applications. On first launch it offers to link the ollama command into /usr/local/bin.
Windows. In PowerShell:
irm https://ollama.com/install.ps1 | iexOr download OllamaSetup.exe. It installs for your user account without administrator rights. To install somewhere else, run OllamaSetup.exe /DIR="d:\some\location".
Linux without the script. The docs also describe a manual install (curl -fsSL https://ollama.com/download/ollama-linux-amd64.tar.zst | sudo tar x -C /usr, plus a ROCm package for AMD GPUs and an arm64 package for ARM). With the manual route you set up your own systemd service.
Piping a script from the internet into your shell gives that script full control. If that bothers you, download install.sh, read it, then run it, or use the manual package.
Check it worked:
ollama -vOn macOS and Windows the app runs the server in the background. On Linux the script sets up a systemd service. If no server is running, start one with ollama serve.
Step 2: Pull and run your first model
Browse models at ollama.com/search. A tag after the colon picks the size. For a first try on a 16 GB laptop, a 3–4 billion parameter model is a safe choice:
ollama pull qwen3.5:4b # about 3.3–4.0 GB depending on build
ollama run qwen3.5:4bollama run downloads the model if needed, then opens an interactive chat. Type /bye to leave. For multi-line input, wrap the text in """. To ask a single question without the interactive session:
ollama run qwen3.5:4b "Explain what a context window is in two sentences."Other small models in the library, with sizes as listed on ollama.com on October 8, 2026:
| Model tag | Download size | Context | License (per model card) |
|---|---|---|---|
llama3.2:1b | 1.3 GB | 128K | Llama 3.2 Community License |
llama3.2:3b | 2.0 GB | 128K | Llama 3.2 Community License |
qwen3.5:4b | 3.3–4.0 GB | 256K | Apache 2.0 |
gemma4:e4b | 6.6–9.5 GB | 128K | Apache 2.0 |
gpt-oss:20b | 14 GB | 128K | Apache 2.0 |
Last verified: October 8, 2026. Licenses differ in important ways: the Llama license, for example, carries an acceptable-use policy and attribution rules. Read our open-weight licenses comparison before you ship anything built on these models.
Some Ollama models can "think" before answering (Qwen 3.5 has thinking on by default, according to its model card). The API's think parameter controls this. ollama show <model> and the /api/show endpoint report which thinking controls a model supports.
Step 3: Manage models
| Task | Command |
|---|---|
| List downloaded models | ollama ls (also ollama list) |
| See what's loaded and on which processor | ollama ps |
| Unload a running model | ollama stop qwen3.5:4b |
| Delete a model | ollama rm qwen3.5:4b |
| Show a model's Modelfile | ollama show --modelfile qwen3.5:4b |
| Server options | ollama serve --help |
ollama ps is the most useful troubleshooting command. Its PROCESSOR column shows 100% GPU, 100% CPU, or a split. A split means the model didn't fit in GPU memory and part of it runs, much more slowly, on the CPU. Models stay loaded for five minutes after the last request by default. You can change that with OLLAMA_KEEP_ALIVE or the per-request keep_alive field.
To see speed, add --verbose to ollama run. After each answer it prints timing statistics, including the prompt eval rate and the generation "eval rate" in tokens per second. This flag is widely used but isn't on Ollama's CLI reference page, so check ollama run --help on your version.
Step 4: Call the model from code
The server listens on http://localhost:11434.
Native API, one-shot generation:
curl http://localhost:11434/api/generate -d '{
"model": "qwen3.5:4b",
"prompt": "Why is the sky blue?",
"stream": false
}'stream defaults to true, which returns a stream of JSON chunks. Setting it to false returns one JSON object. The answer is in response, and the object also includes timing fields: eval_count (output tokens) and eval_duration (nanoseconds). Tokens per second is eval_count / eval_duration × 10⁹.
Native API, chat:
curl http://localhost:11434/api/chat -d '{
"model": "qwen3.5:4b",
"messages": [{"role": "user", "content": "Write a haiku about RAM."}],
"stream": false
}'The answer is in message.content. Useful options include temperature, num_ctx (context length) and num_predict (maximum output tokens).
OpenAI-compatible API. Point any OpenAI SDK at /v1/. The SDK requires an API key, but Ollama ignores it:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama") # key required but ignored
resp = client.chat.completions.create(
model="qwen3.5:4b",
messages=[{"role": "user", "content": "Say this is a test"}],
)
print(resp.choices[0].message.content)Supported endpoints are /v1/chat/completions, /v1/responses (since Ollama v0.13.3), /v1/completions, /v1/models and /v1/embeddings. The docs list gaps: no logprobs, tool_choice, logit_bias or n in Chat Completions; images must be base64, not URLs; and no stateful Responses (previous_response_id). The OpenAI API has no way to set context size, so for that you need a Modelfile (next step). Ollama also exposes an Anthropic-compatible endpoint: point an Anthropic client's base URL at http://localhost:11434.
Official client libraries exist for Python and JavaScript.
Step 5: Customize with a Modelfile
A Modelfile builds a named variant of a model with your own defaults. Save this as Modelfile:
FROM qwen3.5:4b
PARAMETER temperature 0.3
PARAMETER num_ctx 8192
SYSTEM You are a concise technical editor. Answer in plain English.Then:
ollama create editor -f ./Modelfile
ollama run editorOther instructions include TEMPLATE (the prompt template), MESSAGE (seed conversation turns), LICENSE, and REQUIRES (minimum Ollama version). Common PARAMETER names are num_ctx, temperature, top_k, top_p, min_p, repeat_penalty, seed, stop and num_predict.
Step 6: Pick a size that fits, and set the context
Ollama sets the default context length from your GPU memory: 4K tokens under 24 GiB of VRAM, 32K for 24–48 GiB, and 256K at 48 GiB or more, according to its context-length docs. Other doc pages still give older defaults (4,096 in the FAQ; 2,048 for num_ctx on the Modelfile page). Check with ollama ps, whose CONTEXT column shows the value actually in use.
To raise the context for everything:
OLLAMA_CONTEXT_LENGTH=32768 ollama serveOr per request with "options": {"num_ctx": 32768}, or with a Modelfile. A bigger context uses more memory, because the attention cache grows with every token. If ollama ps starts showing a CPU/GPU split, reduce the context or use a smaller model. Setting OLLAMA_KV_CACHE_TYPE to q8_0 (the default is f16) shrinks that cache at some cost in quality. Our laptop guide works through the numbers.
Where models are stored:
| OS | Default location |
|---|---|
| macOS | ~/.ollama/models |
| Linux (service install) | /usr/share/ollama/.ollama/models |
| Windows | C:\Users\%username%\.ollama\models |
Set OLLAMA_MODELS to move them, for example to a larger drive. On Linux, the ollama service user needs read and write access to the new folder (sudo chown -R ollama:ollama <directory>).
How to set environment variables (from Ollama's FAQ):
- macOS app:
launchctl setenv OLLAMA_HOST "127.0.0.1:11434", then restart the Ollama app. - Linux (systemd):
sudo systemctl edit ollama.service, addEnvironment="OLLAMA_HOST=127.0.0.1:11434"under[Service], thensudo systemctl daemon-reload && sudo systemctl restart ollama. - Windows: quit Ollama, edit your user environment variables in Settings, then restart Ollama.
Security: keep the server private
By default Ollama binds to 127.0.0.1:11434, so only programs on your own machine can reach it. The local API has no built-in authentication: Ollama's docs say local requests need no key. If you set OLLAMA_HOST=0.0.0.0:11434, as the FAQ shows to "expose Ollama on your network," anyone who can reach that port can use your models, your hardware and the API's model-management endpoints.
This is a real problem. In January 2026, SentinelOne's SentinelLABS and Censys reported finding 175,108 unique Ollama hosts reachable from the public internet across 130 countries over 293 days of scanning. They said that over 48% advertised tool-calling capabilities. Their conclusion was that such systems "must be treated with the same authentication, monitoring, and network controls" as any other internet-facing service. An earlier Cisco Talos study (September 2025) recommended authentication, firewalls or VPNs, rate limiting and logging.
Practical rules:
- Leave
OLLAMA_HOSTat the default unless you need remote access. - To use Ollama from another machine, tunnel to it instead of opening the port:
ssh -L 11434:localhost:11434 you@your-serverthen usehttp://localhost:11434locally. A private VPN works too. - If you must bind to a network interface, bind to the specific LAN address, not 0.0.0.0. Block port 11434 at the firewall for everything except trusted hosts, and put an authenticating reverse proxy in front for anything shared.
- Be careful with
OLLAMA_ORIGINS. It controls which browser origins can call the API; don't set it to*on a machine that browses the web. - Know what's cloud. Model tags ending in
:cloudrun on Ollama's servers, not yours, and requireollama signin. For strictly local use, setOLLAMA_NO_CLOUD=1.
Troubleshooting
| Symptom | Likely cause and fix |
|---|---|
could not connect to ollama / connection refused | Server isn't running. Start the app or run ollama serve. On Linux: sudo systemctl status ollama. |
| Very slow generation | ollama ps shows CPU or a CPU/GPU split: the model plus context doesn't fit in GPU memory. Use a smaller model or quantization, or a lower num_ctx. |
| Model forgets the start of a long document | Context too short (often 4K by default). Raise num_ctx or OLLAMA_CONTEXT_LENGTH. |
| GPU not used on Linux | Check nvidia-smi (NVIDIA) or ROCm v7 install (AMD). Vulkan can be disabled with OLLAMA_VULKAN=0 if it picks the wrong device. |
| Unstable on a laptop with integrated + discrete GPU (Windows) | Set GGML_VK_VISIBLE_DEVICES to the discrete GPU's index, per Ollama's Windows docs. |
| Downloads fail behind a corporate proxy | Set HTTPS_PROXY for the server. The FAQ warns against HTTP_PROXY. |
| 503 errors under load | Request queue full (OLLAMA_MAX_QUEUE, default 512). Requests are handled one at a time per model by default (OLLAMA_NUM_PARALLEL=1). |
| Where are the logs? | macOS: ~/.ollama/logs/server.log; Windows: %LOCALAPPDATA%\Ollama\server.log; Linux: journalctl -e -u ollama. |
Updating and uninstalling
On Linux, rerun the install script to update. Set OLLAMA_VERSION before the script to pin a specific release. The macOS and Windows apps update themselves. To uninstall, use "Add or remove programs" on Windows. Ollama's macOS and Linux pages list the files and service to remove. Downloaded models are large, so delete the models folder too if you're done.
What's next
- Choosing a model: see how to pick a small model for your laptop. For why labs build these small models at all, see our analysis of mini-model releases.
- Cost: running locally isn't free once you count hardware and your time. Compare it in frontier APIs vs open models: what they really cost.
- Hardware: Ollama runs on CPUs and GPUs. Its GPU documentation doesn't list NPU support, so a laptop's "AI chip" usually sits idle. Here's what an NPU actually does.
About this storyBased on the sources linked below. Editorial standards




