Guide · 11 steps

How to run an open model locally with Ollama (October 2026)

Install Ollama, pull a model, chat with it, call it from code, and keep the server off the public internet. Every command comes from Ollama's official documentation.

By ShajanthanReviewed 8 min read
ByShajanthanFounder & Editor
Published
Reading8 MIN
Versions coveredOllama v0.40.0 (as of Oct 8, 2026)
Diagram of Ollama serving a local model on port 11434 to the command line, apps and scripts on the same computer, with a firewall between it and the internet
In 20 seconds
  1. Ollama v0.40.0 installs with one command on macOS, Linux or Windows and serves models on localhost:11434 by default.
  2. You get a CLI, a native REST API and OpenAI- and Anthropic-compatible endpoints, so existing SDK code can point at a local model.
  3. The local server has no built-in authentication. Researchers have counted more than 175,000 Ollama hosts reachable from the internet, so don't bind it to 0.0.0.0 without a firewall or VPN.
Contents

Ollama is the simplest way to download an open-weight model and run it on your own computer. It wraps model downloads, a local inference server, a command-line chat and an API into one install. This guide covers installation on macOS, Windows and Linux, your first model, calling it from code, customizing it with a Modelfile, and keeping it secure.

How this guide was verified: every command below comes from Ollama's official documentation (docs.ollama.com), its download pages and its GitHub repository, checked on October 8, 2026. The commands were checked against the documentation, not executed, and example output is described from the docs.

Documented against: Ollama v0.40.0, the release marked "Latest" on GitHub on October 8, 2026. Ollama ships updates often, so check ollama -v against the releases page.

Before you start

  • Memory. This matters more than anything else. A model has to fit in your GPU memory, or your system RAM if you're on a CPU or an Apple silicon Mac, with room left for the context. A model of about 4 billion parameters at 4-bit quantization is roughly a 2.5–4 GB download in Ollama's library (2.5 GB for phi4-mini:3.8b, 3.3–4.0 GB for qwen3.5:4b). Our laptop model guide explains the math in detail.
  • Disk. On Windows the program itself needs at least 4 GB, and Ollama's docs warn that models can take "tens to hundreds of GB."
  • Operating system. macOS Sonoma (14) or newer; Windows 10 22H2 or newer; or a mainstream 64-bit Linux (x86-64 or ARM64).
  • GPU (optional). Apple silicon uses the GPU automatically. Intel Macs run on the CPU only. On Windows, NVIDIA needs driver 551.61 or newer. Ollama's GPU page lists NVIDIA cards from compute capability 5.0 upward, AMD Radeon cards via ROCm v7, and Vulkan support (on by default on Windows and Linux) as a fallback for other GPUs.

Step 1: Install Ollama

macOS or Linux. Run the official install script:

BASH
curl -fsSL https://ollama.com/install.sh | sh

On a Mac you can instead download Ollama.dmg from ollama.com/download/mac and drag the app into Applications. On first launch it offers to link the ollama command into /usr/local/bin.

Windows. In PowerShell:

POWERSHELL
irm https://ollama.com/install.ps1 | iex

Or download OllamaSetup.exe. It installs for your user account without administrator rights. To install somewhere else, run OllamaSetup.exe /DIR="d:\some\location".

Linux without the script. The docs also describe a manual install (curl -fsSL https://ollama.com/download/ollama-linux-amd64.tar.zst | sudo tar x -C /usr, plus a ROCm package for AMD GPUs and an arm64 package for ARM). With the manual route you set up your own systemd service.

Piping a script from the internet into your shell gives that script full control. If that bothers you, download install.sh, read it, then run it, or use the manual package.

Check it worked:

BASH
ollama -v

On macOS and Windows the app runs the server in the background. On Linux the script sets up a systemd service. If no server is running, start one with ollama serve.

Step 2: Pull and run your first model

Browse models at ollama.com/search. A tag after the colon picks the size. For a first try on a 16 GB laptop, a 3–4 billion parameter model is a safe choice:

BASH
ollama pull qwen3.5:4b      # about 3.3–4.0 GB depending on build
ollama run qwen3.5:4b

ollama run downloads the model if needed, then opens an interactive chat. Type /bye to leave. For multi-line input, wrap the text in """. To ask a single question without the interactive session:

BASH
ollama run qwen3.5:4b "Explain what a context window is in two sentences."

Other small models in the library, with sizes as listed on ollama.com on October 8, 2026:

Model tagDownload sizeContextLicense (per model card)
llama3.2:1b1.3 GB128KLlama 3.2 Community License
llama3.2:3b2.0 GB128KLlama 3.2 Community License
qwen3.5:4b3.3–4.0 GB256KApache 2.0
gemma4:e4b6.6–9.5 GB128KApache 2.0
gpt-oss:20b14 GB128KApache 2.0

Last verified: October 8, 2026. Licenses differ in important ways: the Llama license, for example, carries an acceptable-use policy and attribution rules. Read our open-weight licenses comparison before you ship anything built on these models.

Some Ollama models can "think" before answering (Qwen 3.5 has thinking on by default, according to its model card). The API's think parameter controls this. ollama show <model> and the /api/show endpoint report which thinking controls a model supports.

Step 3: Manage models

TaskCommand
List downloaded modelsollama ls (also ollama list)
See what's loaded and on which processorollama ps
Unload a running modelollama stop qwen3.5:4b
Delete a modelollama rm qwen3.5:4b
Show a model's Modelfileollama show --modelfile qwen3.5:4b
Server optionsollama serve --help

ollama ps is the most useful troubleshooting command. Its PROCESSOR column shows 100% GPU, 100% CPU, or a split. A split means the model didn't fit in GPU memory and part of it runs, much more slowly, on the CPU. Models stay loaded for five minutes after the last request by default. You can change that with OLLAMA_KEEP_ALIVE or the per-request keep_alive field.

To see speed, add --verbose to ollama run. After each answer it prints timing statistics, including the prompt eval rate and the generation "eval rate" in tokens per second. This flag is widely used but isn't on Ollama's CLI reference page, so check ollama run --help on your version.

Step 4: Call the model from code

The server listens on http://localhost:11434.

Native API, one-shot generation:

BASH
curl http://localhost:11434/api/generate -d '{
  "model": "qwen3.5:4b",
  "prompt": "Why is the sky blue?",
  "stream": false
}'

stream defaults to true, which returns a stream of JSON chunks. Setting it to false returns one JSON object. The answer is in response, and the object also includes timing fields: eval_count (output tokens) and eval_duration (nanoseconds). Tokens per second is eval_count / eval_duration × 10⁹.

Native API, chat:

BASH
curl http://localhost:11434/api/chat -d '{
  "model": "qwen3.5:4b",
  "messages": [{"role": "user", "content": "Write a haiku about RAM."}],
  "stream": false
}'

The answer is in message.content. Useful options include temperature, num_ctx (context length) and num_predict (maximum output tokens).

OpenAI-compatible API. Point any OpenAI SDK at /v1/. The SDK requires an API key, but Ollama ignores it:

PYTHON
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama")  # key required but ignored
resp = client.chat.completions.create(
    model="qwen3.5:4b",
    messages=[{"role": "user", "content": "Say this is a test"}],
)
print(resp.choices[0].message.content)

Supported endpoints are /v1/chat/completions, /v1/responses (since Ollama v0.13.3), /v1/completions, /v1/models and /v1/embeddings. The docs list gaps: no logprobs, tool_choice, logit_bias or n in Chat Completions; images must be base64, not URLs; and no stateful Responses (previous_response_id). The OpenAI API has no way to set context size, so for that you need a Modelfile (next step). Ollama also exposes an Anthropic-compatible endpoint: point an Anthropic client's base URL at http://localhost:11434.

Official client libraries exist for Python and JavaScript.

Step 5: Customize with a Modelfile

A Modelfile builds a named variant of a model with your own defaults. Save this as Modelfile:

CODE · Modelfile
FROM qwen3.5:4b
PARAMETER temperature 0.3
PARAMETER num_ctx 8192
SYSTEM You are a concise technical editor. Answer in plain English.

Then:

BASH
ollama create editor -f ./Modelfile
ollama run editor

Other instructions include TEMPLATE (the prompt template), MESSAGE (seed conversation turns), LICENSE, and REQUIRES (minimum Ollama version). Common PARAMETER names are num_ctx, temperature, top_k, top_p, min_p, repeat_penalty, seed, stop and num_predict.

Step 6: Pick a size that fits, and set the context

Ollama sets the default context length from your GPU memory: 4K tokens under 24 GiB of VRAM, 32K for 24–48 GiB, and 256K at 48 GiB or more, according to its context-length docs. Other doc pages still give older defaults (4,096 in the FAQ; 2,048 for num_ctx on the Modelfile page). Check with ollama ps, whose CONTEXT column shows the value actually in use.

To raise the context for everything:

BASH
OLLAMA_CONTEXT_LENGTH=32768 ollama serve

Or per request with "options": {"num_ctx": 32768}, or with a Modelfile. A bigger context uses more memory, because the attention cache grows with every token. If ollama ps starts showing a CPU/GPU split, reduce the context or use a smaller model. Setting OLLAMA_KV_CACHE_TYPE to q8_0 (the default is f16) shrinks that cache at some cost in quality. Our laptop guide works through the numbers.

Where models are stored:

OSDefault location
macOS~/.ollama/models
Linux (service install)/usr/share/ollama/.ollama/models
WindowsC:\Users\%username%\.ollama\models

Set OLLAMA_MODELS to move them, for example to a larger drive. On Linux, the ollama service user needs read and write access to the new folder (sudo chown -R ollama:ollama <directory>).

How to set environment variables (from Ollama's FAQ):

  • macOS app: launchctl setenv OLLAMA_HOST "127.0.0.1:11434", then restart the Ollama app.
  • Linux (systemd): sudo systemctl edit ollama.service, add Environment="OLLAMA_HOST=127.0.0.1:11434" under [Service], then sudo systemctl daemon-reload && sudo systemctl restart ollama.
  • Windows: quit Ollama, edit your user environment variables in Settings, then restart Ollama.

Security: keep the server private

By default Ollama binds to 127.0.0.1:11434, so only programs on your own machine can reach it. The local API has no built-in authentication: Ollama's docs say local requests need no key. If you set OLLAMA_HOST=0.0.0.0:11434, as the FAQ shows to "expose Ollama on your network," anyone who can reach that port can use your models, your hardware and the API's model-management endpoints.

This is a real problem. In January 2026, SentinelOne's SentinelLABS and Censys reported finding 175,108 unique Ollama hosts reachable from the public internet across 130 countries over 293 days of scanning. They said that over 48% advertised tool-calling capabilities. Their conclusion was that such systems "must be treated with the same authentication, monitoring, and network controls" as any other internet-facing service. An earlier Cisco Talos study (September 2025) recommended authentication, firewalls or VPNs, rate limiting and logging.

Practical rules:

  1. Leave OLLAMA_HOST at the default unless you need remote access.
  2. To use Ollama from another machine, tunnel to it instead of opening the port: ssh -L 11434:localhost:11434 you@your-server then use http://localhost:11434 locally. A private VPN works too.
  3. If you must bind to a network interface, bind to the specific LAN address, not 0.0.0.0. Block port 11434 at the firewall for everything except trusted hosts, and put an authenticating reverse proxy in front for anything shared.
  4. Be careful with OLLAMA_ORIGINS. It controls which browser origins can call the API; don't set it to * on a machine that browses the web.
  5. Know what's cloud. Model tags ending in :cloud run on Ollama's servers, not yours, and require ollama signin. For strictly local use, set OLLAMA_NO_CLOUD=1.

Troubleshooting

SymptomLikely cause and fix
could not connect to ollama / connection refusedServer isn't running. Start the app or run ollama serve. On Linux: sudo systemctl status ollama.
Very slow generationollama ps shows CPU or a CPU/GPU split: the model plus context doesn't fit in GPU memory. Use a smaller model or quantization, or a lower num_ctx.
Model forgets the start of a long documentContext too short (often 4K by default). Raise num_ctx or OLLAMA_CONTEXT_LENGTH.
GPU not used on LinuxCheck nvidia-smi (NVIDIA) or ROCm v7 install (AMD). Vulkan can be disabled with OLLAMA_VULKAN=0 if it picks the wrong device.
Unstable on a laptop with integrated + discrete GPU (Windows)Set GGML_VK_VISIBLE_DEVICES to the discrete GPU's index, per Ollama's Windows docs.
Downloads fail behind a corporate proxySet HTTPS_PROXY for the server. The FAQ warns against HTTP_PROXY.
503 errors under loadRequest queue full (OLLAMA_MAX_QUEUE, default 512). Requests are handled one at a time per model by default (OLLAMA_NUM_PARALLEL=1).
Where are the logs?macOS: ~/.ollama/logs/server.log; Windows: %LOCALAPPDATA%\Ollama\server.log; Linux: journalctl -e -u ollama.

Updating and uninstalling

On Linux, rerun the install script to update. Set OLLAMA_VERSION before the script to pin a specific release. The macOS and Windows apps update themselves. To uninstall, use "Add or remove programs" on Windows. Ollama's macOS and Linux pages list the files and service to remove. Downloaded models are large, so delete the models folder too if you're done.

What's next

Did this guide work for you?

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on Open models & open source

The Week in AI

New guides and explainers, every Friday.

0