What “open” actually means here
A model is a very large file of numbers, the weights. With GPT-5 or Claude those numbers
stay on the company’s servers and you rent access through an API. With an open model the
company publishes the file. You download it and run it, and nobody can switch it off, meter
it, or see what you asked.
Open weights is not the same as open source. Almost all of these publish the
weights under their own licence while keeping the training data and code private. Most allow
commercial use; a few restrict it above a size of business. If you are shipping something
commercial, read the licence rather than assuming.
Why it is worth knowing
- Nothing leaves. The obvious one, and the only answer that satisfies a
compliance team without a procurement cycle.
- No per-token cost. The work is free once the file is on disk, so you can
run it over ten thousand rows without watching a meter.
- It cannot be deprecated. A model you have downloaded behaves the same next
year. Hosted models change underneath you.
- It works offline. Genuinely: turn the wifi off and it still answers.
The families worth knowing
| Family | From | Worth it for |
| Llama | Meta | The default all-rounder. The small ones are fast and good enough for most everyday work. |
| Qwen | Alibaba | Strongest small models right now, and the best multilingual coverage. |
| Gemma | Google | Small, efficient, well-behaved. Good on a modest laptop. |
| Mistral | Mistral AI | European, permissive licensing, strong at its size. |
| Phi | Microsoft | Punches far above its weight at reasoning and code for how small it is. |
| DeepSeek | DeepSeek | Reasoning models at a fraction of the usual size. |
| gpt-oss | OpenAI | OpenAI’s open-weight release. The 20B is the largest thing that fits in plain 16 GB. |
Everything published lives on Hugging Face, which is where these actually get released. It is worth half an hour of browsing.
Three ways to run one
| Tool | Suits |
| Ollama | One command, a desktop window if you want one, and it plugs into coding agents. Start here. |
| LM Studio | A proper app with a model browser. Best if you would rather never see a terminal. |
| Jan | Open source, offline by default, looks like a normal chat app. |
Which size fits your machine
Match the model to your memory first, then worry about which one. The download size is
roughly what it needs in RAM while running, and you need headroom for everything else open.
| You have | Run |
| 4 GB | A 1–2B model. Summarising, extraction, tidying text. |
| 8 GB | A 3–4B model, around 2–2.5 GB on disk. The everyday sweet spot. |
| 16 GB | A 7–8B model. Good enough for most daily work. |
| 16 GB and patient | gpt-oss 20B, about 14 GB. Slow, but it runs with no graphics card. |
Start with the smallest one that could work:
{panel("ollama", "Your first model", "1.4 GB download")}
If it will not load you will see model requires more system memory than is
available. It checks free memory, not installed, so closing a browser with forty tabs
genuinely fixes it. In any model picker, avoid anything with :cloud in the name,
which runs on someone else’s servers and defeats the point.
Be honest about what it cannot do
A small local model is meaningfully worse than a frontier model at hard, multi-step
reasoning. It is entirely good enough for pulling fields out of messy text, summarising,
classifying, redacting, and first drafts you were going to rewrite anyway.
The local model filters. The cloud model finishes.
Run the local one over anything sensitive, strip what matters, send only the remainder to
the better model. Most of your data never leaves, and you still get a frontier answer where it
counts. That is the pattern worth stealing, and the one that gets signed off.