A free, self-contained picker for AI models that run on your own computer (via Ollama — the free engine that runs them). Tell it what machine you have and what you want to do; the chart, scores, and recommendations adapt. No account, no tracking, nothing leaves this page. Researched & fact-checked 2026-07-18.
1 · Your machine
Two numbers decide everything: your graphics-card memory (VRAM — the AI's fast desk) and your regular RAM (the slower overflow table). Plus one housekeeping check: free disk space — these models are big one-time downloads. Windows: Task Manager → Performance → GPU → "Dedicated GPU memory".
Click Memory → the big number at the top is your RAM
For disk: open File Explorer → This PC — each drive shows "x GB free"
Windows — the copy-paste way
Press the Windows key, type powershell, open it, paste this, press Enter. It prints your three numbers — you can send them to whoever is helping you choose.
If Task Manager shows no "GPU 1" with dedicated memory — or the card says "Intel/AMD Graphics" rather than NVIDIA — pick No graphics card above. Plenty still works; the page will show you what.
💭 What about processor speed, cooling, or drive type?
Processor speed — barely matters
When a model fits on the graphics card, the card does virtually all the work — any desktop processor from the last several years is fine, and a faster one wouldn't make answers arrive faster. It only starts to matter on machines with no graphics card, and even there memory speed counts for more than processor speed.
Cooling — nothing special needed
Running a model works the machine about as hard as playing a 3D game. A normal desktop with stock cooling handles it all day — no overclocking, no extra fans. Thin laptops will get warm and loud on long jobs and may quietly slow themselves down to stay safe. Annoying, not harmful.
Disk — the one to actually check
Every model is a one-time download that lives on your drive (the ones here run ~3–24 GB each; a starter set can total 40–60 GB). Keep 15–20 GB of headroom beyond the downloads so the system has breathing room. Drive type barely matters: an SSD loads models quicker, but answering speed is identical once loaded.
Sends someone this page pre-set to a machine and job list — handy when you're helping a friend choose.
2 · What do you want it to do?
Pick up to 4 jobs (click to toggle), then weight how much each matters. Or grab a preset.
3 · The Magic Quadrant
Four dimensions: ↔ smarts for your chosen jobs · ↕ how nicely it runs on your machine · dot size = download size · color = where it lives. Hover any dot for its report card — click it and its full card opens right here beside the chart (one click brings the levers back). Models that can't run on your machine are hidden (noted below the chart).
Best for this mix
Fits fully on the graphics card — fastRuns from regular RAM — big brain, patientRuns, but slowly
4 · 👉 The recommendation
No hedging: for the machine and jobs you picked above, this is the one to install first — and the small supporting cast to add after. It re-computes live when you change anything above.
5 · The full report card
Scores are 1–10 within this local lineup (10 = the best you can run at home, not the best that exists). Speed and fit columns are computed for the machine you picked above. Click headers to sort — your previous sort becomes the tiebreaker. Darker chip = better.
6 · Model cards — with the exact install command
After installing Ollama (ollama.com), each command goes in a terminal: click Copy, paste, press Enter. It downloads the model and starts a chat. The colored badge shows how each runs on YOUR selected machine. Each card also has an 🛠️ Install prompt button — a ready-made brief to paste into Claude Code so it sets everything up for you.
7 · Want a fresh, personalized write-up?
This page is a snapshot (researched 2026-07-18) — the model world moves monthly. If you use Claude, ChatGPT, or any capable AI, click the button: it builds a ready-to-paste prompt containing your hardware and chosen jobs, asking your AI for an up-to-the-minute version of these model cards and recommendations. Costs the page's author nothing; uses your own AI account. You can also save this whole page with Ctrl+S — it works offline.
8 · Straight talk — read before downloading
🔬 The hardware fine print (how the speed estimates work)
The master rule: answer speed ≈ your memory speed divided by how much the model must read per word. A regular model re-reads its whole file for every word; a "mixture-of-experts" (MoE) model reads only a small active slice — that's why 30-billion-brain MoEs run acceptably from ordinary RAM.
Fitting on the graphics card needs the download size + ~2 GB (for a normal conversation). Long documents cost more: at ~32K words allow 4–5 GB extra. Apps often default to short conversations, hiding this — raise the context setting deliberately.
Speeds by card class (7–9B models, standard compression): 24 GB flagships ≈ 90–135 words/sec · 16 GB mid-range (4060 Ti class) ≈ 35–45 · 8–12 GB ≈ 40–60 · Apple base M chips ≈ 14–24 (M1 is the low end) · M Pro/Max ≈ 30–70.
MoE from RAM (any 6 GB+ card + fast DDR5): ≈ 15–35 words/sec. A regular model spilling into RAM instead: 2–5. No graphics card at all: small 2–4B models ≈ 6–15; a MoE on 32 GB+ RAM ≈ 12–15 — the best no-card experience.
Apple Silicon: roughly 66–75% of unified memory is usable for models — a 64 GB Mac acts like a ~48 GB card, letting it run models no consumer PC card can hold. Trade-off: Macs take noticeably longer to read long documents before starting to answer.
Two cheap wins: RAM running at its rated speed (XMP/EXPO on in the BIOS) is worth ~1.2–1.7×; and two RAM sticks instead of one is worth ~2× (single-stick machines run at half speed). The OS itself needs ~4–6 GB of RAM before any model loads.
All speeds assume the standard compressed version of each model, one user, and moderate-length prompts — treat them as honest ballparks, not lab measurements.
9 · ☁️ Beyond home hardware — the giants
These are the biggest downloadable models on Earth. They're free and legal to download — and none of them will run on any machine in the selector above, or any machine you can buy for a normal budget. Here's what they'd actually take, and the sane way to use them instead.
Built from a fact-checked multi-agent research run (web sweeps → scoring → independent skeptics), 2026-07-18. Capability scores are informed judgments grounded in public benchmarks; entries marked "specs partly estimated" carry unverified sizes. Speed figures are computed ballparks from the fine-print rules. Colors validated for colorblind safety in light and dark mode. Verify any install command against ollama.com/library before running it.