What Local AI Models Can a Regular Computer Run?
Token generation speeds were measured for Gemma 4 E2B, Gemma 4 12B, gpt-oss 20B, Gemma 4 26B, Qwen 3.6 35B, and Qwen 3.8 27B using Ollama on a Geekom A9 Max mini PC with 32 GB RAM.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Qwen 3.6 35BModel source: provider262,144 tokens max native context | 35Bactive 3B | 21 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| DDR5 32GB Computer | 32 GBRAM | 89.6 GB/s | no | 20 tok/s4-bit quantization |
| Mac mini M4 | 24 GBUnified RAM | 120 GB/s | no | 26 tok/s4-bit quantization |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 54 tok/s4-bit quantization |
| M4 Pro | 24 GBUnified RAM | 273 GB/s | no | 58 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 58 tok/s4-bit quantization |
| M5 Pro | 64 GBUnified RAM | 307 GB/s | no | 64 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 117 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 148 tok/s4-bit quantization |
| RTX 3090 Ti | 24 GBGPU VRAM | 1008 GB/s | yes | 173 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 197 tok/s4-bit quantization |
| RTX 5090 | 32 GBGPU VRAM | 1792 GB/s | yes | 256 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 256 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 417 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 475 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 475 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
Token generation speeds were measured for Gemma 4 E2B, Gemma 4 12B, gpt-oss 20B, Gemma 4 26B, Qwen 3.6 35B, and Qwen 3.8 27B using Ollama on a Geekom A9 Max mini PC with 32 GB RAM.
Demonstrates offloading experts for Qwen 3.6 35B-A3b on llama-server with 8GB VRAM, covering 12GB corrections, layer offload sweeps, and Bob-bench Wide evaluations.
Demonstrates optimizing a local voice agent in Pithagoras using Qwen3.6-35B-A3B, Breeze-TTS-2, Whisper, and Silero VAD on an RTX 3060 via overlapping pipelines and quantization.
Tests FreeToken against llama-server baseline using OpenCode on 12GB VRAM to evaluate performance constraints.
Demonstrates Qwen3.6-35B-A3B inference on RTX 3060 using llama.cpp MoE expert caching combined with speculative decoding, achieving 70-80 tok/s via --moe-cache-profile flags.
Evaluates Qwen 35B A3B against Ornith 1.5 using llama-server, pi harness, Bob-Bench, and fuzzing across decode speed, wall-clock time, and pareto efficiency graphs.
Qwen 3.8-27B versus Qwen 3.6-35B-A3B at BF16 on a custom exam featuring Tetris, volcano physics, spreadsheets, and SVG tasks, plus Claude Opus 4.6 comparison.
Meta Muse Glimmer, Qwen3.6 35B-A3B, Qwen3-Coder 30B, and GPT-OSS 20B compete in 10 frozen events on GMKtec EVO-X2 via Ollama on ROCm, judging speed and accuracy across tasks like coding, image generat
Compares Qwen 3.6 35B A3B, Qwen 3.6 27B dense, and Qwen 3.6 Fable Fusion by having each recreate the Solar Fortress arcade game using an Arcade Template.
Tests Grug 35B QAT Q4 against Qwen 35B A3B Q4 using Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender, and Godot benchmarks on a 16GB local system.
Regina test set results compare Gemma 26B and Qwen 35B using IQ4 quants from unsloth, identical sampling parameters, prompts, harness, and llama-server versions on Nvidia GPUs.
Qwen 3.6 35B A3B and Gemma 4 26B A4B are tested on updating documentation and verifying tests for mature legacy code using the swing-extras project.
Evaluates KAT Coder V2.5 Dev against Qwen 3.6 35B A3B using performance, memory, agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot tests.
Evaluates Laguna XS 2.1 33B A3B against Qwen 3.6 35B A3B using Performance, Memory, Agency, OpenAI Human Eval, Sand Physics, Kanban, Dungeon Crawler, Blender, and Godot tests.
Compares Laguna S 2.1 and Qwen 3.6 35B-A3B on an RTX 3060 using llama.cpp, evaluating debugging accuracy and animation generation tasks via manual testing.
Compares Laguna XS 2.1 and Qwen3.6 on Apple Silicon via setup, speed, token throughput, response quality, reasoning, and tool-use accuracy for real agentic coding tasks.
Tests Laguna S 2.1's real-world coding capabilities and benchmarks performance against GLM 5.2, Qwen 3.7 Max, Hy3, and Kimi K3 using Woaibench.
Evaluates Qwen 3.6 14B A3B FableVibes against Qwen3.6-35B-A3B using Performance, Memory, Agency, OpenAI Human Eval, Sand Physics and Dungeon Crawler tests on a 16GB local setup.
Tests Bonsai 27B binary and ternary variants against Qwen 3.6 35B-A3B and Gemma 4 12B on an RTX 3060 using llama.cpp, measuring generation speed, total memory footprint, and performance on web design
Compares Qwen 3.7 Max and Qwen 35B A3B using Performance, Memory, Agency, Driving Game, Kanban, and Codebase Prompts tests.
Qwen 3.6 and Gemma 4 variants are tested to design and build a full-screen visualizer using MusicPlayer, with results hosted in the ext-mp-ai-visualizers repo.
Compares Intern Science Agents A1 35B MoE against Qwen 3.6 35B A3B via Performance, Memory, Agency, OpenAI Human Eval, Sand Physics and Dungeon Crawler tests on a 16GB local setup.
Demonstrates building a local deep research agent with Qwen models, implementing sub-agent delegation and tool quotas on AMD Ryzen AI Max Series and Radeon AI PRO R9700 GPUs.
Demonstrates configuring local LLMs like Qwen3.6 MTP and Gemma 4 with GitHub Copilot, Codex, and Claude Code on LM Studio or Ollama, plus setting system prompts in LM Studio to prevent looping.
Claude Fable 5 optimized llama.cpp for Qwen3.6 35B-A3B on an RTX 3060, achieving a 64.5% prefill speedup via four patches, with results verified through self-benchmarking.
Compares QWEN3.6-35B-A3B inference speed with and without Multi-Token Prediction on an Apple M5 Max using Pi Coding Agent and LM Studio.
Compares Ornith 1.0 35B and Qwen3.6-35B-A3B-GGUF on a 16GB VRAM system using Driving Game, Fake Desktop, and Agent Maze coding tests.
Compares Ornith 35B and Qwen 3.6 35B-A3B using llama.cpp, llama-swap, and OpenCode on dual RTX 3090s while building a street-racing car OS, race-control interface, and live race simulator UI.
Compares Ornith 1 35B MoE against Qwen 3.6 35B A3B using Performance, Memory, Agency, OpenAI Human Eval, Sand Physics and Dungeon Crawler tests on a 16GB VRAM system.
Tests Qwen-AgentWorld-35B-A3B against Qwen3.6-35B-A3B using Performance, Memory, Agency, OpenAI Human Eval, Sand Physics, and Dungeon Crawler benchmarks on a system with 16GB VRAM and 32GB DDR4 RAM.
Cohere Labs North Mini and Qwen 3.6 35B A3B are tested via 2D Driving Game, Fake Desktop, and Slime Maze tasks on a local system with 16GB VRAM and 32GB DDR4 RAM.
Connects local Qwen3 35B via LM Studio with TestSprite MCP Server in VS Code to execute complete front-end automation tests entirely offline.
Compares inference speed of QWEN3.6-35B-A3B with and without Multi-Token Prediction on Apple M5 Max using Pi Coding Agent and LM Studio.
Gemma 4 12B with QAT and QWEN 3.6 are tested as Coding Agents in VS Code on an M5 Max (128GB) using real-world coding tasks instead of benchmarks or synthetic tests.
Tests Qwen 3.6 35B A3B versus Qwopus 3.6 35B A3B using an Expense Tracker, Card Matching Game, and Breakout Game coding test suite on a system with 16GB VRAM and 32GB DDR4 RAM.
Tests Gemma 4 26B A4B versus Qwen 3.6 35B A3B on a 16GB VRAM system using a six-test suite covering adherence, memory, implementation, architecture, reasoning, and agency.