Run Claude Code Free on Your Own GPU with a Local Model on Ollama
Configures Claude Code to use Nemotron 3.5 Lightning via Ollama on an RTX 5090 through .claude/settings.local.json, verified by generating and modifying a Python palindrome function.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Nemotron 3.5 LightningModel source: provider1,048,576 tokens max native context | 30Bactive 3B | 18 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| DDR5 32GB Computer | 32 GBRAM | 89.6 GB/s | no | 20 tok/s4-bit quantization |
| Mac mini M4 | 24 GBUnified RAM | 120 GB/s | no | 26 tok/s4-bit quantization |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 54 tok/s4-bit quantization |
| M4 Pro | 24 GBUnified RAM | 273 GB/s | no | 58 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 58 tok/s4-bit quantization |
| M5 Pro | 64 GBUnified RAM | 307 GB/s | no | 64 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 117 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 148 tok/s4-bit quantization |
| RTX 3090 Ti | 24 GBGPU VRAM | 1008 GB/s | yes | 173 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 197 tok/s4-bit quantization |
| RTX 5090 | 32 GBGPU VRAM | 1792 GB/s | yes | 256 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 256 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 417 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 475 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 475 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
Configures Claude Code to use Nemotron 3.5 Lightning via Ollama on an RTX 5090 through .claude/settings.local.json, verified by generating and modifying a Python palindrome function.
Tests Qwen 3.8 27B, Nemotron 3.5 Lightning, and Muse Glimmer on an RTX 5090 using 16 hard math, coding, and reasoning problems via Ollama with Q4_K_M quantization, measuring accuracy, time, VRAM, and
Evaluates NVIDIA Nemotron 3.5 Lightning 30B A3B via Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot tests on a 16GB local setup.
Tests Nemotron 3.5 Lightning performance on browser workflows, C++ coding, 3D CAD modeling, frontend design, long-context recall, niche knowledge, and roleplay tasks.
Benchmark compares Nemotron 3.5 Lightning and Muse Glimmer 30B on an RTX 5090 using needle-in-haystack, math, logic, JSON, knowledge, and coding tests at 64,000 tokens, measuring speed, VRAM, and thin
Installs and tests NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4, a large language model trained by NVIDIA.