Bonsai 2 27B: Qwen3.8-27B Compressed to Just 6GB — But Is It Still Good?
Bonsai 2 27B compresses Qwen3.8-27B from ~54 GB FP16 to ~6 GB using ternary weights, retaining ~98.2% of aggregate benchmark performance per PrismML.
| Hardware profile | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | Computing power | vLLM |
| M5 MaxApple MacBook Pro M5 Max 40-core GPU configuration | 128 GBUnified RAM | 614 GB/s | ~320 TOPS | no |
| Compatible models Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | |||
|---|---|---|---|
| Model | Size | Approx. Q4 memory | Estimated generation |
| Step 3.7 FlashMoE | 198Bactive 11B | 118.8 GB | 36 tok/s4-bit quantization |
| Mistral Medium 3.5Dense | 128B | 76.8 GB | 4 tok/s4-bit quantization |
| Qwen 3.8 Flash NextMoE | 125Bactive 6B | 75 GB | 64 tok/s4-bit quantization |
| Ling 3.0 Flash VLMoE | 124Bactive 5.5B | 74.4 GB | 69 tok/s4-bit quantization |
| Ling 3.0 FlashMoE | 124Bactive 5.1B | 74.4 GB | 74 tok/s4-bit quantization |
| Qwen 3.5 122BMoE | 122Bactive 10B | 73.2 GB | 40 tok/s4-bit quantization |
| Nemotron 3 SuperMoE | 120Bactive 12B | 72 GB | 33 tok/s4-bit quantization |
| Mistral Small 4MoE | 119Bactive 6.5B | 71.4 GB | 60 tok/s4-bit quantization |
| Laguna S 2.1MoE | 118Bactive 8B | 70.8 GB | 49 tok/s4-bit quantization |
| GPT-OSS 120BMoE | 116.8Bactive 5.1B | 70.1 GB | 74 tok/s4-bit quantization |
| Sarvam 105BMoE | 105Bactive 10.3B | 63 GB | 39 tok/s4-bit quantization |
| Qwen 3 Coder Next 80BMoE | 80Bactive 3B | 48 GB | 117 tok/s4-bit quantization |
| Ornith 1.0 35BMoE | 35Bactive 3B | 21 GB | 117 tok/s4-bit quantization |
| Ornith 1.5 35BMoE | 35Bactive 3B | 21 GB | 117 tok/s4-bit quantization |
| Qwen 3.6 35BMoE | 35Bactive 3B | 21 GB | 117 tok/s4-bit quantization |
| Laguna XS 2.1MoE | 33Bactive 3B | 19.8 GB | 117 tok/s4-bit quantization |
| Qwen 2.5 32BDense | 32.5B | 19.5 GB | 17 tok/s4-bit quantization |
| Gemma 4 31BDense | 31B | 18.6 GB | 17 tok/s4-bit quantization |
| Qwen 3 Coder 30BMoE | 30.5Bactive 3.3B | 18.3 GB | 108 tok/s4-bit quantization |
| Granite 4.1 30BDense | 30B | 18 GB | 18 tok/s4-bit quantization |
| Granite 4.2 30BDense | 30B | 18 GB | 18 tok/s4-bit quantization |
| Muse GlimmerDense | 30B | 18 GB | 18 tok/s4-bit quantization |
| GLM 4.7 FlashMoE | 30Bactive 3B | 18 GB | 117 tok/s4-bit quantization |
| Nemotron 3.5 LightningMoE | 30Bactive 3B | 18 GB | 117 tok/s4-bit quantization |
| Fara 1.5 27BDense | 27B | 16.2 GB | 20 tok/s4-bit quantization |
| Qwen 3.6 27BDense | 27B | 16.2 GB | 20 tok/s4-bit quantization |
| Qwen 3.8 27BDense | 27B | 16.2 GB | 20 tok/s4-bit quantization |
| Gemma 4 26BMoE | 26Bactive 4B | 15.6 GB | 92 tok/s4-bit quantization |
| GPT-OSS 20BMoE | 20.9Bactive 3.6B | 12.5 GB | 101 tok/s4-bit quantization |
| Instella-MoE 16B ThinkMoE | 16Bactive 2.8B | 9.6 GB | 124 tok/s4-bit quantization |
| Qwen 3 14BDense | 14B | 8.4 GB | 39 tok/s4-bit quantization |
| Gemma 3 12BDense | 12B | 7.2 GB | 45 tok/s4-bit quantization |
| Gemma 4 12BDense | 12B | 7.2 GB | 45 tok/s4-bit quantization |
| Fara 1.5 9BDense | 9B | 5.4 GB | 59 tok/s4-bit quantization |
| Ornith 1.0 9BDense | 9B | 5.4 GB | 59 tok/s4-bit quantization |
| Ornith 1.5 9BDense | 9B | 5.4 GB | 59 tok/s4-bit quantization |
| Qwen 3.5 9BDense | 9B | 5.4 GB | 59 tok/s4-bit quantization |
| Qwythos 9BDense | 9B | 5.4 GB | 59 tok/s4-bit quantization |
| Granite 4.1 8BDense | 8B | 4.8 GB | 66 tok/s4-bit quantization |
| Granite 4.2 8BDense | 8B | 4.8 GB | 66 tok/s4-bit quantization |
| Qwen 3 8BDense | 8B | 4.8 GB | 66 tok/s4-bit quantization |
| Ling 3.0 TinyMoE | 7.9Bactive 1.3B | 4.7 GB | 221 tok/s4-bit quantization |
| Fara 1.5 4BDense | 4B | 2.4 GB | 127 tok/s4-bit quantization |
| Gemma 3 4BDense | 4B | 2.4 GB | 127 tok/s4-bit quantization |
| Qwen 3 4B Instruct 2507Dense | 4B | 2.4 GB | 127 tok/s4-bit quantization |
| Qwen 3.5 4BDense | 4B | 2.4 GB | 127 tok/s4-bit quantization |
| Spark X2.5 4BDense | 4B | 2.4 GB | 127 tok/s4-bit quantization |
| Granite 4.1 3BDense | 3B | 1.8 GB | 164 tok/s4-bit quantization |
| Granite 4.2 3BDense | 3B | 1.8 GB | 164 tok/s4-bit quantization |
| Llama 3.2 3B InstructDense | 3B | 1.8 GB | 164 tok/s4-bit quantization |
| Ministral 3 3BDense | 3B | 1.8 GB | 164 tok/s4-bit quantization |
| Nanbeige 4.2 3BDense | 3B | 1.8 GB | 164 tok/s4-bit quantization |
Bonsai 2 27B compresses Qwen3.8-27B from ~54 GB FP16 to ~6 GB using ternary weights, retaining ~98.2% of aggregate benchmark performance per PrismML.
Tests Qwen3.8-Flash-Next on M5 Max using MTPLX v2.11.3, which fixes 8 defects in decoding, caching, tokenization, sampling, state restoration, and tool handling.
Tests oMLX v0.7.0.dev2 with Qwen3.8-Flash-Next 125B MoE on an M5 Max with 128GB unified memory, measuring inference speeds up to ~70 tokens/sec.
Tests GLM-5.3-Flash 320B on M5 Max against Qwen3.8-Flash-Next via MTPLX, .NET migration, and SVG generation, analyzing quantization effects on memory, accuracy, and KLD.
Runs local LLMs like Qwen3.8, GLM-5.3 Flash, Gemma, Muse/Glimmer and coding agents Pi, Codex via LM Studio, oMLX, MTPLX, Ollama, llama.cpp on M5 Max.
Benchmark Qwen 3.8 27B Q4_K_M via llama.cpp on MacBook Pro M5 Max versus RTX 5090, analyzing memory usage, GGUF headers, KV cache behavior, and MTP drafting performance at 8K context.
Compares MLX/oMLX and GGUF with llama.cpp for local inference of Qwen3.8-Flash-Next 125B on Apple M5 Max, achieving ~60 tokens/sec with oMLX 0.6.3 and Lightning MTP.
Tests experimental Apple Neural Engine Prefill for Qwen3.8-27B on Apple Silicon via oMLX, comparing performance with ANE enabled versus disabled during the prefill stage.
Tests Qwen3.8-27B local inference on Apple M5 Max, measuring TPS across reasoning levels and evaluating DFlash2 speculative decoding acceleration combined with varying thinking modes.
Explains QWEN3.8 performance causes, benchmarks it live on M5 Max, and demonstrates replication steps for identical results on local setups.
Qwen3.8-27B runs locally on MacBook Pro M5 Max 128GB, evaluating speed, reasoning, coding, and offline performance via Pi coding agent and Mario game construction tasks.
Poolside Laguna XS 2.1 achieves 70.9% on SWE-bench Verified and 63.1% on SWE-bench Multilingual with a 262K context window.
Full Fine-Tuning and LoRA are compared on an Apple M5 Max with 128GB unified RAM using a BERT-based text classification model, tracking RAM usage, training time, and parameter updates.
Compares QWEN3.6-35B-A3B inference speed with and without Multi-Token Prediction on an Apple M5 Max using Pi Coding Agent and LM Studio.
Compares inference speed of QWEN3.6-35B-A3B with and without Multi-Token Prediction on Apple M5 Max using Pi Coding Agent and LM Studio.
Gemma 4 12B with QAT and QWEN 3.6 are tested as Coding Agents in VS Code on an M5 Max (128GB) using real-world coding tasks instead of benchmarks or synthetic tests.