Qwen 3.8 27B vs Qwen 3.6: Why Qwen 3.8 Wins
Compares Qwen 3.8, 3.6 and 3.5 using 397 generations on an RTX 5090, analyzing MTP layer activation, split vision encoder, post-training effects, Terminal-Bench scores and token speeds.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Qwen 3.6 27BModel source: provider262,144 tokens max native context | 27B | 16.2 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| DDR5 32GB Computer | 32 GBRAM | 89.6 GB/s | no | 3 tok/s4-bit quantization |
| Mac mini M4 | 24 GBUnified RAM | 120 GB/s | no | 4 tok/s4-bit quantization |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 8 tok/s4-bit quantization |
| M4 Pro | 24 GBUnified RAM | 273 GB/s | no | 9 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 9 tok/s4-bit quantization |
| M5 Pro | 64 GBUnified RAM | 307 GB/s | no | 10 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 20 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 27 tok/s4-bit quantization |
| RTX 3090 Ti | 24 GBGPU VRAM | 1008 GB/s | yes | 33 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 39 tok/s4-bit quantization |
| RTX 5090 | 32 GBGPU VRAM | 1792 GB/s | yes | 58 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 58 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 145 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 204 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 204 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 225 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 225 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 225 tok/s4-bit quantization |
Compares Qwen 3.8, 3.6 and 3.5 using 397 generations on an RTX 5090, analyzing MTP layer activation, split vision encoder, post-training effects, Terminal-Bench scores and token speeds.
Compares Qwen3.8 27B and Qwen3.6 27B using Regina on Minesweeper, VMs, py2c calc, encryption and memcached, analyzing reasoning tokens, degeneracy and failures.
Evaluates Qwen 3.8 27B against Qwen 3.6 27B on psim-blackbox using pi harness 0.79.3 across low, medium, and xhigh reasoning settings, showing performance varies significantly by mode.
Benchmark comparisons of Muse Glimmer against Qwen and Kimi K3 using Inferencer App v2.3.2 on M3 Ultra 512 GiB hardware.
Compares Qwen 3.6 35B A3B, Qwen 3.6 27B dense, and Qwen 3.6 Fable Fusion by having each recreate the Solar Fortress arcade game using an Arcade Template.
Demonstrates using Qwen 3.6 27B Q6_K to reduce technical debt in ExtPackager and ImageViewer by fixing code problems while verifying no existing functionality breaks.
Evaluates Grug 27B QAT Q4 against Qwen 27B Q4 via Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot tests on a 16GB system.
Tests Traycer multi-agent collaboration using Claude Fable 5 and local Qwen 27B on a coding task, comparing results against Qwen 27B solo performance.
Evaluates Fable Fusion 711 Uncensored Heretic NM DAU NEO MAX MTP versus Base Qwen 3.6 27B using Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot
Evaluates Skywork’s capabilities for deep research, advanced coding design, writing a PhD thesis, and creating educational games using Qwen-3.6 on Inferencer App with M4 Max 128GB hardware.
Gemma 4 31B and Qwen 3.6 27B are compared across multiple stages in a dense LLM model battle, featuring code reviews and unexpected plot twists leading to confusing results.
Tests Tess-4 27B against unsloth/Qwen3.6-27B-MTP-GGUF using Performance, Memory, Agency, OpenAI Human Eval, Sand Physics, Dungeon Crawler, Blender and Godot benchmarks.
Demonstrates tuning a dual-GPU local LLM setup using llama-server startup options, layer split, tensor split, and multi-token prediction to increase tokens per second.
Qwen, Ornith, Qwythos, and Gemma 4 compete in a two-round battle using easy and challenging prompts to demonstrate the importance of writing a good prompt.
Evaluates Ternary Bonsai 27B against Qwen using Performance, Memory, Agency, OpenAI Human Eval, Dungeon Crawler, Blender, and Sand Physics tests on a 16GB local system.
Evaluates BottleCap AI's ThinkingCap Qwen 3.6 27B MTP against Base Qwen using Performance, Memory, Agency, OpenAI Human Eval, Sand Physics and Dungeon Crawler tests on a 16GB local setup.
Tests n-gram speculative decoding on Qwen 3.6 27B via llama.cpp against MATH-500, LiveCodeBench, and an 18-prompt iterative coding suite using AMD Ryzen 9 9950X and NVIDIA RTX PRO 6000 Blackwell.
Builds a desktop app extracting text from images and PDFs using Qwen 3.6 locally, demonstrating development workflow and OCR testing without cloud APIs or subscriptions.
Qwen 3.6 and Gemma 4 variants are tested to design and build a full-screen visualizer using MusicPlayer, with results hosted in the ext-mp-ai-visualizers repo.
Benchmarking DFlash speculative decoding with Qwen 3.6 27B via llama.cpp on RTX PRO 6000 Blackwell using aiperf and MATH-500, measuring throughput gains and accuracy against baselines.
Tests Ornith-1.0-35B against Qwen3.6-27B using a Splunk Cyber CTF question and an agentic Pi extension build judged by Opus 4.8.
Compares OpenRouter Fusion and Qwen3.6 27B building the same web app from scratch on fresh Linux VPSes across phased builds to evaluate planning, adaptation, and shipping capabilities.
Demonstrates Luce KVFlash recalling a hidden fact within a 256K token novel-length prompt while paging most context off the GPU to conserve VRAM.
Qwen 3.6 27B non-MTP versus MTP variants are compared for performance, memory usage, agency capabilities, and coding skills on a system featuring 16GB VRAM and 32GB DDR4 RAM.
Demonstrates running DFlash with SGLang's Spec V2 overlap scheduler on Qwen3.6-27B, achieving 160 tok/s on a single H100 via installation, benchmarking, and performance measurement.
Qwopus3.6-27B-Coder-MTP and Qwen3.6-27B-MTP build a multi-phase LAN whiteboard app with drag-and-drop images, real-time chat, and advanced features on separate fresh VPSs.
Demonstrates fitting a model's full 256K context on a small GPU using Luce KVFlash by keeping a tiny KV pool on the card and paging the rest to RAM.
Tests Qwen 3.6 27B MTP versus Qwopus 3.6 27B v2 MTP using an Expense Tracker, Card Matching Game, Breakout Game, and 2D Driving Game on a system with 16GB VRAM and 32GB DDR4 RAM.
Qwen3.6 27B and Claude Fable 5 build an LLM benchmark dashboard on a fresh local VPS connecting to llama-swap endpoints to run benchmarks and display results.
Compares Qwen3.6 27B Q8 and MiMo-v2.5 via iPhone Replica, Ragdoll Physics Simulator, and 3D Orbital Earth Explorer tests evaluating instruction adherence, UI cleanliness, functionality, and interactiv
Coding challenges test MiniMax-M3 versus Qwen3.6 27B using an iPhone-style UI, ragdoll physics simulator, and a fake Qwen3.7 30B release page to evaluate usability and polish.
Evaluates instruction adherence, UI cleanliness, functionality, and handling of complex browser-based coding prompts using identical single-file HTML, CSS, and JavaScript tasks for an iPhone Mockup, H
Compares Gemma4 12B against Qwen3.6 27B using the same GPU setup.
Compares Qwen3.6 27B Q6 locally versus Step3.7 Flash via OpenRouter on speed, coding ability, creativity, interactivity, UI polish, and physics while building an iPhone-style UI, Sled Sketch, and Volt
Compares Qwen3.6 27B Q8 MTP on consumer hardware versus GPT-5.5 across speed, coding, creativity, interactivity, UI polish, game logic, physics, and single-file HTML handling via Sled Sketch, Cubicle