Qwen 3.8 Flash-Next: Is Halogen Worth the Switch?
Qwen 3.8 Flash-Next ran on Halogen Flash Server via 72 coding jobs across OpenCode, dsh, and Pi harnesses on a 128 GB Strix Halo system.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Qwen 3.8 Flash NextModel source: provider262,144 tokens max native context | 125Bactive 6B | 75 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 28 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 30 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 64 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 83 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 115 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 159 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 304 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 369 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 369 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 388 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 388 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 388 tok/s4-bit quantization |
Qwen 3.8 Flash-Next ran on Halogen Flash Server via 72 coding jobs across OpenCode, dsh, and Pi harnesses on a 128 GB Strix Halo system.
Evaluates Claude Code, OpenCode and native Inferencer using Qwen4 Exp on Mac Studio M3 Ultra via Tic-Tac-Toe generation, modifications, image inferencing and token usage comparisons.
Tests Qwen3.8-Flash-Next on M5 Max using MTPLX v2.11.3, which fixes 8 defects in decoding, caching, tokenization, sampling, state restoration, and tool handling.
Tests Qwen 3.8 Flash-Next generation speed across OpenCode, Pi and DeepSeek Harness using 18 benchmark jobs at xhigh thinking level on a GMKtec EVO-X2 with Strix Halo hardware.
Tests Qwen3.8-Flash-Next on RTX PRO 6000 with SGLang, measuring speed, BFCL, tau^2-bench, long context, strategic reasoning, SVG drawing, video editing and motion design capabilities.
Tests oMLX v0.7.0.dev2 with Qwen3.8-Flash-Next 125B MoE on an M5 Max with 128GB unified memory, measuring inference speeds up to ~70 tokens/sec.
Tests performance, memory, reasoning, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot for unsloth/Qwen3.8-Flash-Next-GGUF on a 16GB local setup.
Compares Halogen Flash Server, EngramHalo and ROCm/Vulkan llama.cpp for Qwen3.8-Flash-Next on AMD Strix Halo via Terminal Bench Mini agentic tasks and speed benchmarks.
Compares Qwen3.8-Flash-Next inference via llama.cpp, SGLang, and FreeToken on RTX PRO 6000 Blackwell using AIPerf, measuring first-token latency, decode speed, accuracy on GSM8K/MATH-500, and startup
Demonstrates Qwen3.8-Flash-Next on 4x RTX 3090s using DeepSeek Harness with MiniMax H3 for autonomous music video creation, comparing throughput against Qwen3.8-27B.
Benchmarks unsloth/Qwen3.8-Flash-Next-GGUF on RTX 3060 using llama.cpp and thecodacus/spec-wins coding lab with contradictory specs and decoys.
Tests MTPLX v2.10.0 optimizations for local Qwen3.8-Flash-Next coding agents via .NET migration and SVG image generation comparisons against oMLX.
Benchmarked Qwen3.8-Flash-Next in llama.cpp across 0–96GB VRAM using AMD Ryzen 9 9950X and NVIDIA RTX PRO 6000 Blackwell, measuring decode speeds, prefill gains, and configuration impacts.
Compares DeepSeek V4 Flash and Qwen 3.8 Flash Next using Unsloth UD-IQ1_S 1-bit builds on 19 hard problems via CPU-only and RTX 5090 expert layer offload benchmarks.
Hands-on comparison of GLM-5.3-Flash vs Qwen3.8-Flash evaluating bug fixing, vision coding capabilities, and AI flirting performance through three distinct tests.
Compares MLX/oMLX and GGUF with llama.cpp for local inference of Qwen3.8-Flash-Next 125B on Apple M5 Max, achieving ~60 tokens/sec with oMLX 0.6.3 and Lightning MTP.
Tests Qwen 3.8-Flash-Next 1-bit to 4-bit IQ4_XS via llama.cpp on GMKtec EVO-X2 against Qwen 3.8-27B references using four one-shot tests, reporting ~20 tok/s and peak memory usage.
Benchmarks Qwen 3.8 Flash Next on Intel Core Ultra 9 285K CPU versus RTX 5090 offloading using llama.cpp, measuring throughput and perplexity for 1-bit and 2-bit quantizations.
Evaluates Qwen3.8-Flash-Next-MLX-Q9 and Q4 variants using Inferencer App v2.3.4 on an M3 Ultra 512 GiB system via local testing.
Tests Qwen3.8-Flash-Next coding and Splunk Cyber CTF performance, investigates hallucinations caused by Qwen Sparse Attention, and demonstrates fixing issues via sparse attention settings.
Reviews Qwen 3.8-Flash-Next specs and Unsloth GGUF sizes.
Tests Qwen3.8 Flash Next on browser workflows, C++ games, Blender, Godot, local Q4, FPS generation, and 3D scenes.
Locally installs and tests Qwen3.8-Flash-Next.
Explores Qwen3.8-Flash-Next architecture and runs local inference via Unsloth GGUF and llama.cpp on Apple Silicon, covering 1-bit and 4-bit quantization setups.
Compares Qwen 3.8 Flash Next and Qwen 3.8 27B architectures, detailing MoE, sparse attention, gated residuals, 51B n-gram lookup tables, and RTX 5090 MTP measurements.