LLM Battle 7 - Muse vs Qwen Rematch!
Muse Glimmer is rematched against Qwen 3.8 27B using TalkWithMe, with Qwen’s reasoning level increased to high for comparison.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Muse GlimmerModel source: provider131,072 tokens max native context | 30B | 18 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| DDR5 32GB Computer | 32 GBRAM | 89.6 GB/s | no | 2 tok/s4-bit quantization |
| Mac mini M4 | 24 GBUnified RAM | 120 GB/s | no | 3 tok/s4-bit quantization |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 7 tok/s4-bit quantization |
| M4 Pro | 24 GBUnified RAM | 273 GB/s | no | 8 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 8 tok/s4-bit quantization |
| M5 Pro | 64 GBUnified RAM | 307 GB/s | no | 9 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 18 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 24 tok/s4-bit quantization |
| RTX 3090 Ti | 24 GBGPU VRAM | 1008 GB/s | yes | 30 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 35 tok/s4-bit quantization |
| RTX 5090 | 32 GBGPU VRAM | 1792 GB/s | yes | 52 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 52 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 132 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 186 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 186 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 206 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 206 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 206 tok/s4-bit quantization |
Muse Glimmer is rematched against Qwen 3.8 27B using TalkWithMe, with Qwen’s reasoning level increased to high for comparison.
Tests Qwen 3.8 27B, Muse Glimmer 30B and Gemma 4 26B on an RTX 5090 using Ollama, evaluating consistency across 12 accounting questions run 10 times each.
Side-by-side coding test using Claude Code on Ollama to generate Space Invaders, Breakout, and Tetris HTML canvas games with Qwen 3.8 27B, Muse Glimmer, and Gemma 4 on an RTX 5090.
Muse Glimmer and Qwen 3.8 27B implement two new features in SuperAsteroids, scored by code review with merging or deleting branches based on results.
Side-by-side coding test using Claude Code on Ollama to generate Space Invaders, Breakout, and Tetris HTML canvas games with Qwen 3.8 27B, Muse Glimmer, and Gemma 4 on an RTX 5090.
Tests Qwen 3.8 27B, Nemotron 3.5 Lightning, and Muse Glimmer on an RTX 5090 using 16 hard math, coding, and reasoning problems via Ollama with Q4_K_M quantization, measuring accuracy, time, VRAM, and
Meta Muse Glimmer, Qwen3.6 35B-A3B, Qwen3-Coder 30B, and GPT-OSS 20B compete in 10 frozen events on GMKtec EVO-X2 via Ollama on ROCm, judging speed and accuracy across tasks like coding, image generat
A private benchmark suite evaluates Muse Glimmer, finding it does not measure up against Qwen.
Benchmark comparisons of Muse Glimmer against Qwen and Kimi K3 using Inferencer App v2.3.2 on M3 Ultra 512 GiB hardware.
Meta Muse Glimmer implements a Solar Fortress clone using a 7-stage development plan, replicating the LLM Battle 5 challenge previously applied to three Qwen models.
Benchmark compares Nemotron 3.5 Lightning and Muse Glimmer 30B on an RTX 5090 using needle-in-haystack, math, logic, JSON, knowledge, and coding tests at 64,000 tokens, measuring speed, VRAM, and thin
Evaluates Muse Glimmer 30B versus Qwen 3.6 27B via woaibench.ai, covering benchmarks, agentic capabilities, tool use, coding, efficiency, DFlash speculative decoding, and local hardware requirements.
Evaluates Muse Glimmer performance on an AMD Ryzen AI Max machine via Ollama and Open WebUI, comparing Meta’s metrics against independent benchmarks including admitted hallucination rates.
Locally tests Muse Glimmer 30B via Muse Code with SGLang and vLLM against Qwen3.6-27B and Gemma4-31B on tool use, agent workflows, vision, coding, reasoning, long context, and Estonia benchmark.
Muse Glimmer 30B performance, memory, agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot are tested on a single GPU.
Installs and tests Muse Glimmer in GGUF format using speculative decoding.
Programming test of Meta's Muse Glimmer 30B model on RTX 3090 Ti using Q4_K_XL and Q4 KV cache quantization, reporting 60 tok/s decode speed and 75% MTP accept rate.
Installs and tests Muse Glimmer, a 30-billion-parameter causal language model with a dedicated perception encoder, distilled from Muse Spark.
Compares Muse Glimmer 30B and Qwen3.6-27B in an agentic coding session using Pi Coding Agent, evaluating actual tool calls, tokens and results.
Tests Meta Muse Glimmer on browser workflows, C++ coding, FPS development, 3D CAD modeling, frontend design, creative writing, multimodal coding, and simulations.