LLM Battle 2 - Qwen, Ornith, Qwythos, Gemma 4
Qwen, Ornith, Qwythos, and Gemma 4 compete in a two-round battle using easy and challenging prompts to demonstrate the importance of writing a good prompt.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Ornith 1.0 35BModel source: derived262,144 tokens max native context | 35Bactive 3B | 21 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| DDR5 32GB Computer | 32 GBRAM | 89.6 GB/s | no | 20 tok/s4-bit quantization |
| Mac mini M4 | 24 GBUnified RAM | 120 GB/s | no | 26 tok/s4-bit quantization |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 54 tok/s4-bit quantization |
| M4 Pro | 24 GBUnified RAM | 273 GB/s | no | 58 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 58 tok/s4-bit quantization |
| M5 Pro | 64 GBUnified RAM | 307 GB/s | no | 64 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 117 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 148 tok/s4-bit quantization |
| RTX 3090 Ti | 24 GBGPU VRAM | 1008 GB/s | yes | 173 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 197 tok/s4-bit quantization |
| RTX 5090 | 32 GBGPU VRAM | 1792 GB/s | yes | 256 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 256 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 417 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 475 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 475 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
Qwen, Ornith, Qwythos, and Gemma 4 compete in a two-round battle using easy and challenging prompts to demonstrate the importance of writing a good prompt.
Compares Ornith 1.0 35B and Qwen3.6-35B-A3B-GGUF on a 16GB VRAM system using Driving Game, Fake Desktop, and Agent Maze coding tests.
Tests Ornith-1.0-35B performance in OpenClaw, Hermes, and Pi agent harnesses versus Qwen3.6 27B dense using its mixture-of-experts architecture.
Compares Ornith 35B and Qwen 3.6 35B-A3B using llama.cpp, llama-swap, and OpenCode on dual RTX 3090s while building a street-racing car OS, race-control interface, and live race simulator UI.
Compares Ornith 1 35B MoE against Qwen 3.6 35B A3B using Performance, Memory, Agency, OpenAI Human Eval, Sand Physics and Dungeon Crawler tests on a 16GB VRAM system.
Compares closed-source Sonnet 5 against open-model Ornith 35B to evaluate whether a local model can outperform its proprietary counterpart.
Locally installs and tests Ornith-1.0, a self-improving family of open-source models for agentic coding.
Tests Ornith 1.0 local coding models via browser OS, subway scene/FPS generation, frontend design, C++, multimodal code, and agentic coding scenarios.
Tests Ornith-1.0-35B against Qwen3.6-27B using a Splunk Cyber CTF question and an agentic Pi extension build judged by Opus 4.8.