LLM Battle 8
Coding competition evaluating Ornith-1.5-9B, Spark-X2.5-4B, and IBM's Granite-4.2-8B on automating mundane, repetitive tasks using OpenCode GPU sidebar plugin and VoxCtrl.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Ornith 1.5 9BModel source: derived262,144 tokens max native context | 9B | 5.4 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| DDR4 16GB Computer | 16 GBRAM | 51.2 GB/s | no | 5 tok/s4-bit quantization |
| DDR5 16GB Computer | 16 GBRAM | 89.6 GB/s | no | 9 tok/s4-bit quantization |
| DDR5 32GB Computer | 32 GBRAM | 89.6 GB/s | no | 9 tok/s4-bit quantization |
| Mac mini M4 | 24 GBUnified RAM | 120 GB/s | no | 12 tok/s4-bit quantization |
| RTX 2000 Ada | 16 GBGPU VRAM | 224 GB/s | yes | 22 tok/s4-bit quantization |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 25 tok/s4-bit quantization |
| M4 Pro | 24 GBUnified RAM | 273 GB/s | no | 27 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 27 tok/s4-bit quantization |
| RTX 4060 Ti 16GB | 16 GBGPU VRAM | 288 GB/s | yes | 28 tok/s4-bit quantization |
| M5 Pro | 64 GBUnified RAM | 307 GB/s | no | 30 tok/s4-bit quantization |
| RTX 3060 12GB | 12 GBGPU VRAM | 360 GB/s | yes | 35 tok/s4-bit quantization |
| RTX 5060 Ti 16GB | 16 GBGPU VRAM | 448 GB/s | yes | 44 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 59 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 78 tok/s4-bit quantization |
| RTX 5080 | 16 GBGPU VRAM | 960 GB/s | yes | 91 tok/s4-bit quantization |
| RTX 3090 Ti | 24 GBGPU VRAM | 1008 GB/s | yes | 95 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 111 tok/s4-bit quantization |
| RTX 5090 | 32 GBGPU VRAM | 1792 GB/s | yes | 160 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 160 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 357 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 468 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 468 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 505 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 505 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 505 tok/s4-bit quantization |
Coding competition evaluating Ornith-1.5-9B, Spark-X2.5-4B, and IBM's Granite-4.2-8B on automating mundane, repetitive tasks using OpenCode GPU sidebar plugin and VoxCtrl.
Benchmark compares Qwen 3.5 9B, Ornith 1.5 9B, Defiant Fable, Heretic Qwen, and Hauhau Qwen on 12GB GPUs like 3060, 4070, or 1080 Ti.
Benchmark IBM Granite 4.2, Qwen 3.8 27B, Gemma 4, and Ornith 1.5 on an RTX 5090, measuring KV cache cost, usable context, generation speed, coding, math, and tool calling via Python tests.
Benchmarking Ornith 1.5 9B and 35B, Qwen 3.8 27B, and Gemma 4 on an RTX 5090 using llama-server with 4-bit weights and FP16 KV cache via code-graded metrics.
Speed benchmarks for Ornith 1.5 9B dense and 35B MoE models running on a MacBook M4 Pro with 24 GB unified memory.
Locally installs and tests the Ornith-1.5-9B model using resources from huggingface.co/ornith-ai/Ornith-1.5-9B.