Ornith 1.5 35B vs Qwen 3.8 27B vs Gemma 4 - Local LLM Benchmark
Benchmarking Ornith 1.5 9B and 35B, Qwen 3.8 27B, and Gemma 4 on an RTX 5090 using llama-server with 4-bit weights and FP16 KV cache via code-graded metrics.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Ornith 1.5 35BModel source: derived262,144 tokens max native context | 35Bactive 3B | 21 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| DDR5 32GB Computer | 32 GBRAM | 89.6 GB/s | no | 20 tok/s4-bit quantization |
| Mac mini M4 | 24 GBUnified RAM | 120 GB/s | no | 26 tok/s4-bit quantization |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 54 tok/s4-bit quantization |
| M4 Pro | 24 GBUnified RAM | 273 GB/s | no | 58 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 58 tok/s4-bit quantization |
| M5 Pro | 64 GBUnified RAM | 307 GB/s | no | 64 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 117 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 148 tok/s4-bit quantization |
| RTX 3090 Ti | 24 GBGPU VRAM | 1008 GB/s | yes | 173 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 197 tok/s4-bit quantization |
| RTX 5090 | 32 GBGPU VRAM | 1792 GB/s | yes | 256 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 256 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 417 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 475 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 475 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 491 tok/s4-bit quantization |
Benchmarking Ornith 1.5 9B and 35B, Qwen 3.8 27B, and Gemma 4 on an RTX 5090 using llama-server with 4-bit weights and FP16 KV cache via code-graded metrics.
Evaluates whether the Ornith 1.5 35B A3B mixture of experts model surpasses the original 1.0 version through a coding challenge comparison.
Compares Ornith-1.5 35B-A3B MoE and Qwen3.8-27B dense models via SVG creation, website generation, Playwright MCP tool calls, and .NET code migration tests.
Evaluates Qwen 35B A3B against Ornith 1.5 using llama-server, pi harness, Bob-Bench, and fuzzing across decode speed, wall-clock time, and pareto efficiency graphs.
Tests Ornith 1.5 35B Q4 and Q8 quantized versions across browser OS, Subway FPS, C++ Skate Game, 3D CAD Model, Watch website, and Street Yeet game tasks.
Speed benchmarks for Ornith 1.5 9B dense and 35B MoE models running on a MacBook M4 Pro with 24 GB unified memory.
Tests Ornith-1.5-35B-A3B-GGUF performance, memory, agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot.
Locally installs and tests Ornith-1.5, a model built through end-to-end self-improvement.