Video
Fahd Mirza
2026-09-10
Ling-3.0-flash-VL: Free Vision Model Standing on Kimi's Shoulders
Tests Ling-3.0-flash-VL, described as a native multimodal model.
| Model profile | ||
|---|---|---|
| Model | Size | Approximate size |
| Ling 3.0 Flash VLModel source: provider131,072 tokens max native context | 124Bactive 5.5B | 74.4 GB4-bit quantization |
| Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone. | ||||
|---|---|---|---|---|
| Hardware | Memory | Bandwidth | vLLM | Estimated generation |
| Radeon 8060S 96GB | 96 GBUnified RAM | 256 GB/s | yes | 31 tok/s4-bit quantization |
| DGX Spark | 128 GBUnified RAM | 273 GB/s | yes | 33 tok/s4-bit quantization |
| M5 Max | 128 GBUnified RAM | 614 GB/s | no | 69 tok/s4-bit quantization |
| M3 Ultra | 96 GBUnified RAM | 819 GB/s | no | 90 tok/s4-bit quantization |
| M5 Ultra | 256 GBUnified RAM | 1200 GB/s | no | 124 tok/s4-bit quantization |
| RTX PRO 6000 | 96 GBGPU VRAM | 1792 GB/s | yes | 169 tok/s4-bit quantization |
| DGX H200 | 1128 GBGPU HBM3e | 4800 GB/s | yes | 318 tok/s4-bit quantization |
| DGX Station | 748 GBCoherent Memory | 7100 GB/s | yes | 383 tok/s4-bit quantization |
| ET900N G3 | 748 GBCoherent Memory | 7100 GB/s | yes | 383 tok/s4-bit quantization |
| DGX B200 | 1440 GBGPU HBM3e | 8000 GB/s | yes | 402 tok/s4-bit quantization |
| GB200 NVL72 | 13400 GBGPU HBM3e | 8000 GB/s | yes | 402 tok/s4-bit quantization |
| GB300 NVL72 | 20000 GBGPU HBM3e | 8000 GB/s | yes | 402 tok/s4-bit quantization |
Tests Ling-3.0-flash-VL, described as a native multimodal model.