LLM model: Qwen 3.5 122B

Model profile
Model Size Approximate size
Qwen 3.5 122BModel source: provider262,144 tokens max native context 122Bactive 10B 73.2 GB4-bit quantization
Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone.
Hardware Memory Bandwidth vLLM Estimated generation
Radeon 8060S 96GB 96 GBUnified RAM 256 GB/s yes 17 tok/s4-bit quantization
DGX Spark 128 GBUnified RAM 273 GB/s yes 18 tok/s4-bit quantization
M5 Max 128 GBUnified RAM 614 GB/s no 40 tok/s4-bit quantization
M3 Ultra 96 GBUnified RAM 819 GB/s no 52 tok/s4-bit quantization
M5 Ultra 256 GBUnified RAM 1200 GB/s no 74 tok/s4-bit quantization
RTX PRO 6000 96 GBGPU VRAM 1792 GB/s yes 105 tok/s4-bit quantization
DGX H200 1128 GBGPU HBM3e 4800 GB/s yes 223 tok/s4-bit quantization
DGX Station 748 GBCoherent Memory 7100 GB/s yes 284 tok/s4-bit quantization
ET900N G3 748 GBCoherent Memory 7100 GB/s yes 284 tok/s4-bit quantization
DGX B200 1440 GBGPU HBM3e 8000 GB/s yes 304 tok/s4-bit quantization
GB200 NVL72 13400 GBGPU HBM3e 8000 GB/s yes 304 tok/s4-bit quantization
GB300 NVL72 20000 GBGPU HBM3e 8000 GB/s yes 304 tok/s4-bit quantization
1 video found Showing 1-1

Ai4Users.com

Copyright © 2026 Piotr Szawdyński. All rights reserved.