LLM model: DeepSeek V4.1 Flash

Model profile
Model Size Approximate size
DeepSeek V4.1 FlashModel source: provider1,048,576 tokens max native context 552Bactive 16B 331.2 GB4-bit quantization
Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone.
Hardware Memory Bandwidth vLLM Estimated generation
DGX H200 1128 GBGPU HBM3e 4800 GB/s yes 159 tok/s4-bit quantization
DGX Station 748 GBCoherent Memory 7100 GB/s yes 211 tok/s4-bit quantization
ET900N G3 748 GBCoherent Memory 7100 GB/s yes 211 tok/s4-bit quantization
DGX B200 1440 GBGPU HBM3e 8000 GB/s yes 229 tok/s4-bit quantization
GB200 NVL72 13400 GBGPU HBM3e 8000 GB/s yes 229 tok/s4-bit quantization
GB300 NVL72 20000 GBGPU HBM3e 8000 GB/s yes 229 tok/s4-bit quantization
7 videos found Showing 1-7
Video AI Search 2026-09-18new

Deepseek just did the impossible

Explains Deepseek V4.1 Flash architecture covering prefill versus decode, causal encoder decoder, global versus local cache, sliding window attention, CSA2, hierarchical sparse indexer, single pass mH

Ai4Users.com

Copyright © 2026 Piotr Szawdyński. All rights reserved.