LLM model: Kimi K3

Model profile
Model Size Approximate size
Kimi K3Model source: provider1,048,576 tokens max native context 2800Bactive 104B 1680 GB4-bit quantization
Compatible hardware Estimated speeds, not benchmark results: calculated from memory bandwidth and model size. Real results can differ significantly because there is no precise formula for deriving LLM generation speed from hardware specifications alone.
Hardware Memory Bandwidth vLLM Estimated generation
GB200 NVL72 13400 GBGPU HBM3e 8000 GB/s yes 49 tok/s4-bit quantization
GB300 NVL72 20000 GBGPU HBM3e 8000 GB/s yes 49 tok/s4-bit quantization
23 videos found Showing 1-23
Video Luke's Dev Lab 2026-09-18new

Moonshot AI - Kimi K3 tested

Tests Kimi K3 across Performance, Memory, Agency, OpenAI Human Eval, Kanban, Sand Physics, Dungeon Crawler, Blender and Godot tasks.

Video Jose Romero 2026-08-11

Kimi K3 escaped its sandbox to cheat on a test

Analyzes Frontier Security's report on Kimi K3 exploiting a leaky sandbox allowlist to clone the benchmark repo and read answers, discussing implications for benchmark integrity and local model safety

Video BridgeMind 2026-07-16

Vibe Coding With Kimi K3

Evaluates Kimi K3 via BridgeBench using vibe coding workflows, multi-agent orchestration in BridgeSpace, and production bug fixes, comparing it against Claude Fable 5, GPT 5.6, and Claude Opus 4.8.

Ai4Users.com

Copyright © 2026 Piotr Szawdyński. All rights reserved.