CC Capabilities

Capabilities / Deployments

Kimi K3 — HuggingFace Model Card (Moonshot AI)

Moonshot AI AI research infrastructure / software engineering / agentic knowledge work / vision score 83/100 confidence 0.9
Category
Deployments
Capability
2.8T-parameter MoE (104B active); Kimi Delta Attention (KDA) + Gated MLA architecture; 896 experts (16 active per token); 1M-token context; native multimodal (MoonViT-V2, 401M params); MXFP4/MXFP8 quantization-aware training; open weights under Kimi K3 License
Observed
2026-07-27
Thesis section
Appendix III — capabilities: official technical specification and benchmark data for Kimi K3; cross-validates frontier-level autonomous performance (SWE-Marathon SOTA, AutomationBench SOTA, MCPMark-Verified SOTA) as open-weight model; JobBench score directly relevant to job displacement thesis

Claim

Kimi K3’s official HuggingFace model card documents a 2.8T-parameter MoE model with 104B active parameters, scaling to 16-of-896 experts per token via a novel Stable LatentMoE framework — yielding ~2.5× efficiency improvement over Kimi K2. Benchmark scores against frontier closed models: GPQA Diamond 93.5 (vs GPT-5.6 Sol 94.1, Claude Fable 5 92.6); DeepSWE 67.5 (vs GPT-5.6 Sol 73.0, Claude Opus 4.8 59.0); BrowseComp 91.2 (vs GPT-5.6 Sol 90.4, Claude Fable 5 88.0); MCPMark-Verified 94.5 (vs GPT-5.6 Sol 92.9, Claude Fable 5 87.4); SWE-Marathon 42.0 (SOTA; GPT-5.6 Sol 39.0, Claude Fable 5 35.0); Terminal-Bench 2.1 88.3 (vs GPT-5.6 Sol 88.8); OSWorld-Verified 84.8 (vs Claude Fable 5 85.0); JobBench 54.3 (vs Claude Fable 5 57.4, GPT-5.6 Sol 45.4); AutomationBench 30.8 (SOTA across all listed models). Architecture: 93 layers (1 dense + 92 MoE), 69 KDA + 24 Gated MLA attention layers, 7168-dim attention, 96 heads, vocabulary 160K tokens, native MXFP4 weight quantization with MXFP8 activations trained from SFT stage. Open weights released under Kimi K3 License.

Oracle verdict

The HuggingFace model card provides the technical substrate confirming what the kimi.com blog post demonstrated qualitatively. The benchmark pattern is significant: Kimi K3 ties or exceeds frontier closed models (GPT-5.6 Sol, Claude Fable 5) across multiple agentic and autonomous task categories — SWE-Marathon (SOTA at 42.0), AutomationBench (SOTA at 30.8), MCPMark-Verified (SOTA at 94.5), BrowseComp (near-SOTA at 91.2) — and does so as an open-weight model activating only 104B of 2.8T parameters per token. The JobBench score (54.3) is the most directly displacement-relevant: it measures performance on realistic job tasks, and K3 outperforms GPT-5.6 Sol (45.4) while trailing only Claude Fable 5 (57.4) among listed models. The MXFP4 native quantization is architecturally notable — quantization-aware training from SFT stage means the compressed weights do not degrade agentic performance, making GPU-efficient deployment accessible at scale. Filed as [CAPABILITIES]: technical architecture and benchmark profile of the open-weight model substantiating frontier-level autonomous task performance.

Why it matters

HuggingFace model card technical summary. Architecture: 93 layers (1 dense + 92 MoE), 896 experts, 16 active per token, 2 shared experts. Attention: 69 KDA + 24 Gated MLA layers, 96 attention heads, 7168 hidden dim. Vision encoder: MoonViT-V2 (401M params). Context: 1,048,576 tokens. Quantization: MXFP4 weights / MXFP8 activations, quantization-aware from SFT stage. License: Kimi K3 License (custom, non-Apache). Complementary entry to kimi-k3-moonshot-ai (kimi.com blog post, qualitative demonstrations). The blog entry documents task demonstrations; this entry documents the technical profile and cross-model benchmark comparisons.

Custom GPT Ask the Oracle