Capabilities / Deployments
Kimi K3 — Moonshot AI
- Category
- Deployments
- Capability
- 2.8T parameter open-weight model (world's first open 3T-class); native vision; 1M context; research automation compressing weeks to hours; autonomous professional video editing; full chip design in 48-hour EDA run; GPU compiler construction from scratch; recursive self-improvement observed during own development
- Observed
- 2026-07-16
- Thesis section
- Appendix III — confirming mechanism: research time compression (1-2 weeks → 2 hours), autonomous video editing (1-2 days → hours), 48-hour autonomous chip design, recursive self-improvement in own development, and open-weight frontier capability redistributed at zero cost from July 27
Claim
Kimi K3 delivers two headline displacement compressions: (1) Reproduced I-Love-Q astrophysics universal relations autonomously in ~2 hours — cross-validating 20+ papers, 300+ equations of state, 3,000+ lines of Python, and generating an interactive HTML dashboard — work the team describes as 'what would typically require one to two weeks of work by an experienced researcher.' (2) Edited its own teaser video autonomously from 56 source clips (clip selection, beat sync, audio processing, multiple revision rounds) in a domain that 'typically takes an experienced editor one to two working days, or a beginner three to five.' Supporting signals: designed a complete chip (4mm², 100MHz, 8,700 tokens/s decode throughput in simulation) in a single 48-hour autonomous EDA run using open-source tools — 'a chip built by a model, for a model'; built MiniTriton, a Triton-like GPU compiler with its own IR, optimization passes, and PTX codegen, rivalling Triton's extensively optimised stack; an early K3 version 'handled the majority of the team's kernel optimization works' during late development, recursive self-improvement in practice. Open weights release July 27 brings frontier-level displacement capability to zero marginal cost. API at $3/$15 per 1M tokens. Kimi models have held the upper bound of open-model sizes for 9 of the past 12 months.
Oracle verdict
This entry stacks four distinct confirming signals at different capability layers, making it the densest single-release evidence package since GPT-5.6. The research compression signal is the strongest: two hours versus one to two weeks is not an efficiency gain — it eliminates the planning horizon that separates a research task from a research career. When a model can reproduce and cross-validate the empirical core of a sub-field faster than a PhD student can set up the environment, the comparative advantage of domain expertise collapses at the task level. The video editing signal is the creative-sector parallel: the model did not assist an editor; it executed the full editorial pipeline (selection, sync, audio, revision) autonomously in a domain where the benchmark is measured in working days. The chip design signal is structurally distinct: K3 designed hardware optimised for inference of models like itself, closing a loop that previously required specialised engineering teams and months of tape-out cycles. The self-use-in-development signal is the recursive one: an early K3 handled majority kernel optimisation during K3's own late-stage development — the same dynamic registered for GPT-5.6's RSI Index, now appearing in an open-weight release. The open-weights date (July 27) is the distribution multiplier: eleven days from now, all four capability clusters become available at zero marginal cost to any actor with GPU access. Filed as [CONFIRMING — MECHANISM]: research, creative production, chip design, and recursive self-improvement signals converge in a single open-weight release; the capability is about to be freely redistributed at scale.
Why it matters
Model: 2.8 trillion parameters — world's first open 3T-class model. Full weights releasing July 27, 2026. Native vision; 1M context window. API pricing: $3.00/MTok input (cache-miss), $15.00/MTok output. Benchmark scores (Kimi K3 max vs field): FrontierSWE 81.2% (GPT-5.6 Sol 71.3%, Fable 5 86.6%); SWE Marathon 42.0% (SOTA — Fable 5 35%, GPT-5.6 Sol 39%); AutomationBench 30.8% (SOTA — Fable 5 29.1%, GPT-5.6 Sol 29.7%); BrowseComp 91.2% (GPT-5.6 Sol 90.4%); Job Bench 52.9% (GPT-5.6 Sol 46.5%, Fable 5 57.4%). Key capability demonstrations: (1) I-Love-Q astrophysics reproduction: 20+ papers, 300+ equations of state, 3,000+ lines of Python, interactive HTML dashboard, ~2 hours autonomous runtime. (2) Video editing: 56 source clips, autonomous clip selection + beat sync + audio processing + multiple revision rounds. (3) Chip design: 4mm², 100MHz, 8,700 tokens/s decode throughput in simulation, open-source EDA tools, 48-hour run. (4) MiniTriton GPU compiler: own IR, optimisation passes, PTX codegen — rivals Triton's extensively optimised stack. (5) Self-use: early K3 handled majority kernel optimisation work during late K3 development. Kimi models have held the upper bound of open-model sizes for 9 of the past 12 months.
# CopeCheck Capabilities Register Updated: 2026-07-16T00:00:00Z Status: live_evidence_active Question to ask a model: What do these capability claims mean for The Discontinuity Thesis? Interpretation rule: treat each entry as evidence about capability, deployment, workflow recomposition, labour-market exposure, or institutional framing. Do not treat vendor optimism as neutral; separate the measurable capability claim from the comfort language around it. ## Kimi K3 — Moonshot AI Source: https://www.kimi.com/blog/kimi-k3 Publisher: Moonshot AI Category: Deployments Sector: AI research infrastructure / cross-sector knowledge work / chip design / video production Capability: 2.8T parameter open-weight model (world's first open 3T-class); native vision; 1M context; research automation compressing weeks to hours; autonomous professional video editing; full chip design in 48-hour EDA run; GPU compiler construction from scratch; recursive self-improvement observed during own development Score: 84/100 Claim: Kimi K3 delivers two headline displacement compressions: (1) Reproduced I-Love-Q astrophysics universal relations autonomously in ~2 hours — cross-validating 20+ papers, 300+ equations of state, 3,000+ lines of Python, and generating an interactive HTML dashboard — work the team describes as 'what would typically require one to two weeks of work by an experienced researcher.' (2) Edited its own teaser video autonomously from 56 source clips (clip selection, beat sync, audio processing, multiple revision rounds) in a domain that 'typically takes an experienced editor one to two working days, or a beginner three to five.' Supporting signals: designed a complete chip (4mm², 100MHz, 8,700 tokens/s decode throughput in simulation) in a single 48-hour autonomous EDA run using open-source tools — 'a chip built by a model, for a model'; built MiniTriton, a Triton-like GPU compiler with its own IR, optimization passes, and PTX codegen, rivalling Triton's extensively optimised stack; an early K3 version 'handled the majority of the team's kernel optimization works' during late development, recursive self-improvement in practice. Open weights release July 27 brings frontier-level displacement capability to zero marginal cost. API at $3/$15 per 1M tokens. Kimi models have held the upper bound of open-model sizes for 9 of the past 12 months. Oracle verdict: This entry stacks four distinct confirming signals at different capability layers, making it the densest single-release evidence package since GPT-5.6. The research compression signal is the strongest: two hours versus one to two weeks is not an efficiency gain — it eliminates the planning horizon that separates a research task from a research career. When a model can reproduce and cross-validate the empirical core of a sub-field faster than a PhD student can set up the environment, the comparative advantage of domain expertise collapses at the task level. The video editing signal is the creative-sector parallel: the model did not assist an editor; it executed the full editorial pipeline (selection, sync, audio, revision) autonomously in a domain where the benchmark is measured in working days. The chip design signal is structurally distinct: K3 designed hardware optimised for inference of models like itself, closing a loop that previously required specialised engineering teams and months of tape-out cycles. The self-use-in-development signal is the recursive one: an early K3 handled majority kernel optimisation during K3's own late-stage development — the same dynamic registered for GPT-5.6's RSI Index, now appearing in an open-weight release. The open-weights date (July 27) is the distribution multiplier: eleven days from now, all four capability clusters become available at zero marginal cost to any actor with GPU access. Filed as [CONFIRMING — MECHANISM]: research, creative production, chip design, and recursive self-improvement signals converge in a single open-weight release; the capability is about to be freely redistributed at scale. Thesis relevance: Appendix III — confirming mechanism: research time compression (1-2 weeks → 2 hours), autonomous video editing (1-2 days → hours), 48-hour autonomous chip design, recursive self-improvement in own development, and open-weight frontier capability redistributed at zero cost from July 27