CC Capabilities

Capabilities / Deployments

GPT-5.6 — OpenAI

OpenAI AI research infrastructure / cross-sector knowledge work score 88/100 confidence 0.94
Category
Deployments
Capability
Long-horizon agentic work across 55 professional fields (Agents' Last Exam SOTA 52.7%), recursive self-improvement (RSI Index 57.9%), computer use (OSWorld 2.0 62.6% at 85% fewer output tokens than competitors), production-ready knowledge work outputs across presentations, financial models, and legal documents
Observed
2026-07-09
Thesis section
Appendix III — confirming mechanism: RSI Index at 57.9% (was 41.7% for GPT-5.5) and AI-majority research code authorship inside OpenAI — direct evidence of recursive capability compression accelerating the discontinuity timeline

Claim

GPT-5.6 Sol achieves new SOTA on Agents' Last Exam (52.7%), Coding Agent Index (80), and BrowseComp (92.2%). The RSI Index — measuring recursive self-improvement capability — scores 57.9% vs 41.7% for GPT-5.5; AI now writes a majority of research code inside OpenAI itself. Knowledge work outputs across presentations, financial models, and legal documents are described by early adopters as ready for production without human polish. At $5/$30 per 1M tokens, frontier-level professional displacement is commercially accessible at scale. The RSI signal is the landmark: the model can accelerate its own successor’s development, compressing the timeline further.

Oracle verdict

GPT-5.6 is the first entry in this register where the confirming signal is not about what AI does to human work, but about what AI does to its own successor’s development. The RSI Index — a direct measure of recursive self-improvement capability — moves from 41.7% to 57.9% in a single generation; OpenAI reports that AI now writes a majority of research code inside the organisation developing the next model. This is the thesis mechanism at the infrastructure level, not the application layer: the system accelerating toward the discontinuity is now measurably contributing to that acceleration. The secondary signals reinforce the case at the displacement layer. Agents’ Last Exam (52.7% SOTA, long-horizon professional workflows across 55 fields), Coding Agent Index v1.1 (80, SOTA), and BrowseComp Ultra (92.2% SOTA) establish frontier capability across the knowledge-work spectrum. OSWorld 2.0 at 62.6% — surpassing Claude Opus 4.8 while using 85% fewer output tokens — means computer-use efficiency has crossed a threshold where cost, not capability, was the remaining constraint. Early-adopter testimony that presentations, financial models, and legal documents are ‘ready for production without human polish’ is the deployment signal this thesis has been predicting. Filed as [CONFIRMING — MECHANISM]: the RSI Index movement is not a benchmark gain; it is evidence that the capability curve is being steepened by the capability itself.

Why it matters

Model family: Sol ($5/$30 per 1M tokens, flagship), Terra (balanced, lower cost), Luna (most cost-efficient). Available via ChatGPT, Codex, and OpenAI API. Benchmark scores (Sol unless noted): Agents’ Last Exam 52.7% (SOTA, long-horizon professional workflows across 55 fields); Coding Agent Index v1.1 80 (SOTA); OSWorld 2.0 62.6% (computer use — surpasses Claude Opus 4.8 at 85% fewer output tokens); BrowseComp Ultra 92.2% (SOTA); GeneBench Pro 28.7% (was 12% for GPT-5.5, more than doubled); RSI Index 57.9% (was 41.7% for GPT-5.5); Internal Research Debugging 68.3%; Management Consulting Tasks 43.2% (vs 31.3% for GPT-5.5). Ultra mode coordinates 4 agents in parallel. The RSI Index directly measures progress toward recursive self-improvement; this register treats it as the primary signal for this entry.

Custom GPT Ask the Oracle