Capabilities / Deployments
GPT-5.6 — OpenAI
- Category
- Deployments
- Capability
- Long-horizon agentic work across 55 professional fields (Agents' Last Exam SOTA 52.7%), recursive self-improvement (RSI Index 57.9%), computer use (OSWorld 2.0 62.6% at 85% fewer output tokens than competitors), production-ready knowledge work outputs across presentations, financial models, and legal documents
- Observed
- 2026-07-09
- Thesis section
- Appendix III — confirming mechanism: RSI Index at 57.9% (was 41.7% for GPT-5.5) and AI-majority research code authorship inside OpenAI — direct evidence of recursive capability compression accelerating the discontinuity timeline
Claim
GPT-5.6 Sol achieves new SOTA on Agents' Last Exam (52.7%), Coding Agent Index (80), and BrowseComp (92.2%). The RSI Index — measuring recursive self-improvement capability — scores 57.9% vs 41.7% for GPT-5.5; AI now writes a majority of research code inside OpenAI itself. Knowledge work outputs across presentations, financial models, and legal documents are described by early adopters as ready for production without human polish. At $5/$30 per 1M tokens, frontier-level professional displacement is commercially accessible at scale. The RSI signal is the landmark: the model can accelerate its own successor’s development, compressing the timeline further.
Oracle verdict
GPT-5.6 is the first entry in this register where the confirming signal is not about what AI does to human work, but about what AI does to its own successor’s development. The RSI Index — a direct measure of recursive self-improvement capability — moves from 41.7% to 57.9% in a single generation; OpenAI reports that AI now writes a majority of research code inside the organisation developing the next model. This is the thesis mechanism at the infrastructure level, not the application layer: the system accelerating toward the discontinuity is now measurably contributing to that acceleration. The secondary signals reinforce the case at the displacement layer. Agents’ Last Exam (52.7% SOTA, long-horizon professional workflows across 55 fields), Coding Agent Index v1.1 (80, SOTA), and BrowseComp Ultra (92.2% SOTA) establish frontier capability across the knowledge-work spectrum. OSWorld 2.0 at 62.6% — surpassing Claude Opus 4.8 while using 85% fewer output tokens — means computer-use efficiency has crossed a threshold where cost, not capability, was the remaining constraint. Early-adopter testimony that presentations, financial models, and legal documents are ‘ready for production without human polish’ is the deployment signal this thesis has been predicting. Filed as [CONFIRMING — MECHANISM]: the RSI Index movement is not a benchmark gain; it is evidence that the capability curve is being steepened by the capability itself.
Why it matters
Model family: Sol ($5/$30 per 1M tokens, flagship), Terra (balanced, lower cost), Luna (most cost-efficient). Available via ChatGPT, Codex, and OpenAI API. Benchmark scores (Sol unless noted): Agents’ Last Exam 52.7% (SOTA, long-horizon professional workflows across 55 fields); Coding Agent Index v1.1 80 (SOTA); OSWorld 2.0 62.6% (computer use — surpasses Claude Opus 4.8 at 85% fewer output tokens); BrowseComp Ultra 92.2% (SOTA); GeneBench Pro 28.7% (was 12% for GPT-5.5, more than doubled); RSI Index 57.9% (was 41.7% for GPT-5.5); Internal Research Debugging 68.3%; Management Consulting Tasks 43.2% (vs 31.3% for GPT-5.5). Ultra mode coordinates 4 agents in parallel. The RSI Index directly measures progress toward recursive self-improvement; this register treats it as the primary signal for this entry.
# CopeCheck Capabilities Register Updated: 2026-07-16T00:00:00Z Status: live_evidence_active Question to ask a model: What do these capability claims mean for The Discontinuity Thesis? Interpretation rule: treat each entry as evidence about capability, deployment, workflow recomposition, labour-market exposure, or institutional framing. Do not treat vendor optimism as neutral; separate the measurable capability claim from the comfort language around it. ## GPT-5.6 — OpenAI Source: https://openai.com/index/gpt-5-6/ Publisher: OpenAI Category: Deployments Sector: AI research infrastructure / cross-sector knowledge work Capability: Long-horizon agentic work across 55 professional fields (Agents' Last Exam SOTA 52.7%), recursive self-improvement (RSI Index 57.9%), computer use (OSWorld 2.0 62.6% at 85% fewer output tokens than competitors), production-ready knowledge work outputs across presentations, financial models, and legal documents Score: 88/100 Claim: GPT-5.6 Sol achieves new SOTA on Agents' Last Exam (52.7%), Coding Agent Index (80), and BrowseComp (92.2%). The RSI Index — measuring recursive self-improvement capability — scores 57.9% vs 41.7% for GPT-5.5; AI now writes a majority of research code inside OpenAI itself. Knowledge work outputs across presentations, financial models, and legal documents are described by early adopters as ready for production without human polish. At $5/$30 per 1M tokens, frontier-level professional displacement is commercially accessible at scale. The RSI signal is the landmark: the model can accelerate its own successor’s development, compressing the timeline further. Oracle verdict: GPT-5.6 is the first entry in this register where the confirming signal is not about what AI does to human work, but about what AI does to its own successor’s development. The RSI Index — a direct measure of recursive self-improvement capability — moves from 41.7% to 57.9% in a single generation; OpenAI reports that AI now writes a majority of research code inside the organisation developing the next model. This is the thesis mechanism at the infrastructure level, not the application layer: the system accelerating toward the discontinuity is now measurably contributing to that acceleration. The secondary signals reinforce the case at the displacement layer. Agents’ Last Exam (52.7% SOTA, long-horizon professional workflows across 55 fields), Coding Agent Index v1.1 (80, SOTA), and BrowseComp Ultra (92.2% SOTA) establish frontier capability across the knowledge-work spectrum. OSWorld 2.0 at 62.6% — surpassing Claude Opus 4.8 while using 85% fewer output tokens — means computer-use efficiency has crossed a threshold where cost, not capability, was the remaining constraint. Early-adopter testimony that presentations, financial models, and legal documents are ‘ready for production without human polish’ is the deployment signal this thesis has been predicting. Filed as [CONFIRMING — MECHANISM]: the RSI Index movement is not a benchmark gain; it is evidence that the capability curve is being steepened by the capability itself. Thesis relevance: Appendix III — confirming mechanism: RSI Index at 57.9% (was 41.7% for GPT-5.5) and AI-majority research code authorship inside OpenAI — direct evidence of recursive capability compression accelerating the discontinuity timeline