Capabilities Dashboard
Highest-signal quantitative metrics. Absolute scores are not meaningful in isolation.
Epoch Capabilities Index (ECI)
ECI Frontier Trend
Illustrative placeholder data, not for quantitative use.
The ECI frontier has advanced roughly linearly since the introduction of reasoning models. Absolute scores are arbitrary; relative trends and the distinction between reasoning and non-reasoning models are more informative.
As of October 10, 2026 · Source: Epoch AI
METR Time Horizon
METR Time Horizon (log scale)
Illustrative placeholder data, not for quantitative use. Log scale used because horizons grow approximately exponentially.
50% time horizon measures the human expert task duration at which models succeed half the time on software engineering tasks. Recent frontier models reach low double-digit hours, but measurements above approximately 16 hours are unreliable with the current task suite.
As of September 8, 2026 (Time Horizon 1.1) · Source: METR
Hard Unsaturated Benchmarks
Currently tracking a small number of benchmarks that still differentiate frontier models according to the criteria in Methodology (clear score gaps between leading models and not yet near human ceiling).
| Benchmark | Current frontier (approx.) | Why still discriminative |
|---|---|---|
| ARC-AGI-2 / ARC-AGI-3 | Low-to-mid double digits | Large gap to human performance on novel visual reasoning |
| FrontierMath (Tier 4 / Erdős) | Low single-digit % | Research-level problems; most remain unsolved by models |
As of October 2026 · Source: ARC Prize / Epoch AI