Capabilities Dashboard

Highest-signal quantitative metrics. Absolute scores are not meaningful in isolation.

Epoch Capabilities Index (ECI)

ECI Frontier Trend

2024 Q22025 Q42026 Q4ECI

Illustrative placeholder data, not for quantitative use.

The ECI frontier has advanced roughly linearly since the introduction of reasoning models. Absolute scores are arbitrary; relative trends and the distinction between reasoning and non-reasoning models are more informative.

As of October 10, 2026 · Source: Epoch AI

Limitations: Absolute ECI scores have no intrinsic meaning. Focus on slopes and relative changes. Reasoning models show a steeper trend than non-reasoning models.

View update history →

METR Time Horizon

METR Time Horizon (log scale)

2024 Q22025 Q42026 Q2Hours (log)
50% horizon (hrs)80% horizon (hrs)

Illustrative placeholder data, not for quantitative use. Log scale used because horizons grow approximately exponentially.

50% time horizon measures the human expert task duration at which models succeed half the time on software engineering tasks. Recent frontier models reach low double-digit hours, but measurements above approximately 16 hours are unreliable with the current task suite.

As of September 8, 2026 (Time Horizon 1.1) · Source: METR

Limitations: Task suite has a practical ceiling. Cheating and reliability issues have been reported on longer tasks. 80% horizons are substantially shorter than 50% horizons.

View update history →

Hard Unsaturated Benchmarks

Currently tracking a small number of benchmarks that still differentiate frontier models according to the criteria in Methodology (clear score gaps between leading models and not yet near human ceiling).

BenchmarkCurrent frontier (approx.)Why still discriminative
ARC-AGI-2 / ARC-AGI-3Low-to-mid double digitsLarge gap to human performance on novel visual reasoning
FrontierMath (Tier 4 / Erdős)Low single-digit %Research-level problems; most remain unsolved by models

As of October 2026 · Source: ARC Prize / Epoch AI

Limitations: Scores are approximate placeholders. Benchmarks can saturate quickly. These are included while they still differentiate models. Replacement criteria are documented in Methodology.

View update history →

← Back to homepage