Methodology
This page documents how the AI Trajectory Observatory selects, presents, and updates information. The goal is to keep claims traceable, limitations explicit, and revisions evidence-based.
Core Approach
- Prioritize quantitative, publicly documented metrics over narrative summaries.
- Separate measured current capabilities from longer-term possibilities and constraints.
- Record the source, date, and known limitations for every major claim or chart.
- Revise when new evidence appears; document the reason for each change in the Update Log.
Analysis of ultimate possibilities and physical/economic constraints will be added in a later phase.
Primary Data Sources
- Epoch Capabilities Index (ECI) — Composite capability scale from Epoch AI. Used for relative trend tracking.
- METR Time Horizon — Human-equivalent task duration at which models succeed at a given rate on software engineering tasks.
- Selected hard benchmarks — A small number of evaluations that still differentiate frontier models (see criteria below).
All charts and numbers link back to the original source where possible.
Hard Benchmark Selection Criteria
Benchmarks are included only while they retain useful discriminative power. The following rules guide inclusion and exit:
Inclusion
- Clear score gaps remain between leading frontier models.
- Performance is not yet near the human ceiling or fully saturated.
- The benchmark is publicly documented with reproducible evaluation conditions.
Exit / Replacement
- When top models cluster tightly and the benchmark no longer distinguishes them.
- When saturation or contamination makes further gains uninformative.
- Replacement benchmarks are chosen according to the same inclusion criteria and noted in the Update Log.
Current placeholders (ARC-AGI series, FrontierMath) will be replaced or confirmed once specific versions and scores are selected.
Known Limitations of Core Metrics
Epoch Capabilities Index (ECI)
- Absolute scores are arbitrary; only relative changes and trends are meaningful.
- The index aggregates multiple benchmarks; individual component saturation can affect interpretation.
- Reasoning models and non-reasoning models show different slopes; comparisons should respect this distinction.
METR Time Horizon
- Based primarily on software engineering and related tasks; does not directly measure open-ended or physical-world performance.
- Current task suite has a practical ceiling; measurements above approximately 16 hours are unreliable.
- Cheating and reliability issues have been observed on longer tasks.
- 50% and 80% horizons can differ substantially; both are reported when available.
Update Process
- Key metrics are reviewed at least monthly, or sooner after major model releases.
- Gap analysis and major interpretive judgments are revised quarterly or when material new evidence appears.
- Every material change is recorded in the Update Log with the date, what changed, and the reason or evidence.
- Placeholder values are clearly labeled until replaced by sourced data.
What This Site Does Not Do
- It does not provide investment advice or personalized recommendations.
- It does not claim that reaching any single signpost constitutes AGI.
- It does not treat speculative timelines as measured facts.