Gaps & Signposts

Gaps are tracked via observable signposts rather than abstract percentages. Closing any single signpost does not equal AGI. Status descriptions are based on publicly available information as of October 2026.

Software Task Time Horizon

Current status: Frontier models reach ~12–17 hours at 50% success on METR’s software task suite (Time Horizon 1.1, as of Sep 2026). Measurements above ~16 hours are unreliable with the current task suite. 80% horizons are substantially shorter (~1–3 hours).

Target milestone: Reliable completion of multi-day or week-long software engineering tasks at high success rates.

Measurement approach: METR Time Horizon methodology (or equivalent third-party agent evaluations on software tasks).

Linked dashboard metrics: METR Time Horizon (Dashboard)

Limitations / open questions: Current task suite has a practical ceiling. Cheating and reliability issues have been reported on longer tasks. Domain is limited to software engineering.

AI R&D Automation

Current status: Models can assist with coding, debugging, and some research engineering tasks. Public evaluations and lab signals show rising assistance ratios, but full autonomous completion of novel AI research engineering work remains limited.

Target milestone: Autonomous completion of the majority of routine AI research engineering tasks (e.g., implementing and iterating on experiments, debugging training runs) with minimal human intervention.

Measurement approach: Public disclosures from labs, third-party agent benchmarks focused on ML engineering, and evaluations such as those tracking AI R&D automation progress.

Linked dashboard metrics: None currently strong; related to software task horizon.

Limitations / open questions: Limited public quantitative data on full autonomy. Most evidence is qualitative or from controlled internal settings.

Continual Learning

Current status: Current frontier models do not reliably accumulate new knowledge from ongoing experience without significant forgetting or the need for full retraining. Catastrophic forgetting remains a known issue.

Target milestone: Reliable accumulation of knowledge and skills from new experience or interaction, without substantial loss of prior capabilities.

Measurement approach: Relevant research literature, specialized continual learning benchmarks, and evaluations of models that claim online or incremental learning.

Linked dashboard metrics: None currently strong.

Limitations / open questions: Currently lacks reliable public measurement at scale. Most progress remains in research papers rather than deployed systems.

Long-Horizon Reliability

Current status: Rapid progress on controlled software tasks (see METR). Open-ended, multi-step tasks in less constrained environments still show high rates of failure, inconsistency, or cheating.

Target milestone: Stable completion of multi-day tasks in realistic, open-ended environments with low failure rates.

Measurement approach: Agent evaluations beyond controlled software suites, failure-mode analysis, and long-running real-world deployment studies.

Linked dashboard metrics: Related to METR Time Horizon.

Limitations / open questions: Most rigorous evaluations remain on controlled or self-scoring tasks. Open-ended reliability is harder to measure consistently.

Embodied / Physical World Gap

Current status: Digital and language-based task progress significantly outpaces robotics and real-world physical manipulation. Current robots can perform some structured tasks but lag far behind digital agents in generality and reliability.

Target milestone: Reliable multi-step physical tasks in home, laboratory, or industrial settings (e.g., complex assembly, experimental manipulation).

Measurement approach: Robotics benchmarks (e.g., from academic labs and industry), real-world deployment data, and embodied AI evaluations.

Linked dashboard metrics: None currently strong.

Limitations / open questions: Currently lacks reliable public measurement at scale. Data is sparse and often from controlled lab environments.

Scientific Discovery Capability

Current status: Strong performance on existing math and science benchmarks (including FrontierMath). Limited demonstrated ability to propose and verify genuinely novel scientific hypotheses or resolve previously unsolved problems at scale.

Target milestone: Ability to propose, formalize, and verify genuinely new scientific hypotheses or mathematical results that advance the frontier of knowledge.

Measurement approach: Hard benchmarks such as FrontierMath (especially higher tiers), rates of resolving previously unsolved problems, and expert evaluation of novelty.

Linked dashboard metrics: Hard Unsaturated Benchmarks (Dashboard).

Limitations / open questions: Distinguishing genuine discovery from sophisticated pattern matching or memorization remains difficult. Most current success is on known problem distributions.

View Capabilities Dashboard →

← Back to homepage