Skip to main content
Why ARL

Capability gets a headline.
Readiness has to be measured.

ARL did not invent a measurement discipline. It ports working ones from fields that solved the measurement of stochastic, safety-critical systems decades ago. This page is the why: the trends that make a shared readiness yardstick necessary, and the reasoning behind ARL’s four axes. It takes no position on who is ahead or on where the technology is going — only on what a readiness claim has to contain to mean something.

Part 1 — The trend

Five shifts in how AI is built and reported, each pointing at the same missing thing: a yardstick that survives translation across labs, languages, and years.

01

No shared yardstick

New models often arrive with new benchmarks, and benchmarks are replaced every few months. A score reported by one effort in one year cannot be honestly compared to a score from another effort years later — and capability numbers are commonly published without error bars, without a methodology fixed before the claim, without disclosing the hardware, without energy attribution, and without saying which exact version of the system was measured.

02

Reliability moves slower than headlines

Research on operational reliability and on task time-horizons documents a widening distance between rising headline scores and how consistently systems behave under operational variation. A single-shot accuracy number does not capture variance, and variance is what determines whether a system can be relied on.

03

Energy is a first-order constraint

Global data-centre electricity grew 17% in 2025; AI-specifically grew 50%, with the IEA projecting 945 TWh by 2030, and grid + generation additions taking 5–10 years to land. Joules-per-token is now proposed as a standard efficiency metric (peer of FLOPs and latency) in current work — and reasoning queries cost roughly 13× a standard query, an order of magnitude invisible to systems that don't measure. A readiness score that omits joules omits the part of the system the physical world actually bills for. Tokens-per-second compresses; joules don't.

04

Systems are agentic, not just models

Outputs of one step feed the inputs of the next; the harness, the tools, and the scaffolding are part of the deployed system, not accessories to the model. New attack surfaces come with that — schema-level techniques such as the Constrained Decoding Attack reach documented success rates of 94–99% by embedding intent in grammar rules while the prompt stays benign. Measurement has to cover the whole configured system.

05

Governance is converging, methodology is deferred

Risk-based AI frameworks are converging across jurisdictions, but most defer the technical how-to-measure to standards still being drafted. There is a methodology-shaped gap between “this system is high-risk” and “here is how its readiness was measured” — a gap a technical readiness framework can fill.

Part 2 — The response

Four axes, each closing a gap the others leave open — and each borrowed from a field that already does this.

No single number characterizes a stochastic system. Pharmaceutical efficacy, radar, communications, and aircraft certification all report several quantities at once, because any one of them alone is gameable. ARL is built on the same principle: Validation Depth (how thoroughly tested), Convergence Class (how stochastic, and whether the stochasticity is bounded), Energy Profile (the physical cost, in joules), and Security Class (measured resistance to an adversary). Hardware is documented alongside every claim for reproducibility, but it is not a peer axis. Three axes leave gaps; five add redundancy; four is the minimum that works.

Validation Depth
Technology Readiness Levels
NASA (Sadin 1974; Mankins 1995; ISO 16290:2013)

Measures how thoroughly a claim has been tested, not how impressive it is. A proven toaster and a proven spacecraft both rate “proven.”

Convergence Class
Pharmaceutical efficacy, radar Pd/Pfa, bit error rate
Stochastic-system characterization practice

Every field that deploys stochastic systems abandoned “it’s probabilistic, so it can’t be measured” decades ago. Variance gets characterized, not waved away.

Energy Profile
ENERGY STAR & PUE
EPA / The Green Grid

Physical equipment energy has been measured and verified across product categories for thirty years. The methodology and the verification practice already exist.

Security Class
Confidentiality / integrity / auditability
NIST SP 800-53, ISO/IEC 27001, Common Criteria EAL

Fields that deploy against an adversary measure resistance to attack as a first-class property, with assurance levels, rather than asserting it.

Part 3 — Trends with structure

Structure is the difference between an anecdote and a trajectory.

A capability score by itself is a snapshot. The improvements people point to only become a trajectory — something you can measure, compare, and extrapolate — once a structured yardstick holds still underneath them. METR’s task-horizon measurement is a clean example: by fixing the methodology, it turns “models feel more capable” into a measured cadence. The length of task a frontier agent completes at 50% reliability has roughly doubled every seven months since 2019 — with recent doublings closer to four — and the same shape now appears across nine domains. That trajectory is only visible because the measurement is structured. ARL is that discipline applied to readiness: fix the axes, and improvement becomes legible and comparable instead of asserted.

And structure is itself on a steep trajectory. Over the last few years, groups in every region have converged on disciplined evaluation and readiness frameworks — independently, and toward the same shape. ARL sits in that lineage as a readiness-level structure for deployed AI systems.

  1. 1974 → 2013
    Technology Readiness Levels

    NASA's readiness-level scale (Sadin 1974, Mankins 1995) is codified as ISO 16290 — the original structure ARL adapts.

  2. 2019
    First national AI governance frameworks

    OECD AI Principles; Singapore's Model AI Governance Framework — the earliest attempts to structure how AI is assessed.

  3. 2020 → 2022
    Structured, holistic evaluation

    Technology Readiness Levels adapted for machine learning (MLTRL); Stanford HELM and BIG-Bench move evaluation from single scores toward multi-metric discipline.

  4. 2022 → 2023
    Accountability + open testing

    China's CAC algorithm and generative-AI registry; Singapore open-sources AI Verify; ISO/IEC 42001 (AI management systems) is published.

  5. 2023 → 2024
    National evaluation bodies + binding law

    UK and US AI Safety Institutes stand up; the EU AI Act enters into force; Korea passes its AI Basic Act; METR begins measuring task time-horizons.

  6. 2025
    Technical standards + reliability discipline

    China's TC260 generative-AI security and content-labeling standards take effect; reliability and time-horizon research matures into measured trajectories.

  7. 2026
    Structure reaches agentic systems

    Singapore's Agentic AI framework; METR's Time Horizon 1.1 corroborates the trend across nine domains; readiness-level structures (ARL) extend the discipline to deployed AI systems end-to-end.

This is a trend, not a scoreboard — the point is that structured measurement is becoming the global default, the same way UL marks, SOC 2, and ISO 9001 did, because comparable results let each one build on the last instead of restarting every cycle.

Built to be useful

A standard earns its place by being needed, not by being mandated.

UL safety marks, SOC 2 audits, ISO 9001, CE marking — each became a commercial necessity before it was ever universally required, because buyers, insurers, and procurement offices needed a defensible basis for a decision. A readiness score is for the same moment: when someone has to choose a system and be able to explain the choice later.

That is why ARL is anchored in math and physics rather than opinion. Each axis rests on a foundation that does not drift across time, languages, or institutions: a claim of “ARL 6, Class D, 12.3 kJ/task, S2” means the same thing everywhere, while a claim of “PhD-level reasoning” does not survive the trip.

Composes, not competes

ARL is a technical readiness measure. It fits underneath the governance frameworks and alongside the benchmarks that already exist.

EU — AI Act Defers conformity-assessment methodology to harmonized standards still in progress — the kind of technical measure ARL provides.
United States — NIST AI RMF A voluntary risk-management frame; ARL supplies a measurable readiness score that a risk process can cite.
China — CAC registry + TC260 The CAC algorithm and generative-AI filing gives provider accountability and a national inventory of deployed systems; the TC260 technical-standards stack (content-labeling GB 45438-2025, generative-AI security and training-data standards effective late 2025) addresses content security and data discipline. ARL measures per-system technical readiness — a different layer of the same goal: deployment that is disclosed and accountable.
UK · Singapore — AISI / AI Verify Their evaluation platforms (Inspect, AI Verify) plug into the ARL Sandbox as Harnesses — the task lives in the Harness, ARL measures around it.
Japan · South Korea — METI / AI Basic Act Risk-tiered guidance and law (Japan's METI AI Guidelines and J-AISI; Korea's AI Basic Act and KAISI) set obligations; ARL supplies the readiness scoring such obligations can reference.
International — ISO/IEC JTC 1 / SC 42 The venue where the AI standards the world interoperates with are written (ISO/IEC 42001 and the rest); ARL is the kind of technical readiness method such standards reference.
Any benchmark MMLU, C-Eval, KMMLU, SWE-Bench, GAIA, HCAST, and language-specific suites worldwide — every benchmark is a task that plugs into ARL-S as a Harness. ARL does not replace benchmarks; it puts discipline around them.