FC / NETWORK
Registry 05 · Project
Continuous Evaluation Infrastructure for Frontier AI
An idea becomes real the moment someone agrees to carry it.
Closing the gap between pre-deployment testing and live system behavior at population scale. Existing audit regimes assume static technology; frontier systems are adaptive.
Classification
Translation work
$4–6M · Multi-year
Brief / 01
Context record
"Continuous Evaluation Infrastructure for Frontier AI" is a live Research-to-Reality opportunity, an explicit bridge between work being done in research environments and the institutions that would have to adopt, fund, or implement it for the work to actually change anything in the world.
The challenge: Closing the gap between pre-deployment testing and live system behavior at population scale. Existing audit regimes assume static technology; frontier systems are adaptive. Futurecraft is convening the specific combination of stakeholders required, Policy researcher, ML safety lab, Foundation underwriter, Standards body, under conditions of trust, with MIRA maintaining continuity between sessions and surfacing adjacent work that strengthens the case.
Estimated funding sits at $4–6M · Multi-year, on a 24 months to v1; 36 months to multi-jurisdiction adoption arc. Policy implications: Establishes precedent for continuous-evaluation authority in technology regulation, with implications for medical devices, autonomous systems, and critical infrastructure. Progress is measured against concrete, falsifiable outcomes, not activity, so the institution underwriting the work can see whether it is on track.
02 / DELIVERY
Implementation record
24 months to v1; 36 months to multi-jurisdiction adoption
Stakeholders
- , Policy researcher
- , ML safety lab
- , Foundation underwriter
- , Standards body
Funding
Policy implications
Measurable outcomes
- , Open-source evaluation harness adopted by ≥3 frontier labs
- , Reference policy text introduced in two jurisdictions
- , Independent evaluation institution incubated