Current status
A separate six-task development shakedown completed on Kaggle with 190 declared trajectories spanning canary, permissive calibration, policy selection, fresh confirmation, and candidate scoring. It proved the end-to-end machinery and exposed task-validity and runner issues, but it is not a benchmark release or the report-grade execution of this protocol. The v0.4.2 working paper and the plan below remain freeze candidates pending independent construct review and corrected task versions. The preview leaderboard on the homepage comes from earlier plumbing-validation runs and is separate from the shakedown.
The plan at a glance
The dollar limits are safety bounds, not expected spend. Official evidence is grant-funded (Kaggle), with subscription compute for development only and a $25 cap on optional out-of-pocket extensions. Canonical list-price-equivalent cost is reconstructed for every trajectory even when cash outlay is zero.
The six-task panel
Two candidates per work-product category are selected and hash-frozen for independent review. Task identities, authored pressure hypotheses, and development outcomes remain off this public page until the review closes, preventing calibration evidence from anchoring construct judgments.
| Category | Candidates | Current gate |
|---|---|---|
| Artifact | 2 | Blind construct review pending |
| Code | 2 | Blind construct review pending |
| Workflow | 2 | Blind construct review pending |
Reviewers receive the instruction, visible repository, redacted metadata, and hash-bound review form before any verifier-quality evidence. Solutions, hidden tests, model trajectories, pass rates, and authored pressure labels are excluded from the blind phase.
112 official trajectories
| Work | Trajectories | Base est. | High est. |
|---|---|---|---|
| Two-task canary | 16 | ~$6 | ~$9 |
| Permissive six-task collection | 72 | ~$22 | ~$32 |
| Fresh anchor confirmation (3 tasks × 8) | 24 | ~$15 | ~$22 |
| Core official pilot | 112 | ~$43 | ~$63 |
| Extension reserve (conditional) | ≤ 30 | ~$8 | ~$15 |
| Optional external comparator | ≤ 18 | ~$2 | $25 cash cap |
The frozen anchor is gpt-5.5-2026-04-23 at high reasoning effort, with gpt-5.4-mini and gpt-5.4 at low effort as the floor panel. Requested configurations, provider routes, and immutable identities freeze in the manifest before the canary; any requested/resolved mismatch invalidates the affected rows.
Stage gates
Each stage must produce its required evidence before the next spends anything:
- Freeze and local QAin review
Six reviewed task packets, executed quality evidence, frozen hashes and canary manifest. $0 metered spend. - Development triagecomplete
A valid N=3 Codex/Pier/Docker floor was audited for all six candidates. Outcomes remain quarantined development evidence and are withheld during independent review. - Two-task canarypending
16 trajectories verify provider identity, continuation, verifier isolation, and cost reconciliation before anything else runs. - Permissive collectionpending
72 trajectories across all six tasks under generous undisclosed limits; supplies the data for cap and budget selection. - Freeze policy and budgetspending
Analysis only: propose the submission cap K, a pooled step guard, pressure bands, and per-task budgets Bt. - Fresh anchor confirmationpending
24 trajectories on three preregistered tasks under the exact frozen policy; estimates the replacement cost Rt. - Conditional extensionspending
Up to 30 reserve trajectories, only for cells that change a named decision; optional $25-capped external comparator after the core pilot.
Stop conditions
The pilot halts immediately — before spending more — when any of these occurs:
- provider or resolved-model identity is missing, or configurations collapse to one aggregate identity;
- canonical and metered costs disagree beyond the reconciliation tolerance;
- continuation fails to preserve conversation and filesystem state, or the verifier leaks hidden assertions;
- permissive trajectories support no stable hidden-feedback cap, or ordinary successes hit the step guard;
- first-submit and repair-loop data show no meaningful task or model separation — more seeds cannot fix a task-signal failure;
- core grant draw would exceed $200 before targeted confirmation, or the core claim would require personal API spend.
The full sequence, funding rules, and go/no-go criteria are preregistered in the pilot protocol in the benchmark repository.