Research · pilot

Six-task protocol-validation pilot

Protocol v0.3 · freeze candidateKaggle shakedown completeReport-grade pilot pending

A six-task Kaggle development shakedown has exercised the full calibration and scoring machinery. The report-grade pilot remains a separate, construct-reviewed study under the protocol below and makes no leaderboard claim until those gates close.

Contents

Current status

A separate six-task development shakedown completed on Kaggle with 190 declared trajectories spanning canary, permissive calibration, policy selection, fresh confirmation, and candidate scoring. It proved the end-to-end machinery and exposed task-validity and runner issues, but it is not a benchmark release or the report-grade execution of this protocol. The v0.4.2 working paper and the plan below remain freeze candidates pending independent construct review and corrected task versions. The preview leaderboard on the homepage comes from earlier plumbing-validation runs and is separate from the shakedown.

The plan at a glance

6
tasks, two per category
112
official trajectories
$43 / $63
base / high estimate
$200 + $100
core limit + reserve

The dollar limits are safety bounds, not expected spend. Official evidence is grant-funded (Kaggle), with subscription compute for development only and a $25 cap on optional out-of-pocket extensions. Canonical list-price-equivalent cost is reconstructed for every trajectory even when cash outlay is zero.

The six-task panel

Two candidates per work-product category are selected and hash-frozen for independent review. Task identities, authored pressure hypotheses, and development outcomes remain off this public page until the review closes, preventing calibration evidence from anchoring construct judgments.

CategoryCandidatesCurrent gate
Artifact2Blind construct review pending
Code2Blind construct review pending
Workflow2Blind construct review pending

Reviewers receive the instruction, visible repository, redacted metadata, and hash-bound review form before any verifier-quality evidence. Solutions, hidden tests, model trajectories, pass rates, and authored pressure labels are excluded from the blind phase.

112 official trajectories

WorkTrajectoriesBase est.High est.
Two-task canary16~$6~$9
Permissive six-task collection72~$22~$32
Fresh anchor confirmation (3 tasks × 8)24~$15~$22
Core official pilot112~$43~$63
Extension reserve (conditional)≤ 30~$8~$15
Optional external comparator≤ 18~$2$25 cash cap

The frozen anchor is gpt-5.5-2026-04-23 at high reasoning effort, with gpt-5.4-mini and gpt-5.4 at low effort as the floor panel. Requested configurations, provider routes, and immutable identities freeze in the manifest before the canary; any requested/resolved mismatch invalidates the affected rows.

Stage gates

Each stage must produce its required evidence before the next spends anything:

  1. Freeze and local QAin review
    Six reviewed task packets, executed quality evidence, frozen hashes and canary manifest. $0 metered spend.
  2. Development triagecomplete
    A valid N=3 Codex/Pier/Docker floor was audited for all six candidates. Outcomes remain quarantined development evidence and are withheld during independent review.
  3. Two-task canarypending
    16 trajectories verify provider identity, continuation, verifier isolation, and cost reconciliation before anything else runs.
  4. Permissive collectionpending
    72 trajectories across all six tasks under generous undisclosed limits; supplies the data for cap and budget selection.
  5. Freeze policy and budgetspending
    Analysis only: propose the submission cap K, a pooled step guard, pressure bands, and per-task budgets Bt.
  6. Fresh anchor confirmationpending
    24 trajectories on three preregistered tasks under the exact frozen policy; estimates the replacement cost Rt.
  7. Conditional extensionspending
    Up to 30 reserve trajectories, only for cells that change a named decision; optional $25-capped external comparator after the core pilot.

Stop conditions

The pilot halts immediately — before spending more — when any of these occurs:

  • provider or resolved-model identity is missing, or configurations collapse to one aggregate identity;
  • canonical and metered costs disagree beyond the reconciliation tolerance;
  • continuation fails to preserve conversation and filesystem state, or the verifier leaks hidden assertions;
  • permissive trajectories support no stable hidden-feedback cap, or ordinary successes hit the step guard;
  • first-submit and repair-loop data show no meaningful task or model separation — more seeds cannot fix a task-signal failure;
  • core grant draw would exceed $200 before targeted confirmation, or the core claim would require personal API spend.

The full sequence, funding rules, and go/no-go criteria are preregistered in the pilot protocol in the benchmark repository.

Synced from lydakis/ShallowSWE @ 65d7516f27af on 2026-07-11. The benchmark repository is the source of truth.

docs/six-task-pilot-protocol-v0.3.md · sha256 9634f2c6d09a3f96