Research · data

Data & downloads

Preview snapshotGenerated 2026-07-18

Everything on this site is computed from the files below — no number appears on a chart that you cannot recompute from a download. The current snapshot is a machinery-validation preview, not a report-grade release; the difference is spelled out at the bottom.

Contents

Current preview snapshot

Snapshot shallowswe-v0.2-website-preview-plus-inkling-high-2026-07-18: 864 bounded repair-loop rows across 18 tasks at N=3 seeds per task × model config. Fixed preview caps apply to every row: $5 spend, 20 verifier submissions, 120 agent steps. Failures keep their actual spend in the bill — this is realized CPSC, the preview view of the metric family.

The preview is nearly saturated by design (853 successes in 864 scored loops): it validates the repair-loop machinery, accounting, and export path rather than discriminating model capability.

Downloads

FileContentsSize
rollouts.jsonRepair-loop rows. Every scored bounded repair-loop row: outcome, spend, tokens, submissions, steps.3.3 MB
aggregate-by-model.jsonAggregate by model. Model-level CPSC, solve rate, and diagnostics.31 KB
aggregate-by-task-model.jsonAggregate by task × model. The per-cell numbers behind every chart.548 KB
workload-index.jsonWorkload index. Declared basket weights for the suite views.7 KB
deepswe-comparison.jsonDeepSWE comparison. The published DeepSWE v1.1 leaderboard rows used as clearly-labeled context.42 KB
run-manifest.jsonRun manifest. Snapshot metadata: protocol caps, panel, seeds, and budget gate.5 KB

Price sheets

Dollars are always token counts × a dated OpenRouter price sheet, so every cost is repriceable. The charts use the newest sheet; older sheets are kept so past numbers stay reproducible.

Provenance & runner

  • Runner: pier-private-repair-loop-pilot, agent shallowswe-resumable-mini-swe-agent.
  • Protocol invariants: one model config per row, no fallbacks, conversation and filesystem continuation across verifier feedback, coarse feedback classes only (generic_failure, runtime_error, missing_required_artifact, output_mismatch).
  • Research documents: the working paper and pilot protocol on this site are synced from lydakis/ShallowSWE @ 65d7516f27af with per-document hashes recorded in a content manifest.
  • DeepSWE context rows come from the published DeepSWE v1.1 leaderboard and are never blended into ShallowSWE values.

License

All ShallowSWE data files are released under CC BY 4.0 — use them freely with attribution. See LICENSE.txt for the exact terms and the attribution line to use.

What changes at v1

A calibrated snapshot is a different artifact from this preview, and its files will say so. When the pilot completes and a report-grade run happens, the exports gain:

  • calibrated per-task budgets Bt and anchor replacement costs Rt, with sample sizes and intervals;
  • charged-spend fields for all three metric views — reference-budget, realized, and escalation CPSC — beside actual spend;
  • immutable model_config_id and agent_policy_id identities with full provider provenance;
  • an explicit claim tier on every artifact (protocol_validation or report_grade), plus the candidate funnel, exclusions, and rank-stability results.

Preview files will remain downloadable after that, clearly separated from calibrated snapshots. The full field list is in paper appendix C.

Synced from lydakis/ShallowSWE @ 65d7516f27af on 2026-07-11. The benchmark repository is the source of truth.