Current preview snapshot
Snapshot shallowswe-v0.2-website-preview-plus-inkling-high-2026-07-18: 864 bounded repair-loop rows across 18 tasks at N=3 seeds per task × model config. Fixed preview caps apply to every row: $5 spend, 20 verifier submissions, 120 agent steps. Failures keep their actual spend in the bill — this is realized CPSC, the preview view of the metric family.
The preview is nearly saturated by design (853 successes in 864 scored loops): it validates the repair-loop machinery, accounting, and export path rather than discriminating model capability.
Downloads
| File | Contents | Size |
|---|---|---|
| rollouts.json | Repair-loop rows. Every scored bounded repair-loop row: outcome, spend, tokens, submissions, steps. | 3.3 MB |
| aggregate-by-model.json | Aggregate by model. Model-level CPSC, solve rate, and diagnostics. | 31 KB |
| aggregate-by-task-model.json | Aggregate by task × model. The per-cell numbers behind every chart. | 548 KB |
| workload-index.json | Workload index. Declared basket weights for the suite views. | 7 KB |
| deepswe-comparison.json | DeepSWE comparison. The published DeepSWE v1.1 leaderboard rows used as clearly-labeled context. | 42 KB |
| run-manifest.json | Run manifest. Snapshot metadata: protocol caps, panel, seeds, and budget gate. | 5 KB |
Price sheets
Dollars are always token counts × a dated OpenRouter price sheet, so every cost is repriceable. The charts use the newest sheet; older sheets are kept so past numbers stay reproducible.
Provenance & runner
- Runner: pier-private-repair-loop-pilot, agent
shallowswe-resumable-mini-swe-agent. - Protocol invariants: one model config per row, no fallbacks, conversation and filesystem continuation across verifier feedback, coarse feedback classes only (generic_failure, runtime_error, missing_required_artifact, output_mismatch).
- Research documents: the working paper and pilot protocol on this site are synced from lydakis/ShallowSWE @ 65d7516f27af with per-document hashes recorded in a content manifest.
- DeepSWE context rows come from the published DeepSWE v1.1 leaderboard and are never blended into ShallowSWE values.
License
All ShallowSWE data files are released under CC BY 4.0 — use them freely with attribution. See LICENSE.txt for the exact terms and the attribution line to use.
What changes at v1
A calibrated snapshot is a different artifact from this preview, and its files will say so. When the pilot completes and a report-grade run happens, the exports gain:
- calibrated per-task budgets Bt and anchor replacement costs Rt, with sample sizes and intervals;
- charged-spend fields for all three metric views — reference-budget, realized, and escalation CPSC — beside actual spend;
- immutable
model_config_idandagent_policy_ididentities with full provider provenance; - an explicit claim tier on every artifact (
protocol_validationorreport_grade), plus the candidate funnel, exclusions, and rank-stability results.
Preview files will remain downloadable after that, clearly separated from calibrated snapshots. The full field list is in paper appendix C.