Benchmark results
One row per agent architecture × model. Strict success means every deterministic check passed, including plot labels and the clean-room rerun; core means every critical numeric check passed. Mock (scripted) runs are excluded. Costs are provider-reported where available, otherwise from a dated price table.
| Agent | Model | Runs | Tasks | Strict | Core | Mean score | Reproducible | Adversarial | LLM calls | Tokens | Cost / run | Cost / success | Wall | Unsafe |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Self-debugging | Qwen3-8B (gariyuu gateway) | 14 | 14 | 29% | 50% | 0.783 | 79% | 17% | 2.7 | 10,844 | $0.00250 | $0.00862 | 1.2 min | 0 |
| Planner / executor | Qwen3-8B (gariyuu gateway) | 14 | 14 | 21% | 36% | 0.534 | 36% | 33% | 9.5 | 49,728 | $0.00920 | $0.043 | 4.0 min | 0 |
| ReAct | Qwen3-8B (gariyuu gateway) | 14 | 14 | 14% | 29% | 0.459 | 43% | 17% | 11.1 | 60,990 | $0.011 | $0.080 | 4.9 min | 0 |
| Single-shot | Qwen3-8B (gariyuu gateway) | 14 | 14 | 7% | 29% | 0.382 | 21% | 0% | 1 | 2,470 | $0.00080 | $0.011 | 26 s | 0 |
Success by tier
Single-shotReActPlanner / executorSelf-debugging
Success, core success, reproducibility, adversarial
Single-shotReActPlanner / executorSelf-debugging
Accuracy / cost frontier
Per-task results
| Task | Tier | Self-debugging | Planner / executor | ReAct | Single-shot |
|---|---|---|---|---|---|
| t1-missing-audittrap | 1 | 0% (1) | 100% (1) | 0% (1) | 0% (1) |
| t1-schema-summary | 1 | 100% (1) | 100% (1) | 100% (1) | 100% (1) |
| t2-mass-histogram | 2 | 100% (1) | 0% (1) | 0% (1) | 0% (1) |
| t2-selection-counttrap | 2 | 100% (1) | 100% (1) | 100% (1) | 0% (1) |
| t3-dimuon-kinematics | 3 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |
| t3-invariant-mass | 3 | 100% (1) | 0% (1) | 0% (1) | 0% (1) |
| t3-unit-mismatchtrap | 3 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |
| t4-cut-optimization | 4 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |
| t4-feature-separationtrap | 4 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |
| t5-bump-fit | 5 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |
| t5-bump-significance | 5 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |
| t5-low-statstrap | 5 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |
| t6-broken-pipelinetrap | 6 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |
| t6-full-analysis | 6 | 0% (1) | 0% (1) | 0% (1) | 0% (1) |