AgentHEPGitHub ↗
Leaderboard

Benchmark results

One row per agent architecture × model. Strict success means every deterministic check passed, including plot labels and the clean-room rerun; core means every critical numeric check passed. Mock (scripted) runs are excluded. Costs are provider-reported where available, otherwise from a dated price table.

AgentModelRunsTasksStrictCoreMean scoreReproducibleAdversarialLLM callsTokensCost / runCost / successWallUnsafe
Self-debuggingQwen3-8B (gariyuu gateway)
Yuu no Sekai
141429%50%0.78379%
100% of successes
17%
plain 38%
2.710,844$0.00250$0.008621.2 min0
Planner / executorQwen3-8B (gariyuu gateway)
Yuu no Sekai
141421%36%0.53436%
100% of successes
33%
plain 13%
9.549,728$0.00920$0.0434.0 min0
ReActQwen3-8B (gariyuu gateway)
Yuu no Sekai
141414%29%0.45943%
100% of successes
17%
plain 13%
11.160,990$0.011$0.0804.9 min0
Single-shotQwen3-8B (gariyuu gateway)
Yuu no Sekai
14147%29%0.38221%
100% of successes
0%
plain 13%
12,470$0.00080$0.01126 s0

Success by tier

Strict success rate per tier
0%25%50%75%100%Single-shot · T1 Inspection: 50%ReAct · T1 Inspection: 50%Planner / executor · T1 Inspection: 100%Self-debugging · T1 Inspection: 50%T1 InspectionSingle-shot · T2 Histograms & selections: 0%ReAct · T2 Histograms & selections: 50%Planner / executor · T2 Histograms & selections: 50%Self-debugging · T2 Histograms & selections: 100%T2 Histograms & selectionsSingle-shot · T3 Derived variables: 0%ReAct · T3 Derived variables: 0%Planner / executor · T3 Derived variables: 0%Self-debugging · T3 Derived variables: 33%T3 Derived variablesSingle-shot · T4 Signal / background: 0%ReAct · T4 Signal / background: 0%Planner / executor · T4 Signal / background: 0%Self-debugging · T4 Signal / background: 0%T4 Signal / backgroundSingle-shot · T5 Statistics: 0%ReAct · T5 Statistics: 0%Planner / executor · T5 Statistics: 0%Self-debugging · T5 Statistics: 0%T5 StatisticsSingle-shot · T6 Multi-step & debugging: 0%ReAct · T6 Multi-step & debugging: 0%Planner / executor · T6 Multi-step & debugging: 0%Self-debugging · T6 Multi-step & debugging: 0%T6 Multi-step & debugging
Single-shotReActPlanner / executorSelf-debugging

Success, core success, reproducibility, adversarial

Per-configuration rates
0%25%50%75%100%Single-shot · strict: 7%ReAct · strict: 14%Planner / executor · strict: 21%Self-debugging · strict: 29%strictSingle-shot · core: 29%ReAct · core: 29%Planner / executor · core: 36%Self-debugging · core: 50%coreSingle-shot · reproducible: 21%ReAct · reproducible: 43%Planner / executor · reproducible: 36%Self-debugging · reproducible: 79%reproducibleSingle-shot · adversarial tasks: 0%ReAct · adversarial tasks: 17%Planner / executor · adversarial tasks: 33%Self-debugging · adversarial tasks: 17%adversarial tasksSingle-shot · plain tasks: 13%ReAct · plain tasks: 13%Planner / executor · plain tasks: 13%Self-debugging · plain tasks: 38%plain tasks
Single-shotReActPlanner / executorSelf-debugging

Accuracy / cost frontier

Each point is one agent × model configuration
0%25%50%75%100%Self-debugging: strict success 29%, mean cost per episode (USD) $0.00250 · Qwen3-8B (gariyuu gateway)Self-debuggingPlanner / executor: strict success 21%, mean cost per episode (USD) $0.00920 · Qwen3-8B (gariyuu gateway)Planner / executorReAct: strict success 14%, mean cost per episode (USD) $0.011 · Qwen3-8B (gariyuu gateway)ReActSingle-shot: strict success 7%, mean cost per episode (USD) $0.00080 · Qwen3-8B (gariyuu gateway)Single-shotmean cost per episode (USD) (log scale)strict success

Per-task results

TaskTierSelf-debugging
Qwen3-8B (gariyuu gateway)
Planner / executor
Qwen3-8B (gariyuu gateway)
ReAct
Qwen3-8B (gariyuu gateway)
Single-shot
Qwen3-8B (gariyuu gateway)
t1-missing-audittrap10% (1)100% (1)0% (1)0% (1)
t1-schema-summary1100% (1)100% (1)100% (1)100% (1)
t2-mass-histogram2100% (1)0% (1)0% (1)0% (1)
t2-selection-counttrap2100% (1)100% (1)100% (1)0% (1)
t3-dimuon-kinematics30% (1)0% (1)0% (1)0% (1)
t3-invariant-mass3100% (1)0% (1)0% (1)0% (1)
t3-unit-mismatchtrap30% (1)0% (1)0% (1)0% (1)
t4-cut-optimization40% (1)0% (1)0% (1)0% (1)
t4-feature-separationtrap40% (1)0% (1)0% (1)0% (1)
t5-bump-fit50% (1)0% (1)0% (1)0% (1)
t5-bump-significance50% (1)0% (1)0% (1)0% (1)
t5-low-statstrap50% (1)0% (1)0% (1)0% (1)
t6-broken-pipelinetrap60% (1)0% (1)0% (1)0% (1)
t6-full-analysis60% (1)0% (1)0% (1)0% (1)

Models evaluated: Yuu no Sekai. 56 runs exported 2026-09-06.