AgentHEPGitHub ↗
Model comparison

Models

Results pooled over agent architectures for each model. Only models with real credentials at run time appear; presets without credentials are listed so the gap is explicit rather than hidden. The exact model string requested and the string the API reported are both stored with every run.

ModelPresetRunsStrict success (pooled)Best architectureTotal costReported as
Qwen3-8B (gariyuu gateway)
Yuu no Sekai
gariyuu-qwen3-8b5618%Self-debugging (29%)$0.335Yuu no Sekai

Pooled success by tier

Strict success per tier, all architectures pooled
0%25%50%75%100%Qwen3-8B (gariyuu gateway) · T1 Inspection: 63%T1 InspectionQwen3-8B (gariyuu gateway) · T2 Histograms & selections: 50%T2 Histograms & selectionsQwen3-8B (gariyuu gateway) · T3 Derived variables: 8%T3 Derived variablesQwen3-8B (gariyuu gateway) · T4 Signal / background: 0%T4 Signal / backgroundQwen3-8B (gariyuu gateway) · T5 Statistics: 0%T5 StatisticsQwen3-8B (gariyuu gateway) · T6 Multi-step & debugging: 0%T6 Multi-step & debugging
Qwen3-8B (gariyuu gateway)

Provider presets

credentials are read from the environment only
PresetKindModel stringContextEvaluated
gariyuu-qwen3-8bopenai_compatYuu no Sekai8,192yes
openai-gpt-4o-miniopenai_compatgpt-4o-mini128,000no credentials at run time
openai-gpt-4.1openai_compatgpt-4.11,000,000no credentials at run time
anthropic-sonnet-5anthropicclaude-sonnet-5200,000no credentials at run time
anthropic-haiku-4.5anthropicclaude-haiku-4-5-20251001200,000no credentials at run time
mockmockmock-scripted100,000tests only

Adding a model is a registry entry plus an API key in the environment; no benchmark code changes. Runs are never fabricated for presets without credentials.