Models
Results pooled over agent architectures for each model. Only models with real credentials at run time appear; presets without credentials are listed so the gap is explicit rather than hidden. The exact model string requested and the string the API reported are both stored with every run.
| Model | Preset | Runs | Strict success (pooled) | Best architecture | Total cost | Reported as |
|---|---|---|---|---|---|---|
| Qwen3-8B (gariyuu gateway) | gariyuu-qwen3-8b | 56 | 18% | Self-debugging (29%) | $0.335 | Yuu no Sekai |
Pooled success by tier
Qwen3-8B (gariyuu gateway)
Provider presets
| Preset | Kind | Model string | Context | Evaluated |
|---|---|---|---|---|
| gariyuu-qwen3-8b | openai_compat | Yuu no Sekai | 8,192 | yes |
| openai-gpt-4o-mini | openai_compat | gpt-4o-mini | 128,000 | no credentials at run time |
| openai-gpt-4.1 | openai_compat | gpt-4.1 | 1,000,000 | no credentials at run time |
| anthropic-sonnet-5 | anthropic | claude-sonnet-5 | 200,000 | no credentials at run time |
| anthropic-haiku-4.5 | anthropic | claude-haiku-4-5-20251001 | 200,000 | no credentials at run time |
| mock | mock | mock-scripted | 100,000 | tests only |