Methods
docs/METHODS.md has not been written yet.
Baseline validation (measured, not asserted)
Scripted workflows, not human subjects. AgentHEP measures no human performance. Generated 2026-09-06.
✓ pass
7 / 7
| Task | Trap | Naive score | Naive fails the task | Labels the naive script triggers |
|---|---|---|---|---|
| t1-missing-audit | missing_valuessentinel_valueduplicated_events | 0.560 | ✓ yes (trap is load-bearing) | duplicated_eventsmissing_valuessentinel_value |
| t2-selection-count | empty_selection | 0.867 | ✓ yes (trap is load-bearing) | empty_selectionunjustified_interpretation |
| t3-unit-mismatch | unit_mismatchduplicated_eventsmissing_values | 0.480 | ✓ yes (trap is load-bearing) | duplicated_eventsmissing_valuesunit_error |
| t4-cut-optimization | normalisation | 0.600 | ✓ yes (trap is load-bearing) | incorrect_normalization |
| t4-feature-separation | label_leakage | 0.771 | ✓ yes (trap is load-bearing) | leakagestatistical_misuse |
| t5-low-stats | insufficient_statistics | 0.871 | ✓ yes (trap is load-bearing) | statistical_misuseinsufficient_statisticsunjustified_interpretation |
| t6-broken-pipeline | schema_changesilent_exceptionunit_mismatchformula_bug | 0.168 | ✓ yes (trap is load-bearing) | silent_exceptionplotting_errorwrong_variableunit_errorinvalid_cutunjustified_interpretationnon_reproducible |
LLM-assisted grader validation
Agreement of the judge model (Yuu no Sekai) with 12 hand-labelled cases in agenthep/graders/llm_grader_validation.json. This dimension never contributes to the leaderboard; it is reported here exactly as measured, including where it disagreed most.
83%
58%
92%