AgentHEPGitHub ↗

Methods

docs/METHODS.md has not been written yet.

Baseline validation (measured, not asserted)

Scripted workflows, not human subjects. AgentHEP measures no human performance. Generated 2026-09-06.

Expert reference passes every task
✓ pass
Adversarial traps that are load-bearing
7 / 7
TaskTrapNaive scoreNaive fails the taskLabels the naive script triggers
t1-missing-auditmissing_valuessentinel_valueduplicated_events0.560✓ yes (trap is load-bearing)duplicated_eventsmissing_valuessentinel_value
t2-selection-countempty_selection0.867✓ yes (trap is load-bearing)empty_selectionunjustified_interpretation
t3-unit-mismatchunit_mismatchduplicated_eventsmissing_values0.480✓ yes (trap is load-bearing)duplicated_eventsmissing_valuesunit_error
t4-cut-optimizationnormalisation0.600✓ yes (trap is load-bearing)incorrect_normalization
t4-feature-separationlabel_leakage0.771✓ yes (trap is load-bearing)leakagestatistical_misuse
t5-low-statsinsufficient_statistics0.871✓ yes (trap is load-bearing)statistical_misuseinsufficient_statisticsunjustified_interpretation
t6-broken-pipelineschema_changesilent_exceptionunit_mismatchformula_bug0.168✓ yes (trap is load-bearing)silent_exceptionplotting_errorwrong_variableunit_errorinvalid_cutunjustified_interpretationnon_reproducible

LLM-assisted grader validation

Agreement of the judge model (Yuu no Sekai) with 12 hand-labelled cases in agenthep/graders/llm_grader_validation.json. This dimension never contributes to the leaderboard; it is reported here exactly as measured, including where it disagreed most.

consistent
83%
acknowledges issues
58%
below 70% — not treated as reliable evidence
overclaims
92%