the.ai

Evaluation / Methods

verified

Statistical Significance

Whether a difference between two numbers is bigger than the noise that produced them. Benchmark tables are full of comparisons where it has not been checked, and a great many reported improvements are inside the range two identical configurations would differ by anyway.

Viz primitive · budget-splitcomparisons-made = 4

comparisons-made holds 17% of the budget; rest holds the remaining 83%.

Comparisons a results table makes at once, against the twenty a 5% threshold tolerates before one false positive is expected, in comparisons. Drag the comparison count up to watch it pass that mark — at twenty, one significant-looking result is what chance alone produces.

4

Reviewed by opendroid · 2026-08-18

  • arXiv:1803.09578 — Why Comparing Single Performance Scores Does Not Allow to Draw Conclusions About Machine Learning Approaches
  • arXiv:2103.03098 — Accounting for Variance in Machine Learning Benchmarks