the.ai

Code / Verification

verified

Execution Verification

Run it and see. Code is checkable in a way essays are not, and that changes what is possible: you can generate a hundred candidates, keep the ones that pass, and never ask a human. But a test suite is not a specification — passing every test you have is evidence, not proof, and the gap between them is where this gets interesting.

Viz primitive · threshold-sweepseparation = 1.4 · threshold = 0.8 · base-rate = 0.35
let throughcutflagged

390 of 1000 flagged. 66% of them were right and 133 were false alarms; 73% of what should have been caught was, leaving 93 missed.

Generated programs that are actually correct, against ones that are not, scored by how thoroughly the tests exercise them. Drag the suite's thoroughness up to watch them separate — at the left a great deal of what passes is wrong, and no threshold helps, because a test has none.

1.4

Reviewed by opendroid · 2026-08-18