the.ai

Evaluation / Metrics

verified

pass@k

For code, whether a single answer is right is the wrong question — a developer can generate several attempts and run the tests. pass@k asks whether at least one of k samples passes, which measures something closer to how the model is actually used. It also rises with k for free, so the k has to be reported or the number means nothing.

Viz primitive · budget-splitk = 10

k holds 10% of the budget; rest holds the remaining 90%.

Share of problems solved by at least one of k samples against those still unsolved. Drag k to watch the metric climb without the model changing.

10

Reviewed by opendroid · 2026-08-04