Reading the evaluation artifact
Reading the evaluation artifact
A spec differ compares two documents and knows nothing about your code. This measures the gap: the compiler says which generated call sites break, and a classifier that never sees compiler output has to predict it from the diff and a parsed read of the source.
94 compiler-verified call-site breakages across 3 vendors, against a threshold of 60 across 3.
| Measure | Development corpus | Held-out slice |
|---|---|---|
| Compiler-verified breakages | 94 | 265 |
| Admitted call sites | 1,338 | 1,342 |
| Precision | 95.8% | 100.0% |
| Recall | 96.8% | 19.6% |
| F1 | 96.3% | 32.8% |
| False negatives | 3.2% | 80.4% |
| False positives | 0.3% | 0.0% |
| Predictor | FN rate ↓ | FP rate ↓ | Precision | Recall | F1 | TP | FP | TN | FN | Abstention |
|---|---|---|---|---|---|---|---|---|---|---|
| SystemPredicts from the spec diff and a parsed read of the call-site source. Never sees compiler output. | 3.2% | 0.3% | 95.8% | 96.8% | 96.3% | 91 | 4 | 1,240 | 3 | 2.2% |
| Baseline · touches changed operationEvery call site that touches a changed operation is IMPACTED. What a team does today with a spec differ and grep. | 0.0% | 64.4% | 10.5% | 100.0% | 19.0% | 94 | 801 | 443 | 0 | 0.0% |
| Baseline · ERR-level changes onlyThe same move with the differ's own severity filter on. | 19.1% | 58.2% | 9.5% | 80.9% | 17.0% | 76 | 724 | 520 | 18 | 0.0% |
strict: Strict counts an UNKNOWN as a miss wherever the compiler found a breakage: the system did not report a breakage that exists. An UNKNOWN on a clean call site is not counted as a false positive, because a false positive is a call site reported IMPACTED and an abstention is not that.
The answer key is the TypeScript compiler, not a human and not a model: call sites are generated from revision A and typecheck clean against it by construction, then recompiled unchanged against revision B, where tsc emits the labels. A label proves a type-level incompatibility at a generated call site against a generated client — it does not prove production breakage, it does not prove runtime behaviour, and a clean compile does not prove safety. Changes of a class the type system cannot express — maxLength, pattern, minimum, format, enum semantics, auth, rate limits, behaviour — are returned as UNKNOWN, because no call site could be made to fail on them and reporting one as safe would be worse than reporting it as broken. The corpus is synthetic call sites against real vendor specs: the specifications, the diffs and the compiler are real, and the client code is generated.
The rows behind these rates — the actual TypeScript, the compiler diagnostics and the change each verdict is anchored to — are under spec pairs. The bytes every specification was read from are under provenance.