# Scoring rules actually applied (deviations from the brief are marked ⚠)

This is the scoring methodology for the Spark judgment-seat bench. It is published as the "how we measured"
record for the study. Cell-level defect detail and the ground-truth answer key are held privately, because
they describe unreleased fixes in a private codebase; only the aggregate scores (see `scores.csv`) and this
methodology are public.

1. **Match** = semantic: same defect, cite landing in the same function/block. Anchor-independent.
2. **weighted_recall** = Σ weight(matched musts) / Σ weight(all musts). Weights 3/2/1.
3. **opus_crit_cov** = matched opus-critical / total opus-critical.
4. **precision** = true findings / total findings. TRUE = matches a ground-truth must, OR the overseer
   verified it as a real, previously-unlisted defect (those earn precision, never recall).
5. **anti_hits** = findings matching a pre-registered anti-hit claim, plus FABRICATED cites.
6. ⚠ **Citation policy — THREE buckets, because the brief's single rule proved ambiguous once arms were read.**
   Every arm reads the LIVE working tree, so cites are judged against the LIVE tree, not the merge commit.
   - **fabricated** (counts as an anti-hit): the claimed code does not exist in the cited file at all, or the
     claim misdescribes what the code does.
   - **anchor drift** (tracked separately, NOT an anti-hit): the quoted code exists verbatim in the cited
     file but at a different line.
   Both counts are reported so the reading intended by the owner can be applied.
7. ⚠ **`false_findings`** — a claim the overseer verified as NOT real that is also not on the pre-registered
   anti-hit list. The brief's formula has no bucket for these; they are counted false for precision and
   reported in their own column, never silently dropped.
8. **severity_fidelity** = fraction of matched musts where the arm's severity ≥ the ground-truth band
   (opus-critical → critical/major, should → major/minor, nice → any).

## ⚠ Ground-truth citation anchor correction (disclosed)

Ground-truth cites were first written against each cell's merge commit. The repository's working tree has
since moved past every cell's merge, and every arm reads and cites the WORKING TREE. No ground-truth
finding, tag, weight or anti-hit was changed after any arm launched — only the understanding of which tree
the line anchors refer to. Matching is semantic (rule 1), so this does not move any match decision; it only
prevents mis-scoring a correct arm cite as fabricated.

### Citation buckets (final — three, not two)
- **fabricated** — the claimed code is absent from the cited file, or the claim misdescribes it. ANTI-HIT.
- **out_of_range** — the cited line number is past the file's end. Reported separately: it is the strongest
  available evidence that a number was generated rather than read, even when the quoted code is real.
- **drift** — right file, real code, wrong line. Not an anti-hit.
