Each file you dropped was parsed and its Ed25519 signature checked against the published key table, exactly as plan_d/src/verify_passport.py does — same checks, same order, same exit codes. A candidate whose signature does not verify, or whose evidence has expired, is excluded from every number on the page rather than ranked last.
The numbers are read straight out of each passport’s computed block. Nothing is recomputed and nothing is combined: there is no composite score anywhere on this page, and no row averages two tasks. Models are grouped by the task they were evaluated on and ranked only inside a group; where two candidates share no task, the page says so in words instead of leaving a cell empty. The per-task composite index that the passports do carry is deliberately not shown here — on a procurement screen the largest bold figure gets read as “the score” and compared across rows it does not span.
Rows run worst-case first — worst subgroup, then worst acquisition setting, then cross-site retention — with headline accuracy last, because the floor is what a hospital is buying. The vendor’s own declarations sit in their own band, styled so they can never be mistaken for a measurement.
The pass/fail column is the Trust Runtime evaluator run with no inference envelope: the same engine that gates a single request answers “may we buy this at all” at procurement binding time. Its thresholds are an example and are labelled as one.
Benchmarks