Peer-Reviewed Benchmarks
Our validation methodology, the metrics we measure, and the status of peer-reviewed publications from customer sites.
Status: Published peer-reviewed benchmarks forthcoming
We do not currently link to peer-reviewed studies here because none have reached publication from customer sites. We will not publish fabricated, pre-print, or selectively-reported numbers in their place. The framework below describes what we measure, how we measure it, and where to get the latest early-pilot validation data for evaluation.
Validation Methodology
Radiologist-Supervised
Every metric is computed against a board-certified radiologist's signed report, never against another model's output. The radiologist is the reference standard.
Prospective, Not Just Retrospective
Pilot sites run on production volume during the evaluation window. Retrospective benchmarking alone is insufficient for clinical reporting claims.
Site-Stratified
Metrics are reported per site, per modality, per study type, and per radiologist. A single aggregated number hides clinically meaningful variance.
Edit-Distance, Not Just Acceptance
"Did the radiologist accept the draft?" is a weak signal. We measure edit-distance per section, per finding, and per critical-result mention.
Critical-Result Recall
We track whether AI-surfaced critical findings reach the radiologist's attention and trigger the site's documented critical-results pathway.
Audit-Trail Verification
Every pilot-signed report is reconciled end-to-end against the ORU^R01 round-trip with the RIS. Audit integrity is a hard gate, not a metric.
Forthcoming Peer-Reviewed Publications
This page will link to each publication here as soon as it is in print or online-first at a peer-reviewed venue. Until then, the slots below remain placeholders.
Customer-site study #1 — TAT & Edit-Distance
Prospective pilot at a single radiology site evaluating turnaround time and edit-distance on AI-drafted reports across a defined modality mix.
Customer-site study #2 — Critical-Result Workflow
Evaluation of critical-finding detection, escalation fidelity, and audit-trail integrity in a supervised AI-assisted reporting workflow.
Customer-site study #3 — Multi-Site Rollout
Cross-site comparison of edit-distance, radiologist satisfaction, and operational metrics after broader deployment.
How We Will Not Report Numbers
- No selective reporting. Every metric we publish will include denominators, confidence intervals where applicable, and the per-site variance.
- No reference to "consensus" of other AIs. The reference standard is always a board-certified radiologist's signed report.
- No pre-prints labeled as peer-reviewed. Pre-prints will be clearly labeled pre-print, and the peer-reviewed version will replace them when available.
- No published numbers without named site and IRB-equivalent oversight. We will not publish customer results without the customer's review and approval.
- No citation of internal eval without provenance. Internal evaluations are explicitly distinguished from externally-published studies.
Get the Latest Validation Data
If you are evaluating ApertureAI for a 30-day pilot or broader rollout, we'll share the most recent internal validation summary, including edit-distance and TAT distributions from the current pilot cohort. No fabricated numbers — just the latest real ones.
Request Validation Data