0→1 radiology platform, multi-tenant, FDA SaMD-compliant. We pivoted the MVP from speed-first to interpretability-first, and adoption followed.
Post-scan, radiologists at partner hospitals were spending hours structuring narrative findings into reports that downstream clinicians, billing, and QA all needed. Existing AI assistants were fast, and ignored. The assumption was that speed was the win. Discovery told a different story.
We ran structured interviews with radiologists, referring clinicians, and QA leads. The pattern was unambiguous: they didn't trust opaque outputs. A model that suggested a finding without showing why got dismissed. A model 15% slower but that surfaced source imagery regions, cited prior scans, and flagged uncertainty got adopted.
We pivoted the MVP. Interpretability first, speed second. The product spec changed in week four, and that was the week the project started working.
The rebuilt MVP shipped with: region-level evidence highlighting, confidence intervals per finding, one-click model-feedback capture, and an audit log aligned to FDA SaMD requirements. We drove NLP accuracy from 68% to 88% through structured feedback loops: radiologists flagged errors inline; those corrections flowed back into prompt tuning and edge-case handling within the sprint.
The platform launched across the client's partner network. ARR surpassed $500K within six months, faster than the original velocity-first spec projected, because adoption was the constraint, not feature count. Downstream, the structured data pipeline we designed eliminated 4,000+ manual hours per year of post-processing across partner sites.
Interpretability beats speed in clinical AI. Radiologists didn't adopt the faster version, they adopted the one they could defend to a patient. The MVP that shipped a week later but showed its reasoning is the one that got used.
Compliance and adoption aren't separate asks. The audit log and clinician sign-off that FDA SaMD required were also what built the trust that drove adoption. Building for compliance built the case for the product.
Feedback loops beat one-time accuracy pushes. 68% to 88% didn't come from a bigger model, it came from making every clinician correction traceable back into prompt tuning within the same sprint.