A pilot that starts with two approvals

For the first 100 patients in Nolla’s Utah pilot, a prescription reaches the patient only after two licensed physicians independently review and approve it. In the next stage, the system issues prescriptions directly and a physician reviews every case afterwards, at least weekly. That shift in timing is the pilot’s central experiment. A mistake caught in the first stage can be intercepted before delivery; the same kind of mistake in the second stage is discovered inside a decision already issued. Measuring agreement between the system and physicians provides a starting point for understanding how these two arrangements work for patients.[1]

The pilot’s boundaries make the comparison easier to interpret. Nolla selects among preapproved topical plans for adults in Utah with mild to moderate acne. It uses structured questions and a five-angle facial scan, without a free-text prescribing interface. Severe acne and pregnancy are referred to a physician or in-person care. A narrow starting point can be a reasonable basis for gradually reducing review: the decision space is constrained, excluded groups are identified and there is a route to human assessment. That rationale does not require this pilot to represent every medical decision.[1]

The denominator behind agreement

Nolla’s announced progression thresholds are at least 95% physician agreement, zero serious adverse events and written state approval. The first stage lasts at least four weeks and the second at least eight. Together these requirements form an advancement rule, but each counts something different. Physician agreement measures how closely a proposed plan matches clinical judgment. Adverse-event measurement concerns what happens to the patient during treatment. Whether the agreement denominator includes all applicants or only those receiving prescriptions determines how difficult cases referred elsewhere affect the success calculation.[1]

The two independent physician assessments in the first stage provide a useful baseline for that reason. Disagreement between the physicians should also enter the comparison; otherwise one human decision becomes a supposedly perfect reference. I would report the AI choice, each physician’s choice and the final decision delivered to the patient separately. That prevents a small change in topical strength and a missed exclusion from disappearing into one disagreement category. Narrow protocols may indeed produce high agreement; an alternative explanation is the constrained task design, rather than broad clinical competence.[1]

Review that reaches the patient

In stage three, monthly review samples at least 10% of prescriptions, with every side-effect and escalation case reviewed as well. Sampling design becomes consequential here. Randomly selected routine cases can produce a different picture from patients who volunteer feedback. Outcomes among people who never send a message, stop scanning or seek care elsewhere differ from reports received inside the app. Showing those routes separately in the pilot’s evaluation would help distinguish low reporting from loss to follow-up.[1]

The most useful results table for this pilot would show applicants, referrals, prescriptions changed by a physician and the time to each correction, separately for every stage. Follow-up duration and completion should appear beside those counts. That reporting design would allow a direct comparison of access after pre-prescription review is removed and of how errors reaching patients are caught. Nolla’s specified weekly and monthly review schedules offer concrete starting points for this comparison. The pilot’s scientific value rests on interpreting its agreement figure within that care pathway.[1]