What the result says

The three-arm, open, randomised trial enrolled 637 patients, of whom 633 entered the analysis, across 33 sites, most of them community-based. The primary outcome was time to healing as determined by blinded assessors, with follow-up of up to 12 months. The hazard ratio was 0.78 (0.61–1.00; p = 0.046) for compression wraps against established evidence-based treatment and 0.79 (0.61–1.01; p = 0.056) for wraps against the two-layer bandage. 12-month healing was 76.2%, 79.4% and 80.1% respectively.[1]

The conclusion on whether the two-layer bandage is inferior to established treatment takes two forms. Under the treatment-policy strategy the hazard ratio is 1.01 (0.79–1.28) and meets the prespecified margin of 1.33; under the hypothetical strategy it is 1.16 (0.86–1.58) and misses that margin. The trial plan prespecified these two estimands before the data were seen. The first asks what follows from offering a method in real practice with treatment switching; the second asks what might have happened without those changes. The non-inferiority interpretation changes with the clinical question being answered.[1]

Where the power went

Two limitations the authors report themselves are decisive here. 455 of the participants, that is 71%, received a compression treatment other than the one they were allocated at some point during follow-up. Healing also occurred less often than the sample size calculation assumed, so the number of observed healing events fell well short of what 80% statistical power required, and slight under-recruitment came on top of that. When low power meets high crossover, a result whose confidence interval touches exactly 1.00 becomes hard to read.[1]

The failure mode should be named: the result is narrow rather than wrong. The difference between p = 0.046 and p = 0.056 reflects the number of events, not a difference between the two treatments. The main message, that wraps do not speed healing, stands, because the 12-month healing rates in the three arms are close together and all point the same way. What does not stand is any precise statement about the size of the difference.[1]

What the next trial has to fix

Fairness is due: a trial run in community settings, with blinded assessors and follow-up of up to 12 months, is hard to do, and the team reported its shortfalls plainly. Three priorities follow for the next trial: base sample size on the healing incidence observed here, measure the direction and reason for treatment changes over time, and make clear from the outset which estimand corresponds to which clinical decision. The estimands in this study were already prespecified. The remaining priority is power and adherence data strong enough to answer both questions securely.[1]