Nobody in SURMOUNT-5 was blinded. That is the first thing to know about the trial the whole market quotes, and it is stated in its own design: phase 3b, open label, no masking of participants or investigators.[1][2] The molecule question itself — which drug, at what price, for whom — belongs to the comparison article. What is below is the trial: how it was built, which of its numbers are observed and which are modeled, and the four things it was never constructed to settle.
The design, and the one thing it could not hide
751 adults with obesity but without type 2 diabetes were randomly assigned 1:1 to the maximum tolerated dose of subcutaneous tirzepatide (10 mg or 15 mg) or the maximum tolerated dose of subcutaneous semaglutide (1.7 mg or 2.4 mg), once weekly for 72 weeks. The preregistered primary endpoint was a single measure: the percent change in body weight from baseline to week 72. Key secondary endpoints were weight reductions of at least 10%, 15%, 20% and 25%, and the change in waist circumference.[1]
Open label is a real limitation and a survivable one here, because the primary endpoint is a scale reading rather than a judgment. A participant who knows which drug they received cannot talk a balance into a different number. It matters far more for the secondary and post hoc endpoints that depend on how someone says they feel, and the trial could not have been blinded anyway: the two products ship in different devices with different escalation schedules, and the registry records injection-site reactions in 32 of 374 tirzepatide participants against 1 of 376 on semaglutide.[2] A blind that thirty-two people could see through their own skin was never on offer.
What that costs is specific rather than general. It does not threaten the primary endpoint, and it does not threaten the threshold proportions below, which are also scale readings. It does threaten anything a participant reported about themselves — the quality-of-life scores, the symptom counts, the decision to keep going — because a person who believes they were assigned the stronger drug is not a neutral observer of their own experience. The right posture is to treat this trial as settled on weight and provisional on everything a person had to say out loud.
The headline, with its interval
At week 72 the least-squares mean percent change in weight was −20.2% (95% CI, −21.4 to −19.1) with tirzepatide and −13.7% (95% CI, −14.9 to −12.6) with semaglutide (P < 0.001). Waist circumference fell 18.4 cm (95% CI, −19.6 to −17.2) against 13.0 cm (95% CI, −14.3 to −11.7).[1] The registry states the contrast directly: a least-squares mean difference of 6.5 percentage points (95% CI, 4.9 to 8.1) on weight and 5.4 cm (95% CI, 3.6 to 7.1) on waist.[2]
An interval running from 4.9 to 8.1 points is the useful part. It does not include zero, so the direction is settled; it is also four and a half points wide, so a reader who quotes “six and a half points” as though it were a constant is quoting the midpoint of a range the trial itself did not narrow further. The individual spread underneath that mean is wider still, which is why the same trial produced participants who lost a third of their body weight and participants who did not reach 10%.
The dose objection, and the analysis that answers it
The standing objection to any tirzepatide-versus-semaglutide result is that semaglutide was underdosed — the objection that makes SURPASS-2 unusable for obesity, as the semaglutide weight article sets out. SURMOUNT-5 anticipated it. Both arms titrated to the maximum each participant tolerated, and a prespecified secondary analysis estimated the comparison at the top of both ladders: −21.8% against −15.4%, a difference of 6.4 percentage points (95% CI, 4.7 to 8.0).[2]
Read what that analysis is before using it. The registry describes it as an effect “evaluated assuming that participants had stayed on treatment and reached the highest dose of treatment” — a mixed-model estimate of a counterfactual, not a subgroup of people who actually reached 15 mg and 2.4 mg.[2] It is the most misquoted number in this trial, routinely presented as though it were an observed head-to-head at label doses. Its value is narrower and still real: the gap does not close when you model both drugs at their ceilings, so the underdosing objection does not rescue the smaller figure.
Where the difference actually lives
The threshold ladder above is where the two arms separate most. At 10% or more the arms were 87.7% against 66.7%; at 20% or more, 55.0% against 31.1%; at 30% or more — a threshold the publication’s abstract does not report at all — 23.0% against 8.2%.[2] The relative gap roughly triples as the bar rises. These are observed proportions posted without confidence intervals, so they carry less precision than the primary contrast, and they should be read as shape rather than as point estimates.
The same asymmetry shows up in how fast people got there. A post hoc analysis defined a rapid responder as someone reaching 15% reduction by week 24: 44% of the tirzepatide arm and 21% of the semaglutide arm qualified, and 32.3% of the whole trial. Rapid responders reported numerically more gastrointestinal and hepatobiliary events in both arms, and completed treatment at rates similar to everyone else.[6]
Two drugs, one side-effect table
This is the result nobody markets. Over 72 weeks the registry records nausea in 163 of 374 tirzepatide participants and 167 of 376 on semaglutide; constipation 101 against 107; diarrhea 88 against 88. Vomiting and reflux ran the other way — 56 against 80, and 23 against 40. Serious adverse events occurred in 18 against 13, and no participant in either arm died. Discontinuation for an adverse event was 6 against 6. Overall, 319 of 375 and 319 of 376 completed.[2]
A reader who assumes the more potent drug must cost more tolerability does not get that from this trial. Both arms produced roughly the same volume of gut symptoms, and the same handful of people walked away from each. What the class costs generally is in the side-effect article; what this trial adds is that between these two molecules, at tolerated doses, the cost was not the variable that separated them.
Attrition says the same thing from the other side. 56 of 375 tirzepatide participants and 57 of 376 on semaglutide did not complete, and the reasons line up almost one for one: 27 against 26 withdrew consent, 16 against 18 were lost to follow-up, 6 against 6 left for an adverse event, one in each arm stopped for non-compliance, and 5 against 3 were recorded as having been assigned treatment by mistake.[2] Very little of that is pharmacology. It is what a 72-week commitment costs in ordinary life, and it fell on both arms equally.
Beating a comparator is not reaching a target
A post hoc analysis applied proposed treat-to-target thresholds to the trial. Reaching a waist-to-height ratio below 0.53 or a body-mass index below 27 was achieved by 23.1% to 33.9% of tirzepatide participants and 14.2% to 20.7% of semaglutide participants. Of those who reached the waist-to-height threshold, about 77% also met goals on at least four of five cardiometabolic risk parameters, an odds ratio of 2.31 against those who did not (p < 0.001); the body-mass index threshold was not statistically associated with the physical-function outcome assessed.[4]
Two thirds of the winning arm did not reach either target. Superiority on a mean and arrival at a clinical goal are different claims, and a seller quoting the first while implying the second is trading on the gap.
Which analysis set produced which number
SURMOUNT-5’s figures move depending on who is counted. The primary −20.2% and −13.7% come from all randomized participants who received a dose. A post hoc analysis restricted to the 425 participants with prediabetes at baseline, using the efficacy analysis set, reports −21.5% against −14.5% (estimated treatment difference −7.1%; 95% CI, −9.1 to −5.0), with mean HbA1c falling 0.60% against 0.48% (difference −0.12%; 95% CI, −0.18 to −0.06; p < 0.001) and 89.9% against 76.2% reverting to normoglycemia.[3]
Those are larger weight numbers from the same trial, and both sets are honest. They differ because the population is narrower and the estimand censors differently. Anyone comparing a figure they have read somewhere against a figure here should check which of the two it came from before concluding that one of them is wrong.
What the trial did not measure
No cardiovascular, renal or mortality endpoint was among its primary or key secondary outcomes, and with zero deaths across 750 treated people it could not have addressed one. It enrolled nobody with type 2 diabetes. It ran for 72 weeks, so it says nothing about year three, and nothing about what happens after stopping. On patient-reported outcomes the gains were mostly shared: physical component scores improved with both treatments (p < 0.001), mental component scores improved with neither (p > 0.05 in both arms), and the one domain separating the drugs was General Health, 5.45 against 4.20 (p = 0.003).[5]
It also describes one country. All 750 treated participants enrolled in the United States; mean age was 44.7 years, 485 were women, and 571 were White, 144 Black or African American and 18 Asian.[2] That is a narrower base than the multinational obesity trials this result is usually set beside, including the trial in the SURMOUNT-1 article.
Finally, both arms received branded, FDA-approved product, supplied without interruption, escalated on a protocol rather than a budget, with no dose lapsing between shipments. Sellers on the tirzepatide board overwhelmingly dispense compounded preparations, which are not FDA-approved and are not reviewed by the FDA for safety, efficacy or quality before they are dispensed. A head-to-head between two approved products is evidence about those two products, and the distance between them and a compounded vial is set out in the compounded-versus-brand article.