Nearly every semaglutide sales page in this market is quoting a single trial, usually without naming it. The number is fifteen percent, the trial is STEP 1, and the distance between those two facts is where the misreading happens. The distribution underneath that mean — how many participants reached 10%, how many reached 15% — is set out in the semaglutide weight article. What follows is the machinery instead: how the trial was built, who was allowed in, what its headline figure is an estimate of, and what happened to the participants once it ended.
Randomization came before tolerance, not after
STEP 1 was a double-blind trial in 1,961 adults with a body-mass index of 30 or greater, or 27 or greater in the presence of at least one weight-related coexisting condition, and without diabetes. They were assigned in a 2:1 ratio to 68 weeks of once-weekly subcutaneous semaglutide at 2.4 mg or to placebo, with a lifestyle intervention in both groups.[1]
The detail worth holding is when the assignment happened. Participants were randomized at week 0, before anyone had taken a dose, and the escalation to the maintenance dose ran inside the 68 weeks rather than ahead of them — the schedule described in the titration article.
That is the structural opposite of STEP 4, where 902 people took semaglutide for twenty weeks first and only the 803 who had reached and held 2.4 mg were randomized. STEP 1 therefore contains the participants who could not tolerate the climb, and STEP 4 does not. The two trials' discontinuation rates describe different populations and cannot be stacked.
The 2:1 split, and what it was for
An even allocation is the statistically efficient choice, and STEP 1 did not make it. Two-thirds of the 1,961 participants — roughly 1,300 people — went to semaglutide and about 650 to placebo.[1]
The trade is deliberate. An uneven split costs a little precision on the treatment difference and buys a much larger safety database on the drug itself, which is what a regulator reviewing a new dose of an existing molecule most needs. It also means the placebo estimate rests on the smaller of the two groups, so the comparator arm is the noisier half of the comparison even though it is the one nobody scrutinizes.
Two consequences follow for anyone reading adverse-event percentages out of this trial. Rates in the semaglutide arm are the better-measured ones, because that arm is twice the size. And a raw count on the semaglutide side will always look larger than its placebo counterpart before any adjustment, because there were twice as many people available to report it. Percentages carry that adjustment; counts lifted out of a table do not.
The cohort was narrower than the market quoting it
Across the five trials in the program, participants had a mean age of 46.2 to 55.3 years, were mostly female (mean 74.1% to 81.0%), and carried a mean body-mass index of 35.7 to 38.5 with a mean waist circumference of 113.0 to 115.7 cm.[2] This was a cohort with class II obesity on average, four-fifths women, and no diabetes.
A later post hoc analysis reported the racial and ethnic composition. Pooling STEP 1 with STEP 3, participants reported race as White (75.3%), Asian (10.6%), Black (8.8%) or another racial group (5.3%), and ethnicity as Hispanic or Latino (13.9%). There were no significant interactions between the treatment effect and race (p ≥ 0.07) or ethnicity (p ≥ 0.40).[5]
The absence of an interaction is the useful half: the drug behaved the same way across those subgroups. The composition still matters for a different reason, which is that a reader with type 2 diabetes is not in this trial at all. The figure that applies to that reader came from a separate trial and is roughly a third smaller, which is the subject of the STEP 2 article.
What the headline number is an estimate of
STEP 1 stated its estimand in the paper itself: the primary estimand assessed effects regardless of treatment discontinuation or rescue interventions.[1] That single sentence changes what the percentage means. It is the answer to “what happens to a group of people prescribed this drug for 68 weeks,” with the people who stopped taking it still counted in the denominator.
A 2025 methods paper uses STEP 1 as its worked example and sets out the alternatives. The original analysis applied a treatment policy strategy, which views nonadherence as an aspect of the treatment regimen and makes no adjustment for it. A supplementary analysis used a hypothetical strategy, targeting the effect that would have been realized had all participants adhered. The paper proposes an instrumental variable method that avoids the unconfoundedness assumption the hypothetical strategy relies on, and its estimates in STEP 1 suggest a sustained, slowly decaying treatment effect on weight.[3]
The commercial consequence runs against the usual framing. A figure built on a treatment policy strategy already has real-world dropout inside it, so quoting it as a best case available only to a perfect adherer inverts what it measures.
Two endpoints, and the thing neither of them counted
The coprimary endpoints were the percentage change in body weight and a weight reduction of at least 5%.[1] In absolute terms the change from baseline to week 68 was −15.3 kg against −2.6 kg on placebo, an estimated treatment difference of −12.7 kg (95% CI, −13.7 to −11.7). Cardiometabolic risk factors improved and participant-reported physical functioning rose more than on placebo.[1]
What STEP 1 never counted was an event. No heart attack, stroke or cardiovascular death was an endpoint, because the trial was neither sized nor run long enough to accumulate them. Answering that question took a separate trial of 17,604 adults who already had cardiovascular disease and a body-mass index of 27 or greater without diabetes.[6] Those results, and their limits, are in the cardiovascular article.
What ended treatment
Nausea and diarrhea were the most common adverse events with semaglutide. They were typically transient, mild to moderate in severity, and subsided with time. More participants on semaglutide than on placebo discontinued because of gastrointestinal events: 59 (4.5%) against 5 (0.8%).[1]
Roughly one participant in twenty-two left for that reason, in a trial with study staff, scheduled visits and a protocol governing the pace of escalation. The timeline those symptoms follow is in the nausea article, and the fuller safety record is in the side-effect article.
The year after the drug and the program both stopped
At week 68 both treatment and lifestyle intervention were discontinued, and an off-treatment extension followed a subset of 327 participants for a further year. Within that subset, mean weight loss from week 0 to week 68 was 17.3% (SD 9.3) with semaglutide and 2.0% (SD 6.1) with placebo. By week 120 the two groups had regained 11.6 (SD 7.7) and 1.9 (SD 4.8) percentage points, leaving net losses from baseline of 5.6% (SD 8.9) and 0.1% (SD 5.8). Cardiometabolic improvements reverted toward baseline for most variables.[4]
That 17.3% is not a correction to the trial's 14.9%. It describes the 327 people who entered the extension, reported as an observed mean rather than as the trial's primary estimand — a different denominator and a different statistic. Anyone stacking the two has built a number that neither paper contains.
The standard deviation of 8.9 around a net 5.6% is the more honest summary of the year off. That interval spans participants who held nearly everything and participants who finished above their starting weight, and the pricing consequence of that spread is the subject of the stopping article.
What was actually in the syringe
Every figure above belongs to branded, FDA-approved semaglutide at a labeled dose, escalated under supervision, with a supply that never lapsed and a lifestyle program running in both arms. Most sellers on the semaglutide board dispense compounded preparations instead. Compounded drugs are not FDA-approved and are not reviewed by the FDA for safety, efficacy or quality before they are dispensed, and the distinction is set out in the compounded-versus-brand article.
STEP 1 is a strong trial and its result is real. It is also a description of one molecule, at one dose, in one carefully bounded population, measured under a stated estimand and reported with the dropouts left in. Each of those qualifiers is missing from the version that reaches a checkout page.