Semaglutide has two US prescribing labels, and they report the same adverse reaction at 20% and at 5.7%. The weight-management label records abdominal pain in 20% of 2,116 treated adults against 10% of 1,261 on placebo.[1] The diabetes label, for the same molecule, records 5.7% of 261 patients at 1 mg against 4.6% of 262 on placebo.[2] A quarter of the dose does not explain a gap that size, and it does not explain why the placebo arms differ too. What explains it is printed in a footnote most readers never open.
The word is not the same word
On the weight label, the row headed “Abdominal Pain” carries a footnote stating that it includes abdominal pain, abdominal pain upper, abdominal pain lower, gastrointestinal pain, abdominal tenderness, abdominal discomfort and epigastric discomfort — seven preferred terms collapsed into one line.[1] The diabetes label’s row carries no footnote at all, because it counts one term.[2]
Tirzepatide does the same thing in the same direction. Its weight label folds abdominal discomfort, abdominal pain, abdominal pain lower, abdominal pain upper and abdominal tenderness into a single row reading 5%, 9%, 9% and 10% across placebo (n=958) and the 5, 10 and 15 mg maintenance doses.[3] Its diabetes label uses the bare term and reads 4%, 6%, 5% and 5% across placebo (n=235) and the same three doses.[4] Two documents, one molecule, one dose, and roughly double the rate on the one that bundles.
The liraglutide label demonstrates the arithmetic by refusing to bundle. It prints Abdominal Pain at 5.4% against 3.1% and Upper Abdominal Pain at 5.1% against 2.7% as two separate rows in the same table of 3,384 treated adults against 1,941 on placebo.[5] Those are two of the seven terms the semaglutide weight label merges. Anyone setting liraglutide’s 5.4% beside semaglutide’s 20% is comparing one preferred term against seven, and the oral semaglutide label — a bare term again, at 10% and 11% for the 7 and 14 mg doses against 4% on placebo — sits somewhere in between.[6] The neighboring row for distension is counted consistently across all six and is a much weaker signal, which is a useful reminder that these tables are not interchangeable even inside one document.
The placebo arms settle it
The strongest evidence that these percentages are not comparable is the column nobody reads. Untreated abdominal pain across the six labels runs 3.1%, 4%, 4.6%, 5%, 6%, 7% and 10% — a threefold spread among people receiving nothing at all.[1][2] [3][4][5][6] Trial populations, questioning methods and coding conventions move the control arm by more than most of these drugs move their own treatment arm. A comparison table that ranks molecules by the drug column, which is what most of them do, is ranking the paperwork.
The definition does not explain everything, and the honest version says so. The pediatric section of the semaglutide weight label uses the bare term, with no composite footnote, and still records 15% of 133 adolescents against 6% of 67 on placebo.[1] A nine-point gap on an unbundled term is a real drug effect, and nothing in the labeling accounts for why the adult diabetes pool produced a one-point gap on the same term.
The dose curve that is not there
One table supplies its own control. In the tirzepatide diabetes pool, across placebo and the 5, 10 and 15 mg doses, nausea runs 4%, 12%, 15% and 18% — the clean ascending curve a dose-driven reaction produces. In the same table, in the same patients, abdominal pain runs 4%, 6%, 5% and 5%.[4] The highest dose sits one point above placebo and below the lowest dose.
The pattern repeats wherever doses are separated. Semaglutide for diabetes reports 7.3% at 0.5 mg and 5.7% at 1 mg — the higher dose lower.[2] Tirzepatide for weight moves from 9% to 9% to 10% across a threefold dose range.[3] Semaglutide for weight, in the trial that tested 7.2 mg against 2.4 mg and placebo, reports 7%, 9% and 12% using the identical seven-term composite — a three-point rise for a threefold dose, from a 2.4 mg figure of 9% that the pooled table puts at 20%.[1] Nausea and vomiting escalate with exposure. This does not, which is why the escalation schedule in the titration planner is a poor predictor of when pain arrives.
The randomized estimates rank the molecules differently
A dose-response network meta-analysis of 39 reports covering 33,354 adults with overweight or obesity and without diabetes pooled the two pain terms separately. Across the 12 trials reporting abdominal pain, the relative risks were 2.34 (95% CI, 1.41 to 3.89) for semaglutide, 2.08 (1.06 to 4.07) for liraglutide and 4.36 (1.29 to 14.78) for tirzepatide, with orforglipron at 1.46 (0.26 to 8.18) and exenatide at 0.33 (0.01 to 7.95), neither significant.[7]
Two things follow. Liraglutide, whose label prints the smallest percentages of the group, is statistically indistinguishable from semaglutide on relative risk — the intervals overlap almost entirely. And tirzepatide’s point estimate is the highest while its interval spans a factor of eleven, which is what a handful of events looks like once it is expressed as a ratio.
The 13 trials reporting upper abdominal pain behave better. Semaglutide came out at 2.14 (1.66 to 2.77) and liraglutide at 1.69 (1.23 to 2.32), both significant, with intervals roughly a third as wide.[7] One symptom family, one population, a similar number of trials, and a tenfold difference in precision — because the trials that report a bundled term are not the trials that report a specific one. A label that merges the two merges a well-measured endpoint with a poorly measured one. The molecule-level comparison people actually want is laid out in the head-to-head article, and abdominal pain is not one of the outcomes that separates them.
The reporting data reverse the order again
A disproportionality analysis of 81,752 adverse event reports in the FDA system, 21,281 of them gastrointestinal, ranked the class by reporting odds ratio. Semaglutide carried the largest signal for nausea (ROR 7.41, 95% CI 7.10 to 7.74), vomiting (6.67), constipation (6.17) and diarrhea (3.55). It did not lead upper abdominal pain. Liraglutide did, at 4.63 (4.12 to 5.21).[8] The same paper put overall gastrointestinal reporting at 3.00 for semaglutide, 2.39 for liraglutide and 1.39 for dulaglutide, and found liraglutide carried the highest proportion of serious gastrointestinal reports at 23.31% against 12.2% for dulaglutide.[8]
So three instruments produce three orderings. The labels put semaglutide highest, the randomized pooling cannot separate semaglutide from liraglutide, and the spontaneous reports put liraglutide highest for the specific term. A reporting database has no denominator and counts reports rather than patients, so it cannot settle the question — but it can show that the label ranking is not robust to changing the instrument.
Timing, and the line the label draws
The same analysis found most gastrointestinal reactions in this class occurred within the first month of treatment.[8] That is the window in which pain is ordinary: arriving during escalation, tied to meals, easing between them, fading as the dose holds. Permanent discontinuation for a gastrointestinal reaction occurred in 4.3% of treated adults against 0.7% on placebo in the semaglutide weight program — a sixfold difference, and still a figure meaning 95 of every 100 people continued.[1]
The label’s own warning language describes something else entirely. It instructs monitoring for “persistent or severe abdominal pain (sometimes radiating to the back), and which may or may not be accompanied by nausea or vomiting,” and instructs stopping the drug if pancreatitis is suspected.[1] The distinguishing features are persistence and escalation rather than site or intensity at a single moment: pain that does not remit between meals, that wakes a person, that builds across days rather than settling. An emergency medicine review of this class lists the presentations that belong in that category — pancreatitis, biliary disease, renal injury following severe gastrointestinal losses, and hypoglycemia where a sulfonylurea or insulin is also being taken.[9] The biliary and pancreatic event rates behind that warning are in the gallbladder and pancreatitis article, and pain accompanied by a distended abdomen that will not pass stool or gas is the separate picture handled in the obstruction article. Reaching for an anti-inflammatory to manage it carries its own problem, set out in the NSAID article.
What a compounded vial adds to the picture
Every percentage above was generated by an FDA-approved product at a labeled dose under trial supervision. Most sellers covered here dispense compounded semaglutide or tirzepatide, which is not FDA-approved and is not reviewed by the FDA for safety, efficacy or quality before it is dispensed, and no surveillance system reports abdominal pain incidence for those products separately.
The nearest measurement comes from a statewide poison center. Across 1,047 human exposure cases to this drug class between December 2017 and December 2023, abdominal pain was the reported effect in 54 cases (5.1%), behind nausea (28.0%) and vomiting (25.5%). The dominant reason for contact was not an adverse reaction at all: 80.0% of exposures were unintentional therapeutic errors. Two hundred twenty patients (21.0%) were treated in an emergency department and 46 (4.4%) were admitted. Thirty-six exposures involved a compounded product, and among those, administration errors accounted for 33 of 36 (91.7%).[10] A separate analysis of 3,348 accidental overdose reports found significant disproportionality across every agent in the class, with reporting odds ratios ranging from 2.64 to 61.12.[11]
That is the specific hazard a multi-dose vial and a syringe introduce, and it is not the hazard the incidence tables describe. Pain following a dose drawn at four times the intended volume is not the labeled adverse reaction; it is a measurement error with a clinical consequence. The broader tolerability picture from the trials themselves is in what the trials recorded.
What this leaves
Abdominal pain is a real and reasonably common effect of this class: significant in pooled randomized data at roughly 2.1 to 2.3 times placebo for the two molecules with enough trials to measure, and clearly elevated on the unbundled adolescent term at 15% against 6%. What it is not is comparable across products, predictable from dose, or reliably ranked by the numbers printed on the labels. A reader deciding between two molecules on the strength of a 20% against a 5.4% is reading a difference in drafting convention.