Frontier scope (kept peripheral, not deepened) — a class efficacy-and-limitations comparison (surrogate-vs-hard-outcome inversion, benefit-harm coupling, lean-mass cost, durability) against the lifestyle reference is IN under the Pharmacotherapy taper; the per-agent which-to-prescribe ranking across nineteen drugs — several of them non-standard/investigational (CagriSema, orforglipron, retatrutide) — is the pharmacopeia depth that stays frontier. Read this page for the class-level decision lessons; the per-agent ranking / discontinuation / GI tables are the peripheral edge and are not deepened (no further head-to-head-ranking acquisition).
Nineteen drugs, 262 RCTs, 99 791 participants, 24 outcomes, GRADE-rated — the single most comprehensive comparison the wiki holds of the anti-obesity drug class (Nong et al., 2026). Because most of these drugs have never been tested head-to-head, the comparison is a network meta-analysis against a common reference (lifestyle modification alone), and its value is the one thing no single-drug page can carry: a ranking across the class on several outcomes at once. The headline is not the ranking on weight — it is that the ranking on weight is not the ranking on the outcomes people actually care about.
The source frames its own reason for looking past weight: «Obesity is increasingly recognised as a complex chronic disease and reliance on weight loss alone as a measure of treatment success may oversimplify benefits and harms while reinforcing stigma» (Nong et al., 2026). That is this wiki’s Surrogate Outcomes discipline stated from inside the guideline machinery — weight is the surrogate; mortality, cardiovascular and kidney events, quality of life, and function are the patient-important outcomes.
The weight-loss ranking (moderate-high certainty, vs lifestyle alone, at one year)
The drugs supported by moderate-to-high certainty split into a clear order on percentage body-weight change (Nong et al., 2026):
| Drug | MD % body weight (95% CI) | Category |
|---|---|---|
| Tirzepatide | -14.9 (-16.0 to -13.9) | among most effective |
| CagriSema (cagrilintide-semaglutide) | -14.8 (-16.9 to -12.7) | among most effective |
| Oral semaglutide | -10.9 (-12.7 to -9.1) | intermediate |
| Orforglipron | -9.9 (-12.4 to -7.5) | intermediate |
| Subcutaneous semaglutide | -9.8 (-10.6 to -9.1) | intermediate |
| Phentermine-topiramate | -8.1 (-9.7 to -6.5) | intermediate |
| Liraglutide, naltrexone-bupropion, orlistat, exenatide, SGLT-2i, dulaglutide, metformin | -1.9 to -4.7 | not convincingly different from lifestyle |
Emerging triple/dual agents (retatrutide, ecnoglutide, mazdutide) show similar or larger effects (13.1-14.6%) but at low-to-very-low certainty — «reflecting early stage development rather than absence of effect», an explicit insufficient-evidence state, not a null (Nong et al., 2026). For context, tirzepatide and CagriSema put «approximately a third to a half of treated people» past a >=20% loss within a year versus «about 2% with lifestyle modification alone» — an effect the authors say «approach those reported after some bariatric procedures» (Nong et al., 2026).
The benefit-harm coupling — bigger loss is bought with more harm and dropout
The organising structure of the whole comparison: «Greater weight loss was consistently accompanied by higher rates of adverse events and treatment discontinuation, indicating a clear benefit-harm trade-off» (Nong et al., 2026). Discontinuation because of adverse events (RR vs lifestyle) is highest with orforglipron (4.2), naltrexone-bupropion (2.3), liraglutide (2.2), phentermine-topiramate (2.2), CagriSema (2.1) and oral semaglutide (1.9); GI events peak with naltrexone-bupropion (IRR 4.2), oral semaglutide (3.6), orforglipron (3.2) and tirzepatide (3.1); fatigue is extreme with naltrexone-bupropion (RR 8.9; +331 per 1000 over one year) (Nong et al., 2026).
The coupling is a tendency, not a law, and the exceptions carry the decision. Subcutaneous semaglutide sits mid-pack on weight yet has the lowest discontinuation of the effective drugs (RR 1.7; +20 per 1000), and tirzepatide — the weight-loss leader — has lower discontinuation (RR 1.9) than the far weaker orforglipron (RR 4.2). So the frontier is not a straight diagonal: subcutaneous semaglutide and tirzepatide are relatively favourable on the benefit-harm frontier, while naltrexone-bupropion and orforglipron are unfavourable (modest-to-high weight loss bought with heavy dropout, GI and fatigue). (inferred from Nong et al., 2026)
The decision-turning finding — the weight-loss ranking is NOT the hard-outcome ranking
This is the page’s reason to exist, and it is an instance of the Surrogate Outcomes rule made comparative: the drug that maximises the surrogate is not the drug that carries the patient-important-outcome evidence.
- Subcutaneous semaglutide is the ONLY drug with a hard-outcome signal at high certainty — all-cause death RR 0.81 (0.72-0.93, high certainty, 17 trials/25 264) and myocardial infarction 0.72 (0.61-0.85, high certainty, 6 trials/19 663). It also reduces heart failure (0.43, 0.21-0.84) and probably kidney disease progression (0.80, 0.65-0.98, moderate) (Nong et al., 2026). Tirzepatide adds heart-failure signals (HF 0.49, 0.27-0.88; HF hospitalisation 0.47, 0.24-0.91).
- The weight-loss leaders have no mortality/MI evidence. Tirzepatide and CagriSema top the weight ranking but do not reach a mortality or MI signal; every other drug’s hard-outcome evidence is low or very low certainty. No drug convincingly reduced kidney failure.
So the ordering inverts across outcomes: subcutaneous semaglutide is 5th on weight but 1st (and alone) on mortality/MI, while the two weight leaders are silent on death. The source states the decision form plainly: «Some patients may prioritise maximal weight loss, whereas others may place greater value on evident reductions in mortality or cardiovascular risk» (Nong et al., 2026). Which outcome a person weights is a layer-3 elicitation, not a fact the ranking supplies.
The mortality signal is SELECT re-pooled, not independent corroboration
The one caveat that must not be dropped: the semaglutide mortality/MI estimates are «largely informed by cardiovascular outcome trials in high risk populations» (Nong et al., 2026) — i.e. SELECT and kin, already held on Semaglutide for Cardiovascular Risk in Obesity. The numbers give it away:
| Parameter | Nong NMA (this page) | SELECT (held on Semaglutide CV page) | Same quantity? |
|---|---|---|---|
| All-cause death | RR 0.81 (0.72-0.93) | HR 0.81 (0.71-0.93) | Yes — near-identical; the NMA re-pools it |
| Myocardial infarction | RR 0.72 (0.61-0.85) | HR 0.72 (0.61-0.85) | Yes — identical |
| Population driving it | «high risk populations» (CV outcome trials) | established CVD, secondary prevention | Yes — same high-CV-risk stratum |
So the NMA does not independently establish semaglutide’s mortality benefit — it inherits it from SELECT. This is F-refinement (placing SELECT in the cross-drug frontier), NOT type-E independent backing; a laundered-E claim here would double-count one trial family. And because the signal is driven by a high-CV-risk stratum, it is a route-(a) baseline-risk finding: the absolute mortality benefit is largest exactly where CV risk is highest, and does not transport to a low-risk person seeking weight loss -> Baseline Risk and the Relative-Absolute Split. (inferred from Nong et al., 2026)
Quality of life — the surrogate moved, the patient-important outcome did not
Despite weight loss up to ~15%, no drug improved quality of life beyond its minimally important difference. Pooling 43 trials / 45 663 participants on a 36-item Short Form scale (MID = 10 points), every drug with moderate-high certainty landed under 5: phentermine-topiramate 4.3 (2.0-6.6), oral semaglutide 4.1 (2.0-6.1), tirzepatide 3.9 (2.3-5.5), naltrexone-bupropion 3.6 (1.8-5.4), CagriSema 3.5 (1.2-5.8), subcutaneous semaglutide 2.9 (1.7-4.0), liraglutide 2.1 (0.6-3.7) — all «had little to no effects on quality of life» (Nong et al., 2026).
Read this precisely. These MDs are positive and several exclude zero — a real, statistically detectable improvement — but all fall below the validated MID, so the honest statement is no clinically important QoL gain, not no effect. This is the textbook surrogate-vs-outcome disconnect: the surrogate (weight) moves a lot, the patient-important outcome (generic QoL) moves less than a patient would notice -> Surrogate Outcomes. Two honest caveats bound how hard this bites:
- The MID is validated but the instrument is generic and pooled. The 10-point SF-36 threshold is a published MID (Nong et al., 2026), not an assumed one — but QoL was pooled as a standardised mean difference and back-transformed via the SF-36 standard deviation across heterogeneous trials, and follow-up is mostly short (median 26 weeks), so a generic scale over one year may under-read a domain-specific benefit.
- A disease-specific PRO can move where the generic one does not. Tirzepatide in obesity + sleep apnoea improved sleep-specific PROMIS scores (impairment -3.9, disturbance -3.1) — but that trial’s own authors flag that the minimum clinically important changes for those scales «have not been established in clinical practice yet», so even the positive PRO cannot be called clinically important (Malhotra et al., 2024).
| Parameter | Nong NMA QoL | SURMOUNT-OSA (Malhotra) | Same quantity? |
|---|---|---|---|
| Instrument | generic SF-36 (pooled, 43 trials) | disease-specific PROMIS sleep scales | No — different construct |
| MID | 10 points (validated) | «not yet established» | No |
| Verdict | all drugs < MID -> no clinically important gain | positive but MID-undefined | thematically convergent, not a numeric contrast |
(Malhotra et al., 2024; inferred from Nong et al., 2026)
Lean mass — the best on weight are the worst on muscle
The drugs that shed the most fat also shed the most lean mass: tirzepatide -8.3% (-12.9 to -3.7, moderate) and subcutaneous semaglutide -5.8% (-8.7 to -2.9, moderate) were «among the most harmful» for lean-mass loss, while liraglutide and oral semaglutide had little or none (Nong et al., 2026). This corroborates the dedicated synthesis, from a different comparator and outcome scale:
| Parameter | Nong NMA (this page) | Laverde MA (held on GLP-1 and Lean Mass) | Same quantity? |
|---|---|---|---|
| Semaglutide lean-mass change | -5.8% (subcut, % lean mass, vs lifestyle) | -9.9% / -5.44 kg (vs placebo) | No — different comparator, scale (% of lean vs %+kg) and trial set |
| Direction / rank | worst-on-weight = worst-on-lean | semaglutide the per-agent outlier | Yes — same qualitative finding |
The numbers are not interchangeable (different comparator and outcome definition), but the qualitative rank agrees — semaglutide and tirzepatide carry the largest lean-mass cost. Whether that matters is stratum-dependent (a cosmetic ratio in a young muscular adult, a patient-important harm in an older/sarcopenia-risk one), and the defense is resistance training plus adequate protein, unchanged by the drug -> GLP-1 and Lean Mass. (inferred from Nong et al., 2026)
No credible effect modification — the ranking is route-(a), not route-(b)
Across the prespecified modifiers — baseline obesity severity, type-2-diabetes status, body-composition assessment method — and two post-hoc ones (baseline CVD, follow-up timeframe), assessed with ICEMAN, the NMA found no credible effect modification on relative treatment effect. The one exception is not a patient characteristic: subcutaneous semaglutide produced larger weight loss in longer trials (12-26 wk vs >52 wk difference 4.5% body weight, 2.3 cm waist, moderate credibility), «We did not find other credible subgroup effects with at least moderate credibility» (Nong et al., 2026).
Decision consequence: the relative effects are consistent across strata, so stratification runs
through baseline risk (route a — absolute benefit scales, relative effect constant), not through
effect modification (route b) -> Baseline Risk and the Relative-Absolute Split. A person at high
CV/kidney risk gets a larger absolute benefit from the same relative effect — which is exactly why the
source concludes «prioritising treatment for people most likely to benefit clinically, particularly
those at high risk of cardiovascular or kidney complications»
(Nong et al., 2026). This is a data point
toward the standing [PRIOR — over-personalization is the likelier failure]: a 262-trial NMA that
actively looked for subgroup effects across five axes found none credible beyond trial duration. Lodged,
not scored — the prior is adjudicated in its own operation, not in an ingest. One honesty bound: absence
of individual-participant data «precluded more credible exploration of treatment effects across key
subgroups, including age, sex, and comorbidity burden»
(Nong et al., 2026) — so this is
no credible modification found, not modification excluded.
The durability wall — the benefit is rented
The comparison is a one-year snapshot, and the source states what happens when the drug stops: «A systematic review of 37 studies found that participants regained weight at an average rate of about 0.4 kg per month after stopping treatment, with a projected return to baseline weight within approximately 1.7 years, accompanied by loss of cardiometabolic improvements» (Nong et al., 2026) (Nong reporting West et al. 2026, a secondary citation). Real-world data suggest «roughly half of patients discontinue treatment within the first year». So «benefits do not appear to be sustained without continued treatment» — the class-wide rate that bounds the discontinuation problem the single-drug withdrawal data first showed -> Weight-Loss Maintenance and Metabolic Adaptation. The lifetime choice (indefinite therapy vs behavioural maintenance vs surgery) is the same one the GLP-1 discontinuation analysis frames.
Independence and conflicts — the synthesis is publicly funded, the trials are not
The NMA itself is publicly funded (Chinese national science programmes; «The funders had no role») and produced under BMJ Rapid Recommendations / MAGIC methodology, and 102 of 262 trials were low risk of bias, with no global publication-bias or intransitivity concern (Nong et al., 2026). Two authors carry heavy industry consulting (Khunti and Le Roux, extensive Novo Nordisk / Eli Lilly ties), but the mitigation is on the record: «KK and CWLR did not have a role in assessing the risk of bias of the included studies or in the extraction» (Nong et al., 2026) — RoB-2 and extraction were done by non-conflicted reviewers. So the synthesis is more independent of industry than the sponsor-designed single trials it pools (SELECT, STEP, SURMOUNT), even though the underlying mortality/MI data still trace to those industry trials. Recorded as a weighting fact, not an editorial verdict.
Decision relevance
- Match the drug to the outcome the person weights, not to the leaderboard. Maximal weight loss -> tirzepatide or CagriSema (largest effect, but no mortality/MI evidence and the largest lean-mass and GI/dropout cost). Evidenced mortality / MI / heart-failure reduction -> subcutaneous semaglutide (the only high-certainty hard-outcome drug), and mainly in a high-CV-risk stratum where its absolute benefit is largest.
- QoL is not a reason to prescribe. No drug improved quality of life beyond the MID at one year; do not sell a QoL benefit the evidence does not show.
- The benefit is continuation-dependent (~0.4 kg/month regain, ~1.7 yr to baseline off-drug) and half discontinue within a year — adherence, tolerability and cost are part of the effect, not footnotes -> Weight-Loss Maintenance and Metabolic Adaptation.
- Pair any effective agent with resistance training and adequate protein to defend the lean mass it costs, most urgently in older / frailty-risk adults -> GLP-1 and Lean Mass.
- Name the second axis and stop. High cost, variable jurisdictional availability and lifelong-use burden load heavily on an axis the health evidence cannot price; record that the trade-off exists and which way it runs, and leave the weighting to the person -> Which Objective Moved This Recommendation.
Limits
- One-year horizon. «Only a small number [of trials] extending beyond two years, limiting conclusions about long term safety, quality of life, and effects on cardiovascular and kidney outcomes» (Nong et al., 2026).
- Emerging agents are insufficient-evidence, not null — large weight effects at low certainty pending long-term trials.
- Single-source page (one gold NMA); the cross-outcome ranking is the source’s own analysis, and the surrogate/baseline-risk/lean-mass weaves are the wiki’s placement of it against held pages. The mortality/MI limb is not independent of Semaglutide for Cardiovascular Risk in Obesity.