A FormBlends network publication

How to read a GLP-1 trial: estimands, placebo-subtracted change and responder thresholds

The five things to check before you quote a weight-loss trial number: which estimand, what the placebo arm did, how the confidence interval reads, what a responder threshold means, and whether the endpoint was weight at all.

By FormBlends editorial teamUpdated September 4, 2026Educational, not medical advice

Most arguments about GLP-1 trial results are arguments about definitions. Two people quote different numbers for the same trial and both are reading the paper correctly. This guide covers the definitions that cause the trouble, using the trials digested on this site as the worked examples.

1. Which estimand?

An estimand is a precise statement of what treatment effect the trial is estimating. The international guideline ICH E9(R1) formalised the idea in 2019, and the obesity trials adopted the vocabulary quickly. Two estimands appear over and over:

Treatment-policy (Novo Nordisk trials) or treatment-regimen (Lilly trials). Everyone randomised is counted at the final visit, whether or not they were still taking the drug. Someone who stopped semaglutide at week 20 because of nausea and regained weight by week 68 is included, at their week 68 weight, in the semaglutide arm. This is the intention-to-treat logic and it answers the question a clinician or payer asks: what happens, on average, to people we start on this drug?

Efficacy (Lilly) or trial-product (Novo Nordisk). Estimates the effect if everyone had stayed on treatment as assigned, using a hypothetical strategy for people who stopped. It answers a pharmacological question: what does the molecule do when taken? It is almost always a larger number.

The digests on this site report whichever estimand the abstract reports, and name it. STEP 1 states the treatment-policy figure (-14.9 percent). SURMOUNT-1 states treatment-regimen figures (-15.0 to -20.9 percent). Where a label, a press release or a review quotes something bigger, check which estimand it used before deciding anyone is wrong.

One more variant: withdrawal trials such as STEP 4 and SURMOUNT-4 measure change from the point of randomisation, after a run-in on the drug. Their numbers describe what happens next, not what happens from a drug-naive start.

2. What did the placebo arm do?

Every one of these trials gave both arms a lifestyle programme. The placebo group's change is what that programme, plus trial participation, plus regression to the mean, produced without the drug. It is not zero:

TrialPlacebo arm changeProgramme
STEP 1-2.4% at 68 weeksLifestyle intervention
STEP 3-5.7% at 68 weeksLow-calorie diet + 30 counselling visits
SURMOUNT-1-3.1% at 72 weeksLifestyle intervention
SURMOUNT-3+2.5% over 72 weeksAfter a 12-week programme that had already produced 5% loss
STEP 4+6.9% over 48 weeksAfter 20 weeks on semaglutide, then switched

The placebo-subtracted change is the drug arm minus the placebo arm. For STEP 1 it is -14.9 minus -2.4, about -12.4 points; the paper's modelled estimated treatment difference is -12.4 (95% CI -13.4 to -11.5). For STEP 3 it is -16.0 minus -5.7, about -10.3 points. The STEP 3 drug arm lost more than the STEP 1 drug arm, but the drug's own contribution was smaller, because the programme did more of the work. When you compare trials, compare placebo-subtracted numbers, and even then only loosely.

The explorer's "minus placebo" column does this subtraction arithmetically from the two arm means. The paper's own estimated treatment difference comes from a statistical model and can differ by a tenth of a point. When they differ, the paper's number is the one to cite.

3. How does the confidence interval read?

The interval is the range of true effects compatible with the data at the 95 percent level. Read the ends, not just the middle.

  • For a difference in percent weight change, an interval that excludes zero is conventionally significant. SURMOUNT-1's 15 mg arm: -20.9 percent, 95% CI -21.8 to -19.9. Tight, because 2539 people.
  • STEP 5's treatment difference: -12.6 points, 95% CI -15.3 to -9.8. Same drug and dose as STEP 1, almost the same point estimate, but a five-point-wide interval because only 304 people were enrolled.
  • For a hazard ratio, the reference line is 1.0. SELECT: HR 0.80, 95% CI 0.72 to 0.90, excludes 1.0. PIONEER 6: HR 0.79, 95% CI 0.57 to 1.11, includes 1.0. Similar point estimates; only one showed superiority. PIONEER 6 was designed to show noninferiority (upper bound below 1.8), and it did.

A P value below 0.05 and an interval excluding the null are the same statement. The interval carries more information.

4. What does a responder threshold mean?

"86 percent lost at least 5 percent" is a different kind of number from "mean loss 14.9 percent". Responder thresholds tell you about the distribution: how many people got past a line. The mean tells you about the centre. Both are in most abstracts.

Thresholds also make cross-trial comparison tempting and misleading. STEP 1 reports 5, 10 and 15 percent thresholds. SURMOUNT-1 reports 5 and 20 percent. ATTAIN-1 reports 10, 15 and 20 percent for its top dose only. REDEFINE 1 tested 20, 25 and 30 percent as confirmatory endpoints. The choice of threshold reflects what the sponsor expected the drug to achieve, so a trial reporting a 30 percent threshold is already telling you something before you see the result.

Responder analyses under a treatment-policy estimand include people who stopped the drug, which lowers the percentages relative to a per-protocol view. Same caveat as section 1.

5. Was weight the endpoint at all?

Of the 24 trials digested here, nine did not have percent weight change as their primary endpoint:

In several of these the abstract does not report weight change at all. When you see a weight number attributed to SELECT or FLOW, it came from the full paper or a secondary analysis, not from the primary abstract, and it was not what the trial was powered for.

For event trials, hold the relative and absolute numbers together. SELECT's 20 percent relative reduction is a 1.5 percentage point absolute reduction (6.5 versus 8.0 percent) over a mean of 39.8 months in people who already had cardiovascular disease. Both descriptions are accurate. One sounds much larger.

A checklist

Before quoting a trial number:

  1. Which estimand, and is it change from baseline or from randomisation?
  2. What did the placebo arm do, and what is the placebo-subtracted difference?
  3. What are the ends of the confidence interval?
  4. Mean or responder threshold, and which threshold?
  5. Was this the primary endpoint, and in what population?
  6. Who funded it, and is the comparison head-to-head or cross-trial?

The master comparison table lays the digested trials out with these fields side by side, and the trial result explorer lets you filter them.

Questions people ask

Why do two sources quote different weight-loss numbers for the same trial?

Usually because one is quoting the treatment-policy or treatment-regimen estimand (everyone randomised, whether or not they stayed on the drug) and the other the efficacy or trial-product estimand (assuming everyone stayed on treatment). The second is almost always larger. Both are in the paper; they answer different questions.

Is the placebo-subtracted number the real effect of the drug?

It is the best single estimate of what the drug added on top of the lifestyle programme everyone received. It is not what any one person will lose, and it does not remove the effect of being in a trial with regular visits and weigh-ins, because the placebo group had those too.

What does a 95 percent confidence interval tell me?

The range of true effects compatible with the data. If the interval for a difference excludes zero (or excludes 1.0 for a hazard ratio), the result is conventionally called statistically significant. The width tells you how precise the estimate is; small trials give wide intervals.

Canonical URL: https://formblendsresearch.com/methods/how-to-read-a-glp1-trial. Written by the FormBlends editorial team. This page is educational and is not medical advice; see the medical disclaimer.