No food is healthy or unhealthy on its own; it is only healthier or less healthy than whatever would have taken its place. So the useful question is rarely is butter bad for you? but bad compared with what? Butter is better than the trans-fat margarine it replaced a generation ago and worse than the olive oil that could replace it now — and both comparisons are true at once. Name a different alternative and the answer changes with it.

Naming that alternative takes two steps most advice skips. First, be clear what the food actually is: “dairy” and “carbs” are baskets of items so unalike that their average describes nothing on a real plate, so eat less of X means little until X is specified. Second, name what the person would genuinely eat instead — the swap they would actually make, not the ideal one a trial assigns.

Even then, the effect is usually smaller than the biology predicts. The body compensates, routines slip, resolve fades; a well-run trial shows the ceiling, not the ordinary week. And when a single swap moves several things that matter — easing the heart, say, while costing money or effort — this page sets them side by side rather than collapsing them into one grade, because how much each outcome weighs is yours to decide, not ours to compute.

So the honest answer is a direction and a range, not a single figure — and which end of the range to act on depends on what you would rather be wrong about. Where a real harm is in play, lean cautious; where the prize is a likely benefit, there is little to gain from overshooting. In short: choose carefully where to act, and settle for roughly right on how much. One limit outlasts the rest — this page can judge only whether a well-informed advisor would frame a swap this way, never whether anyone who followed it was better off. That loop stays open.

Compared to what? Why an effect has no sign until the alternative is named

An effect estimate is never absolute. It contrasts an exposure with the alternative that exposure displaces, so every “X is beneficial” or “X is harmful” carries a hidden second arm: the comparator. Change the comparator and both the sign and the size of the answer change with it. The trouble is not that comparisons are hard — it is that the comparator is usually left unstated, and the unstated default silently does the work. When a study removes a nutrient, its calories are replaced by something; when a person stops a behaviour, the freed time and energy go somewhere. The omitted arm is a premise, not a footnote. -> The Comparator Problem

A comparator-blind claim is therefore not necessarily wrong — it is unfinished, and naming the omitted arm can flip its truth. So state the counterfactual and ask whether the estimate survives a realistic one: what the person would actually do instead, not an idealized substitution the evidence never tested.

The same move resolves disputes that look unrelated, and it resolves them differently each time — which is why the comparator must ride along with every illustration:

  • Saturated fat — the harm is the replacement: cutting it against polyunsaturated fat reads one way, against refined carbohydrate another, so the recommendation splits by what replaces it -> Saturated Fat Intake and Replacement.
  • Free sugars — the isocaloric swap (against other carbohydrate at equal energy) is null on weight, which surfaces only once the comparator is pinned at equal calories; a comparator-blind “sugar is fattening” mistakes the molecule for the energy it carries -> Free Sugars Intake.
  • Dietary patterns — none is best in the absolute; each evidenced benefit is “better than” a specified alternative, so “which diet is best?” has no comparator-free answer -> Dietary Patterns.

Each magnitude, dose-response and curve shape stays on its own lead above; this section names only the framing. The companion checks — what the effect is measured on, and how a comparator-relative benefit must not be promoted into an absolute target — are deferred: surrogate-versus-target and baseline risk live on Metrics for Targeted Health Guidance.

One misuse shadows the whole move, and it is the most common way the comparator frame launders a real harm: because everything carries some cost and some benefit, you can turn the frame to relativize a genuine harm away — it is all a little hormetic, who is to say? That is not a comparison but the unfalsifiable over-generalization the demarcation line rejects. Use the frame to rank among reasonable options; a large, well-evidenced harm survives every comparator and is never readmitted by it — the big-rock stays a big rock -> Big Rocks (Median).

Naming the counterfactual to X, though, presupposes that X is a single, specified thing — which is often the first place the analysis breaks.

What is “X”, exactly? Specifying the exposure before its comparator

You cannot say “instead of what” until you pin down “what”. A food or nutrient label is only a usable exposure if the boundary it draws carries information — and often it does not. Where within-category variance exceeds between-category variance, the grouping has no explanatory power and the category-level estimate describes no actual food: skinless chicken, fatty pork and lean pasture-raised beef share one label, and the distance within it can exceed the distance to the next label. “Replace X” is undefined until X names something real. -> Is the Food Category Doing Any Work

Two further failures block the comparator before it can be named:

  • Design asymmetry can flip the grade without changing the advice. An isolated component can be dosed and blinded where the whole food can only be observed, so the isolate may out-grade the food on study quality alone. SACN grades fibre isolates as showing an effect, but bounds the finding in the same clause — the benefit is «demonstrated at intakes achieved through supplementation» (Scientific Advisory Committee on Nutrition, 2015). The better-graded object is the one that could be randomised, which is a fact about study design, not about what to eat.
  • A food’s identity drifts under a constant name. Breeding, processing and reformulation move the thing on the plate while the word stays fixed, so evidence attached to a label may not transport to the current exposure. Specify implementations, not labels. This is not just tidy practice but a condition for the question to be well-posed: a causal effect is defined only against a sufficiently well-defined intervention, and a vague label bundles several versions of treatment — weight lost by diet versus by illness, fibre from an isolate versus from the whole food — whose effects need not agree. Until the version is pinned there is no single effect to estimate. -> The Target Trial (Emulation and the Well-Defined Intervention)

The specify-the-exposure move is the whole of this section; the verdicts on each category belong to their leads:

Once X and its comparator are both fixed, a second gap opens: the effect the mechanism predicts for that contrast is not the effect a person realizes.

Intended vs realized: what compensates, and does it survive leaving the RCT arm?

A mechanism earns a direction, never a magnitude. The body is a closed loop, not an open one, so the naive arithmetic — add this, subtract that — is routinely wrong, and the correction it needs is signed. Usually the whole-organism response works against the intervention, leaving a net smaller than the mechanism promised; sometimes it redirects the effect onto a different outcome; and occasionally it works with the intervention, leaving a net larger. So the standing question is not the narrower what compensates? — which already assumes the response subtracts — but the signed what does the whole system do: damp, redirect, or reinforce? -> Net Effect vs Intended Effect. The evidence behind the two directions is lopsided: several worked cases show the response damping an effect, only one shows it amplifying, so treat the response can add as a real caution, not a symmetric law.

One clause each. Added exercise energy is partly offset by eating more and moving less the rest of the day -> Exercise Energy Compensation; energy restriction meets a defended set-point that actively regains -> Weight-Loss Maintenance and Metabolic Adaptation; a benefit that lasts only while a drug is taken, or is cancelled by an offsetting harm on another pathway so the all-cause net is null -> Inflammation as a Modifiable Lever; a meal-timing schedule whose effect may be calorie intake under another name -> Time-Restricted Eating. The amplifying mirror runs the other way: a drink’s own calories go un-compensated and it provokes extra eating in the same sitting, so the net surplus exceeds the drink alone -> Alcohol and Mortality and Vascular Disease.

This is why efficacy is not effectiveness. Efficacy is the effect of an assigned, idealized exposure inside a trial; effectiveness is that same parameter after compensation, execution drift, and adherence have each taken their cut. The two are not rival estimates of one quantity — the trial’s efficacy is an upper structural bound on field effectiveness, the ceiling the mechanism could deliver if nothing leaked, not a competing number. Report the parameter through that transformation: neither reify the trial mean as the decision (scientism) nor discard it because it will not land exactly (science-scepticism). -> The Estimate-to-Action Gap.

Adherence is part of the effect, not a footnote to it: an intervention not done has no effect. Under diminishing returns this can invert the naive ranking — a smaller sustained dose beats a larger abandoned one. But that is a condition, not a slogan, and the condition is curve shape: it holds in a flat or diminishing-returns region, and fails on a steep rising arm or below a deficiency floor, where the larger dose is what clears the outcome and abandonment is the real failure to fix. Which regime applies is the floor-and-direction question the region section takes up. -> Net Effect vs Intended Effect.

One refinement keeps adherence is part of the effect honest, because “the effect” is really two. An as-recommended estimate — the effect of assigning the swap, adherence and all — is the effectiveness-side quantity, and it is the conservative one for a benefit (a swap nobody keeps up shows little). But for a harm it can mislead in the opposite direction: a near-null as-recommended result can hide a real risk concentrated in the fraction who actually take the exposure up, so it is not a safety clearance. The discipline is to name which effect a figure is — the effect of being recommended the swap, or the effect of making it — since they answer different questions and, for harms, can point opposite ways. -> The Target Trial (Emulation and the Well-Defined Intervention)

This transformation needs one input the wiki does not hold. Adherence-probability for a given person is thinly evidenced — effectiveness data are sparse — so the wiki names the gap and infers no number for it: the direction (a realized effect below efficacy) is licensed, the size is not. Why an estimate is a range at all, and what evidence structurally cannot show, is deferred -> Limits of Evidence; each compensation magnitude stays on its exposure-lead above.

Reversibility belongs to the same net-effect accounting and is rarely as free as assumed: can I stop is not does stopping restore baseline. Where restoration is unevidenced, you fall back on detectability and surveillance, not an assumed undo — and how far the evidence bar itself should relax for a cheap, reversible choice is deferred -> Limits of Evidence.

Even a correctly-realized net effect is rarely a single number, because a substitution usually moves more than one outcome that matters.

When outcomes compete: laying out the axes instead of summing them

When a substitution moves more than one patient-important outcome, there is no unique optimum without weights across those outcomes — and the weights are not an empirical quantity anyone can look up. But “you cannot optimize” overstates the bind. Dominance rules out any option worse on every axis than an available alternative, and the Pareto frontier — the set where improving one axis necessarily worsens another — is definable, both with no weights at all. You need weights only to order options within the frontier. So the tangle is real but small: it binds on the ordering of the survivors, not on the whole field. -> The Weighting Problem - Why Population Guidance Is Ill-Posed and Individual Advice Is Not.

The wiki’s role here is bounded and specific. It supplies each axis’s evidenced transmission to its outcome — the health coordinate, the one an evidence fabric can actually contest — and where a substitution loads a non-health axis (carbon, welfare, affordability) it names that the axis exists and which direction it runs, then stops. It never prices the axis, never nets it against the health finding, and asserts no carbon, water, or welfare magnitude, because it holds no such data. Naming the axis is the discipline; supplying a cross-axis exchange rate would be the false objectivity the no-scalar-maximand rule forbids.

This is why the individual’s problem is well-posed where the population’s is not. A guideline must issue one recommendation to people who weight the axes differently, and aggregating heterogeneous weights into a single ordering is ill-posed independent of anyone’s bad faith; one person has their own weights. So the division of labour is exact: the wiki supplies the coordinate (each axis’s evidenced transmission) and the interval (the region on each axis); the person supplies the cross-axis weight, elicited at layer 3, never estimated. Several population axes — scalability, communicability, palatability — do not even transfer to an audience of one.

At guidance level the competing objectives are often already on the page. GRADE names four determinants of a recommendation, only one of which is the evidence, and a body publishing an Evidence-to-Decision table lists its non-evidence considerations by name — WHO’s acceptability domain openly includes «potential impact on national economies» (World Health Organization, 2023). So which objective moved this recommendation is a reading skill, not an accusation — and its direction is indeterminate until read: feasibility and acceptability push toward stringency as readily as toward laxity (an environmental load pushes red-meat guidance harder; a staple’s economics push the other way). What bodies disclose is the considerations; what stays unpublished is the weight. -> Which Objective Moved This Recommendation.

Two guards. Watch the health-halo running across axes: a favourable health score must not buy a food exemption from its carbon or welfare column — the same asymmetric scrutiny that lets motivated reasoning survive an evidence-based framing, one level out. And keep a trade within the health axis distinct from a cross-axis one — child-neurodevelopment benefit against methylmercury harm is two patient-important health outcomes, priced by stratum, not a health-versus-environment tangle -> Fish and Seafood Consumption. Surrogate-versus-target and the relative/absolute baseline-risk split are deferred -> Metrics for Targeted Health Guidance.

Even after you lay out the axes, each axis still carries an estimate that is a region — and a recommendation has to say what to actually do with a region.

From estimate to substitution: a region and a direction, not a point

In this domain the evidence structurally yields a floor, a direction within the studied range, and — where the curve bounds it — a harm-ceiling, but not a point optimum: a study can show below here is deficiency, more in this range still helps, past here it harms, essentially never this exact intake is best -> The Underivable Optimum. Why the estimate is a band rather than a point — measurement error and the statistical reason an interval is the honest object — is not re-derived here; state the conclusion and see -> Limits of Evidence. So frame a substitution to clear the floor and move the right way within the region, not to land on a peak that was never identified -> The Estimate-to-Action Gap.

That is not a counsel of despair. It only reads as one if you collapse two jobs into one. The fabric optimizes ALLOCATION — Layer 1 ranks levers by effect size x certainty and spends attention on the largest remediable gap (which lever, magnitudes deferred to the Big Rocks deliverables) — and it satisfices DOSE — per lever, clear the floor, move in the evidenced direction, stay in range rather than chase a point this domain cannot locate. These are a category distinction, not a contradiction: satisfice is a claim about the dose, never about where to spend attention, and neither bleeds into the other.

The division of labour is the same one the competing-axes section drew: the wiki supplies the coordinate (each axis’s evidenced transmission to its outcome) and the interval (the region on that axis); the person supplies the cross-axis weight and — in the next section — the loss tail. Coordinate and interval are supplied; weight and tail are elicited. Neither is an estimate the wiki computes.

One discipline guards the region against being read as more than it is. A comparator-relative benefit (“better than the refined carbohydrate it replaced”) and a floor (“covers the deficiency requirement”) are descriptive objects; promoting either into an absolute eat this target is the descriptive->normative category error -> The Descriptive-Normative Category Error. A floor is where a benefit begins, not where it is complete: read as a target it tells anyone with an objective above deficiency that they are done at the start. A body can name the trap to pre-empt it — EFSA on dietary sugars states that «The Panel wishes to clarify that a UL is not a recommended level of intake» (European Food Safety Authority, 2022). The estimate travels only with its comparator attached.

Two riders keep this honest. First, the earlier point that a smaller sustained dose can beat a larger abandoned one is a condition on curve shape, not a slogan (established above): it holds only near a flat/diminishing-returns region -> Net Effect vs Intended Effect. Second, structural leverage — an exposure that reshapes physiology or the choice environment (so future choices are easier) can outrank a point optimization of similar expected size. This is a telos provision with no dedicated claim page; it is carried here on that provision and on the structural note in The Estimate-to-Action Gap, and flagged as a named gap — a candidate concept page routed to Weave/ingest as residual, not a settled fabric claim.

Once you name a region, one question remains that the region alone cannot answer: which end of it to act on.

Which end of the interval? Asymmetric loss and the conservative default

Even a perfectly symmetric statistical interval has an asymmetric decision-relevant summary, and the two do not conflict — they are about different objects. The interval describes where the parameter plausibly sits; the decision minimizes expected loss, which is rarely symmetric about the estimate. So the action’s summary is not the mean or the midpoint but a quantile — a tail — its position set by the ratio of the cost of under-shooting to that of over-shooting. Where over-estimating a harm is cheap and under-estimating it expensive, the decision reads off the upper (bad) tail; where over-shooting a benefit’s dose is cheap and forgoing it is the loss, off the lower (conservative) tail. The precautionary posture is thus derived from the cost structure, not asserted — the decision-side twin of the statistical asymmetry, which lives on -> Limits of Evidence and is not re-derived here -> The Estimate-to-Action Gap.

Three consequences follow.

  • The decision number is a tail, and WHICH tail is elicited, not estimated. Same routing as the cross-axis weight above: the wiki supplies the coordinate and the interval; the person supplies the loss tail. Which quantile to act on is a layer-3 loss-function fact about this person’s costs, not a quantity the evidence computes.
  • You cannot naively net a harm interval against a benefit interval into one symmetric band. Collapsing both to their midpoints and subtracting throws away exactly the asymmetry that should drive the choice. Weight each side to its own tail, per person, first — then net the loss-weighted objects, not the raw intervals.
  • Symmetric-standards guard: harm-up / benefit-down is a DEFAULT, not a law. Always weight the harm tail is the exact machinery by which manufactured alarm survives an evidence-based framing, and its mirror — always weight the benefit’s upside — is how false hope does. The conservative direction is the reasoned default because harm is often convex and irreversible; but a cheap, reversible exposure with a real upside can rationally be weighted to a benefit’s upper tail. The discipline is to name whose loss sets the direction and never present the weighting as though it were in the data -> The Descriptive-Normative Category Error (the don’t-smuggle guard).

The loss function is direction-agnostic, which is the proof it is bias away from the costly tail rather than a disguised always-conservative rule: the same rule biases protein up toward the upper end of its range (overshoot is low-harm for healthy kidneys, under-dosing forfeits the objective) -> Protein and Resistance Training for Muscle and Strength — and biases further up for older adults, where the cost of under-dosing is higher -> Protein Intake for Older Adults — while it biases training intensity down toward the margin (overshoot loads an often-irreversible injury tail). Opposite directions from one rule is the signature of loss-appropriate bias, not of smuggled precaution -> The Estimate-to-Action Gap.

Caveats and boundaries

  • This is an open loop. No operation grades a decision here against a realized outcome; the wiki verifies only the would-form — whether a well-informed advisor would frame the substitution this way — never whether anyone was better off.
  • This deliverable carries no exposure estimates. Every food, drug or behaviour above is an illustration that links its exposure-lead; the dose-response, magnitude and curve shape live there, never here.
  • Deferrals. What the evidence can and cannot show -> Limits of Evidence; surrogate-versus-target and how baseline risk enters (relative vs absolute) -> Metrics for Targeted Health Guidance; which levers matter by magnitude x certainty -> the Big Rocks deliverables.
  • No scalar maximand. Where an exposure moves competing outcomes, each axis’s evidenced transmission is laid out and any non-health axis is named by direction only; the cross-axis weight and the loss tail are the person’s, elicited at layer 3, never computed here.
  • Named gaps (routed to Weave/ingest, not silently filled): structural leverage has no dedicated claim page (a telos provision only); adherence-probability for a given person is not held as a quantified input (effectiveness data are sparse) — direction licensed, magnitude not.

Evidence box

Question’How does the choice of comparator (replace X with what?) change the estimated effect of a dietary or lifestyle exposure, and how should a recommendation frame substitutions when objectives compete, evidence is partial, and the choice is made in a real environment rather than an RCT arm?‘
Evidence included3 sources — 3 gold
Overall certaintyMedium (see Rating Certainty of Evidence)
Source-selection noteAll sources are gold or high tier.
Last updated2026-08-31 · Independently reviewed: No · Full edit history

References

European Food Safety Authority. (2022). Tolerable upper intake level for dietary sugars. EFSA Journal, 20(2). https://doi.org/10.2903/j.efsa.2022.7074
Scientific Advisory Committee on Nutrition. (2015). Carbohydrates and Health. https://www.gov.uk/government/publications/sacn-carbohydrates-and-health-report
World Health Organization. (2023). Saturated fatty acid and trans-fatty acid intake for adults and children: WHO guideline. https://www.who.int/publications/i/item/9789240073630