NASEM is a standard-of-record body that had every incentive to sound an alarm and deliberately did not. Its calibrated position is itself the finding: non-replication is expected, mostly uninterpretable in aggregate, and confidence comes from the web of evidence, not from any two studies matching. Represent this measured stance; do not import a “science is broken” narrative the report resists. (inferred from National Academies of Sciences Engineering and Medicine, 2019)

The report refuses the “crisis” frame

  • «there seems to be an emerging consensus that it is not helpful, or justified, to refer to psychology as being in a state of “crisis.”»
  • The aggregate rate is unknowable, so a failure rate is uninterpretable: «no one knows the expected level of non-replicability in a healthy science» — «some failures to replicate are inevitable».
  • The ambient alarm is itself thin: media coverage over a decade was low (<200 on-topic articles), with limited evidence it moved public opinion — cited as reason to resist, not amplify, the narrative. (National Academies of Sciences Engineering and Medicine, 2019)

This is the wiki’s symmetric-standards rule applied to a scary story: a crisis claim gets the same scrutiny as any other, and here the evidence for a measured “crisis rate” does not exist. It also cuts the other way — the report does not say science is fine; publication bias, p-hacking, and misaligned incentives are real avoidable degraders -> Sources of Non-Replicability. Uniformity in either direction (all-broken or all-fine) would be the tell.

Where confidence actually comes from — the web, and triangulation

«CONCLUSION 7-2: … The goal of science is to understand the overall effect or inference from a set of scientific studies, not to strictly determine whether any one study has replicated any other.» The engine is convergence across independent weaknesses: «Replicability of an effect is reflected in the consistency of effect sizes across the studies, especially when a variety of methods, each with different weaknesses, converge on the same conclusion.» (National Academies of Sciences Engineering and Medicine, 2019)

This restates the wiki’s own type-E / triangulation criterion — and its volume is not independence rule: what raises confidence is methods with different weaknesses agreeing, not many studies sharing one weakness -> Upgrading Observational Evidence, Rating Certainty of Evidence. NASEM reaches it from the replication literature rather than from the meta-method corpus, but by re-reasoning to the same rule — a convergent restatement, not a confidence-raising independent proof (the triangulation rule is method-layer; a source re-deriving it does not corroborate it).

And at the synthesis level, the trustworthy-SR machinery IS the operational antidote — the point where this concept meets evidence appraisal. The replication crisis’s mechanisms — p-hacking, selective publication, non-reproducible methods (P-Hacking and Researcher Degrees of Freedom, Publication Bias and Selective Reporting) — are precisely what a trustworthy systematic review is built to catch and down-weight — a pre-registered protocol, dual independent screening, risk-of-bias assessment, and an explicit small-study/publication-bias check (What a Trustworthy Systematic Review Requires, Is This Actually a Systematic Review, Risk of Bias Assessment Tools). So “confidence without a crisis” is not passive: it is bought by holding the evidence synthesis to those standards. But the Guard below bounds how much it buys — a pre-registered protocol is interpretability, not a quality stamp, and an SR of a skewed literature still inherits some of that bias; the machinery is where the crisis mechanisms are detected and discounted, not a guarantee they are erased. A body appraised without those steps inherits the crisis unexamined; one appraised with them at least prices it in.

The decision-maker rule, and its symmetry (Rec 7-3)

«Anyone making personal or policy decisions based on scientific evidence should be wary of making a serious decision based on the results, no matter how promising, of a single study. Similarly, no one should take a new, single contrary study as refutation of scientific conclusions supported by multiple lines of previous evidence.» (National Academies of Sciences Engineering and Medicine, 2019)

Both limbs matter, and the second is the one usually forgotten: a single dramatic new study neither establishes an effect nor overturns a well-supported body. This is a direct Layer-1 / weave discipline — the wiki weights a body of evidence, and a lone contrarian source revises the web where it disrupts least, not wherever it is loudest.

Guard: what does not buy confidence

Decision relevance

  • A non-replication updates confidence a little, not to zero — and only after asking which source (Sources of Non-Replicability) plausibly produced it.
  • Trust tracks convergence of independent methods, not study count — the same criterion the fabric already uses for its own claims, now sourced from a domain-external consensus body.
  • The honest register is calibrated, not alarmed — which is the wiki’s own the loop is open, never let a clean audit read as validation stance, stated for the replication question.

(inferred from National Academies of Sciences Engineering and Medicine, 2019)

References

National Academies of Sciences Engineering and Medicine. (2019). Reproducibility and Replicability in Science. National Academies Press. https://doi.org/10.17226/25303