Reference

Definitions

Two dimensions, and the indices that measure them. Every worked value below comes from the two trials on this site: CREST and Jover.

Work in progress. This page is being built out. The canonical, versioned definitions live in the specification at fragilitymetrics.org.

Fragility

fr · a dimension, not an index

The instability of the p-value to cross the alpha threshold of 0.05 with small perturbations in the data.

Fragility asks how securely a result holds its side of the threshold. It is direction-agnostic: the same measure applies whether the result starts significant and would be lost, or starts non-significant and would be gained. A fragile classification is not a wrong one — it is one that a handful of reclassified patients would reverse.

Measured by an index (a count of patients) or a quotient (that count as a proportion of N). The quotient is what gets compared across trials.

Robustness

nb · a dimension, not an index

The geometric distance from therapeutic neutrality.

Robustness asks how far the observed result sits from no effect. It is a property of the point estimate alone — it does not depend on sample size, precision, or any significance test. That separation is deliberate: it is what lets the triplet tell a large-but-imprecise effect apart from a genuinely null one.

Published bands: weak below 0.075, moderate from 0.075 to 0.227, strong at 0.227 and above.

Heston TF. The Neutrality Boundary Framework: Quantifying Statistical Robustness Geometrically. arXiv:2511.00982. 2025.

Fragility Index

FI · count · fragility

The smallest number of patients whose outcome must be reclassified, within a single arm, to move the p-value across alpha.

The original fragility measure. One arm is selected and outcomes are toggled inside it until the classification flips; the arm total stays fixed. Because the search is confined to one arm, the FI is a constrained special case of the global search below.

This site uses the Heston FI: the arm with fewer events is selected, ties broken to the smaller arm, and toggling then runs in whichever direction crosses alpha. Full rule.

CRESTFI = 2 · toggle d → c
JoverFI = 1 · toggle a → b

Walsh M, Srinathan SK, McAuley DF, et al. The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index. J Clin Epidemiol. 2014;67(6):622–628.

Fragility Quotient

FQ · quotient · fragility

The Fragility Index expressed as a proportion of the trial’s total sample.

FQ = FI / N

A raw count grows with trial size, so two FIs from trials of different size cannot be compared. Dividing by N removes that dependence and turns the count into a proportion of patients — which is comparable.

CRESTFQ = 2 / 2502 = 0.000799
JoverFQ = 1 / 33 = 0.030303

Ahmed W, Fowler RA, McCredie VA. Does Sample Size Matter When Interpreting the Fragility Index? Crit Care Med. 2016;44(11):e1142–e1143.

Global Fragility Index

GFI · count · fragility

The smallest number of cell-to-cell reallocations, anywhere in the table, that moves the p-value across alpha.

The search is global: a patient may be moved from any cell to any other, and neither the row nor the column margins are held fixed. Only N stays constant. Because a within-arm toggle is itself a reallocation, every move available to the FI is also available here — so GFI ≤ FI always. That is a mathematical result, confirmed empirically, not a tendency.

CRESTGFI = 1 · move a → c · p 0.029 → 0.062
JoverGFI = 1 · move d → a · p 0.061 → 0.026

Both indices count moves to the same boundary, but they search different move-sets, so an equal count can still name a different toggle and a different resulting p-value.

Heston TF. Fragility Metrics Toolkit. Zenodo. doi:10.5281/zenodo.17254763

Global Fragility Quotient

GFQ · quotient · fragility

The Global Fragility Index expressed as a proportion of the trial’s total sample.

GFQ = GFI / N

GFQ is the value the framework reports as fr. Read it as a proportion, not a count: it answers what proportion of patients would have to be reclassified before the classification changes. Below 0.05 the classification is called fragile — the same cutpoint as the p-value it mirrors.

CRESTGFQ = 1 / 2502 = 0.000400 · fragile
JoverGFQ = 1 / 33 = 0.030303 · fragile

Heston TF. Fragility Metrics Toolkit. Zenodo. doi:10.5281/zenodo.17254763

Neutrality Distance Index

NDI · count · robustness

The smallest number of paired, fixed-margin moves that brings the table to the reachable point closest to therapeutic neutrality.

NDI = round(|ad − bc| / N) = round(N · RQ / 4)

A paired move relocates two patients, one in each arm, so both margins stay fixed. Each such move shifts the cross-product difference by exactly N, which is why reachable values are spaced N apart. NDI invokes no significance test at all — it depends only on ad − bc. Its target is RR = 1, not alpha.

CRESTNDI = round(17976 / 2502) = 7
JoverNDI = round(99 / 33) = 3

NDI is a count, so it grows with trial size and is not comparable between trials. CREST’s NDI of 7 exceeds Jover’s 3, yet CREST is weak and Jover is strong. Use RQ to compare trials; use NDI to state one trial’s distance in patients. Worked step by step.

Heston TF. Fragility Metrics Toolkit. Zenodo. doi:10.5281/zenodo.17254763

Risk Quotient

RQ · quotient · robustness

The normalised geometric distance of a 2×2 table from therapeutic neutrality, on a bounded 0–1 scale.

RQ = |ad − bc| / (N² / 4)

RQ is the value the framework reports as nb for a 2×2 table, and it is the measure to use when comparing trials. It is scale-invariant: multiplying every cell by a constant leaves it unchanged. Neutrality is ad = bc, which is RR = 1, so RQ = 0 means the table sits exactly on the neutrality boundary.

CRESTRQ = 0.011486 · weak
JoverRQ = 0.363636 · strong

RQ and NDI express the same distance in different units: RQ as a bounded proportion, NDI as an integer count of paired moves. RQ is not NDI / N.

Heston TF. The Neutrality Boundary Framework: Quantifying Statistical Robustness Geometrically. arXiv:2511.00982. 2025. · Heston TF. Fragility Metrics Toolkit. Zenodo. doi:10.5281/zenodo.17254763

Compute any of these free at fragilitymetrics.org. Spotted an error? Tell us.

I endorse the Complete Evidence Standard.

Endorsement is a public statement of practice, not a membership fee. There is nothing to pay and nothing to renew.

Your email is used to confirm your endorsement and is never published or shared.

Signatories

1

researcher or clinician has endorsed the Standard.

Thomas F. Heston, MD, MSc