Exhibit two

A large effect, published as absent.

Relapse fell from 83% on placebo to 47% on methotrexate. The p-value came in at 0.06, and the trial entered the literature as negative. The same paper, in the same table, also reports the analysis at p = 0.019.

Nothing about the treatment changed between those two numbers.

The trial

Methotrexate in giant-cell arteritis

A randomised, double-blind, placebo-controlled trial asked whether methotrexate prevents relapse in giant-cell arteritis. This is the completion-of-treatment analysis, as published.

Jover JA, et al. Ann Intern Med. 2001 · Relapse, completion-of-treatment analysis · N = 33
ArmRelapseNo relapseRate
Methotrexate 7 a 8 b 46.7%
Placebo 15 c 3 d 83.3%

Fisher’s exact, two-sided: above 0.05, so the classification is not significant. A 36.7-point difference in relapse rate is reported as no difference.

Fisher’s exact, two-sidedp = 0.061

Fragility — fr

One reclassified patient moves this trial across the threshold.

Fragility measures the stability of the statistical classification: what proportion of patients would have to be reclassified before the classification changes. Here that proportion is 1 in 33 — GFQ = 0.0303, well below the 0.05 cutpoint, so the classification is fragile.

Take one placebo patient who did not relapse. Count that patient instead as a methotrexate patient who did. That is a single reallocation out of 33 — the toggle d → a — and the classification changes.

Methotrexate

7 relapses of 15

Placebo

15 relapses of 18
After one global movep = 0.061

The same count, two different move-sets

The classic fragility index may only toggle outcomes inside one arm, holding that arm’s total fixed. The Global Fragility Index searches every cell-to-cell reallocation in the table, with neither margin fixed. Both find a single move here — but not the same move, and not the same resulting p-value.

Classic FI = 1

a → b

Confined to one arm, selected by the rule that takes the arm with fewest events — methotrexate, 7 against 15. That arm’s total stays at 15. One patient who relapsed instead does not.

(6, 9, 15, 3)

p = 0.014274

GFI = 1

d → a

The whole table, neither margin fixed. One placebo patient without relapse is instead a methotrexate patient with relapse.

(8, 8, 15, 2)

p = 0.025510

A within-arm toggle is itself a cell-to-cell reallocation, so every move available to the classic index is also available to the global search. The classic procedure is a constrained special case, which makes GFI ≤ FI a theorem, not an observation. When the two counts agree, as they do here, they still name different toggles reaching different tables. When they disagree, the classic index reports a trial as more stable than it is.

Robustness — nb

The effect is nowhere near neutrality.

Robustness measures how far the observed result sits from therapeutic neutrality, on a bounded 0–1 scale. It is a property of the point estimate, not of its precision — which is exactly why it can disagree with a p-value.

nb (RQ) 0.363636 Strong. The published bands are weak below 0.075, moderate from 0.075 to 0.227, strong at 0.227 and above.
NDI 3 The same distance stated as a patient count: 3 coupled fixed-margin moves would bring this table to the point closest to no effect. NDI = round(|ad − bc|/N) = round(99/33).

NDI is a count, so it grows with trial size and is not comparable between trials. CREST’s NDI is 7 against this trial’s 3, yet CREST is weak and this is strong. Use RQ to compare trials; use NDI to state one trial’s distance in patients.

The triplet

Complete evidence for this trial

pattern (not significant, fragile, strong)

Fragile nonsignificance, yet far from neutrality. The published interpretation of this cell is unambiguous: likely false negative — the effect exists but was not detected. The trial did not show that methotrexate fails. It showed that this trial was too small to say.

The demonstration

The same trial, analysed twice, in the same table.

Table 2 of the paper reports two analysis populations. They straddle the significance threshold. Watch which number moves and which number does not.

Completion of treatmentCompletion of follow-up
2×2 7 / 8 vs 15 / 3 9 / 11 vs 16 / 3
N3339
p 0.0613 — not significant 0.0187 — significant
GFI12
fr (GFQ) 0.0303 — fragile 0.0513 — stable
nb (RQ) 0.3636 — strong 0.3918 — strong
Pattern not significant, fragile, strong significant, stable, strong
Published action Likely false negative — increase power Complete statistical evidence — trust for clinical use

The p-value moved from 0.061 to 0.019 and carried the verdict with it. The robustness moved from 0.364 to 0.392. The size of the effect was never in question. Only the classification was — and the classification is the thing the literature recorded.

This is what the triplet is for. Reported alone, the first analysis says “negative trial”. Reported as p–fr–nb, it says “large effect, unstable classification, underpowered” — which is both true and actionable.

Against exhibit one

Identical fragility. Opposite failures.

CREST and this trial each sit one reallocation from the opposite classification. GFI = 1 for both. What separates them is robustness.

CREST 2011 — myocardial infarctionJover 2001 — relapse
p 0.0290 — significant 0.0613 — not significant
GFI11
fr (GFQ) 0.0004 — fragile 0.0303 — fragile
nb (RQ) 0.0115 — weak 0.3636 — strong
Pattern significant, fragile, weak not significant, fragile, strong
Published action Thin evidence — reject for clinical use Likely false negative — increase power

One trial reported a negligible effect as established. The other reported a substantial effect as absent. Both passed peer review at a major journal. Both are a single reclassified patient from the opposite conclusion, and neither p-value carried that information.

Check these numbers.

Every figure on this page is reproducible from the two 2×2 tables above. The calculators and the full specification of GFI, GFQ, RQ and NDI are free.

Read the specification →

Trial data from Jover JA, et al. Ann Intern Med. 2001, Table 2 (doi:10.7326/0003-4819-134-2-200101160-00010). Both analysis populations are as published; the p-values here are recomputed by Fisher’s exact test and agree with the paper’s reported 0.06 and 0.018.

I endorse the Complete Evidence Standard.

Endorsement is a public statement of practice, not a membership fee. There is nothing to pay and nothing to renew.

Your email is used to confirm your endorsement and is never published or shared.

Signatories

1

researcher or clinician has endorsed the Standard.

Thomas F. Heston, MD, MSc