Exhibit two
Relapse fell from 83% on placebo to 47% on methotrexate. The p-value came in at 0.06, and the trial entered the literature as negative. The same paper, in the same table, also reports the analysis at p = 0.019.
Nothing about the treatment changed between those two numbers.
The trial
A randomised, double-blind, placebo-controlled trial asked whether methotrexate prevents relapse in giant-cell arteritis. This is the completion-of-treatment analysis, as published.
| Arm | Relapse | No relapse | Rate |
|---|---|---|---|
| Methotrexate | 7 a | 8 b | 46.7% |
| Placebo | 15 c | 3 d | 83.3% |
Fisher’s exact, two-sided: above 0.05, so the classification is not significant. A 36.7-point difference in relapse rate is reported as no difference.
Fragility — fr
Fragility measures the stability of the statistical classification: what proportion of patients would have to be reclassified before the classification changes. Here that proportion is 1 in 33 — GFQ = 0.0303, well below the 0.05 cutpoint, so the classification is fragile.
Take one placebo patient who did not relapse. Count that patient instead as a methotrexate patient who did. That is a single reallocation out of 33 — the toggle d → a — and the classification changes.
The classic fragility index may only toggle outcomes inside one arm, holding that arm’s total fixed. The Global Fragility Index searches every cell-to-cell reallocation in the table, with neither margin fixed. Both find a single move here — but not the same move, and not the same resulting p-value.
a → b
Confined to one arm, selected by the rule that takes the arm with fewest events — methotrexate, 7 against 15. That arm’s total stays at 15. One patient who relapsed instead does not.
(6, 9, 15, 3)
p = 0.014274
d → a
The whole table, neither margin fixed. One placebo patient without relapse is instead a methotrexate patient with relapse.
(8, 8, 15, 2)
p = 0.025510
A within-arm toggle is itself a cell-to-cell reallocation, so every move available to the classic index is also available to the global search. The classic procedure is a constrained special case, which makes GFI ≤ FI a theorem, not an observation. When the two counts agree, as they do here, they still name different toggles reaching different tables. When they disagree, the classic index reports a trial as more stable than it is.
Robustness — nb
Robustness measures how far the observed result sits from therapeutic neutrality, on a bounded 0–1 scale. It is a property of the point estimate, not of its precision — which is exactly why it can disagree with a p-value.
| nb (RQ) | 0.363636 | Strong. The published bands are weak below 0.075, moderate from 0.075 to 0.227, strong at 0.227 and above. |
|---|---|---|
| NDI | 3 | The same distance stated as a patient count: 3 coupled fixed-margin moves would bring this table to the point closest to no effect. NDI = round(|ad − bc|/N) = round(99/33). |
NDI is a count, so it grows with trial size and is not comparable between trials. CREST’s NDI is 7 against this trial’s 3, yet CREST is weak and this is strong. Use RQ to compare trials; use NDI to state one trial’s distance in patients.
The triplet
Fragile nonsignificance, yet far from neutrality. The published interpretation of this cell is unambiguous: likely false negative — the effect exists but was not detected. The trial did not show that methotrexate fails. It showed that this trial was too small to say.
The demonstration
Table 2 of the paper reports two analysis populations. They straddle the significance threshold. Watch which number moves and which number does not.
| Completion of treatment | Completion of follow-up | |
|---|---|---|
| 2×2 | 7 / 8 vs 15 / 3 | 9 / 11 vs 16 / 3 |
| N | 33 | 39 |
| p | 0.0613 — not significant | 0.0187 — significant |
| GFI | 1 | 2 |
| fr (GFQ) | 0.0303 — fragile | 0.0513 — stable |
| nb (RQ) | 0.3636 — strong | 0.3918 — strong |
| Pattern | not significant, fragile, strong | significant, stable, strong |
| Published action | Likely false negative — increase power | Complete statistical evidence — trust for clinical use |
The p-value moved from 0.061 to 0.019 and carried the verdict with it. The robustness moved from 0.364 to 0.392. The size of the effect was never in question. Only the classification was — and the classification is the thing the literature recorded.
This is what the triplet is for. Reported alone, the first analysis says “negative trial”. Reported as p–fr–nb, it says “large effect, unstable classification, underpowered” — which is both true and actionable.
Against exhibit one
CREST and this trial each sit one reallocation from the opposite classification. GFI = 1 for both. What separates them is robustness.
| CREST 2011 — myocardial infarction | Jover 2001 — relapse | |
|---|---|---|
| p | 0.0290 — significant | 0.0613 — not significant |
| GFI | 1 | 1 |
| fr (GFQ) | 0.0004 — fragile | 0.0303 — fragile |
| nb (RQ) | 0.0115 — weak | 0.3636 — strong |
| Pattern | significant, fragile, weak | not significant, fragile, strong |
| Published action | Thin evidence — reject for clinical use | Likely false negative — increase power |
One trial reported a negligible effect as established. The other reported a substantial effect as absent. Both passed peer review at a major journal. Both are a single reclassified patient from the opposite conclusion, and neither p-value carried that information.
Every figure on this page is reproducible from the two 2×2 tables above. The calculators and the full specification of GFI, GFQ, RQ and NDI are free.
Trial data from Jover JA, et al. Ann Intern Med. 2001, Table 2 (doi:10.7326/0003-4819-134-2-200101160-00010). Both analysis populations are as published; the p-values here are recomputed by Fisher’s exact test and agree with the paper’s reported 0.06 and 0.018.
Endorsement is a public statement of practice, not a membership fee. There is nothing to pay and nothing to renew.
Signatories
1
researcher or clinician has endorsed the Standard.