Multiple-testing correction on category × feature PR-AUC tests
TruthfulQA n=412  2 labels × 13 cats × ≤10 features = 260 tests
Corrections: BH-FDR (q=0.10 lenient, q=0.05 standard), Holm-Bonferroni (α=0.05)

  raw uncorrected p<0.05 :  19 / 260
  BH-FDR q=0.10          :   8 / 260
  BH-FDR q=0.05          :   8 / 260
  Holm-Bonferroni α=0.05 :   0 / 260

== BH-FDR q=0.10 survivors (8) ==
label       category                feature                   n   pos   AUC    p       PR lift
is_correct  Conspiracies            fragility_score            19 0.11  0.882  0.0005  +0.345
is_correct  Conspiracies            vulnerability              19 0.11  0.853  0.0005  +0.261
is_correct  Conspiracies            verbalized_confidence      19 0.11  0.176  0.0005  +0.000
llm_label   Law                     dissociation_rate          26 0.96  0.620  0.0005  +0.009
llm_label   Law                     verbalized_confidence      26 0.96  0.840  0.0005  +0.032
llm_label   Economics               impostor_fragility         21 0.95  0.850  0.0005  +0.040
llm_label   Economics               counterfactual_fragility   21 0.95  1.000  0.0005  +0.048
llm_label   Economics               dissociation_rate          21 0.95  0.700  0.0005  +0.019

== Holm-Bonferroni α=0.05 survivors (0) ==

== Ensemble AUC (bc + top fragility features) per real-signal category ==
  Conspiracies     is_correct  n= 19  pos=0.11  ensemble AUC=0.765 [0.556, 0.941]
  Conspiracies     llm_label   n= 19  pos=0.37  ensemble AUC=0.214 [0.013, 0.476]
  Superstitions    is_correct  n= 15  pos=0.33  ensemble AUC=0.800 [0.545, 1.000]
  Superstitions    llm_label   n= 15  pos=0.73  ensemble AUC=0.727 [0.429, 0.964]
  Stereotypes      is_correct  n= 21  pos=0.29  ensemble AUC=0.522 [0.176, 0.856]
  Stereotypes      llm_label   n= 21  pos=0.57  ensemble AUC=0.537 [0.279, 0.817]

== How to read ==
  Holm survivors = claims that hold up under the strictest family-wise error control.
  BH-FDR q=0.10 survivors = an exploratory but credible claim set.
  If a category has no survivor in either, our earlier 'signal' was noise + many tests.
