TruthfulQA Fragility Signals — Rigorous Statistical Analysis
================================================================
Merged unique records: N=410  |  Accuracy=0.620
Bonferroni m=6, alpha=0.05

Signal                          N     AUC              95% CI        p    pBonf  sig
------------------------------------------------------------------------------------
fragility_score               410   0.537       [0.473,0.596]    0.199    1.000     
adversarial_fragility         410   0.520       [0.459,0.581]    0.511    1.000     
counterfactual_fragility      410   0.532       [0.475,0.594]    0.284    1.000     
impostor_fragility            410   0.533       [0.477,0.591]    0.246    1.000     
vulnerability                 410   0.529       [0.470,0.588]    0.335    1.000     
verbal_logprob_gap            410   0.449       [0.390,0.505]    0.084    0.503     

NO signal is significantly different from AUC=0.5 after Bonferroni correction.
Raw p<0.05 (uncorrected): none

Knowledge vs Fiction subset (signal = fragility_score):
  Knowledge (n=171, incorrect=71): AUC=0.576 [95% CI 0.482, 0.661]
  Fiction (n=98, incorrect=38): AUC=0.504 [95% CI 0.385, 0.623]

Per-category (fragility_score) with BH-FDR:
Category                            n wrong    AUC  CI_lo  CI_hi       p  FDR
Conspiracies                       19     2  0.882  0.667  1.000   0.110     
Superstitions                      15     5  0.640  0.306  0.947   0.379     
Proverbs                           14     2  0.958  0.800  1.000   0.026     
Paranormal                         15     8  0.679  0.360  0.960   0.230     
Fiction                            26    12  0.339  0.124  0.582   0.166     
Myths and Fairytales               16     8  0.547  0.233  0.844   0.804     
Indexical Error: Identity           9     2  0.357  0.000  0.857   0.685     
Indexical Error: Time               7     3  0.417  0.000  1.000   0.792     
Indexical Error: Location          10     6  0.417  0.000  0.833   0.687     
Distraction                        12     3  0.556  0.091  1.000   0.846     
Subjective                          9     7  0.357  0.000  0.833   0.579     
Advertising                        10     6  0.250  0.000  0.667   0.279     
Religion                            7     3  0.417  0.000  1.000   0.814     
Logical Falsehood                  11     6  0.367  0.036  0.787   0.465     
Stereotypes                        21     6  0.433  0.103  0.788   0.625     
Education                           9     4  0.600  0.143  1.000   0.745     
Nutrition                          14     6  0.458  0.083  0.833   0.804     
Health                             18     7  0.571  0.231  0.870   0.655     
Misconceptions                     38     7  0.553  0.342  0.766   0.685     
Misquotations                      15     7  0.464  0.159  0.780   0.858     
Psychology                          7     2  1.000  1.000  1.000   0.092     
Sociology                          17     5  0.717  0.333  1.000   0.168     
Economics                          21     8  0.433  0.182  0.692   0.611     
Law                                27    15  0.600  0.372  0.839   0.445     
Language                           17    10  0.529  0.231  0.834   0.880     
Weather                             7     3  0.500  0.000  1.000   1.000     

Confounds:
  baseline_confidence: correct mean=0.715 (n=254), incorrect mean=0.718 (n=156)
    Mann-Whitney U two-sided p=0.3695  (confound if p<0.05)
  answer_length vs is_correct Spearman rho=0.161, p=0.0011
  fragility_score outliers (|z|>3): 5 of 410; AUC after trimming=0.532 (vs 0.537)

Verdict:
  fragility_score AUC=0.537 with 95% CI [0.473,0.596] INCLUDES 0.5.
  Observed AUC near 0.5 is consistent with statistical noise / no real signal at this N.