=== Regime analysis on merged TruthfulQA (n=413) ===
Overall hallucination rate: 0.385
Baseline AUC (fragility_score, full set): 0.541
Baseline AUC (1 - baseline_confidence, full set): 0.483

=== Task 2: bc threshold × fragility threshold grid ===
bc_thr  frag_thr  n     hallu_rate  AUC(frag)   AUC(1-bc)   
0.6     0.01      383   0.368       0.524       0.439       
0.6     0.02      377   0.366       0.530       0.439       
0.6     0.05      228   0.377       0.561       0.418       
0.6     0.08      91    0.451       0.499       0.406       
0.7     0.01      251   0.390       0.529       0.445       
0.7     0.02      247   0.389       0.533       0.448       
0.7     0.05      161   0.404       0.528       0.435       
0.7     0.08      69    0.449       0.500       0.347       
0.8     0.01      58    0.483       0.621       0.373       
0.8     0.02      58    0.483       0.621       0.373       
0.8     0.05      50    0.520       0.574       0.404       
0.8     0.08      32    0.562       0.532       0.389       

=== bc threshold only (no frag filter) ===
bc_thr  n     hallu_rate  AUC(frag)   AUC(1-bc)   
0.6     383   0.368       0.524       0.439       
0.7     251   0.390       0.529       0.445       
0.8     58    0.483       0.621       0.373       

=== Task 3: best regime with n>=30 (by AUC on fragility_score) ===
BEST: bc>0.80, frag>0.000, n=58, hallu=0.483, AUC=0.621
  Bootstrap 95% CI: mean=0.619, [0.461, 0.762]

=== Task 5: verbalized_confidence>0.8 & baseline_confidence subsets ===
vc      bc_thr    n     hallu_rate  AUC(frag)   AUC(1-bc)   
>0.8    0.5       145   0.297       0.544       0.506       
>0.8    0.6       139   0.288       0.557       0.484       
>0.8    0.7       87    0.322       0.570       0.553       
>0.8    0.8       23    0.304       0.420       0.179       

=== Task 6: composite scores in glass-cannon regime (bc>0.8, frag>0.02) ===
n_glass_cannon=58, hallu_rate=0.483
  adaptive               AUC=0.650
  impostor*bc            AUC=0.643
  impostor               AUC=0.635
  frag*bc                AUC=0.632
  fragility_score        AUC=0.621
  weighted_sum           AUC=0.593
  sum_all_frag           AUC=0.585
  vulnerability          AUC=0.581
  adversarial            AUC=0.543
  paraphrase             AUC=0.477
  counterfactual         AUC=0.454
  1-bc                   AUC=0.373
  verbal_logprob_gap     AUC=0.290

=== Task 7: fragility-AUC by baseline_confidence bin ===
bc range       n     hallu_rate  AUC(frag)   AUC(1-bc)   
[0.75,0.85)   120   0.400       0.555       0.514       
[0.85,0.95)   12    0.667       0.594       0.000       
[0.95,1.01)   1      (too few)

=== Scan: ANY subset with AUC>0.70, n>=30 ===
  NONE FOUND with AUC>0.70 at n>=30 via bc×frag thresholds on fragility_score.

=== Scan composites: AUC>0.70, n>=30 ===
  NONE FOUND with AUC>0.70 across composite scores either.
