Feature redundancy on 3-judge majority label (TruthfulQA n=200)
bc-only 5-fold CV AUC: 0.6266

feature                     bc_AUC  bc+feat        Δ      z   judged
fragility_score             0.6266   0.6112  -0.0154  -0.27        —
adversarial_fragility       0.6266   0.6023  -0.0244  -0.44        —
adaptive_fragility          0.6266   0.6083  -0.0184  -0.33        —
impostor_fragility          0.6266   0.6267  +0.0001  +0.01        —
counterfactual_fragility    0.6266   0.6202  -0.0064  -0.12        —
paraphrase_fragility        0.6266   0.5951  -0.0316  -0.56        —
vulnerability               0.6266   0.6220  -0.0047  -0.08        —
dissociation_rate           0.6266   0.6271  +0.0005  +0.00        —
verbalized_confidence       0.6266   0.6063  -0.0203  -0.37        —

== Interpretation ==
  |z| > 2 = significantly different AUC (rare with bc alone strong)
  compare with prior single-label result: all features |z|<0.42

  If some feature shows |z|>2 here but not before, then the previous
  null finding was hidden by label noise. This would re-open the
  perturbation research direction.

  If all still |z|<2, bc-only pivot stands even on clean labels.
