Solo-feature universality check (n≥50 datasets)
Sign convention: higher value = more likely hallucination (features inverted where needed).
Bootstrap 95% CI; ✓ means lower CI bound > 0.5 (feature is statistically informative).

feature                      | TruthfulQA (Cerebr | TriviaQA (n=200)   | NQ-Open (n=75)     | NIM llama-3.1-8b   | Cohere Command-A   | Mistral-Small     
----------------------------------------------------------------------------------------------------------------------------------------------------------
fragility_score              | 0.533[0.48,0.59]   | 0.530[0.44,0.62]   | 0.500[0.36,0.63]   | 0.435[0.32,0.56]   | 0.517[0.39,0.65]   | 0.562[0.44,0.68]  
baseline_confidence          | 0.477[0.42,0.54]   | 0.748[0.67,0.82]✓  | 0.525[0.39,0.66]   | 0.443[0.32,0.57]   | 0.493[0.35,0.64]   | 0.465[0.35,0.59]  
verbalized_confidence        | 0.554[0.50,0.61]   |       —            | 0.523[0.39,0.65]   |       —            |       —            |       —           
verbal_logprob_gap           | 0.445[0.39,0.51]   |       —            | 0.452[0.32,0.58]   |       —            |       —            |       —           
adversarial_fragility        | 0.518[0.46,0.58]   | 0.448[0.36,0.54]   | 0.506[0.37,0.64]   | 0.464[0.35,0.58]   | 0.483[0.34,0.63]   | 0.532[0.40,0.66]  
paraphrase_fragility         | 0.491[0.43,0.55]   | 0.630[0.54,0.72]✓  | 0.476[0.34,0.62]   | 0.402[0.27,0.53]   | 0.509[0.38,0.65]   | 0.519[0.39,0.64]  
adaptive_fragility           | 0.541[0.48,0.60]   | 0.607[0.52,0.69]✓  | 0.463[0.33,0.60]   | 0.441[0.32,0.56]   | 0.521[0.40,0.65]   | 0.516[0.39,0.63]  
impostor_fragility           | 0.533[0.48,0.59]   | 0.610[0.52,0.69]✓  | 0.549[0.41,0.68]   | 0.469[0.35,0.59]   | 0.551[0.42,0.69]   | 0.537[0.42,0.65]  
counterfactual_fragility     | 0.528[0.47,0.58]   | 0.492[0.41,0.57]   | 0.487[0.36,0.62]   | 0.464[0.35,0.58]   | 0.483[0.34,0.63]   | 0.532[0.40,0.66]  
vulnerability                | 0.521[0.46,0.58]   | 0.448[0.36,0.54]   | 0.530[0.40,0.66]   | 0.515[0.39,0.63]   | 0.477[0.33,0.62]   | 0.506[0.38,0.63]  
dissociation_rate            | 0.507[0.46,0.56]   | 0.451[0.37,0.54]   | 0.447[0.33,0.58]   | 0.492[0.44,0.55]   | 0.478[0.39,0.58]   | 0.503[0.46,0.55]  
answer_length_tokens         | 0.434[0.37,0.49]   | 0.584[0.50,0.66]   | 0.479[0.35,0.61]   | 0.482[0.36,0.60]   | 0.445[0.31,0.59]   | 0.340[0.23,0.45]  

== Universality score ==
  fragility_score                0 /  6 datasets with CI lower bound > 0.5
  baseline_confidence            1 /  6 datasets with CI lower bound > 0.5
  verbalized_confidence          0 /  2 datasets with CI lower bound > 0.5
  verbal_logprob_gap             0 /  2 datasets with CI lower bound > 0.5
  adversarial_fragility          0 /  6 datasets with CI lower bound > 0.5
  paraphrase_fragility           1 /  6 datasets with CI lower bound > 0.5
  adaptive_fragility             1 /  6 datasets with CI lower bound > 0.5
  impostor_fragility             1 /  6 datasets with CI lower bound > 0.5
  counterfactual_fragility       0 /  6 datasets with CI lower bound > 0.5
  vulnerability                  0 /  6 datasets with CI lower bound > 0.5
  dissociation_rate              0 /  6 datasets with CI lower bound > 0.5
  answer_length_tokens           0 /  6 datasets with CI lower bound > 0.5

== Interpretation ==
  A 'universal' hallucination predictor should have CI lower bound > 0.5 in
  most datasets. Features that hit in 0-2/6 datasets are unreliable.
