Ultimate Ensemble Hallucination Detector -- 5-fold CV
n=412, features=105, seed=42

== PRIMARY (LLM-judge label)  pos_rate=0.614 ==
  LogReg                                             AUC=0.7304  [95% CI 0.6776,0.7792]
  BLEND_top4(LogReg+GB+CatBoost)                     AUC=0.7235  [95% CI 0.6752,0.7720]
  STACK(logreg)                                      AUC=0.7191  [95% CI 0.6680,0.7700]
  GB                                                 AUC=0.6967  [95% CI 0.6459,0.7491]
  CatBoost                                           AUC=0.6904  [95% CI 0.6401,0.7415]
  RF                                                 AUC=0.6568  [95% CI 0.6034,0.7108]
  MLP                                                AUC=0.6522  [95% CI 0.5980,0.7035]
  LGBM                                               AUC=0.6518  [95% CI 0.5982,0.7079]
  XGB                                                AUC=0.6296  [95% CI 0.5771,0.6840]
  HistGB                                             AUC=0.5879  [95% CI 0.5334,0.6426]
  best=LogReg  AUC=0.7304  Brier=0.2192
  Top-10 features (CatBoost gain):
    answer_length                                 5.113
    inv__verbal_logprob_gap                       4.876
    inv__baseline_confidence                      4.356
    log__verbal_logprob_gap                       4.015
    int__adversarial_fragility__x__impostor_fragility 3.467
    log__baseline_confidence                      3.111
    baseline_confidence                           3.001
    verbal_logprob_gap                            2.926
    int__fragility_score__x__adversarial_fragility 2.358
    verbalized_confidence                         2.307
  Calibration (mean_pred, frac_pos) deciles:
    0.041 -> 0.190
    0.148 -> 0.415
    0.277 -> 0.512
    0.414 -> 0.512
    0.522 -> 0.683
    0.614 -> 0.683
    0.731 -> 0.683
    0.832 -> 0.707
    0.919 -> 0.927
    0.978 -> 0.833

== SECONDARY (flex label)  pos_rate=0.383 ==
  LogReg                                             AUC=0.6166  [95% CI 0.5601,0.6714]
  BLEND_top4(LogReg+LGBM+GB)                         AUC=0.6156  [95% CI 0.5614,0.6684]
  STACK(logreg)                                      AUC=0.6103  [95% CI 0.5554,0.6645]
  LGBM                                               AUC=0.5788  [95% CI 0.5210,0.6365]
  GB                                                 AUC=0.5747  [95% CI 0.5190,0.6279]
  CatBoost                                           AUC=0.5614  [95% CI 0.5050,0.6177]
  RF                                                 AUC=0.5599  [95% CI 0.5047,0.6140]
  XGB                                                AUC=0.5452  [95% CI 0.4910,0.6012]
  MLP                                                AUC=0.5418  [95% CI 0.4843,0.5997]
  HistGB                                             AUC=0.5366  [95% CI 0.4779,0.5945]
  best=LogReg  AUC=0.6166  Brier=0.2593
  Top-10 features (CatBoost gain):
    answer_length                                 4.957
    inv__verbal_logprob_gap                       3.878
    log__verbal_logprob_gap                       3.434
    inv__impostor_fragility                       3.018
    verbal_logprob_gap                            2.991
    counterfactual_fragility                      2.948
    inv__baseline_confidence                      2.835
    int__adversarial_fragility__x__impostor_fragility 2.772
    adversarial_fragility                         2.745
    inv__adversarial_fragility                    2.724
  Calibration (mean_pred, frac_pos) deciles:
    0.041 -> 0.214
    0.144 -> 0.220
    0.242 -> 0.341
    0.346 -> 0.366
    0.430 -> 0.439
    0.503 -> 0.415
    0.573 -> 0.390
    0.650 -> 0.390
    0.751 -> 0.488
    0.909 -> 0.571

Verdict on primary target: GOAL MET (>= 0.70)  (AUC=0.7304)