Cross-model baseline_confidence robustness audit

model                           n    acc     bc-AUC (flex) [95% CI]          p
cerebras/llama-3.1-8b           413  0.62  0.483 [0.424, 0.541]  p=0.546
nvidia_nim/llama-3.1-8b         100  0.62  0.443 [0.321, 0.571]  p=0.370
cohere/command-a                100  0.79  0.493 [0.348, 0.638]  p=0.910
mistral/small                   100  0.63  0.465 [0.348, 0.585]  p=0.573

POOLED (713 rows across families): bc-AUC = 0.468 [0.422, 0.512]  p=0.1410

== Interpretation ==
bc-AUC > 0.6 on every model + non-overlapping CI from 0.5
  → robust 'bc is a signal' claim across families.
If any family shows AUC ≈ 0.5 or CI spanning 0.5,
  → the bc-only pivot has a generalization hole.
