n=200 joined with multi-judge
flex vs majority agreement: 0.530
  unanimous rate: 0.700
  flex+/maj+ = 58  flex+/maj- = 64  flex-/maj+ = 30  flex-/maj- = 48

  full (n=200):
    label=flex       bc AUC=0.455 [0.365,0.544] p=0.336
    label=majority3  bc AUC=0.635 [0.560,0.710] p=0.001
  unanimous-only (n=140):
    label=flex       bc AUC=0.433 [0.326,0.541] p=0.255
    label=majority3  bc AUC=0.654 [0.567,0.739] p=0.001
  non-unanimous (n=60):
    label=flex       bc AUC=0.516 [0.366,0.670] p=0.843
    label=majority3  bc AUC=0.607 [0.454,0.752] p=0.188

== Interpretation ==
  If bc AUC unanimous > bc AUC non-unanimous: label noise was hiding signal
  If both ≈ 0.5: bc is genuinely chance on TruthfulQA
  majority3 > flex AUC: multi-judge is cleaner than surface flex
