# Decoy / Negative-Control Calibration of the Boltz-2 Validation (RP-1)

**Purpose.** The mock peer-review's controlling objection (Reviewer 2, echoed by Reviewer 3) was that
"59/60 passed ipTM>0.5 & complex-pLDDT>0.7" has no null distribution, so the pass rate is
uninterpretable and the single-model co-folding is circular. This memo reports a decoy-control
experiment that supplies the missing null, run under the **identical** Boltz-2 protocol as the
original design validation (MSA server, 3 recycling steps, 3 diffusion samples, complex confidence
read from `confidence_*_model_0.json`).

**Design.** 36 binder–target complexes co-folded on 1×A100-80GB (job wall 24 min, 0 failures):
- **24 scrambled decoys** (12 per target): each is a real design's binder sequence with residues
  randomly permuted — identical length and amino-acid composition, interface geometry destroyed.
- **8 random decoys** (4 per target): length-matched sequences drawn from the natural aa background.
- **4 reproduction controls** (2 per target): original lead + a mid-ranked design, re-folded verbatim
  to confirm the protocol reproduces the reported metrics.

**Reproduction controls check out.** Re-folding the originals recovered their reported metrics within
diffusion-sampling noise (CCL26 lead ipTM 0.942→0.904; POSTN lead 0.901→0.928), so the decoy
comparison is apples-to-apples.

---

## Headline finding: ipTM and the pass threshold do NOT discriminate designs from scrambled decoys

| Metric | CCL26 design | CCL26 scrambled | CCL26 random | POSTN design | POSTN scrambled | POSTN random |
|---|---|---|---|---|---|---|
| median ipTM | 0.852 | 0.797 | 0.909 | 0.780 | 0.804 | 0.943 |
| % pass (ipTM>0.5 & pLDDT>0.7) | 100% | 92% | 0% | 97% | 100% | 75% |

- **Scrambled decoys pass the confidence gate at essentially the same rate as real designs**
  (92% CCL26, 100% POSTN vs 100%/97% for designs).
  Mann–Whitney design>scrambled on ipTM is **not significant** for either target (CCL26 p=0.16,
  POSTN p=0.80 — POSTN scrambled actually scored marginally *higher*).
- **ipTM is even higher for random-composition sequences** (median ~0.89–0.94), confirming that for a
  small compact target, Boltz-2 assigns confident interface scores to almost any polypeptide placed
  against it. ipTM alone is **not** evidence of designed, specific binding.

## What the threshold *does* catch: complex pLDDT excludes random junk, weakly

complex pLDDT separates **random** sequences from designs, but unevenly by target. For CCL26 the effect
is clear (random median pLDDT ~0.55 vs design 0.84; random pass rate 0%). For POSTN it is weak: random
median pLDDT ~0.73 sits just above the 0.7 line and 75% (3/4) of random decoys still pass the full gate,
so pLDDT does **not** meaningfully exclude random sequences for POSTN. The design>random pLDDT
difference is significant for both (CCL26 p=0.018, POSTN p=0.0015), but the threshold only converts that
into exclusion for CCL26. Against **scrambled** designs pLDDT does no work for either target (scrambled
retains composition and folds to design-level pLDDT). Net: the fold-confidence half of the gate filters
unstructured CCL26 sequences and little else.

## The epitope-overlap analysis is the only metric with signal — and it is weak

Focal epitope engagement (≥3 hotspots contacted) is the one axis where designs beat scrambled decoys:

| | CCL26 | POSTN |
|---|---|---|
| designs ≥3 epitope hits | 43% | 17% |
| scrambled ≥3 epitope hits | 25% | 0% |

The direction is correct (designs engage the epitope more focally than scrambled controls), but the
margin is modest and, on this sample, the Mann–Whitney tests on epitope-hit count and on the
on-target composite score do **not** reach significance (all p>0.1). The on-target composite is
therefore better supported than raw ipTM as a *ranking* heuristic, but the decoy set does not
establish that it distinguishes true binders from same-composition noise at the single-design level.

---

## Consequences for the manuscript

1. **"59/60 passed" must be reframed, not just softened.** The correct statement is that 59/60 designs
   cross a confidence threshold that ~92–100% of *composition-matched scrambled sequences also cross*.
   The pass rate is not evidence of design success; it is close to the null pass rate. This is the
   single most important revision and it strengthens the paper's credibility to state it plainly.
2. **The paper's real (weak) evidence for on-target design is the epitope-overlap analysis**, which
   should be promoted to the primary validation readout, with the decoy null shown alongside it, and
   its statistical weakness disclosed.
3. **ipTM/pLDDT thresholds should be presented as fold-plausibility filters, not binding validation.**
   "Validated" → "computationally assessed"; interface confidence cannot establish affinity or
   specificity, and here demonstrably does not distinguish scrambled from designed sequence.
4. **The orthogonal-predictor check (RP-2) is now higher priority**, since single-model ipTM is shown
   to be uninformative for these targets.

---

## Follow-up: orthogonal-predictor concordance (RP-2, completed)

All 60 designs were re-folded with Chai-1 (architecturally distinct co-folder; ESM2 embeddings, MSA
server, 3 trunk recycles, 200 diffusion timesteps, seed 42) under matched settings. The two predictors
disagree sharply:

| | CCL26 (n=30) | POSTN (n=30) | all 60 |
|---|---|---|---|
| Boltz-2 median ipTM | 0.852 | 0.780 | — |
| Chai-1 median ipTM | 0.186 | 0.328 | — |
| Chai-1 ipTM>0.5 pass | 0% | 13% | 7% |
| Spearman ρ (Boltz vs Chai) | −0.14 (n.s.) | +0.38 (p=0.04) | −0.13 (n.s.) |

Both leads collapse in Chai-1 (CCL26 lead 0.942→0.134; POSTN lead 0.901→0.225). Only 4/60 designs reach
Chai-1's ipTM>0.5 line, versus 59/60 in Boltz-2, and per-design rankings are not concordant overall. No
design is confidently a binder in both models. This is the decisive confirmation that single-model
Boltz-2 confidence is not evidence of binding for these targets: the decoy calibration shows ipTM cannot
separate designs from scrambled decoys *within* Boltz-2, and the concordance result shows the designs'
Boltz-2 confidence does not replicate in a second model at all. Any future selection must require
cross-model agreement plus decoy-null-beating epitope engagement as an explicit acceptance gate.

*Data: `chai_boltz_concordance.csv` (per-design Boltz vs Chai), `TableS6_concordance_and_controls.csv`.
Figure: `chai_concordance.png`.*

*Data: `decoy_control_results.csv` (all 36 complexes: ipTM, pLDDT, pTM, epitope hits, on-target score),
`decoy_control_spec.csv` (sequence provenance). Figure: `decoy_calibration.png`.*
