# Response to Reviewers — EoE de novo binder design manuscript

We thank the editor and reviewers. The central methodological objection — that the Boltz-2
"59/60 passed" figure had no null distribution and the single-model co-folding was circular — was
correct, and addressing it changed the paper substantially. We ran two new compute experiments
(a negative-control decoy calibration and an orthogonal-predictor concordance check) and reframed the
validation narrative around what in-silico confidence can and cannot establish. Point-by-point below;
manuscript line references are to `EoE_manuscript_revised.md`.

## Editor's essential revisions

**E1. Calibrate the word "validated."** Done. We removed every unqualified use of "validated" for the
in-silico results (Abstract, §2.4, §2.5 header, Figure 1 legend) and replaced them with "assessed" /
"computationally assessed." The word "validated" now appears only in cautionary constructions ("not an
experimentally validated binder"; "we refrain from any claim of validated binding") or referring to
external drug-target validation.

**E2. Boltz circularity / no controls (decoys + positive control + orthogonal predictor).** Done, and
this is the core of the revision.
- *Decoys (new Fig. 5 / Fig. S13, Table S6):* 36 control complexes folded under the identical protocol
  — 24 scrambled (composition-matched, residues permuted), 8 random-composition, 4 reproduction
  controls. Result: **ipTM does not distinguish designs from scrambled decoys** (scrambled pass the
  gate at 92–100%; Mann–Whitney design>scrambled n.s., CCL26 p=0.16, POSTN p=0.80). Reported plainly in
  Abstract, §2.4, and Limitations.
- *Reproduction/positive control:* the two leads re-folded verbatim recovered their original metrics
  within diffusion noise (0.942→0.904, 0.901→0.928), establishing the comparison is like-for-like.
- *Orthogonal predictor (new Fig. 7 / Fig. S14, Table S6):* all 60 designs re-folded with Chai-1.
  Boltz-2 and Chai-1 disagree sharply (Chai median ipTM 0.19/0.33 vs 0.85/0.78; 7% pass Chai's ipTM>0.5
  vs 98% in Boltz; overall Spearman ρ=−0.13, n.s.). Both leads collapse in Chai-1. Reported in
  Discussion "Orthogonal-predictor concordance."

**E3. Reframe "59/60 passed" as threshold-crossing on a pre-filtered subset.** Done (§2.4). The text now
states the 60 are 30/target hand-selected from the 1,074 that survived complexity filtering out of
1,920 designs, and describes 59/60 as a "threshold-crossing rate on a multiply pre-filtered subset, not
a design-success rate — and … close to the null."

**E4. Show meta-analysis I², cohort count, FDR audit.** Done (§2.1, Methods). We now report the
random-effects model (DerSimonian–Laird), BH-FDR across all genes, per-gene k (CCL26 k=9, POSTN k=8),
I² (CCL26 98%, POSTN 72%), and a leave-one-cohort-out sensitivity check. The apparent "9/9 vs 8/8"
discrepancy is explained as differential per-gene cohort coverage, not an inconsistency.

**E5. Reconcile CCL26 "most discriminating" vs 11th/20 druggability rank.** Done (§2.1). The two
statements measure different things; the composite penalizes CCL26's very high I², and we state that
CCL26 was prioritized on mechanistic grounds, not composite rank.

**E6. Disclose search + druggability-score methodology; fix "2026" typo.** Done (Methods). We disclose
the PubMed query, date restriction (corrected 2015–2025; the "2026" was a typo), trial-set definition,
per-target counting method, and the composite priority-score construction (with weights deferred to
released code). "Zero direct interventional trials" is now explicitly defined.

## Reviewer-specific points (selected)

- *R2 — isolated FAS1-IV domain:* now flagged in both Results (§2.5) and Limitations.
- *R2 — acidic/helical MPNN bias:* retained in Limitations, now connected to the decoy finding (confident
  co-folding may be partly generic electrostatic complementarity).
- *R1/R3 — eotaxin redundancy (CCL11/CCL24) and periostin-as-biomarker:* added to Introduction and
  Discussion as explicit target-specific caveats.
- *R3 — "histologic remission" definition:* clarified (eosinophil-count endpoint) in Discussion.

## Items deferred (stated as future work, not claimed)

- **Buried-SASA for all 60 designs:** the original 60 design structures were not archived (only the two
  leads), so buried SASA is reported for the leads only; interface-residue counts are in Table S4 for all
  60. Re-folding all 60 solely to add this column was judged disproportionate. Noted as future work.
- **Wet-lab validation:** out of scope for this in-silico study, as agreed; the Next-steps section lists
  the specific experiments required.

## New artifacts accompanying this revision

- `EoE_manuscript_revised.md`, `EoE_supplementary_revised.md`
- `decoy_control_memo.md` (standalone write-up of both compute experiments)
- Figures: `decoy_calibration.png` (Fig. 5/S13), `chai_concordance.png` (Fig. 7/S14)
- Tables: `TableS6_concordance_and_controls.csv`, `decoy_control_results.csv`, `decoy_control_spec.csv`,
  `chai_boltz_concordance.csv`
