# Mock Peer-Review Report — EoE de novo binder-design manuscript

**Manuscripts assessed:**
- **Main manuscript — De novo binders vs CCL26 & POSTN in EoE**
- **Supplementary Information**

**Review model:** 3 independent expert reviewers + handling-editor synthesis. Reviewers were briefed on the manuscript framing and scope; every requested revision is achievable within that scope.

**Panel:**
- **Reviewer 1** — EoE clinician-scientist / type-2 immunology lens
- **Reviewer 2** — Computational structural biologist & protein-design methodologist lens
- **Reviewer 3** — Biostatistician + research-methodology/ethics lens

*Generated 2026-07-09 · reviews produced by independent model instances under distinct expert personas. Advisory only — verify any specific factual claim before acting on it.*

---

## Editor's Decision Letter

# Editor's Decision Letter

**Manuscript:** De novo binders vs CCL26 & POSTN in EoE (Main Manuscript + Supplementary Information)
**Submission type:** Preprint / hackathon computational deliverable

---

## Overview

This submission nominates CCL26 and POSTN as convergent, undrugged EoE targets via a nine-cohort transcriptomic meta-analysis cross-referenced against literature/trial data, then executes a de novo binder design pipeline (RFdiffusion → SolubleMPNN → Boltz-2) to produce two computationally-validated lead binder–target complexes. All three reviewers agree the design pipeline itself is well-engineered and transparently documented for a hackathon-timeline deliverable, and that the limitations section is unusually candid. All three also converge, from different angles (clinical biology, structural-design methodology, biostatistics), on the same underlying problem: several headline claims — "undrugged," "validated," "59/60 passed," and the target-nomination statistics themselves — are asserted at a higher confidence level than the presented evidence supports.

## Decision Per Manuscript

- **Main Manuscript: Major Revision**
- **Supplementary Information: Major Revision** (contingent on / parallel to the main manuscript's revisions)

This is consistent with the balance of reviews: Reviewer 1 called minor revision, but Reviewers 2 and 3 independently called major revision on grounds (validation circularity; unaudited statistics) serious enough that the editor weights them as controlling. None of the required fixes fall outside the stated in-scope set — all are computational reporting, disclosure, sensitivity-analysis, or calibration-of-language tasks achievable with data already in hand.

---

## Essential Revisions — Main Manuscript

1. **Stop using "validated" as an unqualified term (R1, R2, R3).** Replace with "computationally validated"/"in-silico assessed" consistently from the Abstract onward. Currently the hedge appears only in the Discussion while "validated" is used unqualified three times within the Abstract and Results (Abstract "we validated 60 binder–target complexes"; §2.4 "We validated the 60 selected binder–target pairs"; §2.5 header "Two validated lead binders"), plus in figure legends (R1 #3, R2 #1).

2. **Address the circularity and lack of controls in the Boltz-2 validation stage (R2, echoed by R3's point on selection cascades).** At minimum:
 - Run scrambled/decoy-sequence and, if feasible, a known-binder positive control through the identical Boltz-2 protocol to calibrate what ipTM/pLDDT thresholds actually distinguish (R2 Major #1–2, Feasible #1).
 - Report the on-target composite score's rank-correlation against raw ipTM across all 60 designs, and justify its weighting constants (R2 Major #3, Feasible #3).
 - Explicitly frame "59/60 passed" as a threshold-crossing rate on a multiply-pre-filtered subset, not a general design-success rate (R2 Major #4, R3 Major #6 — raised independently by both reviewers, this is a convergent must-fix).

3. **Surface the meta-analysis statistical audit trail (R1, R3).** Report per-cohort effect sizes/direction and I² for CCL26 and POSTN specifically (not just "computed" in the abstract), explain the 9/9-vs-8/8 cohort denominator discrepancy for POSTN, and state the total number of FDR-significant genes to contextualize the reported adjusted p-values (R3 Major #1–3, Feasible #1–3).

4. **Engage with why these targets are "undrugged" rather than presenting the trial-pipeline gap as unambiguous opportunity (R1, R3).** Address eotaxin redundancy (CCL11/CCL24 vs CCL26 via shared CCR3), periostin's biomarker-vs-target history and physiological repair role, and disclose the literature/trial corpus search methodology (query strings, inclusion criteria, date-range — including the apparent 2026 typo flagged by R3 Minor #1) so "zero direct trials" is falsifiable rather than asserted (R1 Major #1–2, R3 Major #5).

5. **Reconcile CCL26's druggability rank (11th/20) with its claim as "the single most discriminating EoE transcript" (R1 #6, R3 #4).** Both reviewers independently flagged this internal tension; disclose the shortlist scoring formula/weights (or state it was qualitative) and address the gap in text.

6. **Temper the POSTN causal/clinical framing (R1 Major #4–5).** Label incremental benefit of direct POSTN neutralization over existing IL-13-pathway blockade as an untested hypothesis, and clarify that "histologic remission" is an eosinophil-count endpoint, not evidence of ongoing symptomatic disease.

7. **Qualify the isolated-domain caveat for POSTN in the Results section itself, not only the Discussion (R2 Major #5).** State explicitly in §2.5 that validation used isolated FAS1-IV (5WT7), not full-length POSTN.

8. **Second-opinion structural check (R2 Major #1, Feasible #2).** Run the two leads (and ideally top ~10 by on-target score) through an architecturally distinct co-folding predictor and report agreement — this was listed as "future work" but is computationally feasible now and directly addresses the single-predictor circularity concern.

## Essential Revisions — Supplementary Information

1. **Deposit the full meta-analysis table and per-cohort statistics (R1, R3 — convergent).** Table S1 currently shows only top-60 genes; both reviewers want the full 25,654-gene table (or at minimum per-cohort log2FC/SE/I² for CCL26 and POSTN) to make the pooled statistics auditable.

2. **Disclose the druggability shortlist scoring formula/weights (R1, R3 — convergent).** Both reviewers flagged the same asymmetry: the Boltz-2 on-target score formula is fully disclosed, but Table S2's construction is not.

3. **Specify the literature/trial corpus search methodology in Supplementary Methods (R1, R3 — convergent).** Query strings, inclusion criteria, date-range handling, and how the 45-target Table S3 subset was drawn from the 591-abstract/157-trial corpus.

4. **Extend Table S4 to include buried SASA and interface-residue counts for all 60 designs, and add decoy/control rows (R2).** Currently only the two leads are contextualized.

5. **Version-pin the computational pipeline (R2).** Software commit hashes/weights for RFdiffusion, ProteinMPNN, Boltz-2, and random seeds (or documented run-to-run variance) for reproducibility.

---

## Substantive Disagreements and Adjudication

- **Overall recommendation split (R1: minor vs. R2/R3: major).** R1's assessment focused primarily on the epidemiological/clinical-mechanistic argument and found it fixable with citations and hedging; R2 and R3 independently identified deeper evidentiary gaps — one in the structural-validation logic (circularity, no controls), one in the statistical backbone (unshown I², undisclosed formulas). These are not contradictory so much as complementary blind spots each reviewer's expertise caught that the others didn't emphasize as strongly. The authors should treat this as an instruction to fix **both** the narrative-calibration issues (R1) **and** the harder quantitative-audit issues (R2, R3) — the manuscript cannot be considered adequately revised by addressing only the softer language fixes.
- **No reviewer contests the design pipeline's basic soundness or the appropriateness of AI-disclosure/in-silico framing.** This is not in dispute and should not be re-litigated in revision.
- **Minor tension on emphasis:** R1 treats the "undrugged" framing as a citation/nuance gap; R2/R3 treat validation and statistics as the more consequential gap. Authors should prioritize R2's control experiments and R3's statistical disclosures as the highest-value fixes, since these bear most directly on whether the two headline leads can be trusted at all, before polishing the epidemiological narrative.

## Path to Acceptance

**Main manuscript:** The paper will be ready for acceptance once (a) all "validated" language is calibrated to what a single co-folding model plus composite heuristic actually shows, backed by decoy/positive-control comparisons and, ideally, an orthogonal structural predictor; (b) the meta-analysis statistics (I², cohort-count discrepancy, FDR audit) are shown rather than asserted; (c) the "undrugged" and "single most discriminating transcript" claims are reconciled with the paper's own druggability ranking and the known redundancy/biomarker history of these targets; and (d) hedging is applied consistently from the Abstract onward rather than concentrated in the Discussion. These are all achievable from data and literature the authors already have in hand.

**Supplementary Information:** Acceptance follows directly from closing the reproducibility gaps that parallel the main-text revisions — full gene-level statistics, disclosed scoring formulas, corpus methodology, decoy-control tables, and version/seed pinning. No new computation beyond what the pipeline already supports is required; this is a disclosure and transparency exercise.

We look forward to a revised submission and thank the authors for a genuinely interesting and mostly well-hedged computational deliverable — the core pipeline is a credible contribution; it needs its evidentiary claims brought into tighter alignment with what has actually been shown.

---

## Full Reviews

## Reviewer 1
### Reviewer 1: EoE clinician-scientist, type-2 immunology / therapeutic pipeline lens

## Main manuscript — De novo binders vs CCL26 & POSTN in EoE

**Brief summary**

The authors combine a nine-cohort EoE transcriptomic meta-analysis with literature/trial-pipeline mining to nominate CCL26 and POSTN as "convergent, undrugged" targets, then run a de novo binder design pipeline (RFdiffusion → SolubleMPNN → Boltz-2) against defined functional epitopes on each and report two computationally validated lead binder–target complexes. The paper is explicit throughout that this is entirely in-silico work intended as a starting point for future wet-lab validation.

**Strengths**

- The two-target framing (recruitment axis vs. remodeling axis) is a genuinely sensible way to think about residual disease space not covered by IL-4Rα/IL-5/IL-5Rα/TSLP biologics, and is stated with reasonable care in most places.
- Appropriately distinguishes a "ligand trap" mechanism (occluding a functional surface on a secreted protein) from receptor antagonism, and gives a defensible reason for preferring the ligand-trap route for CCL26 (CCR3 is a GPCR with expression breadth issues; CCL26 is small and tractable).
- The POSTN loop-interruption logic (IL-13→DSG1↓→POSTN↑ feed-forward circuit) is correctly flagged as an *indirect* route to barrier benefit, not a direct desmosomal effect — this is the right level of hedging for that specific claim, and the Discussion explicitly walks it back to "a downstream consequence, not a direct binding effect."
- Limitations section names real, specific technical caveats (SolubleMPNN acidic/helical bias, FAS1-IV-only construct, POSTN's physiological repair role requiring confirmation of neutralizing vs. agonist behavior) rather than a generic "further work is needed" gesture.

**Major concerns**

1. **The "zero direct interventional trials" argument is doing more work than it can bear.** The abstract and Results (2.1) present the absence of CCL26/POSTN-directed trials as evidence of an "unexploited" target. But there is a simpler explanation the manuscript never engages with: eotaxin/CCR3-axis blockade was tried in other eosinophilic diseases (e.g., CCR3 antagonists, anti-IL-5 approaches that work partly by depleting the eosinophils CCL26 recruits) and periostin itself has been explored as a *biomarker*, not a drug target, in asthma/EoE for reasons that may include biological or druggability concerns (e.g., periostin's normal role in tissue repair, redundancy among eotaxins CCL11/CCL24/CCL26 signaling through the same CCR3). The paper needs to at least address *why* these axes have not been drugged despite mechanistic plausibility — otherwise "undrugged" risks being read as "unrecognized" when it may be "tried and deprioritized" or "mechanistically harder than it looks." This is a scoping/citation gap, fully fixable without new experiments.

2. **Redundancy among eotaxins is not addressed and materially weakens the CCL26-trap rationale.** CCR3 is also activated by CCL11 (eotaxin-1) and CCL24 (eotaxin-2). A CCL26-selective trap leaves CCR3 signaling from CCL11/CCL24 untouched. Given that CCL26 is highlighted as "the single most discriminating EoE transcript" (PMID 17900656), the authors should explicitly state whether CCL26 is thought to be non-redundant in EoE tissue (e.g., because CCL11/CCL24 are not co-induced to the same degree in esophageal epithelium) — otherwise a reader is left to assume monotherapy against CCL26 would fully block eosinophil chemotaxis, which the biology does not support. This needs a sentence of nuance in 2.2 and the Discussion, not new data.

3. **The clinical-translation framing overstates what a "computationally-validated lead" means.** The abstract's closing sentence — "we present them as a mechanistically-grounded starting point for two differentiated EoE therapeutic strategies" — is acceptable, but earlier phrasing ("we validated 60 binder–target complexes," "two computationally-validated lead binder-target complexes") uses "validated" in a way that a non-specialist or press-facing reader will misread as biophysical validation. Boltz-2 ipTM/pLDDT is a structure-prediction confidence metric, not a binding affinity or neutralization readout, and the paper itself notes "de novo binder success rates at the bench remain modest even for high-confidence designs" — but that caveat arrives only in the Discussion, after "validated" has been used unqualified three times in the Abstract/Results (plus figure legends). The terminology should be consistently qualified (e.g., "computationally validated" or "in-silico validated" every time, not just "validated") from the abstract onward.

4. **The POSTN barrier-recovery claim is directionally reasonable but under-cited for causality.** The Discussion states DSG1 loss "potentiates inflammation partly *through* POSTN" (PMID 24220297) and periostin is "the top induced gene when DSG1 is lost." This is consistent with the cited Sherrill/Blanchard-lineage literature, but the causal arrow POSTN→remodeling→barrier persistence in the specific sense implied (i.e., that neutralizing POSTN would meaningfully accelerate barrier recovery beyond what IL-13 blockade already achieves) is asserted rather than demonstrated even in the cited literature. Since dupilumab and other IL-13-pathway agents already lower POSTN indirectly and produce histologic and (to a degree) barrier improvement, the incremental benefit of a *direct* POSTN binder over existing upstream blockade is speculative and should be labeled as a hypothesis to be tested, not a mechanistic strategy already differentiated from the existing pipeline. The current wording ("addresses the disease process that persists in histologic remission") implies a clinical benefit that has not been shown for *any* periostin-directed intervention in any indication.

5. **Histologic remission ≠ symptomatic/functional remission, and this distinction is glossed.** The manuscript's central residual-disease argument rests on PMID 39343172 (POSTN/DSG1 persisting in histologic remission). This is a fine citation, but the paper should be explicit that "histologic remission" in EoE is an eosinophil-count-based definition and does not by itself establish ongoing symptomatic disease or need for further therapy — readers less familiar with EoE endpoints could over-read this as "patients in remission are still sick and need this drug," which is not established.

6. **The druggability shortlist ranking (POSTN 3rd/20, CCL26 11th/20) is stated without any indication of what separates rank 3 from rank 11, or what's above CCL26.** Given the abstract leads with CCL26 as "the single most discriminating EoE transcript," a reader will want to know why it ranks only 11th on tractability — is this because of the eotaxin redundancy issue in #2, structural flexibility, or something else? This should be unpacked in text, not left to Fig. S4/Table S2 alone.

**Minor concerns**

1. Abstract states "59/60 passed" but doesn't clarify until Results 2.4 which single design failed (a POSTN design) — a one-clause addition to the abstract would pre-empt a reader wondering if a CCL26 design failed.
2. The epitope-hotspot-hit disparity between targets (13/30 CCL26 ≥3/5 hotspots vs. 5/30 POSTN ≥3/4 hotspots) is attributed to "the flatter POSTN integrin face" — plausible, but this is asserted identically in three places (Results 2.4, Fig. S10 legend, Discussion) without independent structural evidence (e.g., surface curvature metric) supporting the claim beyond visual/qualitative inference.
3. PMID 12061839 is cited for CCL26 being "≈100× more potent" than IL-13 as an inducer — worth double-checking this is IL-4 vs IL-13 induction *potency* and not conflating with eosinophil chemotactic potency, since the sentence structure is slightly ambiguous on which potency is being compared.
4. FAS1 domain IV construct (5WT7) is 140 residues but described as "the isolated FAS1 domain IV" — worth a clause noting whether this fragment's integrin-binding patch conformation is expected to match the full-length, multi-FAS1-domain protein (this is partially addressed in Limitations but could be flagged where the structure is first introduced in 2.3, not just in Discussion).
5. "Clinical pipeline mining" claims 157 EoE trials cross-referenced — no versioning/date-stamp of the ClinicalTrials.gov pull is given in Methods, which matters for reproducibility of the "zero direct trials" claim as the pipeline evolves.

**Feasible revisions**

1. Add a paragraph (Introduction or Discussion) addressing eotaxin redundancy (CCL11/CCL24 vs CCL26) and citing esophageal-tissue-specific expression data if available in the meta-analysis/single-cell dataset already assembled, to justify CCL26 selectivity as non-redundant in this tissue context.
2. Soften/standardize "validated" → "computationally validated" or "in-silico validated" at every occurrence in Abstract and Results, matching the hedge already present in the Discussion.
3. Add explicit discussion of why the eotaxin/CCR3 and periostin axes have historically not been drugged despite long-standing mechanistic interest (redundancy, biomarker vs. target history, physiological repair roles), rather than presenting the trial-pipeline gap as unambiguous opportunity.
4. Rephrase the POSTN barrier-recovery claim to explicitly flag that incremental benefit over existing IL-13-pathway blockade is an untested hypothesis, not an established differentiation.
5. Add one sentence clarifying that "histologic remission" is an eosinophil-count endpoint, to prevent over-reading of the residual-disease argument.
6. Expand on Fig. S4/Table S2 in-text to explain the basis for the POSTN vs. CCL26 tractability rank gap (what pushed CCL26 to 11th).
7. Add ClinicalTrials.gov / PubMed pull date to Methods for reproducibility of the "zero trials" and literature-count claims.

**Recommendation:** **Minor revision.** The target rationale is fundamentally sound and appropriately caveated in the Discussion/Limitations, but several claims in the Abstract/Results (redundancy of eotaxins, causal weight of the POSTN loop-breaking hypothesis, and consistent hedging of "validated") need tightening before the mechanistic story is presented as cleanly as it currently reads.

---

## Supplementary Information

**Brief summary**

The SI provides the complete figure set (S1–S12), five supplementary tables covering meta-analysis, druggability shortlist, literature/trial cross-reference, fold-back validation, and MPNN summaries, plus two rotation-video files, together giving a reasonably complete audit trail of the pipeline described in the main text.

**Strengths**

- Table S4 (per-design ipTM/pLDDT/pTM/epitope-hits for all 60 designs) and Table S5 (MPNN summary across 240 backbone×temperature groups) are exactly the granularity needed to independently re-rank designs and check the lead-selection logic — this is good reproducibility practice.
- Fig. S6 (per-residue SASA with hotspots overlaid) directly substantiates the epitope-exposure claim rather than asserting it, which is the correct standard for a design-hotspot justification.
- Fig. S2's explicit combination of omics effect size, literature attention, and trial count in one plot is a clear and appropriately modest way of making the "convergent but undrugged" case, and it's good that the underlying counts are also tabulated (Table S3) rather than left only as a scatter plot.

**Major concerns**

1. **Table S3 (literature × omics cross-reference) needs its search methodology specified.** "591 EoE abstracts, 2015–2026" is stated in the main Methods, but neither Methods nor SI specifies the actual PubMed query string, inclusion criteria, or how "mentions" of a target were counted (title/abstract text search? MeSH terms? Any disambiguation for common gene-symbol collisions, e.g., "TSLP" vs. other acronyms)? Given that the "zero direct trials" and "high literature attention" claims underpin the entire target-nomination argument, this needs to be reproducible in the SI Methods, not just asserted.
2. **The 20-gene druggability shortlist (Table S2) construction criteria are stated qualitatively ("effect size, direction concordance, single-cell specificity, antibody tractability") but the actual scoring formula/weights are not given anywhere** — this is analogous to the on-target composite score formula, which *is* given explicitly for the Boltz-2 stage. The asymmetry (rigorous formula disclosure for the design-validation stage, none for the target-nomination stage) should be corrected; if a scoring formula exists it should be in Supplementary Methods, and if it was a qualitative/manual rank it should say so.
3. **Heterogeneity (I²) is mentioned as computed ("yielding pooled log2FC, standard errors, heterogeneity (I²)...")** but no I² values are shown for CCL26/POSTN anywhere in the main text, figures, or the described Table S1 columns as presented — despite I² being explicitly listed as a Table S1 column. Given nine cohorts of presumably heterogeneous platforms/cohort designs, an I² for the two lead targets should be surfaced in-text or in a figure legend, since high heterogeneity would qualify the "up in 9/9 cohorts" concordance claim non-trivially.
4. **Multiple-testing correction is stated as Benjamini-Hochberg over "25,654 genes,"** but it is unclear whether this FDR correction was applied within each cohort before pooling, on the pooled random-effects p-values, or both — this matters for the strength of the "adjusted p = 3×10⁻⁴" (CCL26) and "3×10⁻¹⁶" (POSTN) figures reported in the Abstract. A one-line clarification in Supplementary Methods would resolve this.

**Minor concerns**

1. Fig. S3 legend notes single-cell analysis is restricted to "endothelial, basal/suprabasal epithelial, mast, myeloid, T cell" compartments — no eosinophil cluster is listed, which is somewhat notable for an eosinophilic disease; a sentence on why eosinophils are absent (e.g., not well captured by the scRNA-seq platform/digestion protocol used, consistent with known technical difficulty of profiling eosinophils by scRNA-seq) would preempt a reviewer/reader question.
2. Fig. S8's complexity filter thresholds (entropy ≥2.6, max run ≤5, top-aa fraction ≤0.40) are presented without justification for why these specific cutoffs were chosen over alternatives — even a one-sentence rationale (e.g., "based on typical composition of natural globular proteins") would help.
3. Supplementary Media files are referenced only by filename in prose ("CCL26_binder_complex_360.mp4") without confirmation of accessibility/format notes (codec, duration) beyond the Methods table — minor, but worth a consistent cross-reference index as promised ("see Supplementary Information for the full index") which is not fully realized as a standalone index section.
4. Table S1 is described as "top 60 genes by |z|" — worth stating explicitly whether CCL26 and POSTN's rank within that top-60 list is disclosed (i.e., are they #1 and #2, or lower), since the abstract's superlative framing ("single most discriminating EoE transcript") would be best directly supported by showing rank in this specific meta-analysis, not only via the external citation PMID 17900656.

**Feasible revisions**

1. Add PubMed query string, date range handling, and disambiguation method to Supplementary Methods for the literature corpus underlying Table S3/Fig. S2.
2. Disclose the exact scoring formula (or explicitly state it was qualitative/manual) for the Table S2 druggability shortlist, matching the transparency given to the Boltz-2 on-target score.
3. Report I² (and k, already a stated column) for CCL26 and POSTN explicitly in text or a table excerpt, and comment on what heterogeneity level is observed for the two lead targets specifically.
4. Clarify at which stage(s) BH correction was applied (per-cohort vs. pooled) in Supplementary Methods.
5. Add a brief note on eosinophil representation (or lack thereof) in the single-cell dataset used for Fig. S3/2.1.
6. State CCL26/POSTN's actual rank (not just "top 60") within Table S1 to directly substantiate the "single most discriminating transcript" claim from the group's own data rather than relying solely on an external citation.

**Recommendation:** **Minor revision.** The supplementary package is substantially complete and well-organized for a hackathon-scale deliverable, but three specific reproducibility gaps (literature-search methodology, druggability-score formula, and I²/heterogeneity reporting) need to be closed before the target-nomination statistics can be taken at face value.

---

**Cross-cutting comment:** This is a well-structured, honestly-caveated computational pipeline paper whose weakest link is not the design methodology (which is described in good, checkable detail) but the epidemiological/mechanistic argument that CCL26 and POSTN are "undrugged" targets ripe for intervention — that argument needs to engage with eotaxin redundancy, the biomarker-vs-target history of periostin, and heterogeneity/robustness of the underlying meta-analysis before the design work it motivates can be read as addressing a genuine, well-characterized therapeutic gap rather than a plausible but under-interrogated one. None of the requested fixes require new wet-lab or clinical data — they are citations, formula disclosures, and calibration of existing claims, all achievable within the current dataset and literature already assembled by the authors.

## Reviewer 2
# Peer Review

### Reviewer 2: Computational structural biologist specializing in de novo protein design pipelines (RFdiffusion/ProteinMPNN/AF-Boltz co-folding validation)

---

## Main Manuscript — De novo binders vs CCL26 & POSTN in EoE

**Brief summary.** The authors combine a nine-cohort transcriptomic meta-analysis with literature/trial mining to nominate CCL26 and POSTN as undrugged EoE targets, then run an RFdiffusion→SolubleMPNN→Boltz-2 pipeline (80 backbones, 1,920 sequences, 60 folded complexes) to produce one "validated" lead binder per target, reporting ipTM 0.942 (CCL26) and 0.901 (POSTN) with epitope-hotspot engagement. The paper is explicit that all results are computational and frames the leads as starting points for future wet-lab work.

**Strengths**
- Clear, staged design funnel (Fig. 3A) with quantitative attrition at each step (80→1,920→1,074→60→2), which is good reproducibility hygiene.
- Sensible move to look beyond raw ipTM/pLDDT and construct an epitope-coverage-weighted "on-target" score (Fig. 4C, S11) — recognizes explicitly that confident-but-off-target folds are a real failure mode.
- Reasonable, literature-grounded epitope choices (CCR3/GAG basic face on CCL26; FAS1-IV integrin patch on POSTN), with SASA-based confirmation that hotspot residues are surface-exposed.
- Appropriately worded limitations paragraph flags the MPNN acidic/helical compositional bias, use of isolated FAS1-IV vs. full-length POSTN, and the physiological (non-pathogenic) role of POSTN.

**Major concerns**

1. **"Validated" is used well beyond what a single co-folding model can establish.** The abstract, Results (§2.4–2.5), and even the figure legend for Fig. 1 assert the complexes are "computationally-validated." Boltz-2 ipTM/pLDDT/pTM are self-consistency metrics from the *same* model class used to generate hotspot-conditioned inputs; they are a measure of "does this model believe its own fold," not evidence of binding, affinity, or specificity. There is no independent structural predictor (AlphaFold3-class, Chai-1) used to cross-validate any of the 60 complexes, let alone the two leads, despite the Discussion listing this as a "next step." Using a single non-orthogonal predictor for both design conditioning and validation is a circularity risk that should be disclosed as a limitation of "validation" in the Results section itself, not deferred to the Discussion.

2. **No negative/decoy controls and no positive/reference controls.** There is no reported ipTM/pLDDT distribution for (a) scrambled or shuffled binder sequences against the same targets, (b) binders designed against an unrelated/non-epitope surface, or (c) a known true binder–target pair (e.g., an existing anti-CCL26 nanobody or the CCR3 N-terminal peptide complexed with CCL26) run through the identical Boltz-2 protocol. Without a decoy baseline, "ipTM > 0.5, pLDDT > 0.7" cannot be interpreted quantitatively — Boltz-2 is known to sometimes assign moderate-to-high ipTM to non-interacting or weakly-interacting pairs, especially for short designed binders against defined hotspots (the hotspot conditioning itself can inflate apparent interface confidence). A held-out positive control establishes that the pipeline can distinguish "real" affinity-relevant contacts from designed-into-existence ones; a scrambled/negative control establishes the false-positive rate of the pass criteria. Neither is present, so the near-universal 59/60 pass rate is not clearly interpretable as pipeline success versus a permissive threshold.

3. **The on-target composite score is ad hoc and not benchmarked.** ipTM × (0.5 + 0.5 × coverage) × min(pLDDT/0.7,1) is a reasonable-looking heuristic, but its functional form (why 0.5 baseline weight on coverage rather than a different split; why the pLDDT cap at 0.7 rather than the pass threshold value used elsewhere) is not justified against any external outcome (e.g., experimentally known binder rankings, or even just correlation with buried SASA or interface residue count across the 60 designs). As constructed, it can reward a design with perfect epitope coverage and mediocre ipTM over one with excellent ipTM elsewhere; whether that trade-off is actually predictive of biological on-target neutralization is untested. At minimum, report the correlation/rank-agreement between raw ipTM ranking and on-target-score ranking for the 60 designs to show this metric is doing non-trivial work, not just re-ordering the same top hits.

4. **Selection funnel introduces uncontrolled information leakage/optimism bias.** 30 designs per target were chosen for Boltz-2 validation "by a composite of MPNN score and sequence entropy" (§2.3) — i.e., pre-filtered by a proxy for sequence quality before ever seeing structural validation, and then the *same* on-target composite score is used both to rank the 60 validated designs and to pick the lead. This multi-stage best-of-N selection (80→60 backbones scored on hotspot contacts, then MPNN-score/entropy-selected 30/target, then Boltz-2-ranked, then on-target-score-ranked) means the final ipTM=0.942 is the maximum of a curated top percentile, not a representative pipeline output. The manuscript should report the full distribution (already partly done in Fig. S9/S11) but should also explicitly state the effective "N" at each selection stage so the reader can judge the multiple-comparisons-adjusted extremeness of the reported single best value, rather than reading 0.942 as characteristic performance.

5. **Full-length POSTN vs. isolated FAS1-IV domain is a significant, under-weighted caveat.** POSTN is a multi-domain secreted protein (EMI domain + 4 FAS1 repeats + C-terminal region); the design and validation were performed exclusively against the isolated 140-aa FAS1-IV fragment (PDB 5WT7). It is entirely possible that in the full-length, likely oligomeric/multimeric physiological context, this face is occluded by inter-domain packing, glycosylation, or self-association, or that the "exposed integrin patch" identified in the isolated NMR construct is not solvent-accessible in the native protein. This is mentioned once in the Discussion as caveat #2 but should also temper the Results-section framing (§2.5) where the POSTN lead is described as engaging "the FAS1 integrin-interaction patch" without qualification that this was assessed on an isolated domain construct, not the physiological antigen.

6. **Buried SASA and interface-residue-count reporting lack context for interpretation.** 1,220 Å² (CCL26) and 842 Å² (POSTN) buried surface areas are reported without comparison to typical interface sizes for productive protein–protein interactions (~800–1,500 Å² is typical for a mini-binder, so these values are plausible, but the manuscript doesn't state this benchmark) or to the distribution of buried SASA across the other 58 designs. Readers cannot tell if the lead's interface area is typical, small, or unusually large for the design set — this should be added to Fig. S12 or Table S4.

7. **Compositional bias (acidic/helical) is flagged but not quantified or mitigated.** The Discussion states the leads are "strongly helical and acidic" as a "known compositional bias of SolubleMPNN" — but no actual sequence composition data (pI, %Glu/Asp, %helix by secondary-structure prediction) is given for the two leads or the design population as a whole, and no attempt (e.g., re-designing with a charge-balance constraint, or running a naive electrostatic complementarity calculation beyond hand-waving "electrostatically complementary to the basic CCL26 GAG face") was made to test whether this bias is (a) real (b) actually beneficial or a solubility/aggregation liability. A brief quantitative table (net charge, GRAVY, %helix from the RFdiffusion backbone or DSSP on Boltz-2 output) would substantiate or refute this claim rather than asserting it.

8. **Epitope-coverage metric is coarse and possibly gameable.** "Any target atom within 5 Å of a binder atom" for a hotspot residue is a binary yes/no per hotspot; a single transient/glancing contact from one binder side-chain atom counts identically to a fully buried, multi-atom contact. This 5 Å cutoff is generous (typical direct H-bond/vdW contact distances are 3.5–4.5 Å) and may over-count marginal contacts as "engaging" the hotspot. Given that epitope engagement is central to the paper's differentiator claim (on-target vs. merely-confident), a stricter/more informative metric (e.g., number of atom-pairs <4 Å, or fraction of hotspot residue SASA buried) should be reported alongside, or in place of, the binary hit count.

**Minor concerns**

1. Abstract states "we validated 60 binder–target complexes" — validated should be reserved for something closer to experimental confirmation, or the phrase should read "computationally assessed" throughout to keep internal consistency with the (appropriately hedged) Discussion.
2. RFdiffusion "noise scales set to 0" is stated without justification — zero noise scale increases design determinism/success rate on the hotspot but is also known to sometimes reduce backbone diversity; a one-line rationale would help reproducibility-minded readers.
3. Random seeds are not reported for RFdiffusion, ProteinMPNN, or Boltz-2 sampling — for full reproducibility these should be specified (or noted as not fixed, with a comment on run-to-run variance).
4. Table S4/S5 column definitions are described only in prose in the SI table list; a proper data dictionary (or header row description) should accompany the CSVs.
5. Fig. 4A/S9 "marker size proportional to epitope hotspot hits" is a good idea but with only 5 discrete hotspot values per target this is hard to read at figure resolution; consider a categorical color or faceted panel instead.
6. It would help to state explicitly, per lead, whether the reported ipTM/pLDDT is from a single seed or the best/median of multiple Boltz-2 diffusion samples (3 diffusion samples are run per prediction per Methods) — is the reported number the max over samples, or an average? This materially affects interpretation of "0.942."

**Feasible revisions**

1. Run the identical Boltz-2 protocol on (a) 10–20 scrambled/shuffled versions of each lead binder sequence and (b) if available, one published CCL26- or POSTN-binding protein/peptide (e.g., a CCR3-derived peptide or an antibody CDR-loop model) as positive/negative controls, and report where the leads fall relative to these distributions. This is fully computational and directly addresses the circularity concern.
2. Run at least the two lead complexes (ideally the top ~10 by on-target score) through a second, architecturally distinct co-folding predictor (AlphaFold3-class server or Chai-1) and report agreement/disagreement — already flagged as a "next step" in Discussion but should be pulled forward into Results as it is computationally feasible now, not contingent on wet-lab access.
3. Report full-distribution or rank-correlation statistics (raw ipTM vs. on-target score, Spearman's ρ) across all 60 designs to justify that the composite score is not simply relabeling the ipTM ranking, and briefly justify the specific weighting constants chosen.
4. Add a quantitative sequence-composition table (net charge/pI, %acidic, DSSP-assigned %helix) for the two leads and the design population, to substantiate (or qualify) the "acidic/helical bias" claim made in the Discussion.
5. Explicitly state in §2.5 (not only the Discussion) that POSTN validation used the isolated FAS1-IV domain construct, not full-length POSTN, so the Results section is not overstating epitope engagement claims relative to the physiological antigen.
6. Report buried-SASA and interface-residue-count distributions for all 60 designs (not just the two leads) so readers can judge whether the leads are representative or outliers of the design population.
7. Clarify and standardize the epitope-coverage/hotspot-contact distance metric (e.g., report both 4 Å and 5 Å cutoffs, or switch to fractional hotspot-residue SASA burial) and state explicitly whether reported ipTM/pLDDT are single-sample, best-of-3, or mean-of-3 Boltz-2 diffusion outputs.
8. State (or fix) random seeds for RFdiffusion/ProteinMPNN/Boltz-2 runs, or explicitly report run-to-run variance from repeat sampling with different seeds, to support the reproducibility claims made in §5/Methods.

**Recommendation: Major revision.** The pipeline and rationale are sound and the hedging language is largely appropriate, but the central evidentiary claim ("validated" leads, 59/60 pass rate) rests entirely on a single non-orthogonal co-folding model with no negative/positive controls and an unbenchmarked composite score — these are addressable computational analyses, not wet-lab asks, and are necessary before the confidence language in the abstract and Results is warranted.

---

## Supplementary Information

**Brief summary.** The SI provides the full figure set (S1–S12), five data tables, two rotation-video files, and a compact methods-parameter table underpinning every stage of the pipeline described in the main text.

**Strengths**
- Genuinely useful diagnostic figures beyond the headline numbers — S5, S8, S10 show the actual distributions behind the pass/fail summary statistics rather than only reporting the winners, which is good practice.
- The parameter table is a clean, scannable reproducibility reference (checkpoint names, weight files, sampling temperatures, filter thresholds all in one place).
- Table S4 (per-design metrics for all 60 folds) and Table S5 (per-backbone×temperature MPNN summary) are the right granularity to allow independent reanalysis, in principle.

**Major concerns**

1. **No raw structure files, code, or environment specification are listed as directly accessible in the SI itself** — the main text says data are "available as project artifacts," but the SI index does not include the actual PDB/CIF coordinate files for all 60 complexes (only the two lead structures are mentioned as available), the RFdiffusion/ProteinMPNN/Boltz-2 invocation scripts, or software version/commit hashes. For a paper whose central claims are computational reproducibility depends on being able to re-run or at least re-score the exact same complexes; version pinning (RFdiffusion commit, Boltz-2 version/model weights hash) is currently absent.
2. **Table S4 should additionally report, per design, buried SASA and interface residue count** (currently only described for the two leads in the main text) so that outlier-ness of the leads on these dimensions (not just ipTM/on-target score) can be assessed by the reader without rerunning anything.
3. **No decoy/negative-control rows exist in any table** — as noted for the main manuscript, this SI would be the natural place to house a scrambled-sequence or off-target-epitope control table; its absence is consistent with the major concern raised above.
4. **Fig. S5/S10 report backbone-level and coverage-level QC but never connect them to final lead identity** — it would strengthen traceability to overlay/mark, in each relevant SI figure, which point corresponds to the eventual chosen lead design (as is nicely done in Fig. S9/S11), so a reader can follow one design's trajectory through every QC stage.

**Minor concerns**
1. Table S5 description says "240 backbone×temperature groups" (80 backbones × 3 temperatures) but the main text's sequence count is 1,920 = 80×8×3 — worth a one-line reconciliation so it's clear S5 is a per-group summary, not per-sequence.
2. Movie S1/S2 codecs/parameters (libx264, 30 fps, 180 frames) are given but not the camera/rendering script — for full reproducibility this should be versioned alongside the PyMOL session file if being deposited.
3. Table S2's "antibody tractability" and "suggested modality" columns are qualitative/subjective; a brief methods note on how these were scored (rule-based? literature-derived? LLM-assisted?) would help readers judge Fig. S4's ranking.

**Feasible revisions**
1. Add explicit software version/commit hashes and a data/code availability statement listing exactly which artifacts (scripts, structure files, environment) are retrievable, rather than a generic "available as project artifacts."
2. Extend Table S4 to include buried SASA and interface-residue counts for all 60 designs (not just the 2 leads), and add a decoy/control block (scrambled or off-epitope designs) per Major Concern #3 above.
3. Cross-reference the lead design's identity across Figs. S5, S7, S8, S9, S10, S11 (e.g., consistent marker highlighting) so its provenance through each QC gate is traceable in one visual thread.
4. Add a one-line reconciliation note clarifying the S5 "240 groups" vs. main-text "1,920 sequences" relationship.

**Recommendation: Minor revision** (contingent on the main manuscript's major revisions) — the SI is structurally solid and mostly just needs the additional control/version-tracking data generated by the main-manuscript revisions to be reflected here.

---

**Cross-cutting comment:** This is a well-organized, honestly-framed hackathon deliverable whose target-selection logic (omics + literature + trial-gap cross-referencing) is more rigorous than its structural-validation logic. The recurring weakness across both documents is that "validated" is asserted on the strength of a single self-consistency metric (Boltz-2 ipTM/pLDDT) with no decoy, no positive control, and no orthogonal predictor — all of which are purely computational additions achievable within the stated scope and would substantially strengthen the paper's core claim without requiring any wet-lab work. Tempering "validated" to "computationally prioritized" or "in-silico supported" throughout the Abstract/Results, pending those additional controls, would already bring the claims into better calibration with the evidence shown.

## Reviewer 3
### Reviewer 3: Biostatistician & research-methodology reviewer, meta-analysis and computational-claims focus

---

## Manuscript 1: Main manuscript — De novo binders vs CCL26 & POSTN in EoE

**Brief summary**
This paper combines a nine-cohort transcriptomic meta-analysis with literature/trial-pipeline mining to nominate CCL26 and POSTN as "undrugged" EoE targets, then runs an RFdiffusion → SolubleMPNN → Boltz-2 pipeline to design and computationally validate one lead binder per target. The headline claims are a 59/60 pass rate on interface-confidence thresholds and two leads (ipTM 0.942 and 0.901) that engage their intended epitopes.

**Strengths**
- The target-nomination framing (omics + literature + trial-pipeline convergence) is a genuinely useful triangulation strategy, more rigorous than picking a target by tractability alone.
- Explicit epitope-level, on-target scoring (ipTM × epitope-coverage × pLDDT term) rather than relying on raw ipTM/pLDDT alone is a sound methodological instinct that guards against confidently-folded but off-target designs.
- The limitations section names real, specific caveats (helical/acidic compositional bias, isolated-domain caveat for POSTN, neutralizing-vs-agonist risk) rather than a generic disclaimer.
- Reasonable internal consistency between text, figures, and methods for the design/validation funnel (80 → 1,920 → 1,074 → 60 → 2).

**Major concerns**

1. **Meta-analysis statistical machinery is asserted, not shown.** The Methods state "random-effects meta-analysis... yielding pooled log₂FC, SE, heterogeneity (I²), and BH-adjusted p-values," but no I² values, no forest plots, and no per-cohort effect sizes for CCL26/POSTN appear anywhere in the main text, and Table S1 is described only as "top 60 genes by |z|." A reader cannot assess whether the pooled effect is driven by 1-2 outlier cohorts or is genuinely consistent, nor whether heterogeneity is high enough to make a fixed pooled estimate (+4.56, +5.55 log₂FC) misleading. Given 9 cohorts of unspecified platform (RNA-seq vs microarray, unspecified overlap in patient cohorts across GEO accessions), heterogeneity should be expected to be substantial, and its absence from the reporting is a significant gap for a paper whose entire target-nomination argument rests on this analysis.

2. **The "9/9" vs "8/8" cohort discrepancy for POSTN is unexplained.** CCL26 is reported "up in 9/9 cohorts" while POSTN is "up in 8/8 cohorts" — implying POSTN was only evaluated in, or only available in, 8 of the 9 cohorts, with no stated reason (probe absence, low counts, QC failure, missing annotation?). This matters because silently dropping a cohort for one gene and not another is exactly the kind of undisclosed analytic choice that can bias a "reproducible across cohorts" claim. The paper needs to state explicitly why POSTN's denominator differs and confirm the missing cohort's data were not simply discordant and excluded post hoc.

3. **Multiple-testing correction across 25,654 genes is stated but not audited.** BH-FDR is applied, and the reported adjusted p-values (3×10⁻⁴ for CCL26, 3×10⁻¹⁶ for POSTN) are internally plausible for top hits, but there is no accompanying report of the total number of genes surviving FDR 0.05 (needed to judge whether "among the most strongly up-regulated" is a top-10 or top-500 claim), nor any sensitivity analysis (e.g., Benjamini-Yekutieli under dependence, given that genes within a pathway are correlated and BH assumes independence or a specific positive-dependence structure). This is a minor-to-moderate statistical nicety but relevant given the paper leans heavily on "adjusted p" as a headline number.

4. **Druggability-score ranking is a single-point, unreplicated composite with no sensitivity or uncertainty reported.** POSTN ranks 3rd/20 and CCL26 11th/20 on an "omics-derived druggability shortlist" built from effect size, direction concordance, single-cell specificity, and antibody tractability. No weighting scheme is disclosed, no rank-robustness check (e.g., how sensitive is CCL26's rank of 11 to reasonable reweighting?) is shown, and CCL26's middling rank (11/20) sits awkwardly against the headline claim that it is "the single most discriminating EoE transcript" — the paper should reconcile why its own composite score does not rank CCL26 near the top if the univariate evidence is this strong, rather than leaving the tension unaddressed.

5. **The "zero direct interventional trials" / cherry-picking risk in the 591-abstract, 157-trial corpus.** The denominators (591 abstracts, 157 trials) are stated without inclusion/exclusion criteria, search string, date-range justification for the odd "2015–2026" window (2026 is in the future relative to any real submission date, which needs correcting or explaining — see Minor #1), or a description of how "direct interventional trial" was operationalized (e.g., would a trial of an anti-IL-13 agent that secondarily affects CCL26 transcription count?). Without this, "zero direct trials" is unfalsifiable to a reader and vulnerable to the criticism that the corpus was constructed to make the two chosen targets look like a gap.

6. **Headline "59/60 pass" is a threshold-crossing count on a self-selected pool, not an unbiased success rate, and is reported without uncertainty or acknowledgment of the selection cascade.** The 60 validated complexes were already selected (30/target) from 1,074 filtered sequences by MPNN score and entropy, which were themselves selected from 1,920 raw sequences generated against pre-screened backbones. Reporting "59/60 passed" invites the reader to read this as a general success rate of the design approach, when it is actually a success rate conditional on multiple upstream selection/filtering steps that were themselves tuned partly post hoc (e.g., the entropy ≥2.6 threshold, the composite selection metric). The paper should explicitly state this is not an unbiased estimate of design success and should avoid language ("59/60 passed" in the abstract) that reads as a general validation rate.

**Minor concerns**

1. The literature-corpus date range "2015–2026" includes future dates; either this is a typo (2025?) or an artifact of an LLM-generated search window that needs to be explained/corrected, since as written it looks like fabricated or auto-generated coverage.
2. Cohort accession list (GSE250595, GSE148381, etc.) is given without sample sizes, platform, or tissue-source detail (biopsy vs brushing) — a table of per-cohort N, platform, and design (case/control definition) belongs in the supplement even briefly, since cohort heterogeneity in EoE transcriptomics (pediatric vs adult, active vs remission) is well known to matter.
3. "The single most discriminating EoE transcript" for CCL26 rests on one 2007 citation (PMID 17900656) and is presented as established fact in the abstract; given the paper's own 9-cohort re-analysis, this claim should be re-anchored to the new meta-analysis result (e.g., rank by z-score in Table S1) rather than a single 18-year-old paper.
4. Table S3's "45 targets" for the literature×omics cross-reference is a narrower list than the "591 abstracts" / "157 trials" corpus— it's unclear how the 45 were chosen from the larger corpus; this mapping should be stated.
5. No confidence intervals or error bars accompany the pooled log₂FC point estimates in Fig. 2A/S1 — even a simple CI whisker per gene would substantially help readers assess reproducibility.

**Feasible revisions**

1. Report per-cohort effect sizes/directions and I² (with a forest plot, e.g., an added supplementary figure) for CCL26 and POSTN specifically, and add 1–2 sentences interpreting heterogeneity.
2. Explicitly state and justify why POSTN's cohort count is 8/8 versus CCL26's 9/9 (which cohort was excluded and why).
3. Report the total number of FDR-significant genes at q<0.05 and, ideally, a rank-robustness/sensitivity check for the druggability shortlist (e.g., leave-one-criterion-out re-ranking of CCL26 and POSTN).
4. Provide the search strategy (query strings, inclusion criteria, date-range rationale, and correction of the apparent "2026" typo) for the 591-abstract/157-trial corpus, and define "direct interventional trial" operationally.
5. Reframe the "59/60 passed" statistic in both abstract and Results with an explicit caveat that this reflects a pre-filtered/selected subset, not an unbiased design success rate, and avoid presenting it as a general pipeline performance metric.
6. Add a sentence reconciling CCL26's 11/20 druggability rank with its claim to be the single most discriminating transcript.

**Recommendation:** Major revision — the design pipeline and validation logic are sound and well-documented, but the statistical backbone of the target-nomination claim (heterogeneity reporting, cohort-count discrepancy, corpus construction, and calibration of the "59/60" headline) needs to be shown, not asserted, before the paper's central "reproducible, undrugged, high-priority target" narrative can be fully credited.

---

## Manuscript 2: Supplementary Information

**Brief summary**
The SI provides the figures, five data tables, two rotation-video files, and a parameter table underlying the main text's meta-analysis, design, and validation pipeline, intended to support reproducibility of the full workflow.

**Strengths**
- The parameter table (hotspots, checkpoint, temperatures, filter thresholds, scoring formula) is unusually complete for a hackathon-timeline deliverable and would let a technically fluent reader largely reproduce the computational pipeline.
- Figure S9/S10/S11 together give a genuinely useful multi-angle view of the validation funnel (raw confidence, hotspot count, composite on-target score) rather than just one summary statistic.

**Major concerns**

1. **Table S1 only reports the "top 60 genes by |z|"** — this is insufficient for an outside reader to audit the meta-analysis. Full per-gene, per-cohort summary statistics (or at minimum, the full 25,654-gene table with pooled estimates, SE, I², adjusted p) should be deposited; without it, no independent verification of the CCL26/POSTN numbers, the FDR procedure, or the heterogeneity claims is possible from what's provided.
2. **No supplementary table or figure gives per-cohort raw effect sizes for CCL26/POSTN**, which is the single most important robustness check missing from the entire submission (see Major #1/#2 on the main manuscript). This belongs here if not in the main text.
3. **Table S3's "45 targets"** is never reconciled against the 591-abstract/157-trial corpus size mentioned in the main text and Fig. S2 legend — the sampling/selection logic that produced 45 from the larger corpus should be documented in the supplementary methods table, not left implicit.
4. **Table S2 druggability shortlist scoring weights are not disclosed** even in the SI parameter table; the formula/weights used to combine effect size, concordance, single-cell specificity, and antibody tractability into a single priority score should be given explicitly so the 3rd/20 and 11th/20 ranks are auditable.

**Minor concerns**

1. Figure legends (S1–S12) are generally clear, but several (e.g., S5, S9) would benefit from stating exact n per panel directly in the legend rather than requiring cross-reference to the main text.
2. Table S4 and S5 descriptions are adequate, but neither table's column schema is given (e.g., is "confidence" in Table S4 the raw Boltz-2 aggregate score or the on-target composite? this ambiguity should be resolved with a one-line data dictionary).
3. Movie files are referenced but no still-frame thumbnail or checksum/hash is given for verification of correspondence between the movies and the reported lead design IDs.

**Feasible revisions**

1. Deposit the full 25,654-gene meta-analysis table (not just top 60) alongside per-cohort effect estimates for at least CCL26 and POSTN.
2. Add a supplementary table giving per-cohort log₂FC/SE/direction for CCL26 and POSTN specifically, with the I² statistic, to directly support the reproducibility claim.
3. Reconcile and document how the 45-target Table S3 list was drawn from the 591/157 corpus.
4. Disclose the druggability composite-score formula/weights in the parameter table or a supplementary methods paragraph.
5. Add a one-line data dictionary for Table S4/S5 column definitions.

**Recommendation:** Major revision — the SI is unusually thorough for pipeline parameters but is missing exactly the statistical audit trail (full gene table, per-cohort effect sizes, scoring-formula disclosure) that would let a reader independently verify the paper's central quantitative claims.

---

**Cross-cutting comment:** The design/validation pipeline (RFdiffusion → SolubleMPNN → Boltz-2, with an epitope-aware on-target score) is the strongest and most transparently reported part of this submission, and the limitations section is commendably candid about in-silico status. The weaker link is the statistical foundation used to justify target choice in the first place: heterogeneity, cohort-count discrepancies, corpus-construction details, and score-weighting are all asserted rather than shown, which matters more here than in a typical binder-design paper because the entire "mechanistically-grounded, convergent-evidence" framing is the paper's distinguishing contribution over a target picked by tractability alone. None of the requested fixes require new wet-lab or clinical data — they are reporting, disclosure, and sensitivity-analysis additions well within the computational scope already completed.
