# Mock Peer-Review Report — EAC manuscript, Round 1

**Manuscript assessed:** M1 — "Steering into competitive whitespace: a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory" (v1)

**Review model:** 3 independent expert reviewers + handling-editor synthesis. Reviewers were briefed on the manuscript's framing and scope (in-silico, AI-assisted, public-data; no wet-lab data claimed or required); every requested revision is achievable within that scope.

**Panel:**
- **Reviewer 1** — Domain clinician-scientist (GI oncology, translational target/biomarker)
- **Reviewer 2** — Computational biologist & biostatistician
- **Reviewer 3** — Methodology & research-ethics

**Recommendations:** Reviewer 1 — Major revision · Reviewer 2 — Major revision · Reviewer 3 — Major revision

*Generated 2026-07-11. Reviews produced by independent model instances under distinct expert personas (Reviewer 1 recovered on an alternate model instance after a transient generation refusal on the primary; content unaffected). Advisory only — every specific factual claim was treated as a lead to verify, not ground truth.*

---

## Editor's Decision Letter

# Editor's Decision Letter

**Journal of Computational Drug Discovery & Design**
*Manuscript Reference: JCDD-2024-[assigned]*

---

**Re:** "Steering into competitive whitespace" — a public-data, AI-assisted, human-supervised computational discovery-to-design campaign in esophageal adenocarcinoma

---

## Overview

This manuscript reports a fully computational, AI-assisted campaign that characterizes EAC as a copy-number-driven malignancy, maps competitive whitespace across the clinical trial landscape, applies an explicit druggability triage to nominate two leads (GUCY2C, treatment arm; DKK1, interception arm), and delivers de novo protein binders for each via an RFdiffusion→ProteinMPNN→Boltz-2 pipeline validated in-silico by ipTM/pLDDT metrics. The work is transparently scoped — no wet-lab data are claimed or required — and the open-questions table (O-1 through O-5) reflects unusual intellectual honesty about what remains unresolved. Three expert reviewers (a GI oncology clinician-scientist; a computational biologist/biostatistician; and a methodology/research-ethics reviewer) have evaluated the submission. Their assessments converge on the same high-level verdict: the paper's framing, transparency practices, and competitive-landscape contribution are genuine strengths, but several quantitative pillars that underlie the paper's central claims are currently insufficiently supported by disclosed methods, sensitivity testing, or orthogonal validation. All three recommend **Major Revision**.

---

## Decision: **Major Revision**

The submission is not rejected. The competitive-landscape analysis, the trajectory-spanning strategic framing, and the honest acknowledgment of limitations are valuable contributions appropriate for a high-quality computational preprint venue. However, the paper cannot be accepted in its current form because several of its headline quantitative claims — composite triage scores, the "15/16 pass" rate, and top-design ipTM values — rest on undisclosed scoring rubrics, single-model fold-back, and small single-run samples that a rigorous reader cannot yet evaluate or reproduce. Every concern raised by reviewers is addressable by in-silico analysis, expanded reporting of already-generated data, or manuscript editing. No wet-lab experiments are required or expected.

---

## Essential Revisions

The following concerns were raised independently by two or more reviewers and must be addressed. They are listed in priority order.

---

**1. GUCY2C expression in EAC must be established from public data before the treatment-arm target selection stands.**
*(Reviewers 1 and 3)*

The tumor_selectivity score of 5 (highest ordinal) for GUCY2C is the most consequential unsupported quantitative claim in the paper. GUCY2C is canonically a colorectal antigen; the cited clinical precedent is a CRC vaccine. The manuscript assigns it top selectivity without citing any primary evidence that GUCY2C is expressed at therapeutically actionable levels in EAC tumor cells. This is fully addressable in silico: retrieve GUCY2C RNA expression (RSEM/TPM) from TCGA ESCA adenocarcinoma samples via cBioPortal or the GDC and cross-reference available Human Protein Atlas protein-level data in esophageal adenocarcinoma vs. normal esophageal mucosa. Report the result as a supplementary table or an additional figure panel. If expression is low or absent, the tumor_selectivity score and treatment-arm target discussion must be revised accordingly. This is the single highest-priority revision.

---

**2. Orthogonal structural validation of the two named leads must be completed before publication.**
*(All three reviewers)*

All reviewers independently flag that Boltz-2 ipTM is the sole structural validation metric, that this is a single-model, self-referential confidence score from within the same tool family used for generation, and that the paper itself names AF-Multimer or Chai-1 cross-validation as the very next step (O-1, Discussion). Re-folding the two leads (gucy2c_bb2, dkk1_bb1) — and ideally the top 2–4 designs per target — with at least one independent structure predictor (AlphaFold-Multimer or Chai-1) is an in-scope computational task that should be completed before publication, not deferred. Report the orthogonal ipTM and interface residue agreement or discordance in a supplementary table, and revise the Abstract's headline ipTM claims to reflect multi-model consensus or to flag single-model provenance where consensus is absent.

---

**3. A negative/specificity control for binder designs must be added.**
*(Reviewers 2 and 3)*

Several of the top-scoring designs (e.g., dkk1_bb2, dkk1_bb5; and gucy2c_bb2's terminal poly-alanine run) display the hallmarks of a known RFdiffusion/ProteinMPNN failure mode: generic amphipathic helical bundles that fold with high confidence against almost any target. The manuscript reports only a pairwise binder-vs-binder sequence-similarity check — it never tests whether the top binders fold with high pLDDT in isolation (no target) or score well against a non-cognate decoy. Without this control, an ipTM of 0.915 cannot be distinguished from "generically foldable helix that happens to sit near the target in this one prediction." Add: (a) monomer pLDDT of the binder in isolation, and (b) ipTM of each lead against one or two irrelevant decoy proteins. Flag high-alanine designs as potential developability liabilities in Table S1. This is a fully in-scope computational addition.

---

**4. The composite triage scoring rubric must be fully disclosed and sensitivity-tested.**
*(All three reviewers)*

All three reviewers independently identify this as a reproducibility-critical gap. The manuscript states "composite = association + selectivity + tractability + novelty" but does not disclose: the component weight(s), the ordinal scale anchors for tumor_selectivity and novelty (what does "5" mean for each?), the size of the candidate universe from which the 16 targets were drawn, or whether a single undisclosed weighting choice could reorder GUCY2C/DKK1 relative to CDH17, CEACAM5, or ERBB3, which score close behind. Required additions: (a) a complete scoring rubric in Methods or a new Table S0, including scale definitions, ordinal anchors, and weights; (b) a weight-perturbation or rank-robustness analysis (e.g., Monte Carlo over plausible weight vectors, or at minimum explicit re-ranking under 2–3 alternative weighting schemes) showing that GUCY2C and DKK1 remain in the top tier for their respective lanes under reasonable parameter variation; (c) disclosure of the pre-scoring candidate pool size and source (e.g., "top N targets by Open Targets association score for EAC MONDO identifiers"). A paper whose central contribution is an "honest, structured triage" must disclose the structure of the triage.

---

**5. The "15/16 designs cleared ipTM > 0.5" claim must be reframed and statistically grounded.**
*(Reviewers 2 and 3)*

n=8 backbones per target is a pilot-scale sample; presenting "15 of 16 pass" as a headline result without context implies a generalizable success rate that this sample cannot support. Required: (a) explicitly caveat in the Abstract, Results, and Figure 3 legend that this is a pilot-scale observation (n=8/target) from a single generation run, not a representative success rate; (b) report the mean ± range of ipTM/pLDDT across the 5 Boltz-2 diffusion samples per design rather than only the single best value (the data already exist); (c) contextualize the ipTM > 0.5 threshold against published calibration data for RFdiffusion/ProteinMPNN/Boltz pipelines so readers can assess whether 0.5 is a lenient or stringent bar. If additional backbone generation seeds are computationally feasible, running at least one additional seed per target and reporting seed-stability of the pass rate and top-design identity would substantially strengthen this section.

---

**6. The ClinicalTrials.gov "active-EAC-trial intervention mentions" methodology must be documented and the query must be timestamped.**
*(All three reviewers)*

The whitespace argument — the paper's most distinctive contribution — depends critically on the counts in Figure 1b and Table S3 (41 PD-1/PD-L1, 12 HER2, etc.). These are currently unvalidated annotations with no documented extraction protocol: it is unclear whether classification was by intervention name string match, MeSH term, or manual read; whether combination trials are double-counted across axes; and how the stated 106 active trials relate to the axis-level counts (reviewers note the per-axis counts do not readily sum to 106). Required: (a) add a Methods subsection documenting the exact extraction protocol and any manual curation steps; (b) provide a snapshot/query date for ClinicalTrials.gov, Open Targets, and Drugs@FDA; (c) provide the raw trial-ID list as a supplementary table; (d) reconcile the per-axis counts with the total trial count. Without this, the paper's central whitespace claim is not reproducible.

---

**7. The AI-disclosure statement must be specific and non-placeholder.**
*(Reviewer 3, supported by Reviewer 1)*

The Author Contributions section contains a bracketed placeholder: "[To be completed. The discovery-to-design campaign was executed as a human-supervised, AI-assisted workflow...]." This is inadequate and conflates the named ML tools that are objects of study (RFdiffusion, ProteinMPNN, Boltz-2 — methodology) with general-purpose AI assistance in analysis, scoring, drafting, or figure generation. A reader cannot determine which judgments (e.g., composite weighting in Table S2, the "whitespace verdict" column in Table S3) were made by a human analyst versus proposed by an LLM and reviewed. The revision must include a specific, non-placeholder statement identifying: which AI system(s) were used for which steps; what was deterministic bioinformatics versus AI-assisted synthesis or generation; and what human verification was applied at each step. Given that AI-assistance is framed as a feature of the contribution, the disclosure must be commensurate with that claim.

---

**8. DKK1 clinical-precedent evidence must be segregated by patient population.**
*(Reviewer 1)*

The DKN-01 citation (Klempner 2024) is a Phase 2 study in advanced gastric/GEJ adenocarcinoma — a treatment-arm precedent in late-stage patients who accept significant toxicity. The manuscript credits this as clinical precedent for the DKK1 *interception* arm in a predominantly asymptomatic Barrett's surveillance population (annual EAC progression ~0.3%/year). These are radically different clinical contexts with incommensurable safety expectations and endpoint definitions. Required: add a column or note to Table S2 (and corresponding text) explicitly distinguishing clinical-precedent evidence in advanced disease from evidence in the precursor/Barrett's setting — which is essentially absent beyond biomarker studies — and adjust the narrative accordingly. This is an editorial revision that requires no new data.

---

## Substantive Reviewer Disagreements

There is one area of substantive difference in reviewer emphasis that the authors should acknowledge and adjudicate rather than try to satisfy in contradictory directions:

**Scope of structural validation required.** Reviewer 1 would be satisfied by AF-Multimer or Chai-1 fold-back of the two named leads (gucy2c_bb2, dkk1_bb1). Reviewer 3 additionally requires negative/specificity controls (monomer folding, decoy-target folding) as a prerequisite for treating ipTM values as lead-selection evidence. Reviewer 2 emphasizes seed-stability of the generation run. These are compatible but have different cost/priority ordering. **The authors should complete the orthogonal fold-back (satisfying all three reviewers' minimal ask) and the negative control (Reviewer 3's additional ask), report whether the poly-alanine designs fail the monomer/decoy test, and note seed-stability as an acknowledged limitation if a full reseed is computationally prohibitive within the revision timeline.** This ordering satisfies the consensus while being honest about what was and was not completed.

---

## Path to Acceptance

This manuscript is closer to acceptance than the "Major Revision" designation might suggest. The paper's core intellectual contribution — a reproducible, competition-aware, honesty-gated nomination and binder-design template — is sound and well-executed. The revisions required are specific, bounded, and all achievable by in-silico analysis or editorial editing: retrieve and report GUCY2C EAC expression from TCGA/HPA; complete the AF-Multimer or Chai-1 fold-back of the two leads; add a specificity control for the binder designs; disclose and sensitivity-test the triage scoring rubric; reframe the "15/16 pass" claim with appropriate sample-size caveats and diffusion-sample spread; document the trial-counting methodology with query dates and a raw trial-ID supplement; replace the AI-disclosure placeholder with a specific statement; and segregate DKK1 evidence by patient population. A revised manuscript that addresses these eight points — maintaining the intellectual honesty that all three reviewers praised and extending it to the binder-validation and scoring-provenance sections — would meet the bar for acceptance at this venue.

Authors are invited to submit a point-by-point response letter addressing each Essential Revision in order, with clear indication of where in the revised manuscript each change appears. The revised manuscript will be returned to at least two of the original reviewers.

---

*Handling Editor*
*Journal of Computational Drug Discovery & Design*

---

# Full Reviews (Round 1)

## Reviewer 1

### Reviewer 1: Clinician-scientist, GI oncology and early-phase drug development; experience with esophageal/GEJ malignancies, Barrett's surveillance programs, and biologic/bispecific therapeutic development

---

## Manuscript M1: "Steering into competitive whitespace: a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory"

---

### Brief Summary

This manuscript describes an AI-assisted, human-supervised computational campaign that mines publicly available EAC genomics (TCGA, n=182), maps the competitive clinical landscape (ClinicalTrials.gov, Open Targets), applies an explicit druggability triage to nominate two therapeutic leads — GUCY2C for a T-cell-engager treatment arm and DKK1 for a neutralizing-trap interception arm in Barrett's — and delivers de novo protein binders for each using an RFdiffusion→ProteinMPNN→Boltz-2 pipeline, with 15/16 designs clearing an ipTM >0.5 threshold. The work is transparently framed as computationally validated only, with no experimental data, and explicitly names its two make-or-break unresolved questions (GUCY2C therapeutic window; DKK1 surrogate endpoint). The paper's principal value is as a reproducible, competition-aware nomination and design template, not as a claim of experimental proof-of-concept.

---

### Strengths

- **Intellectual honesty is a genuine feature, not a rhetorical pose.** The abstract, main text, figure legends, and supplement consistently and correctly label ipTM as a confidence proxy, not an affinity, and the open-questions table (O-1 through O-5) explicitly names what is unknown. This is unusual and praiseworthy for an early computational study and should be preserved in revision.

- **The competitive-landscape analysis is the paper's most distinctive contribution.** Table S3 constitutes a genuinely useful, structured articulation of where EAC clinical effort is concentrated. The logic of *deliberately routing away* from PD-1/HER2/VEGF — rather than casually asserting those axes are "crowded" — and mapping the resulting whitespace is clear, reproducible (the sources are named), and analytically coherent.

- **The druggability triage (Figure 2, Table S2) is methodologically defensible.** The explicit rejection of GPX7 as the highest raw-composite candidate on the grounds of target class (lost tumor suppressor, not antibody-addressable) demonstrates that the composite score is not just narrative decoration; it is tested against a meaningful filter. This is good scientific discipline.

- **Trajectory-spanning strategy is clinically cogent.** The framing of Barrett's→dysplasia→carcinoma as presenting two distinct intervention opportunities — molecular interception at the precursor stage and directed killing at the invasive stage — is biologically sound and reflects actual clinical unmet need, particularly the absence of any approved molecular agent for Barrett's high-risk surveillance patients.

- **Binder design metrics are reported comprehensively.** Table S1 reports all 16 designs with full sequences, lengths, mpnn_score, ipTM, complex_pLDDT, a combined confidence, and pass/fail, which allows the reader to assess the design landscape rather than only the cherry-picked leads.

- **The DKN-01 citation (Klempner 2024) for DKK1 precedent is appropriate and current**, and grounding the selection in existing clinical experience with a DKK1-neutralizing agent in esophagogastric cancer materially strengthens the biological plausibility argument.

- **AI assistance is disclosed as a feature of the method, not obscured.** This is correct practice.

---

### Major Concerns

**1. The GUCY2C therapeutic-window argument is asserted, not evaluated, and the assertion is clinically fragile.**

The entire treatment-arm thesis rests on the claim that GUCY2C's "apical, luminal-facing restriction in normal gut" provides a sufficient safety margin for a T-cell engager. This is stated in a single sentence in Results, one sentence in Discussion, and labeled as open question O-2. But the clinical risk is asymmetric and deserves far more rigorous treatment than it receives here. GI-directed T-cell engager toxicity is not a hypothetical: the documented experience with EpCAM-directed catumaxomab (peritoneal), the well-characterized cytokine-driven GI toxicity of CD3-bispecifics generally, and the specific issue that *murine* GUCY2C may not be cross-reactive with the human protein (which the authors acknowledge) together mean that preclinical safety data cannot be generated in the most common model system. This is not merely "one arm of a bispecific whose CD3 arm is not done" (Discussion); it is a target-class question that bears on whether a T-cell engager is the *correct format* for a GI-luminal antigen, even assuming excellent tumor selectivity. The manuscript should more explicitly discuss why a T-cell engager was selected over a safer-window format (ADC, CAR, vaccine, as listed in Table S3) for an antigen with acknowledged normal enterocyte expression. The current framing implies the window exists; it has not been established by any data cited here.

**2. The GUCY2C EAC expression evidence is absent and is essential to the target selection argument.**

The genomic characterization in the paper establishes GUCY2C is absent from the most frequent EAC amplicons and is not listed among the top altered genes in Figure 1a. Table S2 assigns it tumor_selectivity = 5 (the highest ordinal) with a note of "luminal-restricted." But GUCY2C is canonically a colorectal antigen: the cited clinical precedent (Ad5-GUCY2C-PADRE, Table S3) is a CRC vaccine. The manuscript does not cite primary evidence that GUCY2C is expressed at therapeutically actionable levels in EAC tumor cells, nor that its expression is maintained or upregulated across the Barrett's→EAC progression it is meant to target. A tumor_selectivity score of 5 without an EAC-specific expression citation is the single largest unsupported quantitative claim in the paper. This is achievable in silico: the Human Protein Atlas (HPA) and the TCGA RNA-seq data available via cBioPortal or the GDC contain esophageal/EAC expression data for GUCY2C. The authors should retrieve and report GUCY2C RNA expression (TPM or RSEM) in EAC tumor vs. normal esophageal tissue from TCGA and cross-reference HPA protein-level data. If GUCY2C is lowly expressed or absent in EAC, the entire treatment-arm target selection requires revision.

**3. The DKK1 interception arm conflates two distinct patient populations in a clinically consequential way.**

Table S2 labels DKK1's trajectory as "advanced + progression," and the competitive-landscape table (S3) places it under "Advanced." The main text and Figure 1c assign it to the interception arm across "Barrett's metaplasia and dysplasia." But the Klempner 2024 DKN-01 citation is a Phase 2 study in *advanced gastric or GEJ adenocarcinoma*, i.e., a treatment-arm precedent, not an interception-arm precedent. These are radically different clinical settings with different safety expectations, endpoint definitions, patient populations, regulatory pathways, and commercial value propositions. A drug that is tolerable in advanced cancer patients (who accept significant toxicity) may be unacceptable in the predominantly asymptomatic, mostly-non-progressing Barrett's surveillance population, where the annual progression rate to EAC is ~0.3% per year. The manuscript's Discussion and Open Questions (O-4) gesture at this — "very high safety bar" — but the conflation persists in the data tables and in the target-selection scoring, where DKK1's clinical-precedent credit appears to have been awarded for the wrong population. The authors should segregate the evidence base for DKK1 in advanced disease vs. DKK1 in the precursor setting and be explicit that the interception indication has essentially no clinical-precedent support beyond biomarker studies.

**4. The composite scoring methodology is opaque and not validated for sensitivity.**

Table S2 presents composite scores (e.g., GPX7 = 13.86, GUCY2C = 13.0, DKK1 = 11.37) that determine candidate rank order, but the methods section does not state the individual component weights, the scoring rubric for each ordinal, or whether the system is linear-additive, weighted, or otherwise. "A composite of disease association, tumor selectivity (ordinal), antibody tractability, and novelty" is the full description. This is insufficient for reproducibility. A reader cannot determine whether the 0.86-point difference between GPX7 and GUCY2C is meaningful, whether CDH17 (12.38, which is higher than GUCY2C 13.0 when corrected — wait: the table shows CDH17 = 12.38 < GUCY2C = 13.0, which is correct) is close enough to GUCY2C that the choice is effectively arbitrary, or whether DKK1 at 11.37 would survive any reasonable parameter perturbation. The paper cannot claim a "structured" triage without disclosing the structure. A sensitivity analysis — showing that the top-2 selections are stable under ±10-20% weight perturbation across components — is achievable computationally and should be added.

**5. The binder design pipeline lacks any orthogonal structural confidence check.**

The paper relies entirely on Boltz-2 ipTM as the validation metric. The Discussion correctly notes this as a limitation and names ESM, AlphaFold-Multimer, and Chai-1 as alternatives, and O-1 lists this as a near-term computational step. However, for a manuscript whose central technical claim is that these are "structurally credible, in-silico-validated leads," relying on a single model's self-reported confidence score without any orthogonal fold-back is a significant methodological gap. Boltz-2's ipTM is a diffusion-model confidence output that may be correlated across samples from the same run; it is not the same as independent model consensus. AF-Multimer and Chai-1 fold-back of the two leads (gucy2c_bb2 and dkk1_bb1) would require only additional compute and would substantially strengthen the "structurally validated" claim. Given that O-1 is already listed as near-term computational work, this should be completed before publication, not deferred.

---

### Minor Concerns

**1. The TCGA cohort size (n=182) and its limitations need a brief note.** The alteration frequencies cited (87% TP53, 39% CDKN2A, 35% CCND1) are well-supported by the TCGA 2017 ESCC/EAC paper, but the cohort is sequenced tissue, not necessarily representative of clinical EAC in terms of stage at presentation, treatment history, or geographic population. A single sentence acknowledging this is appropriate.

**2. The "active EAC trial intervention mentions" metric (Figure 1b, x-axis) is defined only cursorily.** Is this a count of unique trials, of mentions across all arms, or something else? The number 41 for PD-1/PD-L1 is plausible but should be operationally defined so that another researcher could reproduce it. The methods say "the count of active interventional trials referencing each target/axis" — does "referencing" mean the primary intervention, any arm, or any drug known to act on that target? This matters for the whitespace argument.

**3. The interface residue numbering in S4 is ambiguous.** The GUCY2C interface residues are given as [130,132,146,192,325,326,327,330,347,348,349,378,379] with the parenthetical "complex-file numbering," and separately the hotspot input residues are [239,240,267,422,424,425] in "isolated-domain numbering." The relationship between these two coordinate systems is not explained, and a reader cannot determine whether the binder contacts the intended hotspot patch. The authors should either provide a residue-number mapping or confirm explicitly that the interface residues correspond to the hotspot region in canonical UniProt numbering.

**4. The sequence of dkk1_bb2** ("MTREELIAAAGAAGLAYGAALTAALAAAAAAAGADTAAVLALGAAGAAAAAALAAAAAAAAA") has an unusually high alanine content (>50% Ala), and dkk1_bb5 and dkk1_bb3 show similar patterns. While the paper notes diverse solutions, a poly-alanine bias in a subset of designs may indicate aggregation-prone or low-complexity sequences that would not be developable. This should be noted as a liability flag, and the ESM developability screen (O-1) should be flagged as specifically needed for these designs.

**5. The mpnn_score is reported but not defined in the manuscript or supplement.** A brief definition (e.g., negative log-likelihood from ProteinMPNN; lower is better) should be added to Table S1's header or footnote.

**6. The reference list is abbreviated and inconsistently formatted.** The Klempner 2024 citation is cited as "DKN-01 in Combination With Tislelizumab and Chemotherapy" — this is a correct citation for the DisTinGuish Phase 2, but should note that this is a gastric/GEJ study, not an EAC-specific study, since the clinical-precedent argument depends on the distinction.

---

### Feasible Revisions

All of the following are achievable in silico or by manuscript editing; none requires wet-lab experiments.

1. **Retrieve and report GUCY2C expression in EAC from public RNA-seq data.** Query TCGA ESCA adenocarcinoma samples and/or Human Protein Atlas for GUCY2C RNA (RSEM/TPM) and available protein-level data in esophageal adenocarcinoma vs. normal esophageal mucosa. Report this in a supplementary table or an additional panel to Figure 1 or Figure 2. If expression is low or absent, revise the tumor_selectivity score and discussion accordingly. This revision directly addresses Major Concern 2 and is the highest-priority revision.

2. **Run AF-Multimer or Chai-1 fold-back on the two leads (gucy2c_bb2, dkk1_bb1) and report the orthogonal ipTM/interface score.** Add results as a supplementary table and revise the "structurally validated" claim to reflect multi-model consensus or discordance. This addresses Major Concern 5 and converts O-1 from a deferred item to a completed one.

3. **Disclose and sensitivity-test the composite scoring rubric.** Add to Methods (or a new Table S0) the full scoring schema: component definitions, ordinal scales, and weights. Run a ±15% weight perturbation analysis and report whether GUCY2C and DKK1 remain in the top tier for their respective lanes. This addresses Major Concern 4.

4. **Segregate DKK1 evidence by population in Table S2 and the main text.** Add a column or note distinguishing clinical-precedent evidence in advanced disease (DKN-01, Klempner 2024) from evidence in the precursor/Barrett's setting (which is essentially absent). Revise the trajectory label from "advanced + progression" to something like "Barrett's (interception; no clinical-precedent in precursor setting)" and adjust the biologic_druggability note to reflect the safety-bar asymmetry. This addresses Major Concern 3.

5. **Expand the GUCY2C therapeutic-window discussion and provide format-choice justification.** Add a paragraph (or supplementary note) that (a) reviews published safety data from GI-directed T-cell engagers and ADCs in GI cancers to contextualize the luminal-restriction claim; (b) explicitly argues why a T-cell-engager format was preferred over ADC or CAR-T for this antigen despite known enterocyte expression; and (c) identifies what HPA/GTEx expression data already show about GUCY2C in human gut segments relevant to the safety argument. This partially addresses Major Concern 1 within scope.

6. **Clarify the "active EAC trial intervention mentions" operational definition** in the Methods section with enough precision that the count is reproducible, and add a date-of-query for the ClinicalTrials.gov and Open Targets downloads. This addresses Minor Concern 2 and is essential for reproducibility.

7. **Add a residue-numbering mapping table** (UniProt canonical → isolated-domain → complex-file) for both GUCY2C and DKK1 in S4. Confirm or correct whether the predicted interface residues correspond to the intended hotspot patches. This addresses Minor Concern 3.

8. **Flag the high-alanine-content designs (dkk1_bb2, dkk1_bb5, dkk1_bb3) as potential developability liabilities** in Table S1 and in S5/S6. State that ESM screen (O-1) is specifically prioritized for these sequences. This addresses Minor Concern 4.

9. **Define mpnn_score** in Table S1's header and add a brief note on its interpretation. This addresses Minor Concern 5.

---

### Recommendation

**Major Revision**

The paper's competitive-landscape analysis, trajectory-spanning strategic framing, and honest disclosure practices make it a genuinely useful contribution to the computational drug discovery literature as a preprint. However, two issues require substantive revision before the work's central claims are defensible: the absence of any EAC-specific expression evidence for GUCY2C (the treatment-arm target) leaves the most important tumor-selectivity score unsupported, and the single-model structural validation without orthogonal fold-back is insufficient for the "structurally credible leads" claim when O-1 is already identified as near-term and feasible. Both are achievable in silico. The composite-scoring opacity and the DKK1 population conflation are additionally necessary to fix for intellectual honesty. None of the required revisions involve wet-lab experiments.

---

## Cross-cutting Comment

This is a single-manuscript submission, so cross-cutting remarks bear on the campaign as a whole. The work's strongest feature — radical transparency about what is and is not known — is also the lens through which its gaps are most visible: having named O-2 (GUCY2C therapeutic window) and O-4 (DKK1 surrogate endpoint) as make-or-break questions, the paper is obligated to present the best available computational and public-data case for why those questions are *worth asking*, and that case is currently stronger for DKK1 than for GUCY2C. The treatment arm's target nomination would be substantially more credible if EAC-specific expression data were integrated; their absence is the one place where the campaign's stated honesty-gating principle is not fully applied. The precedent of designing for both arms in a single cycle is methodologically sound, and the binder quality metrics are reported with appropriate transparency — the ask is simply that the same rigor be applied to the upstream target-selection evidence as is applied to the downstream design confidence.

---

## Reviewer 2

### Reviewer 2: Computational biologist / biostatistician with a focus on genomics reproducibility and structure-prediction benchmarking

**Brief summary**

The manuscript is a computational discovery-to-design campaign that (1) characterizes EAC genomics from a single TCGA cohort, (2) builds a competitive-landscape map from ClinicalTrials.gov and Open Targets, (3) applies a scoring/triage scheme to nominate GUCY2C (treatment arm) and DKK1 (interception arm), and (4) generates and folds-back 16 de novo binders (RFdiffusion → ProteinMPNN → Boltz-2), reporting ipTM/pLDDT as validation metrics. The authors are explicit that everything is in-silico and flag two "decisive open questions" as unresolved. As a statistics/reproducibility reviewer, my concerns center on the quantitative scoring systems (Fig. 1b, Fig. 2, Tables S2–S3) being under-specified and unvalidated, the single-cohort/single-model nature of every quantitative claim, and the self-referential circularity of using Boltz-2 ipTM as the sole validation for designs whose hotspots were also chosen heuristically from the same AlphaFold structures.

**Strengths**

- Transparent, fully-sourced quantitative claims (each number traced to a named table/file) — genuinely good reproducibility practice for a preprint of this kind.
- Explicit, auditable rejection of the top-scoring raw candidate (GPX7) on mechanistic grounds — a real demonstration that the triage isn't just cherry-picking a foregone conclusion.
- Full disclosure of the complete 16-design table (not just winners), including the one design that failed threshold (gucy2c_bb6), which allows an outside reader to recompute pass rates and inspect the distribution rather than take summary claims on faith.
- Correct, repeated caveating of ipTM as a "docking confidence proxy," not affinity — this is stated in figure legends, results, and discussion, which is unusual honesty for this genre of paper.
- Cross-reactivity/similarity check across the design set is a sensible, cheap robustness step and its negative result is reported plainly rather than oversold.

**Major concerns**

1. **Composite scores in Table S2/Figure 2 have no disclosed weighting, scale, or uncertainty.** "Composite = association + selectivity + tractability + novelty" is given as a formula but none of the four component scales (1–5? what is "novelty=5" anchored to?) are defined, and there is no sensitivity analysis showing whether GUCY2C/DKK1's selection is robust to reasonable re-weighting of these four components. Given that DKK1 (composite 11.37) ranks *below* four rejected/down-weighted candidates (13.86, 12.47, 12.38, 12.29) purely on the "druggability filter" override, the paper needs to show this ranking is not an artifact of arbitrary weights. A simple weight-perturbation or rank-robustness check (e.g., resample weights, recompute rank of DKK1/GUCY2C across a plausible weight simplex) is entirely in scope and would substantially strengthen the "honest triage" claim.

2. **Single-cohort genomic claims presented without any statistical framing.** The 87%/39%/35% alteration frequencies (Fig. 1a) come from n=182 TCGA samples with no confidence intervals, no comparison to an independent EAC cohort (e.g., ICGC OCCAMS/OAC), and no discussion of the fact that TCGA ESCA "adenocarcinoma subtype" filtering itself involves a sample-selection choice that should be stated (how many ESCA samples were excluded as non-adenocarcinoma, and on what histology annotation?). A single-cohort percentage stated as a fact ("TP53 altered in 87% of tumors") should carry at least a binomial CI (which is trivial to compute: 87% of 182 ⇒ ~[81–91%] at 95%) and ideally a cross-check against the published TCGA ESCA paper's own reported frequency (ref. 3) to confirm cBioPortal's current pipeline reproduces the original TCGA numbers.

3. **"Active-EAC-trial intervention mentions" (Fig. 1b, Table S3) is an unvalidated, seemingly manual annotation with no counting methodology.** How was a trial classified as "mentioning" a target/axis — by intervention name string match, by MeSH term, by manual read of the trial description? Is a single trial counted once per target it might touch (pseudoreplication risk if combination trials touch both e.g. PD-1 and HER2)? Without a documented extraction protocol and inter-rater or at least a spot-check reproducibility statement, these counts (41 PD-1, 12 HER2, 8 VEGF, 1 each for the "whitespace" targets) cannot be taken as reproducible measurements — they read as narrative-supporting tallies. This directly affects the paper's central "whitespace" argument.

4. **Boltz-2-only structural validation is a single-model, single-tool result with no orthogonal consensus — and the design and validation loops are not independent.** The hotspot residues were chosen using the *same* AlphaFold model (pLDDT-based heuristic) that Boltz-2 also uses internally for confidence calibration; folding the binder back onto a structure selected because it was already high-confidence in that same modeling family is a mild form of circularity. The manuscript acknowledges the "no orthogonal consensus" limitation in the Discussion, which is appropriate, but Results and Abstract nonetheless state ipTM figures with three-decimal precision (0.915, 0.858) as if they were precise, reproducible measurements, with no report of run-to-run variance (Boltz-2 stochastic diffusion sampling — 5 samples were generated per the Methods; only the best is reported, not the spread). Reporting the mean/SD or range across the 5 diffusion samples per design, rather than a single best value, is a straightforward in-silico robustness addition and should be added to Table S1.

5. **Multiple-comparisons / selection-bias framing is missing for the "15 of 16 pass" headline.** With 16 designs and a threshold of ipTM>0.5, presenting "15/16 pass, one clear failure" without discussing the fact that the pass threshold (0.5) is a soft, community-convention cutoff (not calibrated against any known-binder benchmark set) risks the reader over-interpreting the pass rate as evidence of a well-calibrated pipeline. A brief calibration statement — e.g., citing published ipTM/success-rate correlations from the RFdiffusion/ProteinMPNN benchmark papers, or noting where these designs would fall on that curve — would let readers judge whether 0.5 is a stringent or lenient bar here.

6. **No independent replicate seeds for RFdiffusion/ProteinMPNN generation.** Eight backbones per target were generated once; there is no report of whether repeating the generation with a different random seed set would give a similar top-ipTM winner or count of passers. Given the very large spread in ipTM within DKK1 designs alone (0.524–0.858) and within GUCY2C (0.401–0.915), the "top design" statistic is an extreme-value statistic from a small, single sample of backbones — its stability under reseeding is an easy, in-scope computational check (rerun generation with 1–2 additional seeds, report whether the top-line ipTM leaders are stable) that would materially strengthen confidence in "gucy2c_bb2/dkk1_bb1 are the leads" rather than "the best of one particular sample of 8."

**Minor concerns**

1. Figure 1a legend and Results text should state whether alteration percentages are of the full n=182 or of some subset with available CNA+mutation data (cBioPortal sometimes reports slightly different denominators by data type); please confirm and state the exact denominator used for each gene.
2. "106 active trials" and "351 interventional trials" — please state the ClinicalTrials.gov query date/snapshot, since this is a live, constantly-updated database and reproducibility requires a timestamp.
3. Table S2's "tumor_selectivity" and "novelty" columns are integers 1–5 with no rubric given anywhere (main text, supplement, or methods) — a short scoring rubric (even qualitative, e.g., "5 = expression restricted to a single normal tissue with no oncologic overlap") should be added as a supplementary table.
4. The claim "a sequence-similarity check... detected no high-similarity binder pairs" (Results, S5) should report the actual metric and threshold used (% identity? alignment method? pairwise all-vs-all?) rather than only the qualitative conclusion.
5. Figure 3/S1 conflates "mpnn_score" units without defining scale/sign convention (lower/higher better?) — should be defined once in Methods.
6. The Discussion's claim that this is a "reproducible template" would be more convincing if the manuscript stated explicitly whether the triage scoring code/weights and the trial-counting scripts are included in the supplementary artifacts, or only the output CSVs — reproducibility of a *score* requires the scoring code, not just the resulting numbers.

**Feasible revisions**

1. Add binomial confidence intervals to all cohort-derived percentages in Fig. 1a/Results and explicitly cross-check the 87%/39%/35% figures against the original TCGA ESCA publication's reported frequencies (ref. 3) to confirm consistency between the two data vintages/pipelines.
2. Add a weight-sensitivity/rank-robustness analysis for the Table S2 composite score (e.g., Monte Carlo over plausible weight vectors, or at minimum report rank of GUCY2C/DKK1 under 2–3 alternative reasonable weightings) and report it as a supplementary figure.
3. Document the exact methodology for "active-EAC-trial intervention mentions" (search terms, database snapshot date, whether trials are double-counted across axes, and any manual curation step), and add this as a subsection in Methods with the raw trial-ID list as a supplementary table.
4. For each of the 16 designs, report ipTM/complex-pLDDT as mean ± range across the 5 Boltz-2 diffusion samples (not only the best), in an expanded Table S1, and note explicitly if the reported "lead" values are the max of 5 draws.
5. Rerun RFdiffusion/ProteinMPNN backbone generation with at least one additional random seed for each target and report whether the top-ipTM lead and overall pass rate are stable across seeds (or, at minimum, explicitly flag in Discussion/Limitations that current leads derive from a single generation run and could shift under reseeding).
6. State the exact sequence-similarity metric/threshold used for the S5 cross-reactivity screen, and add the underlying pairwise similarity matrix as a supplementary table.
7. Add a rubric (even a short qualitative one) defining the 1–5 scales for tumor_selectivity and novelty in Table S2, and clarify in Methods whether these were assigned by a single annotator or reconciled across multiple raters.
8. Add a sentence in Methods/Discussion contextualizing the ipTM>0.5 threshold against published calibration data for RFdiffusion/ProteinMPNN/Boltz-type pipelines, so readers can judge stringency rather than take 15/16 as self-evidently a strong result.

**Recommendation:** Major revision — the computational scope and honesty framing are appropriate for a preprint of this kind and the core claims are plausible, but several quantitative pillars (trial-count methodology, composite-score weighting, single-seed/single-model structural validation) currently lack the statistical scaffolding needed to support the paper's central "differentiated, defensible whitespace" argument, and all of the requested fixes are achievable by additional in-silico analysis or expanded reporting of already-generated data, not new experiments.

---

**Cross-cutting comment:** As this is a single-manuscript submission, this section is not applicable beyond noting that the paper's greatest asset — its unusually complete disclosure of raw tables (all 16 designs, full triage table, trial counts) — is also what exposes its weakest seams (unweighted composite scores, single-run structural metrics, undocumented trial-counting) to a statistically-minded reader; tightening the quantitative provenance of those specific artifacts, all achievable without new wet-lab work, would bring the rigor of the analysis fully in line with the honesty of its framing.

---

## Reviewer 3

### Reviewer 3: Methodology & research-ethics reviewer (computational oncology target validation, generative protein design, AI-disclosure norms)

**Manuscript: "Steering into competitive whitespace" — EAC dual-arm campaign (GUCY2C / DKK1)**

**Brief summary**

The manuscript reports a fully computational campaign that (i) profiles TCGA EAC genomics, (ii) maps the ClinicalTrials.gov/Open Targets competitive landscape, (iii) applies a stated "honest druggability filter" to nominate GUCY2C (treatment arm, T-cell engager) and DKK1 (interception arm, neutralizing trap), and (iv) runs an RFdiffusion→ProteinMPNN→Boltz-2 pipeline to generate and self-validate 16 de novo binders (8 per target), reporting top ipTM of 0.915 (GUCY2C) and 0.858 (DKK1). The paper is explicit and largely consistent about labeling every claim as in-silico, and it flags two "decisive open questions" (therapeutic window, interception surrogate endpoint) as unresolved. The core methodological question for this review is whether the exhibited numbers are a fair, non-cherry-picked, and adequately cross-validated representation of what the campaign actually produced.

**Strengths**

- Rejecting the top raw-composite candidate (*GPX7*, 13.86) because it is a lost tumor suppressor rather than an antibody-addressable target (Table S2, Fig. 2) is a genuinely useful demonstration of the "honesty gate" the paper claims to apply — this is exactly the kind of self-critical step that guards against narrative-driven target selection, and it is good that it is shown rather than asserted.
- Every headline number is tied to a named source table (`design_leads.csv`, `target_nomination.csv`, `competitive_landscape.csv`), which is the right posture for reproducibility even before the underlying data can be independently re-run.
- Explicit ipTM/pLDDT caveats ("a docking-confidence proxy...not a measurement of affinity") appear in Results, Discussion, and figure legends rather than once in a footnote — repetition of the caveat where the number is used is good practice, not boilerplate.
- The two open "make-or-break" questions (O-2 GUCY2C therapeutic window, O-4 DKK1 surrogate endpoint) are framed as could-halt-the-arm risks rather than minimized — this is a stronger form of honesty than most target-nomination papers offer.

**Major concerns**

1. **No orthogonal structural validation despite it being clearly achievable in-silico.** All 16 designs are scored by Boltz-2, the same tool family used implicitly downstream of the RFdiffusion/ProteinMPNN generation loop. The paper itself names an AlphaFold-Multimer or Chai-1 cross-check as the very next step (O-1, Discussion) but does not perform it. Given this is squarely in-scope computational work with no wet-lab dependency, presenting single-model ipTM values (0.915, 0.858) as headline results in the Abstract without at least one orthogonal fold-back is a real gap — self-consistency between a design pipeline and its own confidence metric is a known confound in de novo binder work, and the paper's own limitations paragraph concedes this without acting on it. I would not accept the leads as reported without at least the top 2–4 designs per target re-folded with an independent structure predictor.

2. **No negative/specificity controls anywhere in the validation.** Several of the top-scoring sequences are strikingly generic idealized helical bundles with very high alanine content and heptad-repeat character — e.g., dkk1_bb2 ("MTREELIAAAGAAGLAYGAALTAALAAAAAAAGADTAAVLALGAAGAAAAAALAAAAAAAAA", Table S1) and dkk1_bb5/gucy2c_bb2 (which ends "...AAEAAAAAAAA"). This pattern is a recognized failure mode in RFdiffusion/ProteinMPNN binder campaigns, where the model converges on an amphipathic-helix solution that folds with high confidence against almost any target rather than one shape-complementary to the specific epitope. The manuscript reports only a pairwise binder-vs-binder sequence-similarity check (S5) as a "diversity" control — it never asks whether the same binder scores well against a scrambled/irrelevant target, or whether the binder alone (no target) already folds into a high-pLDDT idealized bundle independent of the interface. Without that negative control, "ipTM 0.915" cannot be distinguished from "generically foldable helix that happens to sit near the target in this one prediction." This is fully achievable computationally and should be added before the ipTM numbers are treated as lead-selection evidence.

3. **The composite triage score is not reproducible or sensitivity-tested.** Table S2 gives only the final composite (e.g., GUCY2C 13.0, DKK1 11.37) and four ordinal sub-scores, with no disclosed weighting scheme, scale bounds, or the size of the candidate universe from which these 16 targets were drawn. Without knowing whether DKK1 (11.37) was the 2nd-best interception candidate out of 20 evaluated or out of 200, the reader cannot assess whether the list was pre-narrowed to fit the whitespace narrative before scoring began. A single, undisclosed weighting choice could plausibly reorder GUCY2C/DKK1 relative to CDH17, CEACAM5, or ERBB3, all of which score close behind in Table S2. This is a denominator problem central to this reviewer's brief and needs to be fixed by disclosure, not by re-scoring.

4. **Design-set sample size is too small to support the "15/16 pass" framing as a rate.** Eight backbones per target, one sequence carried forward per backbone, yields n=8 per arm. Reporting "15 of 16 designs cleared ipTM > 0.5" (Results, Fig. 3) reads as a success-rate statistic but is really two n=8 pilot batches; the true expected hit rate at any meaningful design scale (hundreds–thousands of backbones is standard in published RFdiffusion binder campaigns) is unknown from this data. The manuscript should either explicitly caveat that 15/16 is a pilot-scale observation, not a representative success rate, or scale the design sweep (an in-silico step, fully within scope) before quoting a pass fraction.

5. **AI-disclosure is a placeholder, not a statement.** The framing claims "AI assistance is disclosed and is part of the contribution," but the manuscript's only concrete disclosure is the bracketed Author Contributions line ("[To be completed. The discovery-to-design campaign was executed as a human-supervised, AI-assisted workflow...]"). This conflates two very different things: (a) the named ML tools that are the object of study (RFdiffusion, ProteinMPNN, Boltz-2 — these are methodology, not "AI assistance" in the disclosure sense) and (b) whatever general-purpose AI system(s) performed literature synthesis, scoring, figure generation, or drafting. As written, a reader cannot tell which specific judgments (e.g., the composite weighting in Table S2, the "whitespace verdict" column in Table S3) were made by a human analyst versus generated by an LLM and merely reviewed. This needs a specific, non-placeholder statement: what tool(s), what steps, what was human-verified versus AI-proposed.

**Minor concerns**

1. Figure 1b/Table S3 crowding counts (41, 12, 8, 3, 2, 1–2, 1×6) sum to roughly 70–75 of the stated 106 active trials; the remainder (likely non-targeted chemotherapy/other combination arms) is never explained, leaving the reader unsure whether "crowding" undercounts or the 106 total includes many non-axis-specific trials. A one-line reconciliation would remove ambiguity.
2. No access/snapshot date is given for the ClinicalTrials.gov, Open Targets, or Drugs@FDA queries (Methods, "Target-association and competitive-landscape analysis"). These sources change continuously; a query date is needed for reproducibility.
3. The TCGA cohort (n=182, `esca_tcga_pan_can_atlas_2018`) is used without any note on its demographic composition, and EAC/Barrett's epidemiology is known to skew heavily toward specific ancestries/sexes; a brief statement on cohort representativeness (or its absence) would strengthen the equity framing implicit in an "interception in a large, identifiable population" claim (Introduction).
4. Figure 3's legend ("Reproduced from the design campaign; source `design_results.png`...") reads as unedited pipeline boilerplate rather than a finished caption; it should specify what is plotted (e.g., ipTM vs. pLDDT scatter, colored by target) directly rather than by reference to an external file name.
5. Several bracketed placeholders remain (author list, affiliations, correspondence, competing interests) — expected for a preprint draft, but should be flagged as incomplete rather than left silent.

**Feasible revisions**

1. Re-fold the top 2–4 designs per target (at minimum the two named leads, gucy2c_bb2 and dkk1_bb1) with an independent structure-prediction method (AlphaFold-Multimer or Chai-1, both public) and report agreement/disagreement in ipTM and interface residues; this converts O-1 from a stated future step into an actual result and is entirely in-scope.
2. Add a specificity/negative-control analysis: fold each top binder alone (no target) and, ideally, against 1–2 non-cognate decoy targets, and report whether pLDDT/ipTM discriminate the cognate complex from these controls. Flag any binder whose "confidence" appears target-independent (the high-alanine sequences noted above are the most likely candidates to fail this test).
3. Disclose the size and source of the pre-scoring candidate universe for Table S2 (e.g., "top N targets by Open Targets association score for MONDO_0005028/MONDO_0013662"), the exact composite weighting formula, and ideally a brief sensitivity check (e.g., re-rank under two or three alternative weightings) to show GUCY2C/DKK1 selection is not an artifact of one weighting choice.
4. Reword the "15 of 16 designs cleared ipTM > 0.5" claim (Abstract, Results, Fig. 3 legend) to explicitly state the pilot sample size (n=8 backbones/target) and avoid implying a generalizable success rate; alternatively, expand the backbone sweep computationally and report the pass rate at larger n.
5. Replace the placeholder AI-disclosure sentence in Author Contributions with a specific statement identifying which AI system(s) performed which steps (literature/data synthesis, scoring, code generation, figure/text drafting) and which steps were deterministic bioinformatics tools (RFdiffusion/ProteinMPNN/Boltz-2), plus what human verification was applied to each.
6. Add ClinicalTrials.gov/Open Targets/Drugs@FDA query/snapshot dates to Methods, and reconcile the Fig.1b/Table S3 trial-count denominators.
7. Add one sentence noting the demographic composition (or its absence) of the TCGA cohort used, given known epidemiological skew in EAC/Barrett's populations, to properly frame the "large, identifiable high-risk population" claim used to motivate the interception arm.

**Recommendation: Major revision.** The campaign is transparently in-silico, honestly caveated in its prose, and methodologically interesting as a template, but several core numeric claims (design pass-rate, ipTM leads, composite triage ranking) currently rest on single-model, small-n, and undisclosed-weighting evidence that a rigorous reviewer cannot yet take at face value; all requested fixes are computational or editorial and achievable before resubmission.

---

**Cross-cutting comment:** As a single-manuscript submission, the main cross-cutting risk is internal: the paper's professed "honesty gating" is real and demonstrable in the target-triage step (rejecting GPX7) but is not yet extended with equal rigor to the binder-validation step, where self-consistency (Boltz-2 scoring its own design lineage) and small sample size are treated as acknowledged future work rather than fixed now, even though both fixes are within scope. The gap between the rhetorical framing ("explicit honest druggability triage," "reproducible, honesty-gated template") and the actual evidentiary support for the headline ipTM numbers and the 16-candidate shortlist's denominator is the main thing standing between this being a genuinely rigorous methods paper and a well-written but partially under-verified one.
