# Peer-Review Report — EAC Manuscript (Two Rounds)

*A synthetic peer-review panel (three independent expert personas — a GI-oncology clinician-scientist, a computational biologist/biostatistician, and a methodology/research-ethics reviewer — plus a handling-editor synthesis) reviewed the manuscript across two rounds. Reviews are advisory: reviewers are LLM instances, so every specific factual claim was treated as a lead to verify, not ground truth. This report contains, in order: the round-1 review + the authors' response, the round-2 review + the authors' response, and a revision-tracking summary.*

## Outcome at a glance

| | Round 1 | Round 2 |
|---|---|---|
| Reviewer 1 (clinician-scientist) | Major revision | Minor revision |
| Reviewer 2 (comp-bio / biostatistics) | Major revision | Minor revision |
| Reviewer 3 (methodology / ethics) | Major revision | Minor revision |
| Editor decision | **Major revision** | **Minor revision** |

All three reviewers advanced the manuscript by one decision tier between rounds. The editor explicitly endorsed the authors' choice to *scope rather than fabricate* the GPU-dependent orthogonal structural validation.

---

# ROUND 1

# Mock Peer-Review Report — EAC manuscript, Round 1

**Manuscript assessed:** M1 — "Steering into competitive whitespace: a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory" (v1)

**Review model:** 3 independent expert reviewers + handling-editor synthesis. Reviewers were briefed on the manuscript's framing and scope (in-silico, AI-assisted, public-data; no wet-lab data claimed or required); every requested revision is achievable within that scope.

**Panel:**
- **Reviewer 1** — Domain clinician-scientist (GI oncology, translational target/biomarker)
- **Reviewer 2** — Computational biologist & biostatistician
- **Reviewer 3** — Methodology & research-ethics

**Recommendations:** Reviewer 1 — Major revision · Reviewer 2 — Major revision · Reviewer 3 — Major revision

*Generated 2026-07-11. Reviews produced by independent model instances under distinct expert personas (Reviewer 1 recovered on an alternate model instance after a transient generation refusal on the primary; content unaffected). Advisory only — every specific factual claim was treated as a lead to verify, not ground truth.*

---

## Editor's Decision Letter

# Editor's Decision Letter

**Journal of Computational Drug Discovery & Design**
*Manuscript Reference: JCDD-2024-[assigned]*

---

**Re:** "Steering into competitive whitespace" — a public-data, AI-assisted, human-supervised computational discovery-to-design campaign in esophageal adenocarcinoma

---

## Overview

This manuscript reports a fully computational, AI-assisted campaign that characterizes EAC as a copy-number-driven malignancy, maps competitive whitespace across the clinical trial landscape, applies an explicit druggability triage to nominate two leads (GUCY2C, treatment arm; DKK1, interception arm), and delivers de novo protein binders for each via an RFdiffusion→ProteinMPNN→Boltz-2 pipeline validated in-silico by ipTM/pLDDT metrics. The work is transparently scoped — no wet-lab data are claimed or required — and the open-questions table (O-1 through O-5) reflects unusual intellectual honesty about what remains unresolved. Three expert reviewers (a GI oncology clinician-scientist; a computational biologist/biostatistician; and a methodology/research-ethics reviewer) have evaluated the submission. Their assessments converge on the same high-level verdict: the paper's framing, transparency practices, and competitive-landscape contribution are genuine strengths, but several quantitative pillars that underlie the paper's central claims are currently insufficiently supported by disclosed methods, sensitivity testing, or orthogonal validation. All three recommend **Major Revision**.

---

## Decision: **Major Revision**

The submission is not rejected. The competitive-landscape analysis, the trajectory-spanning strategic framing, and the honest acknowledgment of limitations are valuable contributions appropriate for a high-quality computational preprint venue. However, the paper cannot be accepted in its current form because several of its headline quantitative claims — composite triage scores, the "15/16 pass" rate, and top-design ipTM values — rest on undisclosed scoring rubrics, single-model fold-back, and small single-run samples that a rigorous reader cannot yet evaluate or reproduce. Every concern raised by reviewers is addressable by in-silico analysis, expanded reporting of already-generated data, or manuscript editing. No wet-lab experiments are required or expected.

---

## Essential Revisions

The following concerns were raised independently by two or more reviewers and must be addressed. They are listed in priority order.

---

**1. GUCY2C expression in EAC must be established from public data before the treatment-arm target selection stands.**
*(Reviewers 1 and 3)*

The tumor_selectivity score of 5 (highest ordinal) for GUCY2C is the most consequential unsupported quantitative claim in the paper. GUCY2C is canonically a colorectal antigen; the cited clinical precedent is a CRC vaccine. The manuscript assigns it top selectivity without citing any primary evidence that GUCY2C is expressed at therapeutically actionable levels in EAC tumor cells. This is fully addressable in silico: retrieve GUCY2C RNA expression (RSEM/TPM) from TCGA ESCA adenocarcinoma samples via cBioPortal or the GDC and cross-reference available Human Protein Atlas protein-level data in esophageal adenocarcinoma vs. normal esophageal mucosa. Report the result as a supplementary table or an additional figure panel. If expression is low or absent, the tumor_selectivity score and treatment-arm target discussion must be revised accordingly. This is the single highest-priority revision.

---

**2. Orthogonal structural validation of the two named leads must be completed before publication.**
*(All three reviewers)*

All reviewers independently flag that Boltz-2 ipTM is the sole structural validation metric, that this is a single-model, self-referential confidence score from within the same tool family used for generation, and that the paper itself names AF-Multimer or Chai-1 cross-validation as the very next step (O-1, Discussion). Re-folding the two leads (gucy2c_bb2, dkk1_bb1) — and ideally the top 2–4 designs per target — with at least one independent structure predictor (AlphaFold-Multimer or Chai-1) is an in-scope computational task that should be completed before publication, not deferred. Report the orthogonal ipTM and interface residue agreement or discordance in a supplementary table, and revise the Abstract's headline ipTM claims to reflect multi-model consensus or to flag single-model provenance where consensus is absent.

---

**3. A negative/specificity control for binder designs must be added.**
*(Reviewers 2 and 3)*

Several of the top-scoring designs (e.g., dkk1_bb2, dkk1_bb5; and gucy2c_bb2's terminal poly-alanine run) display the hallmarks of a known RFdiffusion/ProteinMPNN failure mode: generic amphipathic helical bundles that fold with high confidence against almost any target. The manuscript reports only a pairwise binder-vs-binder sequence-similarity check — it never tests whether the top binders fold with high pLDDT in isolation (no target) or score well against a non-cognate decoy. Without this control, an ipTM of 0.915 cannot be distinguished from "generically foldable helix that happens to sit near the target in this one prediction." Add: (a) monomer pLDDT of the binder in isolation, and (b) ipTM of each lead against one or two irrelevant decoy proteins. Flag high-alanine designs as potential developability liabilities in Table S1. This is a fully in-scope computational addition.

---

**4. The composite triage scoring rubric must be fully disclosed and sensitivity-tested.**
*(All three reviewers)*

All three reviewers independently identify this as a reproducibility-critical gap. The manuscript states "composite = association + selectivity + tractability + novelty" but does not disclose: the component weight(s), the ordinal scale anchors for tumor_selectivity and novelty (what does "5" mean for each?), the size of the candidate universe from which the 16 targets were drawn, or whether a single undisclosed weighting choice could reorder GUCY2C/DKK1 relative to CDH17, CEACAM5, or ERBB3, which score close behind. Required additions: (a) a complete scoring rubric in Methods or a new Table S0, including scale definitions, ordinal anchors, and weights; (b) a weight-perturbation or rank-robustness analysis (e.g., Monte Carlo over plausible weight vectors, or at minimum explicit re-ranking under 2–3 alternative weighting schemes) showing that GUCY2C and DKK1 remain in the top tier for their respective lanes under reasonable parameter variation; (c) disclosure of the pre-scoring candidate pool size and source (e.g., "top N targets by Open Targets association score for EAC MONDO identifiers"). A paper whose central contribution is an "honest, structured triage" must disclose the structure of the triage.

---

**5. The "15/16 designs cleared ipTM > 0.5" claim must be reframed and statistically grounded.**
*(Reviewers 2 and 3)*

n=8 backbones per target is a pilot-scale sample; presenting "15 of 16 pass" as a headline result without context implies a generalizable success rate that this sample cannot support. Required: (a) explicitly caveat in the Abstract, Results, and Figure 3 legend that this is a pilot-scale observation (n=8/target) from a single generation run, not a representative success rate; (b) report the mean ± range of ipTM/pLDDT across the 5 Boltz-2 diffusion samples per design rather than only the single best value (the data already exist); (c) contextualize the ipTM > 0.5 threshold against published calibration data for RFdiffusion/ProteinMPNN/Boltz pipelines so readers can assess whether 0.5 is a lenient or stringent bar. If additional backbone generation seeds are computationally feasible, running at least one additional seed per target and reporting seed-stability of the pass rate and top-design identity would substantially strengthen this section.

---

**6. The ClinicalTrials.gov "active-EAC-trial intervention mentions" methodology must be documented and the query must be timestamped.**
*(All three reviewers)*

The whitespace argument — the paper's most distinctive contribution — depends critically on the counts in Figure 1b and Table S3 (41 PD-1/PD-L1, 12 HER2, etc.). These are currently unvalidated annotations with no documented extraction protocol: it is unclear whether classification was by intervention name string match, MeSH term, or manual read; whether combination trials are double-counted across axes; and how the stated 106 active trials relate to the axis-level counts (reviewers note the per-axis counts do not readily sum to 106). Required: (a) add a Methods subsection documenting the exact extraction protocol and any manual curation steps; (b) provide a snapshot/query date for ClinicalTrials.gov, Open Targets, and Drugs@FDA; (c) provide the raw trial-ID list as a supplementary table; (d) reconcile the per-axis counts with the total trial count. Without this, the paper's central whitespace claim is not reproducible.

---

**7. The AI-disclosure statement must be specific and non-placeholder.**
*(Reviewer 3, supported by Reviewer 1)*

The Author Contributions section contains a bracketed placeholder: "[To be completed. The discovery-to-design campaign was executed as a human-supervised, AI-assisted workflow...]." This is inadequate and conflates the named ML tools that are objects of study (RFdiffusion, ProteinMPNN, Boltz-2 — methodology) with general-purpose AI assistance in analysis, scoring, drafting, or figure generation. A reader cannot determine which judgments (e.g., composite weighting in Table S2, the "whitespace verdict" column in Table S3) were made by a human analyst versus proposed by an LLM and reviewed. The revision must include a specific, non-placeholder statement identifying: which AI system(s) were used for which steps; what was deterministic bioinformatics versus AI-assisted synthesis or generation; and what human verification was applied at each step. Given that AI-assistance is framed as a feature of the contribution, the disclosure must be commensurate with that claim.

---

**8. DKK1 clinical-precedent evidence must be segregated by patient population.**
*(Reviewer 1)*

The DKN-01 citation (Klempner 2024) is a Phase 2 study in advanced gastric/GEJ adenocarcinoma — a treatment-arm precedent in late-stage patients who accept significant toxicity. The manuscript credits this as clinical precedent for the DKK1 *interception* arm in a predominantly asymptomatic Barrett's surveillance population (annual EAC progression ~0.3%/year). These are radically different clinical contexts with incommensurable safety expectations and endpoint definitions. Required: add a column or note to Table S2 (and corresponding text) explicitly distinguishing clinical-precedent evidence in advanced disease from evidence in the precursor/Barrett's setting — which is essentially absent beyond biomarker studies — and adjust the narrative accordingly. This is an editorial revision that requires no new data.

---

## Substantive Reviewer Disagreements

There is one area of substantive difference in reviewer emphasis that the authors should acknowledge and adjudicate rather than try to satisfy in contradictory directions:

**Scope of structural validation required.** Reviewer 1 would be satisfied by AF-Multimer or Chai-1 fold-back of the two named leads (gucy2c_bb2, dkk1_bb1). Reviewer 3 additionally requires negative/specificity controls (monomer folding, decoy-target folding) as a prerequisite for treating ipTM values as lead-selection evidence. Reviewer 2 emphasizes seed-stability of the generation run. These are compatible but have different cost/priority ordering. **The authors should complete the orthogonal fold-back (satisfying all three reviewers' minimal ask) and the negative control (Reviewer 3's additional ask), report whether the poly-alanine designs fail the monomer/decoy test, and note seed-stability as an acknowledged limitation if a full reseed is computationally prohibitive within the revision timeline.** This ordering satisfies the consensus while being honest about what was and was not completed.

---

## Path to Acceptance

This manuscript is closer to acceptance than the "Major Revision" designation might suggest. The paper's core intellectual contribution — a reproducible, competition-aware, honesty-gated nomination and binder-design template — is sound and well-executed. The revisions required are specific, bounded, and all achievable by in-silico analysis or editorial editing: retrieve and report GUCY2C EAC expression from TCGA/HPA; complete the AF-Multimer or Chai-1 fold-back of the two leads; add a specificity control for the binder designs; disclose and sensitivity-test the triage scoring rubric; reframe the "15/16 pass" claim with appropriate sample-size caveats and diffusion-sample spread; document the trial-counting methodology with query dates and a raw trial-ID supplement; replace the AI-disclosure placeholder with a specific statement; and segregate DKK1 evidence by patient population. A revised manuscript that addresses these eight points — maintaining the intellectual honesty that all three reviewers praised and extending it to the binder-validation and scoring-provenance sections — would meet the bar for acceptance at this venue.

Authors are invited to submit a point-by-point response letter addressing each Essential Revision in order, with clear indication of where in the revised manuscript each change appears. The revised manuscript will be returned to at least two of the original reviewers.

---

*Handling Editor*
*Journal of Computational Drug Discovery & Design*

---

# Full Reviews (Round 1)

## Reviewer 1

### Reviewer 1: Clinician-scientist, GI oncology and early-phase drug development; experience with esophageal/GEJ malignancies, Barrett's surveillance programs, and biologic/bispecific therapeutic development

---

## Manuscript M1: "Steering into competitive whitespace: a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory"

---

### Brief Summary

This manuscript describes an AI-assisted, human-supervised computational campaign that mines publicly available EAC genomics (TCGA, n=182), maps the competitive clinical landscape (ClinicalTrials.gov, Open Targets), applies an explicit druggability triage to nominate two therapeutic leads — GUCY2C for a T-cell-engager treatment arm and DKK1 for a neutralizing-trap interception arm in Barrett's — and delivers de novo protein binders for each using an RFdiffusion→ProteinMPNN→Boltz-2 pipeline, with 15/16 designs clearing an ipTM >0.5 threshold. The work is transparently framed as computationally validated only, with no experimental data, and explicitly names its two make-or-break unresolved questions (GUCY2C therapeutic window; DKK1 surrogate endpoint). The paper's principal value is as a reproducible, competition-aware nomination and design template, not as a claim of experimental proof-of-concept.

---

### Strengths

- **Intellectual honesty is a genuine feature, not a rhetorical pose.** The abstract, main text, figure legends, and supplement consistently and correctly label ipTM as a confidence proxy, not an affinity, and the open-questions table (O-1 through O-5) explicitly names what is unknown. This is unusual and praiseworthy for an early computational study and should be preserved in revision.

- **The competitive-landscape analysis is the paper's most distinctive contribution.** Table S3 constitutes a genuinely useful, structured articulation of where EAC clinical effort is concentrated. The logic of *deliberately routing away* from PD-1/HER2/VEGF — rather than casually asserting those axes are "crowded" — and mapping the resulting whitespace is clear, reproducible (the sources are named), and analytically coherent.

- **The druggability triage (Figure 2, Table S2) is methodologically defensible.** The explicit rejection of GPX7 as the highest raw-composite candidate on the grounds of target class (lost tumor suppressor, not antibody-addressable) demonstrates that the composite score is not just narrative decoration; it is tested against a meaningful filter. This is good scientific discipline.

- **Trajectory-spanning strategy is clinically cogent.** The framing of Barrett's→dysplasia→carcinoma as presenting two distinct intervention opportunities — molecular interception at the precursor stage and directed killing at the invasive stage — is biologically sound and reflects actual clinical unmet need, particularly the absence of any approved molecular agent for Barrett's high-risk surveillance patients.

- **Binder design metrics are reported comprehensively.** Table S1 reports all 16 designs with full sequences, lengths, mpnn_score, ipTM, complex_pLDDT, a combined confidence, and pass/fail, which allows the reader to assess the design landscape rather than only the cherry-picked leads.

- **The DKN-01 citation (Klempner 2024) for DKK1 precedent is appropriate and current**, and grounding the selection in existing clinical experience with a DKK1-neutralizing agent in esophagogastric cancer materially strengthens the biological plausibility argument.

- **AI assistance is disclosed as a feature of the method, not obscured.** This is correct practice.

---

### Major Concerns

**1. The GUCY2C therapeutic-window argument is asserted, not evaluated, and the assertion is clinically fragile.**

The entire treatment-arm thesis rests on the claim that GUCY2C's "apical, luminal-facing restriction in normal gut" provides a sufficient safety margin for a T-cell engager. This is stated in a single sentence in Results, one sentence in Discussion, and labeled as open question O-2. But the clinical risk is asymmetric and deserves far more rigorous treatment than it receives here. GI-directed T-cell engager toxicity is not a hypothetical: the documented experience with EpCAM-directed catumaxomab (peritoneal), the well-characterized cytokine-driven GI toxicity of CD3-bispecifics generally, and the specific issue that *murine* GUCY2C may not be cross-reactive with the human protein (which the authors acknowledge) together mean that preclinical safety data cannot be generated in the most common model system. This is not merely "one arm of a bispecific whose CD3 arm is not done" (Discussion); it is a target-class question that bears on whether a T-cell engager is the *correct format* for a GI-luminal antigen, even assuming excellent tumor selectivity. The manuscript should more explicitly discuss why a T-cell engager was selected over a safer-window format (ADC, CAR, vaccine, as listed in Table S3) for an antigen with acknowledged normal enterocyte expression. The current framing implies the window exists; it has not been established by any data cited here.

**2. The GUCY2C EAC expression evidence is absent and is essential to the target selection argument.**

The genomic characterization in the paper establishes GUCY2C is absent from the most frequent EAC amplicons and is not listed among the top altered genes in Figure 1a. Table S2 assigns it tumor_selectivity = 5 (the highest ordinal) with a note of "luminal-restricted." But GUCY2C is canonically a colorectal antigen: the cited clinical precedent (Ad5-GUCY2C-PADRE, Table S3) is a CRC vaccine. The manuscript does not cite primary evidence that GUCY2C is expressed at therapeutically actionable levels in EAC tumor cells, nor that its expression is maintained or upregulated across the Barrett's→EAC progression it is meant to target. A tumor_selectivity score of 5 without an EAC-specific expression citation is the single largest unsupported quantitative claim in the paper. This is achievable in silico: the Human Protein Atlas (HPA) and the TCGA RNA-seq data available via cBioPortal or the GDC contain esophageal/EAC expression data for GUCY2C. The authors should retrieve and report GUCY2C RNA expression (TPM or RSEM) in EAC tumor vs. normal esophageal tissue from TCGA and cross-reference HPA protein-level data. If GUCY2C is lowly expressed or absent in EAC, the entire treatment-arm target selection requires revision.

**3. The DKK1 interception arm conflates two distinct patient populations in a clinically consequential way.**

Table S2 labels DKK1's trajectory as "advanced + progression," and the competitive-landscape table (S3) places it under "Advanced." The main text and Figure 1c assign it to the interception arm across "Barrett's metaplasia and dysplasia." But the Klempner 2024 DKN-01 citation is a Phase 2 study in *advanced gastric or GEJ adenocarcinoma*, i.e., a treatment-arm precedent, not an interception-arm precedent. These are radically different clinical settings with different safety expectations, endpoint definitions, patient populations, regulatory pathways, and commercial value propositions. A drug that is tolerable in advanced cancer patients (who accept significant toxicity) may be unacceptable in the predominantly asymptomatic, mostly-non-progressing Barrett's surveillance population, where the annual progression rate to EAC is ~0.3% per year. The manuscript's Discussion and Open Questions (O-4) gesture at this — "very high safety bar" — but the conflation persists in the data tables and in the target-selection scoring, where DKK1's clinical-precedent credit appears to have been awarded for the wrong population. The authors should segregate the evidence base for DKK1 in advanced disease vs. DKK1 in the precursor setting and be explicit that the interception indication has essentially no clinical-precedent support beyond biomarker studies.

**4. The composite scoring methodology is opaque and not validated for sensitivity.**

Table S2 presents composite scores (e.g., GPX7 = 13.86, GUCY2C = 13.0, DKK1 = 11.37) that determine candidate rank order, but the methods section does not state the individual component weights, the scoring rubric for each ordinal, or whether the system is linear-additive, weighted, or otherwise. "A composite of disease association, tumor selectivity (ordinal), antibody tractability, and novelty" is the full description. This is insufficient for reproducibility. A reader cannot determine whether the 0.86-point difference between GPX7 and GUCY2C is meaningful, whether CDH17 (12.38, which is higher than GUCY2C 13.0 when corrected — wait: the table shows CDH17 = 12.38 < GUCY2C = 13.0, which is correct) is close enough to GUCY2C that the choice is effectively arbitrary, or whether DKK1 at 11.37 would survive any reasonable parameter perturbation. The paper cannot claim a "structured" triage without disclosing the structure. A sensitivity analysis — showing that the top-2 selections are stable under ±10-20% weight perturbation across components — is achievable computationally and should be added.

**5. The binder design pipeline lacks any orthogonal structural confidence check.**

The paper relies entirely on Boltz-2 ipTM as the validation metric. The Discussion correctly notes this as a limitation and names ESM, AlphaFold-Multimer, and Chai-1 as alternatives, and O-1 lists this as a near-term computational step. However, for a manuscript whose central technical claim is that these are "structurally credible, in-silico-validated leads," relying on a single model's self-reported confidence score without any orthogonal fold-back is a significant methodological gap. Boltz-2's ipTM is a diffusion-model confidence output that may be correlated across samples from the same run; it is not the same as independent model consensus. AF-Multimer and Chai-1 fold-back of the two leads (gucy2c_bb2 and dkk1_bb1) would require only additional compute and would substantially strengthen the "structurally validated" claim. Given that O-1 is already listed as near-term computational work, this should be completed before publication, not deferred.

---

### Minor Concerns

**1. The TCGA cohort size (n=182) and its limitations need a brief note.** The alteration frequencies cited (87% TP53, 39% CDKN2A, 35% CCND1) are well-supported by the TCGA 2017 ESCC/EAC paper, but the cohort is sequenced tissue, not necessarily representative of clinical EAC in terms of stage at presentation, treatment history, or geographic population. A single sentence acknowledging this is appropriate.

**2. The "active EAC trial intervention mentions" metric (Figure 1b, x-axis) is defined only cursorily.** Is this a count of unique trials, of mentions across all arms, or something else? The number 41 for PD-1/PD-L1 is plausible but should be operationally defined so that another researcher could reproduce it. The methods say "the count of active interventional trials referencing each target/axis" — does "referencing" mean the primary intervention, any arm, or any drug known to act on that target? This matters for the whitespace argument.

**3. The interface residue numbering in S4 is ambiguous.** The GUCY2C interface residues are given as [130,132,146,192,325,326,327,330,347,348,349,378,379] with the parenthetical "complex-file numbering," and separately the hotspot input residues are [239,240,267,422,424,425] in "isolated-domain numbering." The relationship between these two coordinate systems is not explained, and a reader cannot determine whether the binder contacts the intended hotspot patch. The authors should either provide a residue-number mapping or confirm explicitly that the interface residues correspond to the hotspot region in canonical UniProt numbering.

**4. The sequence of dkk1_bb2** ("MTREELIAAAGAAGLAYGAALTAALAAAAAAAGADTAAVLALGAAGAAAAAALAAAAAAAAA") has an unusually high alanine content (>50% Ala), and dkk1_bb5 and dkk1_bb3 show similar patterns. While the paper notes diverse solutions, a poly-alanine bias in a subset of designs may indicate aggregation-prone or low-complexity sequences that would not be developable. This should be noted as a liability flag, and the ESM developability screen (O-1) should be flagged as specifically needed for these designs.

**5. The mpnn_score is reported but not defined in the manuscript or supplement.** A brief definition (e.g., negative log-likelihood from ProteinMPNN; lower is better) should be added to Table S1's header or footnote.

**6. The reference list is abbreviated and inconsistently formatted.** The Klempner 2024 citation is cited as "DKN-01 in Combination With Tislelizumab and Chemotherapy" — this is a correct citation for the DisTinGuish Phase 2, but should note that this is a gastric/GEJ study, not an EAC-specific study, since the clinical-precedent argument depends on the distinction.

---

### Feasible Revisions

All of the following are achievable in silico or by manuscript editing; none requires wet-lab experiments.

1. **Retrieve and report GUCY2C expression in EAC from public RNA-seq data.** Query TCGA ESCA adenocarcinoma samples and/or Human Protein Atlas for GUCY2C RNA (RSEM/TPM) and available protein-level data in esophageal adenocarcinoma vs. normal esophageal mucosa. Report this in a supplementary table or an additional panel to Figure 1 or Figure 2. If expression is low or absent, revise the tumor_selectivity score and discussion accordingly. This revision directly addresses Major Concern 2 and is the highest-priority revision.

2. **Run AF-Multimer or Chai-1 fold-back on the two leads (gucy2c_bb2, dkk1_bb1) and report the orthogonal ipTM/interface score.** Add results as a supplementary table and revise the "structurally validated" claim to reflect multi-model consensus or discordance. This addresses Major Concern 5 and converts O-1 from a deferred item to a completed one.

3. **Disclose and sensitivity-test the composite scoring rubric.** Add to Methods (or a new Table S0) the full scoring schema: component definitions, ordinal scales, and weights. Run a ±15% weight perturbation analysis and report whether GUCY2C and DKK1 remain in the top tier for their respective lanes. This addresses Major Concern 4.

4. **Segregate DKK1 evidence by population in Table S2 and the main text.** Add a column or note distinguishing clinical-precedent evidence in advanced disease (DKN-01, Klempner 2024) from evidence in the precursor/Barrett's setting (which is essentially absent). Revise the trajectory label from "advanced + progression" to something like "Barrett's (interception; no clinical-precedent in precursor setting)" and adjust the biologic_druggability note to reflect the safety-bar asymmetry. This addresses Major Concern 3.

5. **Expand the GUCY2C therapeutic-window discussion and provide format-choice justification.** Add a paragraph (or supplementary note) that (a) reviews published safety data from GI-directed T-cell engagers and ADCs in GI cancers to contextualize the luminal-restriction claim; (b) explicitly argues why a T-cell-engager format was preferred over ADC or CAR-T for this antigen despite known enterocyte expression; and (c) identifies what HPA/GTEx expression data already show about GUCY2C in human gut segments relevant to the safety argument. This partially addresses Major Concern 1 within scope.

6. **Clarify the "active EAC trial intervention mentions" operational definition** in the Methods section with enough precision that the count is reproducible, and add a date-of-query for the ClinicalTrials.gov and Open Targets downloads. This addresses Minor Concern 2 and is essential for reproducibility.

7. **Add a residue-numbering mapping table** (UniProt canonical → isolated-domain → complex-file) for both GUCY2C and DKK1 in S4. Confirm or correct whether the predicted interface residues correspond to the intended hotspot patches. This addresses Minor Concern 3.

8. **Flag the high-alanine-content designs (dkk1_bb2, dkk1_bb5, dkk1_bb3) as potential developability liabilities** in Table S1 and in S5/S6. State that ESM screen (O-1) is specifically prioritized for these sequences. This addresses Minor Concern 4.

9. **Define mpnn_score** in Table S1's header and add a brief note on its interpretation. This addresses Minor Concern 5.

---

### Recommendation

**Major Revision**

The paper's competitive-landscape analysis, trajectory-spanning strategic framing, and honest disclosure practices make it a genuinely useful contribution to the computational drug discovery literature as a preprint. However, two issues require substantive revision before the work's central claims are defensible: the absence of any EAC-specific expression evidence for GUCY2C (the treatment-arm target) leaves the most important tumor-selectivity score unsupported, and the single-model structural validation without orthogonal fold-back is insufficient for the "structurally credible leads" claim when O-1 is already identified as near-term and feasible. Both are achievable in silico. The composite-scoring opacity and the DKK1 population conflation are additionally necessary to fix for intellectual honesty. None of the required revisions involve wet-lab experiments.

---

## Cross-cutting Comment

This is a single-manuscript submission, so cross-cutting remarks bear on the campaign as a whole. The work's strongest feature — radical transparency about what is and is not known — is also the lens through which its gaps are most visible: having named O-2 (GUCY2C therapeutic window) and O-4 (DKK1 surrogate endpoint) as make-or-break questions, the paper is obligated to present the best available computational and public-data case for why those questions are *worth asking*, and that case is currently stronger for DKK1 than for GUCY2C. The treatment arm's target nomination would be substantially more credible if EAC-specific expression data were integrated; their absence is the one place where the campaign's stated honesty-gating principle is not fully applied. The precedent of designing for both arms in a single cycle is methodologically sound, and the binder quality metrics are reported with appropriate transparency — the ask is simply that the same rigor be applied to the upstream target-selection evidence as is applied to the downstream design confidence.

---

## Reviewer 2

### Reviewer 2: Computational biologist / biostatistician with a focus on genomics reproducibility and structure-prediction benchmarking

**Brief summary**

The manuscript is a computational discovery-to-design campaign that (1) characterizes EAC genomics from a single TCGA cohort, (2) builds a competitive-landscape map from ClinicalTrials.gov and Open Targets, (3) applies a scoring/triage scheme to nominate GUCY2C (treatment arm) and DKK1 (interception arm), and (4) generates and folds-back 16 de novo binders (RFdiffusion → ProteinMPNN → Boltz-2), reporting ipTM/pLDDT as validation metrics. The authors are explicit that everything is in-silico and flag two "decisive open questions" as unresolved. As a statistics/reproducibility reviewer, my concerns center on the quantitative scoring systems (Fig. 1b, Fig. 2, Tables S2–S3) being under-specified and unvalidated, the single-cohort/single-model nature of every quantitative claim, and the self-referential circularity of using Boltz-2 ipTM as the sole validation for designs whose hotspots were also chosen heuristically from the same AlphaFold structures.

**Strengths**

- Transparent, fully-sourced quantitative claims (each number traced to a named table/file) — genuinely good reproducibility practice for a preprint of this kind.
- Explicit, auditable rejection of the top-scoring raw candidate (GPX7) on mechanistic grounds — a real demonstration that the triage isn't just cherry-picking a foregone conclusion.
- Full disclosure of the complete 16-design table (not just winners), including the one design that failed threshold (gucy2c_bb6), which allows an outside reader to recompute pass rates and inspect the distribution rather than take summary claims on faith.
- Correct, repeated caveating of ipTM as a "docking confidence proxy," not affinity — this is stated in figure legends, results, and discussion, which is unusual honesty for this genre of paper.
- Cross-reactivity/similarity check across the design set is a sensible, cheap robustness step and its negative result is reported plainly rather than oversold.

**Major concerns**

1. **Composite scores in Table S2/Figure 2 have no disclosed weighting, scale, or uncertainty.** "Composite = association + selectivity + tractability + novelty" is given as a formula but none of the four component scales (1–5? what is "novelty=5" anchored to?) are defined, and there is no sensitivity analysis showing whether GUCY2C/DKK1's selection is robust to reasonable re-weighting of these four components. Given that DKK1 (composite 11.37) ranks *below* four rejected/down-weighted candidates (13.86, 12.47, 12.38, 12.29) purely on the "druggability filter" override, the paper needs to show this ranking is not an artifact of arbitrary weights. A simple weight-perturbation or rank-robustness check (e.g., resample weights, recompute rank of DKK1/GUCY2C across a plausible weight simplex) is entirely in scope and would substantially strengthen the "honest triage" claim.

2. **Single-cohort genomic claims presented without any statistical framing.** The 87%/39%/35% alteration frequencies (Fig. 1a) come from n=182 TCGA samples with no confidence intervals, no comparison to an independent EAC cohort (e.g., ICGC OCCAMS/OAC), and no discussion of the fact that TCGA ESCA "adenocarcinoma subtype" filtering itself involves a sample-selection choice that should be stated (how many ESCA samples were excluded as non-adenocarcinoma, and on what histology annotation?). A single-cohort percentage stated as a fact ("TP53 altered in 87% of tumors") should carry at least a binomial CI (which is trivial to compute: 87% of 182 ⇒ ~[81–91%] at 95%) and ideally a cross-check against the published TCGA ESCA paper's own reported frequency (ref. 3) to confirm cBioPortal's current pipeline reproduces the original TCGA numbers.

3. **"Active-EAC-trial intervention mentions" (Fig. 1b, Table S3) is an unvalidated, seemingly manual annotation with no counting methodology.** How was a trial classified as "mentioning" a target/axis — by intervention name string match, by MeSH term, by manual read of the trial description? Is a single trial counted once per target it might touch (pseudoreplication risk if combination trials touch both e.g. PD-1 and HER2)? Without a documented extraction protocol and inter-rater or at least a spot-check reproducibility statement, these counts (41 PD-1, 12 HER2, 8 VEGF, 1 each for the "whitespace" targets) cannot be taken as reproducible measurements — they read as narrative-supporting tallies. This directly affects the paper's central "whitespace" argument.

4. **Boltz-2-only structural validation is a single-model, single-tool result with no orthogonal consensus — and the design and validation loops are not independent.** The hotspot residues were chosen using the *same* AlphaFold model (pLDDT-based heuristic) that Boltz-2 also uses internally for confidence calibration; folding the binder back onto a structure selected because it was already high-confidence in that same modeling family is a mild form of circularity. The manuscript acknowledges the "no orthogonal consensus" limitation in the Discussion, which is appropriate, but Results and Abstract nonetheless state ipTM figures with three-decimal precision (0.915, 0.858) as if they were precise, reproducible measurements, with no report of run-to-run variance (Boltz-2 stochastic diffusion sampling — 5 samples were generated per the Methods; only the best is reported, not the spread). Reporting the mean/SD or range across the 5 diffusion samples per design, rather than a single best value, is a straightforward in-silico robustness addition and should be added to Table S1.

5. **Multiple-comparisons / selection-bias framing is missing for the "15 of 16 pass" headline.** With 16 designs and a threshold of ipTM>0.5, presenting "15/16 pass, one clear failure" without discussing the fact that the pass threshold (0.5) is a soft, community-convention cutoff (not calibrated against any known-binder benchmark set) risks the reader over-interpreting the pass rate as evidence of a well-calibrated pipeline. A brief calibration statement — e.g., citing published ipTM/success-rate correlations from the RFdiffusion/ProteinMPNN benchmark papers, or noting where these designs would fall on that curve — would let readers judge whether 0.5 is a stringent or lenient bar here.

6. **No independent replicate seeds for RFdiffusion/ProteinMPNN generation.** Eight backbones per target were generated once; there is no report of whether repeating the generation with a different random seed set would give a similar top-ipTM winner or count of passers. Given the very large spread in ipTM within DKK1 designs alone (0.524–0.858) and within GUCY2C (0.401–0.915), the "top design" statistic is an extreme-value statistic from a small, single sample of backbones — its stability under reseeding is an easy, in-scope computational check (rerun generation with 1–2 additional seeds, report whether the top-line ipTM leaders are stable) that would materially strengthen confidence in "gucy2c_bb2/dkk1_bb1 are the leads" rather than "the best of one particular sample of 8."

**Minor concerns**

1. Figure 1a legend and Results text should state whether alteration percentages are of the full n=182 or of some subset with available CNA+mutation data (cBioPortal sometimes reports slightly different denominators by data type); please confirm and state the exact denominator used for each gene.
2. "106 active trials" and "351 interventional trials" — please state the ClinicalTrials.gov query date/snapshot, since this is a live, constantly-updated database and reproducibility requires a timestamp.
3. Table S2's "tumor_selectivity" and "novelty" columns are integers 1–5 with no rubric given anywhere (main text, supplement, or methods) — a short scoring rubric (even qualitative, e.g., "5 = expression restricted to a single normal tissue with no oncologic overlap") should be added as a supplementary table.
4. The claim "a sequence-similarity check... detected no high-similarity binder pairs" (Results, S5) should report the actual metric and threshold used (% identity? alignment method? pairwise all-vs-all?) rather than only the qualitative conclusion.
5. Figure 3/S1 conflates "mpnn_score" units without defining scale/sign convention (lower/higher better?) — should be defined once in Methods.
6. The Discussion's claim that this is a "reproducible template" would be more convincing if the manuscript stated explicitly whether the triage scoring code/weights and the trial-counting scripts are included in the supplementary artifacts, or only the output CSVs — reproducibility of a *score* requires the scoring code, not just the resulting numbers.

**Feasible revisions**

1. Add binomial confidence intervals to all cohort-derived percentages in Fig. 1a/Results and explicitly cross-check the 87%/39%/35% figures against the original TCGA ESCA publication's reported frequencies (ref. 3) to confirm consistency between the two data vintages/pipelines.
2. Add a weight-sensitivity/rank-robustness analysis for the Table S2 composite score (e.g., Monte Carlo over plausible weight vectors, or at minimum report rank of GUCY2C/DKK1 under 2–3 alternative reasonable weightings) and report it as a supplementary figure.
3. Document the exact methodology for "active-EAC-trial intervention mentions" (search terms, database snapshot date, whether trials are double-counted across axes, and any manual curation step), and add this as a subsection in Methods with the raw trial-ID list as a supplementary table.
4. For each of the 16 designs, report ipTM/complex-pLDDT as mean ± range across the 5 Boltz-2 diffusion samples (not only the best), in an expanded Table S1, and note explicitly if the reported "lead" values are the max of 5 draws.
5. Rerun RFdiffusion/ProteinMPNN backbone generation with at least one additional random seed for each target and report whether the top-ipTM lead and overall pass rate are stable across seeds (or, at minimum, explicitly flag in Discussion/Limitations that current leads derive from a single generation run and could shift under reseeding).
6. State the exact sequence-similarity metric/threshold used for the S5 cross-reactivity screen, and add the underlying pairwise similarity matrix as a supplementary table.
7. Add a rubric (even a short qualitative one) defining the 1–5 scales for tumor_selectivity and novelty in Table S2, and clarify in Methods whether these were assigned by a single annotator or reconciled across multiple raters.
8. Add a sentence in Methods/Discussion contextualizing the ipTM>0.5 threshold against published calibration data for RFdiffusion/ProteinMPNN/Boltz-type pipelines, so readers can judge stringency rather than take 15/16 as self-evidently a strong result.

**Recommendation:** Major revision — the computational scope and honesty framing are appropriate for a preprint of this kind and the core claims are plausible, but several quantitative pillars (trial-count methodology, composite-score weighting, single-seed/single-model structural validation) currently lack the statistical scaffolding needed to support the paper's central "differentiated, defensible whitespace" argument, and all of the requested fixes are achievable by additional in-silico analysis or expanded reporting of already-generated data, not new experiments.

---

**Cross-cutting comment:** As this is a single-manuscript submission, this section is not applicable beyond noting that the paper's greatest asset — its unusually complete disclosure of raw tables (all 16 designs, full triage table, trial counts) — is also what exposes its weakest seams (unweighted composite scores, single-run structural metrics, undocumented trial-counting) to a statistically-minded reader; tightening the quantitative provenance of those specific artifacts, all achievable without new wet-lab work, would bring the rigor of the analysis fully in line with the honesty of its framing.

---

## Reviewer 3

### Reviewer 3: Methodology & research-ethics reviewer (computational oncology target validation, generative protein design, AI-disclosure norms)

**Manuscript: "Steering into competitive whitespace" — EAC dual-arm campaign (GUCY2C / DKK1)**

**Brief summary**

The manuscript reports a fully computational campaign that (i) profiles TCGA EAC genomics, (ii) maps the ClinicalTrials.gov/Open Targets competitive landscape, (iii) applies a stated "honest druggability filter" to nominate GUCY2C (treatment arm, T-cell engager) and DKK1 (interception arm, neutralizing trap), and (iv) runs an RFdiffusion→ProteinMPNN→Boltz-2 pipeline to generate and self-validate 16 de novo binders (8 per target), reporting top ipTM of 0.915 (GUCY2C) and 0.858 (DKK1). The paper is explicit and largely consistent about labeling every claim as in-silico, and it flags two "decisive open questions" (therapeutic window, interception surrogate endpoint) as unresolved. The core methodological question for this review is whether the exhibited numbers are a fair, non-cherry-picked, and adequately cross-validated representation of what the campaign actually produced.

**Strengths**

- Rejecting the top raw-composite candidate (*GPX7*, 13.86) because it is a lost tumor suppressor rather than an antibody-addressable target (Table S2, Fig. 2) is a genuinely useful demonstration of the "honesty gate" the paper claims to apply — this is exactly the kind of self-critical step that guards against narrative-driven target selection, and it is good that it is shown rather than asserted.
- Every headline number is tied to a named source table (`design_leads.csv`, `target_nomination.csv`, `competitive_landscape.csv`), which is the right posture for reproducibility even before the underlying data can be independently re-run.
- Explicit ipTM/pLDDT caveats ("a docking-confidence proxy...not a measurement of affinity") appear in Results, Discussion, and figure legends rather than once in a footnote — repetition of the caveat where the number is used is good practice, not boilerplate.
- The two open "make-or-break" questions (O-2 GUCY2C therapeutic window, O-4 DKK1 surrogate endpoint) are framed as could-halt-the-arm risks rather than minimized — this is a stronger form of honesty than most target-nomination papers offer.

**Major concerns**

1. **No orthogonal structural validation despite it being clearly achievable in-silico.** All 16 designs are scored by Boltz-2, the same tool family used implicitly downstream of the RFdiffusion/ProteinMPNN generation loop. The paper itself names an AlphaFold-Multimer or Chai-1 cross-check as the very next step (O-1, Discussion) but does not perform it. Given this is squarely in-scope computational work with no wet-lab dependency, presenting single-model ipTM values (0.915, 0.858) as headline results in the Abstract without at least one orthogonal fold-back is a real gap — self-consistency between a design pipeline and its own confidence metric is a known confound in de novo binder work, and the paper's own limitations paragraph concedes this without acting on it. I would not accept the leads as reported without at least the top 2–4 designs per target re-folded with an independent structure predictor.

2. **No negative/specificity controls anywhere in the validation.** Several of the top-scoring sequences are strikingly generic idealized helical bundles with very high alanine content and heptad-repeat character — e.g., dkk1_bb2 ("MTREELIAAAGAAGLAYGAALTAALAAAAAAAGADTAAVLALGAAGAAAAAALAAAAAAAAA", Table S1) and dkk1_bb5/gucy2c_bb2 (which ends "...AAEAAAAAAAA"). This pattern is a recognized failure mode in RFdiffusion/ProteinMPNN binder campaigns, where the model converges on an amphipathic-helix solution that folds with high confidence against almost any target rather than one shape-complementary to the specific epitope. The manuscript reports only a pairwise binder-vs-binder sequence-similarity check (S5) as a "diversity" control — it never asks whether the same binder scores well against a scrambled/irrelevant target, or whether the binder alone (no target) already folds into a high-pLDDT idealized bundle independent of the interface. Without that negative control, "ipTM 0.915" cannot be distinguished from "generically foldable helix that happens to sit near the target in this one prediction." This is fully achievable computationally and should be added before the ipTM numbers are treated as lead-selection evidence.

3. **The composite triage score is not reproducible or sensitivity-tested.** Table S2 gives only the final composite (e.g., GUCY2C 13.0, DKK1 11.37) and four ordinal sub-scores, with no disclosed weighting scheme, scale bounds, or the size of the candidate universe from which these 16 targets were drawn. Without knowing whether DKK1 (11.37) was the 2nd-best interception candidate out of 20 evaluated or out of 200, the reader cannot assess whether the list was pre-narrowed to fit the whitespace narrative before scoring began. A single, undisclosed weighting choice could plausibly reorder GUCY2C/DKK1 relative to CDH17, CEACAM5, or ERBB3, all of which score close behind in Table S2. This is a denominator problem central to this reviewer's brief and needs to be fixed by disclosure, not by re-scoring.

4. **Design-set sample size is too small to support the "15/16 pass" framing as a rate.** Eight backbones per target, one sequence carried forward per backbone, yields n=8 per arm. Reporting "15 of 16 designs cleared ipTM > 0.5" (Results, Fig. 3) reads as a success-rate statistic but is really two n=8 pilot batches; the true expected hit rate at any meaningful design scale (hundreds–thousands of backbones is standard in published RFdiffusion binder campaigns) is unknown from this data. The manuscript should either explicitly caveat that 15/16 is a pilot-scale observation, not a representative success rate, or scale the design sweep (an in-silico step, fully within scope) before quoting a pass fraction.

5. **AI-disclosure is a placeholder, not a statement.** The framing claims "AI assistance is disclosed and is part of the contribution," but the manuscript's only concrete disclosure is the bracketed Author Contributions line ("[To be completed. The discovery-to-design campaign was executed as a human-supervised, AI-assisted workflow...]"). This conflates two very different things: (a) the named ML tools that are the object of study (RFdiffusion, ProteinMPNN, Boltz-2 — these are methodology, not "AI assistance" in the disclosure sense) and (b) whatever general-purpose AI system(s) performed literature synthesis, scoring, figure generation, or drafting. As written, a reader cannot tell which specific judgments (e.g., the composite weighting in Table S2, the "whitespace verdict" column in Table S3) were made by a human analyst versus generated by an LLM and merely reviewed. This needs a specific, non-placeholder statement: what tool(s), what steps, what was human-verified versus AI-proposed.

**Minor concerns**

1. Figure 1b/Table S3 crowding counts (41, 12, 8, 3, 2, 1–2, 1×6) sum to roughly 70–75 of the stated 106 active trials; the remainder (likely non-targeted chemotherapy/other combination arms) is never explained, leaving the reader unsure whether "crowding" undercounts or the 106 total includes many non-axis-specific trials. A one-line reconciliation would remove ambiguity.
2. No access/snapshot date is given for the ClinicalTrials.gov, Open Targets, or Drugs@FDA queries (Methods, "Target-association and competitive-landscape analysis"). These sources change continuously; a query date is needed for reproducibility.
3. The TCGA cohort (n=182, `esca_tcga_pan_can_atlas_2018`) is used without any note on its demographic composition, and EAC/Barrett's epidemiology is known to skew heavily toward specific ancestries/sexes; a brief statement on cohort representativeness (or its absence) would strengthen the equity framing implicit in an "interception in a large, identifiable population" claim (Introduction).
4. Figure 3's legend ("Reproduced from the design campaign; source `design_results.png`...") reads as unedited pipeline boilerplate rather than a finished caption; it should specify what is plotted (e.g., ipTM vs. pLDDT scatter, colored by target) directly rather than by reference to an external file name.
5. Several bracketed placeholders remain (author list, affiliations, correspondence, competing interests) — expected for a preprint draft, but should be flagged as incomplete rather than left silent.

**Feasible revisions**

1. Re-fold the top 2–4 designs per target (at minimum the two named leads, gucy2c_bb2 and dkk1_bb1) with an independent structure-prediction method (AlphaFold-Multimer or Chai-1, both public) and report agreement/disagreement in ipTM and interface residues; this converts O-1 from a stated future step into an actual result and is entirely in-scope.
2. Add a specificity/negative-control analysis: fold each top binder alone (no target) and, ideally, against 1–2 non-cognate decoy targets, and report whether pLDDT/ipTM discriminate the cognate complex from these controls. Flag any binder whose "confidence" appears target-independent (the high-alanine sequences noted above are the most likely candidates to fail this test).
3. Disclose the size and source of the pre-scoring candidate universe for Table S2 (e.g., "top N targets by Open Targets association score for MONDO_0005028/MONDO_0013662"), the exact composite weighting formula, and ideally a brief sensitivity check (e.g., re-rank under two or three alternative weightings) to show GUCY2C/DKK1 selection is not an artifact of one weighting choice.
4. Reword the "15 of 16 designs cleared ipTM > 0.5" claim (Abstract, Results, Fig. 3 legend) to explicitly state the pilot sample size (n=8 backbones/target) and avoid implying a generalizable success rate; alternatively, expand the backbone sweep computationally and report the pass rate at larger n.
5. Replace the placeholder AI-disclosure sentence in Author Contributions with a specific statement identifying which AI system(s) performed which steps (literature/data synthesis, scoring, code generation, figure/text drafting) and which steps were deterministic bioinformatics tools (RFdiffusion/ProteinMPNN/Boltz-2), plus what human verification was applied to each.
6. Add ClinicalTrials.gov/Open Targets/Drugs@FDA query/snapshot dates to Methods, and reconcile the Fig.1b/Table S3 trial-count denominators.
7. Add one sentence noting the demographic composition (or its absence) of the TCGA cohort used, given known epidemiological skew in EAC/Barrett's populations, to properly frame the "large, identifiable high-risk population" claim used to motivate the interception arm.

**Recommendation: Major revision.** The campaign is transparently in-silico, honestly caveated in its prose, and methodologically interesting as a template, but several core numeric claims (design pass-rate, ipTM leads, composite triage ranking) currently rest on single-model, small-n, and undisclosed-weighting evidence that a rigorous reviewer cannot yet take at face value; all requested fixes are computational or editorial and achievable before resubmission.

---

**Cross-cutting comment:** As a single-manuscript submission, the main cross-cutting risk is internal: the paper's professed "honesty gating" is real and demonstrable in the target-triage step (rejecting GPX7) but is not yet extended with equal rigor to the binder-validation step, where self-consistency (Boltz-2 scoring its own design lineage) and small sample size are treated as acknowledged future work rather than fixed now, even though both fixes are within scope. The gap between the rhetorical framing ("explicit honest druggability triage," "reproducible, honesty-gated template") and the actual evidentiary support for the headline ipTM numbers and the 16-candidate shortlist's denominator is the main thing standing between this being a genuinely rigorous methods paper and a well-written but partially under-verified one.


---

# AUTHORS' RESPONSE TO ROUND 1

# Point-by-Point Response to Round-1 Reviews

**Manuscript:** "Steering into competitive whitespace: a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory"
**Decision:** Major revision (unanimous). Below, each of the editor's eight Essential Revisions is answered with the specific change made. All revisions were achievable in silico or by editing; none required wet-lab work.

---

**ER-1 — Establish GUCY2C selectivity/expression from public data (Reviewers 1, 3; highest priority).**
Done, with an honest split between what the data support and what they do not. We retrieved GTEx v8 normal-tissue RNA for GUCY2C and five comparators (new **Figure S1**, **Table S4**). GUCY2C's normal expression is 85% GI-restricted with a top non-GI tissue of only ~1.7 TPM — supporting the tissue-selectivity premise (matched only by CDH17; contrasted with pan-epithelial TROP2). We explicitly added two caveats now in Results and Discussion: (a) GTEx is *normal* tissue and does **not** establish EAC tumor-cell expression (the precedent is colorectal; EAC tumor expression remains to be confirmed — flagged, not assumed); (b) GUCY2C is genuinely expressed in normal small intestine/colon (~28 TPM), so "luminal restriction" is a polarity argument and the normal-GI window for a T-cell engager is exactly open question O-2. We did not inflate the tumor_selectivity score; we contextualized it.

**ER-2 — Orthogonal structural validation of the two leads (all reviewers).**
Partially addressed and honestly scoped. Re-running AlphaFold-Multimer/Chai-1 fold-back requires GPU compute outside this manuscript-preparation environment, so we did not fabricate consensus numbers. Instead we (a) removed all language implying multi-model validation, (b) state throughout (Abstract, Results R5, Discussion, O-1) that structural confidence rests on single-model Boltz-2 ipTM and that orthogonal fold-back + specificity controls are the decisive, not-yet-run next step, and (c) added the composition-based specificity screen below (ER-3) as the check we *could* run rigorously now. The claim strength was lowered to match the evidence rather than the evidence overstated to match the claim.

**ER-3 — Negative/specificity control for binder designs (Reviewers 2, 3).**
Done as a sequence-composition liability screen on the actual ProteinMPNN sequences (new **Table S5**). This directly confirmed the reviewers' predicted failure mode: the top GUCY2C design by ipTM (gucy2c_bb2) is 44% alanine with an 8-residue homopolymer run, and dkk1_bb2/bb3/bb5 exceed 50% alanine. We now report that **dkk1_bb1 is the most robust lead** (ipTM 0.858, 17.5% Ala, max run 2) and that gucy2c_bb2's headline ipTM must be read against its composition liability. Monomer-only and decoy-target fold-back are named as the remaining O-1 controls that require GPU re-runs.

**ER-4 — Disclose and sensitivity-test the triage rubric (all reviewers).**
Done. We reverse-engineered and disclose the exact additive rubric — `composite = tumor_selectivity + novelty + tractability_score + 5 × association` — which reproduces all 16 published scores to within 0.005 (new **Table S0**, and Methods). We ran a 20,000-draw Monte-Carlo weight-perturbation (±50% on all weights): GUCY2C ranks #1 among biologic-addressable treatment targets in 80.7% of draws (top-3 in 94.2%; CDH17 nearest). We disclose candidly that DKK1 is the *only* biologic-addressable interception target after the filter, so its rank is trivial and reflects a thin addressable space — not a contrived win.

**ER-5 — Reframe "15/16 pass" as pilot-scale (Reviewers 2, 3).**
Done. Abstract, Results R5, and the Figure 3 legend now state this is a pilot-scale observation (n=8 backbones/target, single generation seed, single-best-of-five diffusion samples), that ipTM>0.5 is a permissive bar, and that n=8 cannot support a generalizable success rate. Per-design metrics are in Table S1/S5.

**ER-6 — Document trial-count methodology + query date (all reviewers).**
Partially addressed honestly. The upstream artifacts did not record the exact ClinicalTrials.gov query strings or snapshot dates, and we did not invent them. Methods now (a) label the per-axis "active-trial mentions" as relative crowding indicators, (b) state they were not de-duplicated across combination trials and therefore do not sum to the 106 active-trial total, and (c) flag the missing query strings/snapshot date and raw trial-ID list as a reproducibility gap to close before external submission. We also removed the one unsupported point ("Barrett's interception" at x≈1.4) from Figure 1b that had no valid same-axis count.

**ER-7 — Specific, non-placeholder AI-disclosure (Reviewer 3, supported by 1).**
Done. The placeholder is replaced with a three-category disclosure separating (1) deterministic bioinformatics, (2) published ML methods as objects of study (RFdiffusion/ProteinMPNN/Boltz-2), and (3) AI-assisted synthesis/drafting with human review — and explicitly identifies the ordinal Table S2 scores and the "whitespace verdict" column as analyst judgment, not measurement.

**ER-8 — Segregate DKK1 precedent by patient population (Reviewer 1).**
Done. Results R3 now states that the DKN-01 precedent is in *advanced* gastric/GEJ disease (a late-stage treatment population) and is **not** evidence for interception in a pre-malignant Barrett's population, that direct precursor-setting therapeutic evidence is essentially absent, and that DKK1 is carried on mechanistic/druggability grounds with the population mismatch stated as risk O-4.

---

**On the reviewer disagreement (scope of structural validation).** We adopted the editor's recommended ordering: report the composition specificity screen we could run rigorously (ER-3), name orthogonal fold-back + monomer/decoy controls as the not-yet-completed O-1 step rather than claim them, and note single-seed generation as an acknowledged limitation. We chose transparency over a fabricated consensus.

**Net effect on framing.** The revisions lowered several claim strengths (pass-rate, GUCY2C selectivity, DKK1 precedent, single-model confidence) to match the evidence, and the one genuinely new in-silico finding that emerged — that the top-ipTM GUCY2C binder has a developability liability while the DKK1 binder is clean — is now the honest headline of the design section.


---

# ROUND 2 (revised manuscript)

# Mock Peer-Review Report — EAC manuscript, Round 2 (revised submission)

**Manuscript assessed:** M1 v2 — "Steering into competitive whitespace…" (revised after round-1 Major Revision), with the authors' point-by-point response letter.

**Review model:** Same 3-persona panel + handling-editor synthesis. Reviewers were given the round-1 essential revisions and the authors' response, and asked to judge the adequacy of each revision and flag any newly-introduced problems. (Reviews run on claude-sonnet-4-6 after transient generation refusals on the primary reasoning model; content unaffected.)

**Recommendations:** Reviewer 1 — Minor revision · Reviewer 2 — Minor revision · Reviewer 3 — Minor revision  *(all three moved up from Major revision in round 1)*

*Advisory only — every specific factual claim was treated as a lead to verify, not ground truth.*

---

## Editor's Decision Letter (Round 2)

## Decision Letter — Round 2

**Journal of Computational Biology & Drug Discovery (Preprint Track)**
**Manuscript:** "Steering into competitive whitespace — a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory"
**Decision Date:** [current date]
**Handling Editor:** [HE]

---

### Opening Assessment

The revised manuscript represents a substantive and honest response to the round-1 Major Revision decision. The authors successfully resolved five of the eight essential revisions to an adequate standard and made creditable progress on two others, most notably by (a) delivering a genuine new scientific finding — the alanine-scaffold liability of the headline GUCY2C binder — that reranks the leads in a direction that reduces rather than inflates confidence in their most dramatic result, and (b) declining, rather than fabricating, orthogonal fold-back validation numbers they could not responsibly compute. The response letter is internally consistent with the manuscript text throughout. That said, three reviewers converge on a common set of residual problems — one introduced by the revision itself — that are materially important for the manuscript's core biological claims and are all addressable by public-data query and text editing.

---

### Decision: **Minor Revision**

All three reviewers independently recommend Minor Revision. The manuscript is not ready to accept as-is, but none of the remaining issues requires new experimental data, GPU-intensive recomputation, or a full-cycle structural overhaul. A focused author response addressing the items below should allow acceptance without a further formal review round, at editor discretion.

---

### Status of the Eight Essential Revisions

| ER | Status | Notes |
|---|---|---|
| **ER-1** GUCY2C selectivity/expression from public data | **PARTIALLY ADEQUATE** | GTEx normal-tissue selectivity quantified correctly and well-contextualized; the normal-tissue vs. tumor-expression distinction is now explicit. Two residual problems: (i) GUCY2C mRNA expression in the TCGA *esca_tcga_pan_can_atlas_2018* cohort — from the same public repository used for Figure 1a — was not queried; (ii) the DKK1 fibroblast artifact introduced by the ER-1 GTEx panel is unaddressed (see Remaining Items). |
| **ER-2** Orthogonal structural validation | **ADEQUATE for preprint scope** — see separate judgment below | |
| **ER-3** Negative/specificity control | **ADEQUATE** | Composition liability screen is a genuine contribution; gucy2c_bb2 correctly reassigned from headline to provisional; dkk1_bb1 elevated on combined grounds. Residual: binder-diversity metric is undefined (see Remaining Items). |
| **ER-4** Triage rubric disclosure and sensitivity test | **PARTIALLY ADEQUATE** | Full disclosure and reproduction are exemplary; Monte Carlo design is creditable. Residual: the ±50% perturbation preserves the ×5 association-weight premium rather than testing it structurally; a narrow grid sensitivity (association weight ∈ {1,3,5,7,10}) would constitute genuine robustness vs. local noise. |
| **ER-5** Reframe "15/16 pass" as pilot-scale | **ADEQUATE** | Consistently applied across Abstract, Results, and legends. |
| **ER-6** Trial-count methodology and query date | **PARTIALLY ADEQUATE** | The non-de-duplication caveat and honest acknowledgment of missing query strings are appropriate. Residual: the caveat appears in Methods but not in the Figure 1b legend, where readers encounter the counts first. A one-sentence label fix closes this. |
| **ER-7** Specific non-placeholder AI disclosure | **ADEQUATE** | Three-category framework is substantive and operationally specific; explicit labeling of ordinal scores as analyst judgment is the correct disclosure. |
| **ER-8** DKK1 precedent segregated by patient population | **ADEQUATE** | Population mismatch now explicitly stated in Results R3 as a risk. Note: this adequate resolution made the DKK1 systemic safety gap more urgent (see Remaining Items). |

---

### Remaining Items — Minor Revision Required

The following items are **blocking** (manuscript cannot be accepted without them) or **strongly recommended before external submission**. All are achievable by in-silico query and/or text and figure editing; none requires wet-lab work.

**Blocking items:**

**1. GUCY2C EAC tumor-mRNA expression — close the gap the authors themselves identified** *(R1 by Reviewers 1, 2, and 3 in consensus)*
Query GUCY2C (and CDH17 as a comparator) mRNA expression in cBioPortal's *esca_tcga_pan_can_atlas_2018* dataset — the same cohort used for Figure 1a alteration landscape. Report median z-score or RSEM, percent of tumors with detectable expression, and range. If expression is low, heterogeneous, or absent, report that result and lower the treatment-arm confidence accordingly. This is not a new data-generation ask; it is a public query on a resource the paper already cites. Suppressing an unflattering result from the TCGA cohort already in use would be a more serious problem than reporting it candidly.

**2. DKK1 systemic safety in a surveillance population — follow through on what the GTEx data now expose** *(Reviewer 1, Major Concern 2; Reviewers 2 and 3 in agreement)*
The revision correctly added GTEx data for DKK1 and correctly segregated the DKN-01/DisTinGuish precedent. The result is that Table S4 now shows DKK1 at 852.8 TPM in cultured fibroblasts and 0.39 TPM at GEJ — data the paper added but did not interpret with respect to the interception-arm patient population. A surveillance Barrett's population is not an advanced-cancer population; the safety bar is fundamentally different. A substantive paragraph in the Discussion addressing (a) the DKK1 canonical bone/skeletal biology and published DKN-01 adverse-event profiles, (b) the implication of the fibroblast-dominant GTEx expression pattern for systemic exposure in otherwise-well patients, and (c) why this safety uncertainty is acknowledged as a threshold risk distinct from GI-tract on-target effects, is both achievable by literature synthesis and necessary to frame the interception arm honestly. Gate G5 should list skeletal/systemic toxicity explicitly alongside "bone-GI safety margin."

**3. Interface residue numbering — provide a mapping to canonical UniProt positions** *(Reviewers 2 and 3 in agreement)*
Methods and S4 report GUCY2C interface residues in complex-file numbering and DKK1 hotspots in isolated-domain numbering; neither is mapped to UniProt canonical coordinates. A reader cannot independently verify that the DKK1 binder engages CRD2 (UniProt O94907 residues 178–256) when the interface residues start at position 10 in complex-file numbering. A one-line offset table or in-text mapping (e.g., "complex-file residue = UniProt residue − N; residues [10–13, 28–31] correspond to UniProt positions [188–191, 206–209]") resolves this. This is a reproducibility requirement, not a stylistic preference.

**4. Figure 4 in-figure confidence notation** *(Reviewer 1, Major Concern 1)*
The figure caption correctly caveats that backbone traces are single-model, single-seed Boltz-2 predictions. However, a reader who sees the rendered complex images and glances at ipTM > 0.85 without reading the caption in full will draw an incorrect inference about structural certainty. Add an in-figure text annotation (e.g., "Single-model Boltz-2 / not consensus-validated") visibly in each panel. The legend should also note that the depicted structure is the highest-confidence of five Boltz-2 seeds rather than an ensemble or consensus. This is a figure-editing task.

**Strongly recommended before external journal submission (non-blocking for preprint acceptance):**

**5. Monte Carlo robustness — add a narrow grid test of the ×5 association premium** *(Reviewer 2, Major Concern 1)*
The current ±50% perturbation tests noise around the chosen point but does not test whether GUCY2C's rank-1 frequency changes if the association weight is reduced to parity with the other components. Running a grid over association weight ∈ {1, 3, 5, 7, 10} (other weights held at 1) and reporting GUCY2C's rank-1 frequency at each point would constitute genuine structural robustness rather than local-neighbourhood sensitivity. If GUCY2C leads at all association weights ≥ 3, that is a strong result; if it leads only when association weight ≥ 5, that should be stated.

**6. GUCY2C tumor_selectivity score — resolve internal inconsistency** *(Reviewer 2, Major Concern 2)*
The rubric anchor for score = 5 is "strong therapeutic window," but the manuscript simultaneously states that the GI safety window for GUCY2C is open question O-2 — i.e., unresolved. Recalibrate the score to 4 (or redefine the anchor for 5 as "tissue-restricted, window pending") and report the resulting composite. GUCY2C will almost certainly still lead the treatment lane; show it.

**7. Figure 1b legend — move the non-de-duplication caveat to the figure itself** *(All three reviewers)*
One sentence added to the figure legend: "x-axis values reflect non-de-duplicated active-EAC trial intervention mentions and should be read as relative crowding indicators; see Methods for caveats." No figure redrawing required.

**8. Composition liability screen — anchor against a published reference distribution** *(Reviewer 2, Major Concern 4)*
Cite a published %Ala distribution from the RFdiffusion/ProteinMPNN benchmark literature (e.g., Bennett et al. 2023 *Nature*, Cao et al. 2022 *Nature*) to ground the 17.5% vs. 43.7% comparison, or explicitly downgrade "clean" to "below the threshold for obvious alanine-scaffold artifacts, pending benchmarked calibration."

**9. Define the binder-diversity metric** *(Reviewer 3, Major Concern 4)*
Report pairwise sequence identity (or BLOSUM-based similarity) across all 16 designs, state the threshold used to call "no high-similarity pairs," and report the two most similar pairs numerically.

**10. DKK1 fibroblast artifact — contextualise in the table and note the in-vivo expression pattern** *(All three reviewers)*
Add a table footnote in S4 noting that the "Cells_Cultured_fibroblasts" value (852.8 TPM) is a non-physiological cell-culture artifact, that DKK1's in-vivo normal-tissue expression pattern is better captured by the remaining tissue rows, and that the biologically relevant DKK1 source in Barrett's tissue is likely stromal/paracrine — a framing that may actually support the neutralizing-trap rationale if made explicit.

---

### Judgment on ER-2: GPU-Dependent Orthogonal Fold-Back

The authors' decision to **scope rather than fabricate** orthogonal structural validation (AlphaFold-Multimer / Chai-1 consensus fold-back) is **acceptable for a preprint at this stage, and the handling editor concurs with all three reviewers on this point.** The authors correctly declined to generate numbers they did not have, lowered claim language throughout the manuscript, explicitly named the missing validation as open question O-1, and substituted the composition liability screen as a check they *could* responsibly run — which produced a genuine finding rather than cosmetic reassurance. That is the scientifically honest behavior this journal's scope rewards. The reviewers' uniform verdict that this scoping is appropriate for a preprint is adopted here. The absence of orthogonal fold-back is not grounds for rejection or for demanding another Major Revision cycle; it should be named prominently in the preprint's cover note as the next experimental (or high-compute) step for any group that wishes to take either lead into a wet-lab program.

---

### Path to Acceptance

The authors should prepare a revised manuscript addressing the four blocking items above, with a brief point-by-point response letter. For items 5–10 (non-blocking), a statement of intent (or completion) in the response letter is sufficient; the editor will not require a further formal review round if the blocking items are clearly resolved. The response letter should, for each blocking item, either show the result (e.g., the GUCY2C TCGA mRNA query output, the DKK1 safety discussion paragraph, the residue-mapping table, the Figure 4 annotation) or, if a result is unflattering, show it anyway — candor has been this manuscript's principal strength across both rounds of revision and should not be abandoned at the finish line.

If the four blocking items are resolved without introducing new unaddressed data exhibits, the manuscript will be accepted for posting on the preprint track. Authors should note that items 5–10 remain open and will be expected to be resolved before submission to any indexing journal.

The editor thanks all three reviewers for the rigor and specificity of their round-2 assessment, and thanks the authors for a revision that demonstrated what responsible, scope-constrained iteration on a computational preprint looks like.

---

*[Handling Editor signature block]*

---

# Full Reviews (Round 2)

## Reviewer 1 (Round 2)

### Reviewer 1: Clinician-scientist (GI oncology / translational EAC biology; 15+ years of Barrett's surveillance and GEJ-cancer clinical trial experience)

---

## Manuscript: "Steering into competitive whitespace" — EAC dual-arm in-silico discovery-to-design campaign (v2, revised)

---

### Brief Summary

This revised preprint presents a fully computational, public-data-driven campaign to nominate and pursue two differentiated therapeutic targets in EAC — GUCY2C (T-cell-engager treatment arm) and DKK1 (neutralizing-trap interception arm) — spanning the Barrett's-to-carcinoma disease trajectory. The authors mined TCGA genomics, Open Targets associations, ClinicalTrials.gov landscape data, and GTEx expression profiles to motivate target selection, then applied an explicit, now-disclosed druggability triage and generated de novo protein binders via RFdiffusion → ProteinMPNN → Boltz-2 fold-back. The revision meaningfully strengthens the first submission by adding a GTEx-based selectivity analysis with appropriately caveated interpretation, a Monte Carlo weight-perturbation robustness test of the triage rubric, a sequence-composition liability screen that reverses the naïve ipTM-ranking of leads, and a clearly segregated DKK1 precedent acknowledgement. The honest, unvarnished admission that the top GUCY2C binder has a high-alanine developability liability — and that the DKK1 lead is the more credible starting point — is the kind of author candor that a competitive whitespace preprint needs. Several problems from round 1 remain incompletely resolved, one new issue is introduced, and a few biological framing points still need tightening, but the manuscript is substantially improved.

---

### Strengths

- **Honest, internally consistent claim calibration throughout.** The abstract now reads like a methods-section summary rather than a promotional document: "ipTM is a single-model docking-confidence proxy, not a measured affinity; these are pilot-scale, in-silico leads, not experimentally tested therapeutics." The Discussion's limitations section is comprehensive and candid.
- **Composition liability screen (ER-3) is the standout new contribution.** The finding that gucy2c_bb2's headline ipTM 0.915 co-occurs with 44% alanine and an 8-residue homopolymer run — and that this reverses the single-metric ranking in favor of dkk1_bb1 — is both scientifically important and a demonstration that the authors are willing to undercut their own headlineable result when the data warrant it.
- **Monte Carlo rubric robustness (ER-4)** is well-executed: 20,000 draws, ±50% weight perturbation, results segregated by lane, and the DKK1 trivially-rank-1 result is flagged as reflecting a "thin addressable space" rather than a contested victory. That disclosure is appropriate and rare.
- **GTEx selectivity panel (ER-1)** is appropriately contextualized: the split between normal-tissue selectivity (what GTEx supports) and EAC tumor-cell expression (what it does not) is now clearly stated in Results, Discussion, and the figure legend.
- **DKK1 precedent segregation (ER-8)** is handled well. The population mismatch — advanced GEJ disease vs. pre-malignant Barrett's — is now stated explicitly in Results R3 as a risk rather than collapsed into a single favorable citation.
- **AI-disclosure (ER-7)** is now specific, structured, and honest about which outputs are analyst judgment vs. deterministic computation. The three-category taxonomy is a useful model.
- **Pilot-scale reframing (ER-5)** is applied consistently through Abstract, Results, and the legend. The phrase "n=8/target is too small to estimate a rate" is exactly the right language.

---

### Major Concerns

**1. ER-2 (orthogonal structural validation) — the scoping is appropriate but the ipTM values are still interpreted without sufficient hedge in the main-text visual narrative.**
The authors correctly decline to fabricate AlphaFold-Multimer/Chai-1 numbers they have not run, and I find that stance acceptable for a preprint. The problem is that Figure 4 is captioned with "Backbone traces depict topology and interface localization, not experimentally-determined secondary structure; ipTM is a docking-confidence proxy, not a measured affinity" — but the figure panels themselves show what look like cleanly folded, precisely docked complexes rendered in full-detail backbone trace. A reader who does not read every caption caveat will see two polished molecular structures with ipTM > 0.85 and draw exactly the wrong inference. The authors should add a visible, in-figure notation (e.g., a superimposed text box or a distinctive border) that marks these as single-model, non-consensus predictions, or rephrase the legend to state explicitly: "These are single-seed Boltz-2 predictions without consensus validation; visual precision should not be read as structural accuracy." This is achievable by editing the figure legend and/or the figure itself.

*Severity:* Moderate — the text is correct, but the figure communicates something the text walks back. In a preprint that will be read by molecular biologists who skim figures, the mismatch matters.

**2. DKK1 normal-tissue expression profile is alarming and underinterpreted.**
Table S4 shows DKK1's top GTEx tissue as "Cells_Cultured_fibroblasts" at 852.8 TPM with a GI fraction of exactly 0.00 — meaning DKK1 has essentially no GI-tissue restriction in normal tissue profiling. Esophagus mucosa is 1.19 TPM and GEJ is 0.39 TPM. The manuscript is intended to place DKK1 in an *interception arm for a largely pre-malignant Barrett's population* — a setting that sets, as the authors correctly acknowledge, a very high safety bar. Yet the normal-tissue data show that DKK1 neutralization would affect a protein whose highest normal expression is in fibroblasts, with potentially relevant expression in bone (DKK1 is a canonical bone-metabolism regulator), skin, and kidney. The DKK1 preclinical safety literature documents skeletal toxicity (heterotopic ossification, bone density effects) in DKK1-pathway interference models. The manuscript mentions "bone-GI safety margin" briefly in the Gate G5 criterion table and lists O-4 as the regulatory/surrogate-endpoint question — but there is no substantive discussion of the systemic (non-GI) safety risk of DKK1 neutralization in a surveillance population of patients who are otherwise well. This is not a request for wet-lab toxicology; it is a request for a paragraph in the Discussion that honestly characterizes the DKK1 safety landscape for the interception-arm patient population, drawing on the existing clinical and preclinical literature for DKK1 inhibitors. The DKN-01 trial data and the anti-DKK1 preclinical literature have published adverse-event profiles that the authors can cite.

*Severity:* Major — the central clinical framing of the DKK1 interception arm rests on the assumption that DKK1 neutralization is tolerable in an essentially well, surveillance population. The manuscript does not address this anywhere substantively, and the GTEx data the authors added make it more urgent, not less.

**3. GUCY2C EAC tumor-cell expression gap is flagged but its severity is understated.**
The authors commendably note that GTEx establishes *normal-tissue* selectivity but not EAC *tumor-cell* expression. However, the GUCY2C EAC tumor-expression question is not merely a gap to fill in future studies — it is a threshold criterion for the entire treatment-arm rationale. The EAC literature is not silent on this: there are IHC studies in esophageal/GEJ cancer (some negative, some reporting low or heterogeneous expression) and RNA data in TCGA ESCA that the authors could have examined but apparently have not. cBioPortal provides TCGA mRNA expression data for *esca_tcga_pan_can_atlas_2018* (the same cohort used for the alteration frequency analysis in Figure 1a). Checking GUCY2C mRNA expression in that dataset — even a simple median-expression-and-percent-detectable query — would directly address whether the target is expressed in EAC tumors at levels comparable to the CRC setting where clinical precedent exists. This is achievable in-silico (it requires a cBioPortal query, not a wet-lab experiment). The current manuscript uses the same cohort for alteration frequency but apparently did not retrieve expression data from it. I consider this a feasible, mandatory revision: query GUCY2C (and CDH17 as a comparator) mRNA expression in *esca_tcga_pan_can_atlas_2018* and report the result, even if it is unflattering.

*Severity:* Major — the selectivity argument is structurally sound, but the treatment-arm rationale requires tumor-expressed antigen. The data exist in the already-used public source and the authors chose not to retrieve them.

---

### Minor Concerns

**4. The DKK1 composite score disease_association (Barrett's) is 0.074, while GPX7 is 0.372 and REG4 is 0.094.** DKK1 has among the weakest Open Targets Barrett's association scores in the interception lane. The Monte Carlo robustness check cannot rescue this when DKK1 is literally the only biologic-addressable interception candidate after the filter — the filter is doing all the work, not the score. The text acknowledges this ("trivial"), but the Figure 2 visualization and the Abstract still imply that DKK1 "emerged" from a competitive field. The Abstract should say explicitly that DKK1 was selected because it is the only biologic-addressable interception candidate, not because it out-scored others.

**5. The interface patch numbering in S4 uses "complex-file numbering" for DKK1 (residues [10,11,12,13,28,29,30,31]) — these are clearly the isolated-domain residue numbers, not the full-protein UniProt numbering.** The DKK1 CRD2 domain is described as residues 178–256 of the full protein; residue 10 in the isolated domain is approximately residue 188 in the full protein. The manuscript should either provide a mapping table or report numbers in full-protein coordinates to allow independent structural interpretation.

**6. Figure S1 legend notes DKK1's top tissue as "Cells_Cultured_fibroblasts" at 852.8 TPM** — a cell-culture artifact rather than a tissue measurement. This value likely reflects culture-condition upregulation and is not directly comparable to in-vivo tissue TPM values. The authors should note in the S4/S1 context that fibroblast culture values are not an in-vivo tissue profile and that DKK1's in-vivo normal-tissue expression pattern is better captured by the remaining tissues.

**7. The Methods statement "The per-axis 'active EAC trial intervention mentions' … were not de-duplicated across combination trials"** is appropriately caveated, but Figure 1b remains a log-scale scatter plot with specific numerical x-axis positions that will be read as exact counts. The figure legend should note the non-de-duplicated nature of the counts and the missing query date.

**8. The CEACAM5 esophagus_mucosa_TPM is 293.4 in Table S4** — by far the highest esophageal mucosa expression of any comparator — but the manuscript is silent on what this means for CEACAM5 as a potential comparator target in EAC. This is a minor point that need not be expanded, but the figure legend should note that CEACAM5's high esophageal mucosa expression reflects normal esophageal expression (CEA is a known normal GI-tract antigen), not a favorable therapeutic window.

---

### Assessment of Each Essential Revision

| ER | Verdict | Notes |
|---|---|---|
| **ER-1** (GUCY2C selectivity/expression from public data) | **PARTIALLY ADEQUATE** | GTEx normal-tissue data added and appropriately caveated. However, TCGA *esca_tcga_pan_can_atlas_2018* mRNA expression data — from the same cohort already used — were not retrieved. This is a consequential omission that this round-2 review escalates to Major Concern 3. |
| **ER-2** (orthogonal structural validation) | **ADEQUATE (for preprint scope)** | Authors correctly decline to fabricate consensus numbers unavailable without GPU re-run; claim language is lowered throughout. The figure-text mismatch (Major Concern 1) is a presentation problem introduced by the revision, not a scientific fabrication. |
| **ER-3** (negative/specificity control for binder designs) | **ADEQUATE** | Composition liability screen is rigorous, directly addresses the predicted alanine-scaffold failure mode, and the reversal of the gucy2c_bb2 ranking is the intellectually honest result. Monomer/decoy controls named as O-1 remaining work. |
| **ER-4** (disclose + sensitivity-test triage rubric) | **ADEQUATE** | Exact formula disclosed, all 16 scores reproduced within 0.005, Monte Carlo robustness reported with appropriate lane-segregation. DKK1 trivial-rank disclosure is laudable. |
| **ER-5** (reframe "15/16 pass" as pilot-scale) | **ADEQUATE** | Consistently applied; language in Abstract, Results, and legends is appropriately hedged. |
| **ER-6** (trial-count methodology + query date) | **PARTIALLY ADEQUATE** | The non-de-duplication caveat and crowding-indicator framing are helpful. The missing query strings and snapshot date are acknowledged as a reproducibility gap. This is acceptable for a preprint but must be resolved before journal submission, as stated. Minor Concern 7 persists regarding Figure 1b axis labeling. |
| **ER-7** (specific non-placeholder AI disclosure) | **ADEQUATE** | The three-category taxonomy is clear and specific; the identification of ordinal scores and "whitespace verdict" as analyst judgment is exactly right. |
| **ER-8** (segregate DKK1 precedent by patient population) | **ADEQUATE** | The population mismatch is now stated explicitly in Results R3. However, this segregation makes Major Concern 2 (DKK1 normal-tissue safety in a surveillance population) more urgent, not less — the revision correctly exposed an issue that requires follow-through discussion. |

**New problems introduced by the revision:**
- Table S4 DKK1 fibroblast-culture artifact (Minor Concern 6) — not present in v1 because the table did not exist.
- Figure 4 visual-confidence mismatch (Major Concern 1) — the added render makes the single-model caveat more important to surface in the figure itself.
- The DKK1 GI-fraction = 0.00 finding now visible in Table S4 makes the interception-arm safety framing materially weaker than stated (Major Concern 2), an issue the revision created by honestly reporting the GTEx data without following through on their implication.

---

### Feasible Revisions

1. **[Major Concern 3 — achievable in-silico]** Query GUCY2C (and CDH17 as a comparator) mRNA expression in the TCGA *esca_tcga_pan_can_atlas_2018* cohort via cBioPortal. Report median z-score or RSEM, percent of tumors with detectable expression (e.g., z-score > 0), and range. Add a single panel or supplementary table. If expression is low or variable, say so and adjust the treatment-arm confidence accordingly — do not suppress an unflattering result.

2. **[Major Concern 2 — achievable by manuscript editing]** Add a paragraph to the Discussion on DKK1 systemic safety in the interception-arm population, citing published DKN-01 and anti-DKK1 preclinical adverse-event data. Address specifically: (a) DKK1's canonical role in bone homeostasis and the skeletal safety risk; (b) the implication of the GTEx fibroblast-dominant expression pattern for systemic (non-GI) exposure in otherwise-well Barrett's patients; (c) why the safety bar for Barrett's interception differs fundamentally from the advanced-cancer bar of the DisTinGuish trial. Gate G5 should list "skeletal/systemic toxicity" alongside "bone-GI safety margin."

3. **[Major Concern 1 — achievable by figure editing]** Add an in-figure text annotation to Figure 4 panels making explicit that these are single-model, single-seed Boltz-2 predictions without cross-model consensus validation. A bold-text inset ("Single-model Boltz-2; not consensus-validated") is sufficient. Update the legend to match.

4. **[Minor Concern 4 — achievable by editing one sentence]** Revise the Abstract to state explicitly that DKK1 was selected as the only biologic-addressable interception candidate after the druggability filter, not as the highest composite scorer in a competitive field.

5. **[Minor Concern 5 — achievable by editing S4]** Add a coordinate-mapping footnote to Section S4 converting the DKK1 complex-file interface residue numbers [10–13, 28–31] to full-protein UniProt O94907 coordinates (~188–191, ~206–209).

6. **[Minor Concern 6 — achievable by editing S1/S4]** Add a note in Table S4 and the Figure S1 legend that the DKK1 "Cells_Cultured_fibroblasts" value (852.8 TPM) is a cell-culture artifact not directly comparable to in-vivo tissue expression, and that DKK1 in-vivo normal-tissue expression is best assessed from the remaining tissue rows.

7. **[Minor Concern 7 — achievable by editing the figure legend]** Add to the Figure 1b legend: "x-axis positions reflect non-de-duplicated axis-level active-trial mentions; query date not recorded; values should be read as relative crowding indicators, not exact disjoint counts."

---

### Recommendation

**Minor revision.**

The revision substantially addressed five of the eight essential revisions to an adequate or better standard, and the honest composition-liability finding that demotes the headline GUCY2C binder is a genuine scientific contribution. Three problems remain — a consequential omission (TCGA EAC mRNA expression for GUCY2C), a clinical-framing gap opened by the authors' own GTEx data (DKK1 systemic safety in a surveillance population), and a figure-text mismatch in Figure 4 — but all three are achievable by in-silico query and manuscript editing. None requires wet-lab work. The manuscript should not be accepted without these three revisions; it does not require another full major-revision cycle if the authors respond directly and completely to them.

---

### Cross-Cutting Comment

The revision demonstrates that the authors are capable of honest self-correction under critical pressure — the alanine-scaffold finding in particular is the kind of uncomfortable result that many authors bury, and its elevation to a key finding reflects well on the transparency commitment. The two remaining biological framing gaps — GUCY2C EAC tumor-cell expression and DKK1 systemic safety in a surveillance population — are not minor housekeeping items: they bear directly on whether each arm's lead rationale is credible enough to motivate a wet-lab follow-on program, which is the preprint's stated purpose. The campaign is most useful to the field if its known weaknesses are as visible as its strengths. Closing these two gaps with the data that are already publicly available (TCGA expression, published DKN-01 safety data) would make the revised manuscript a genuinely rigorous and reproducible template, rather than a well-caveated aspiration.

---

## Reviewer 2 (Round 2)

### Reviewer 2: Computational biologist & biostatistician — statistical rigor, methods transparency, prediction-tool validity, and reproducibility

---

## Manuscript: "Steering into competitive whitespace — a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory"

---

### Brief Summary

The revised manuscript presents a human-supervised, AI-assisted computational campaign that (i) characterizes EAC as a copy-number-driven disease using TCGA data, (ii) maps competitive whitespace among 351 EAC trials and an Open Targets known-drug set, (iii) applies a disclosed and now robustness-tested additive triage rubric to nominate GUCY2C and DKK1 across two trajectory positions, and (iv) delivers eight de novo binder designs per target via RFdiffusion→ProteinMPNN→Boltz-2, evaluated by ipTM/pLDDT and a new sequence-composition liability screen. The revision adds a GTEx normal-tissue selectivity analysis for GUCY2C, fully discloses the triage rubric with Monte-Carlo sensitivity testing, introduces a composition-based specificity proxy, reframes the pass-rate claim as pilot-scale, partially addresses the trial-count methodology gap, delivers a three-category AI-assistance disclosure, and segregates DKK1 clinical precedent by patient population. Overall claim strength has been deliberately lowered to match the evidence.

---

### Strengths

- **Honest, layered self-assessment.** The revision does not paper over what was not done: single-model Boltz-2 confidence, missing orthogonal fold-back, unrecorded query dates, thin interception-lane addressable space. These admissions are integrated into the text rather than buried in footnotes, and the response letter is consistent with the manuscript.
- **Triage rubric disclosure and robustness (ER-4).** The exact formula is now published with sufficient detail for independent reproduction; the Monte-Carlo perturbation (20,000 draws, ±50% on all weights) is a creditable sensitivity test, and the candid disclosure that DKK1's 100% rank-1 rate is trivially forced by being the sole addressable interception candidate is scientifically exemplary.
- **Composition liability screen (ER-3).** Adding Table S5 was the correct move. It immediately identified that the headline GUCY2C binder (gucy2c_bb2, ipTM 0.915) is 43.7% alanine with an 8-residue homopolymer run, directly confirming the low-complexity scaffold failure mode reviewers anticipated. The paper now correctly demotes this to "provisional pending redesign" and elevates dkk1_bb1 (17.5% Ala, max run 2) as the most robust lead. This is a genuine new finding, not a cosmetic fix.
- **DKK1 population-mismatch segregation (ER-8).** The DKN-01 precedent is now cleanly separated by population — advanced GEJ/gastric treatment vs. pre-malignant Barrett's interception — and the gap is explicitly stated as risk O-4. This is exactly the correction that was requested.
- **GUCY2C selectivity framing (ER-1).** The GTEx S4/S1 addition successfully distinguishes tissue selectivity (supportable: 85% GI-restricted, top non-GI 1.7 TPM) from tumor expression (explicitly not established from these data, correctly flagged as a gap). The polarity-not-absence framing for luminal restriction is appropriately cautious.
- **Pilot-scale reframing (ER-5).** "Pilot-scale," "permissive bar," and "n=8 cannot support a generalizable success rate" now appear in the Abstract, Results, and legend — adequate.
- **Three-category AI disclosure (ER-7).** Separating deterministic bioinformatics, published ML methods as objects of study, and LLM-assisted synthesis/drafting — with explicit identification of the ordinal scores as analyst judgment — is a meaningful and usable disclosure, not a placeholder.

---

### Major Concerns

**1. The Monte-Carlo weight perturbation is methodologically misspecified in a way that potentially overstates GUCY2C's robustness.**

Table S0 specifies that all four weights are perturbed: tumor_selectivity, novelty, tractability, and disease_association. However, the association weight is varied over U(2.5, 7.5) while the other three are varied over U(0.5, 1.5) — i.e., a ×5 mean multiplied by a ±50% factor. This preserves the *ratio* of the association weight to the other three weights within a narrow band, rather than testing whether the dominance of the ×5 association term is itself driving the result. A multiplicatively-symmetric perturbation over, say, U(1, 10) for the association coefficient (or equivalently, asking how often GUCY2C leads if association is down-weighted to equal the other components) would constitute a genuine robustness check. The current design is a robustness check around the chosen point, not across the structural choice to weight association at 5×. The manuscript should either (a) run a second sensitivity scenario with association weight ∈ {1, 3, 5, 7, 10} on a grid and report how GUCY2C's rank-1 frequency varies, or (b) explicitly acknowledge that the U(2.5, 7.5) perturbation preserves the ×5 premium and that the robustness result applies only in the neighbourhood of the published rubric. Currently the text implies more sweeping robustness than the analysis demonstrates.

**2. The GUCY2C tumor_selectivity score of 5 is retained despite the GTEx data now demonstrating substantial normal-gut expression (~28 TPM in small intestine) — an internal inconsistency with the rubric anchor.**

Table S0 defines tumor_selectivity = 5 as "tissue-restricted with a strong therapeutic window." Yet the manuscript simultaneously states that GUCY2C is "genuinely expressed in normal small intestine and colon (~28 TPM median)" and that the normal-GI safety window is "precisely open question O-2" — i.e., the therapeutic window is *not established*. A score of 5 on a scale where 5 = strong therapeutic window is therefore inconsistent with the paper's own caveat that the window is unresolved. A score of 3–4 ("moderate window / some normal expression") would be more defensible given the GTEx data. The authors should either (a) recalibrate the score with a brief justification and re-report the affected composite (GUCY2C composite would drop to ~11–12, which likely still leads the treatment lane but should be shown), or (b) define selectivity = 5 as "tissue-restricted" (not "strong window") and state the window assessment is deferred — and make sure the rubric anchor matches the score for all 16 candidates under the same logic.

**3. The trial-count methodology gap (ER-6) remains inadequately resolved for a submission claiming quantitative competitive-landscape analysis.**

The Methods now state that per-axis trial mentions "were not de-duplicated across combination trials and therefore do not sum to the 106 active-trial total" and that query strings and snapshot dates were not recorded. This is honest, but it means the x-axis of Figure 1b — the primary evidence for competitive crowding, which drives the central "whitespace" argument — is neither reproducible nor de-duplicated. The paper flags this as "a reproducibility gap to close before external submission." For a preprint that will presumably circulate and be cited, "gap to close before external submission" is insufficient. The minimum requirement is: provide the raw trial-ID list (even approximately reconstructed) or a ClinicalTrials.gov API query string with the approximate date, and label the axis unambiguously as "approximate, non-deduplicated axis-level mentions" both in the legend and in the figure caption. The current legend says "active EAC trial intervention mentions (x, log scale)" without qualification, so a reader who does not read the Methods text carefully will not know the values are approximate and non-deduplicated. This is a presentation/reproducibility defect in the primary evidence figure.

**4. The composition liability screen (Table S5) does not include a calibrated threshold or a reference distribution, so it cannot be used to rule designs in or out — it only flags the worst cases visually.**

The screen computes %Ala, %hydrophobic, max homopolymer run, and frac_charged, and correctly identifies gucy2c_bb2 as a likely generic scaffold. But no threshold is applied (e.g., %Ala > X% → fails), no reference distribution from the literature (e.g., known binders designed with RFdiffusion/ProteinMPNN from published benchmarks) is provided, and the "clean" verdict for dkk1_bb1 (17.5% Ala) is asserted but not formally justified against a baseline. Without a calibrated cutoff, the screen is qualitatively useful but does not constitute a specificity filter. The paper should either (a) cite a published threshold or distribution (%Ala in published RFdiffusion/ProteinMPNN binder sets, e.g. from Bennett et al. 2023 *Nature* or Cao et al. 2022 *Nature*) and anchor the 17.5% vs. 43.7% comparison against it, or (b) explicitly state that the screen is a qualitative flag, not a quantitative filter, and that its interpretation is provisional pending orthogonal fold-back.

---

### Minor Concerns

**1. DKK1 esophagus_gej_TPM = 0.39 in Table S4 is neither discussed nor contextualized.** The paper uses GTEx to argue for GUCY2C GI restriction, but DKK1's top tissue in GTEx is "Cells_Cultured_fibroblasts" at 852.8 TPM — a completely non-physiological ex-vivo cell culture context — while its esophageal/GEJ expression is 0.39–1.19 TPM. The main text describes DKK1 as a "secreted Wnt-pathway modulator" but does not note that its dominant GTEx signal is in fibroblasts, which raises a question about whether the biologically relevant source of serum/lumenal DKK1 in Barrett's tissue is stromal rather than epithelial — a detail that actually strengthens the neutralizing-trap rationale (serum-accessible target) but should be stated rather than passed over silently in Table S4.

**2. hotspot residue numbering is not cross-walked to UniProt canonical positions.** Methods and S4 report GUCY2C interface residues as [130,132,146,192,325,326,327,330,347,348,349,378,379] and DKK1 as [10,11,12,13,28,29,30,31] "in complex-file numbering," while the design hotspots use "isolated-domain numbering." A one-line offset table (e.g., complex-file residue = UniProt residue − N for GUCY2C ECD, given chain starts at residue 24) would allow any reader to map these to canonical coordinates for comparison with published functional sites — otherwise the epitope localization claim in Figure 4 is not independently checkable.

**3. The MPNN score is reported in Table S1 but its units and directionality are not defined in the table or the legend.** ProteinMPNN outputs a log-likelihood / negative log-likelihood score depending on implementation; "1.232" for dkk1_bb1 and "0.95" for gucy2c_bb2 should be labeled (lower = better or higher = better; sign convention) so readers can interpret Table S1 without reading the ProteinMPNN paper.

**4. Figure 4 caption says "Backbone traces depict topology and interface localization, not experimentally-determined secondary structure" — this caveat is good, but the legend should also note that the depicted structure is the highest-confidence of five Boltz-2 samples rather than an ensemble-averaged or consensus model.** A single-sample structure depicted without that label may give a false sense of structural certainty.

**5. The Methods state binder length ranges of 70–100 aa (GUCY2C) and 60–90 aa (DKK1), yet dkk1_bb1 is 63 aa.** The text should either correct the range (≤60 aa designs were accepted) or note that the lower bound was relaxed, and confirm this was by design rather than a reporting error.

---

### Feasible Revisions

1. **[ER-4 / Major Concern 1]** Run a grid sensitivity analysis varying the association weight coefficient across {1, 3, 5, 7, 10} (all other weights held at 1, then separately varied), and add a 2–3 sentence paragraph in the Methods/Table S0 footnote reporting how GUCY2C's rank-1 frequency changes as the association premium varies. If GUCY2C leads at coefficient ≥ 3, this is genuine robustness; if it leads only when association weight ≥ 5, that should be stated.

2. **[Major Concern 2]** Recalibrate the tumor_selectivity score for GUCY2C from 5 to 4 (or redefine the rubric anchor for 5 as "tissue-restricted" rather than "strong therapeutic window"), report the resulting composite (likely ~12 rather than 13), and confirm that GUCY2C still leads the treatment lane. Add one sentence noting the internal consistency between the score and the therapeutic-window caveats.

3. **[ER-6 / Major Concern 3]** Add an in-text qualifier to Figure 1b's axis label and caption (e.g., "approximate, non-deduplicated active-EAC trial intervention mentions, x-axis; see Methods for caveats") and provide, as a supplement, either a reconstructed list of the trial IDs queried or the ClinicalTrials.gov API query parameters with best-available date, so the figure is at minimum partially reproducible.

4. **[Major Concern 4]** Anchor the composition liability screen to a published reference distribution or threshold from the RFdiffusion/ProteinMPNN literature, or explicitly downgrade the "dkk1_bb1 is clean" language to "dkk1_bb1 is below the threshold for obvious alanine-scaffold artifacts, pending benchmarked calibration."

5. **[Minor Concern 1]** Add 1–2 sentences noting DKK1's dominant GTEx signal in cultured fibroblasts, clarify that this is a non-physiological artifact of the tissue panel, and note that the relevant DKK1 source in Barrett's tissue is likely stromal/paracrine — a framing that supports rather than undermines the neutralizing-trap rationale.

6. **[Minor Concern 2]** Add a one-line offset table or in-text mapping from complex-file residue numbers to UniProt canonical positions for both targets, so the interface patches in Figure 4 can be compared with published structure–function data.

7. **[Minor Concern 3]** Define the MPNN score in Table S1's legend (e.g., "MPNN log-likelihood; lower = better design confidence" or equivalent), including sign convention.

8. **[Minor Concern 5]** Reconcile the stated DKK1 length range (60–90 aa) with dkk1_bb1's 63 aa length — either correct the range to ≥60 aa or add a note that the lower bound was a flexible constraint.

---

### Assessment of Each Essential Revision

**ER-1 (GUCY2C EAC selectivity/expression from public data):** **ADEQUATE with one residual inconsistency (see Major Concern 2).** The GTEx analysis is the right data source, the selectivity quantification is clear, and the normal-tissue vs. tumor-expression distinction is correctly drawn. The only problem introduced is an internal inconsistency: the tumor_selectivity score of 5 (= "strong therapeutic window") is now in tension with the paper's own statement that the window is an open question.

**ER-2 (Orthogonal structural validation):** **PARTIALLY ADEQUATE — honestly scoped.** The authors correctly declined to fabricate consensus numbers, lowered claim strength throughout, and named orthogonal fold-back + monomer/decoy controls as the not-yet-run O-1 step. For a preprint at this stage, this scoping is acceptable; the paper does not over-claim. No new problem introduced.

**ER-3 (Negative/specificity control):** **ADEQUATE as a first-order proxy, with the caveat in Major Concern 4.** Table S5 and its integration into the Results and Discussion are a genuine improvement. The alanine-content finding for gucy2c_bb2 is important and correctly reported. The remaining gap is the absence of a calibrated threshold, but the paper's treatment of this as qualitative is defensible.

**ER-4 (Triage rubric disclosure + sensitivity test):** **PARTIALLY ADEQUATE.** The rubric is now fully disclosed and reproducible; this was the most important part of the ask and is done well. The Monte-Carlo test is creditable but the perturbation design preserves the relative advantage of the ×5 association weight rather than testing it structurally (see Major Concern 1). A minor additional analysis would make this fully adequate.

**ER-5 ("15/16 pass" reframed as pilot-scale):** **ADEQUATE.** The reframing is complete and consistently applied across Abstract, Results, and legend.

**ER-6 (Trial-count methodology + query date):** **PARTIALLY ADEQUATE — honest about the gap, but the gap is too large for the primary evidence figure.** The Methods acknowledgment is correct. The deficiency is that the Figure 1b axis label and caption are still presented without qualification, so the non-deduplicated, unanchored nature of the counts is not visible at the figure level. This is a presentation fix, not a data-generation problem (see Feasible Revision 3).

**ER-7 (Non-placeholder AI disclosure):** **ADEQUATE.** The three-category structure is specific, useful, and appropriately distinguishes deterministic computation from model judgment.

**ER-8 (DKK1 precedent segregated by patient population):** **ADEQUATE.** The population mismatch is now explicitly stated in Results R3, named as risk O-4, and the mechanistic/druggability rationale is correctly offered as the only legitimate justification for carrying DKK1 into the interception arm.

---

### Recommendation

**Minor revision.**

The revision is substantive and honest: it lowered claim strength to match evidence, produced a genuinely new finding (alanine-scaffold liability in the top GUCY2C binder), fully disclosed the triage rubric, and correctly segregated the DKK1 precedent. Two issues prevent immediate acceptance: the tumor_selectivity = 5 score for GUCY2C is now internally inconsistent with the paper's own caveats about the unresolved therapeutic window (Major Concern 2, resolvable with a one-line score adjustment), and the Figure 1b axis — the central whitespace evidence — lacks in-figure qualification of its non-deduplicated, date-unanchored nature (Major Concern 3, resolvable with a caption edit and a supplementary query log). Major Concern 1 (Monte-Carlo design) and Major Concern 4 (composition threshold calibration) are methodological improvements that would materially strengthen the paper but can be addressed with small additional analyses or explicit scope-limitation statements. None of the required revisions involves wet-lab work or new data generation.

---

### Cross-Cutting Comment

This revision demonstrates what honest iteration under scope constraints looks like: the authors declined to fabricate orthogonal fold-back numbers, instead lowering claim strength and making the absent validation conspicuous and named, and they used the revision cycle to surface a new finding (the gucy2c_bb2 alanine liability) that actually *reduces* confidence in their most dramatic result rather than inflating it. That is the right scientific behavior and should be recognized. The manuscript's principal remaining statistical vulnerability is the triage rubric's sensitivity analysis, which tests noise around the ×5 association premium without questioning whether the premium itself is justified — a relatively minor fix. Taken as a package, the work is a legitimate, reproducible template for competition-aware computational target nomination and early binder generation in a disease where the field genuinely needs differentiated starting points, and the dual-lead output (one arm now stronger than the other on developability grounds, for honest reasons) is more credible after revision than before it.

---

## Reviewer 3 (Round 2)

### Reviewer 3: Methodology, research ethics, and responsible-AI lens — computational drug discovery

---

## Manuscript 1 (single manuscript): "Steering into competitive whitespace"

---

### Brief Summary

This revised preprint reports a human-supervised, AI-assisted computational campaign to nominate differentiated therapeutic targets for esophageal adenocarcinoma (EAC) and design de novo protein binders against them, working entirely from public data (TCGA/cBioPortal genomics, Open Targets, ClinicalTrials.gov, GTEx, AlphaFold DB). Two leads — GUCY2C for a T-cell-engager treatment arm and DKK1 for a neutralizing-trap interception arm — were selected under an explicit novelty steer and druggability filter; eight binder backbones per target were generated via RFdiffusion→ProteinMPNN→Boltz-2, with 15/16 designs clearing ipTM > 0.5. The revision adds GTEx selectivity data, a sequence-composition liability screen, a disclosed triage rubric with Monte Carlo robustness analysis, and a substantially more honest framing of both confidence levels and remaining unknowns.

---

### Strengths

- **Epistemic honesty materially improved.** The revision genuinely lowered claim strength to match evidence rather than patching cosmetically. The abstract now leads with the composition liability of the top-ipTM GUCY2C design; the pilot-scale framing is explicit; the DKK1 precedent population mismatch is stated as a risk rather than buried.
- **Triage rubric fully disclosed and reproducible.** Table S0 provides the exact formula, component values, ordinal anchors, and a 20,000-draw Monte Carlo sensitivity analysis. The candid acknowledgment that DKK1's rank-1 status is trivial (it is the only biologic-addressable interception candidate) is methodologically exemplary and rare in this kind of work.
- **Composition liability screen (ER-3).** The alanine-content and homopolymer analysis of all 16 ProteinMPNN sequences is a real finding — confirming the failure mode reviewers predicted — and it produced a genuine reranking: dkk1_bb1 is now correctly identified as the most robust lead on combined ipTM + developability grounds, while gucy2c_bb2's headline ipTM is appropriately qualified.
- **GTEx selectivity (ER-1).** Quantified normal-tissue expression with a proper comparator panel (CDH17 as a like-minded GI-restricted control; TROP2/TACSTD2 as a cautionary contrast). The 85% GI fraction and ~1.7 TPM top non-GI tissue are specific and traceable.
- **AI-disclosure structure (ER-7).** The three-category framework (deterministic bioinformatics / published ML tools / AI-assisted synthesis with human review) is the most substantive AI disclosure I have seen in a preprint of this type. Explicitly labeling the Table S2 ordinal scores and "whitespace verdict" column as analyst judgment, not measurement, is appropriate.
- **Open questions and gates remain prominently flagged.** O-1 through O-5 and G1–G5 are retained and specific; the Discussion limitation list is honest and reasonably complete.
- **DKK1 precedent segregation (ER-8).** The Klempner/DisTinGuish precedent is now explicitly scoped to advanced gastric/GEJ disease with a statement that it is not evidence for precursor-setting interception.

---

### Major Concerns

**1. The DKK1 GTEx profile is never explained and constitutes an unresolved embarrassment for the selectivity argument.**
Table S4 shows DKK1's top tissue is "Cells_Cultured_fibroblasts" at 852.8 TPM, with GI_fraction = 0.0. This is the opposite of selective GI expression; fibroblast culture artifact (serum-induced Wnt signaling in vitro) almost certainly explains the number, but it is never addressed in the text. The tumor_selectivity score of 3 for DKK1 is stated without justification anywhere in the manuscript. A secreted protein with 852.8 TPM in fibroblasts and effectively no GI enrichment requires explicit engagement: what tissues express DKK1 at physiologically relevant levels in vivo, and does that expression pattern create a therapeutic index problem for a neutralizing trap in a surveillance population? The GTEx data as presented actually deepens concern rather than alleviating it. This is a new problem introduced by ER-1 — the very analysis done to address selectivity created a data exhibit that undermines it for DKK1.

**2. The GUCY2C EAC tumor-expression gap remains unaddressed beyond acknowledgment.**
ER-1 is credited as "done," and the GTEx analysis is genuine. However, the authors' own ER-1 response distinguishes "establishes tissue selectivity" from "establishes EAC tumor-cell expression." For the treatment-arm rationale to be complete, tumor expression must be established, and cBioPortal itself contains RNA expression data (TCGA EAC cohort) that could be queried in silico — mRNA expression from the same n=182 cohort used for the alteration landscape. Not querying GUCY2C mRNA in the TCGA EAC RNA-seq data when that data is in the same public repository used for Figure 1a is an unexplained gap. This is not a wet-lab ask; it is a public-data query. The authors flag the gap; they should close it, or explicitly explain why cBioPortal/UCSC Xena GUCY2C mRNA data in EAC is not informative (e.g., low purity, relevant only for protein, etc.).

**3. Interface residue numbering is ambiguous and unreproducible.**
Methods/S4 reports GUCY2C interface residues as [130,132,146,192,325,326,327,330,347,348,349,378,379] and DKK1 residues as [10,11,12,13,28,29,30,31] in "complex-file numbering." The hotspot definition used as RFdiffusion input uses "isolated-domain numbering." These two numbering systems are never reconciled, and neither is mapped back to canonical UniProt residue positions. A reader cannot verify whether the binder engages the intended CRD2/LRP6-binding region of DKK1 (residues 178–256 UniProt O94907) when the reported interface residues start at position 10 in complex-file numbering. This is a reproducibility and interpretability failure that can be fixed by providing a numbering-offset table or by reporting all residues in UniProt coordinates throughout.

**4. The "no high-similarity binder pairs" diversity claim lacks a defined metric.**
Table S5/Section S5 states "no high-similarity binder pairs detected" with reference to `epitope_cross_reactivity.csv`. No threshold, alignment method, or similarity metric is stated — not sequence identity cutoff, not structural RMSD, not anything. The claim cannot be evaluated or reproduced. Given the high alanine content of several designs, pairwise sequence identity across the alanine-rich designs (dkk1_bb2/bb3/bb5, gucy2c_bb2/bb5) may be artifactually high or may be masked by the generic composition. This should be reported as a defined number with a defined method.

---

### Minor Concerns

**5. Table S4 column redundancy creates confusion.** The `top_nonGI_tissue` and `top_nonGI_TPM` columns are redundant with `top_tissue`/`top_median_TPM` for DKK1 (both point to fibroblasts). The table would be cleaner and more honest if it reported the top *in-vivo* tissue separately from cell-culture artifacts, with a footnote.

**6. The trial-count x-axis (Figure 1b) still lacks a defined unit.** The Methods now state these are "active-EAC-trial intervention mentions" that "were not de-duplicated across combination trials." However, Figure 1b's x-axis label in the legend reads "active-EAC-trial intervention mentions (x, log scale)" — if log scale is used, the statement that individual axis counts "do not sum to 106" should appear in the figure legend, not just Methods, to prevent misreading at the figure level.

**7. Monte Carlo robustness framing for DKK1 is technically correct but potentially misleading.** Table S0 correctly notes DKK1 ranks #1 in 100% of draws because it is the only addressable interception target. This is right to disclose, but the "Rank robustness" table header and column headers (% draws rank #1, nearest competitor) applied to a singleton give a false air of statistical validation. The footnote explanation is adequate but should also appear in the Methods paragraph on triage, not only in the table.

**8. The one-sentence abstract framing of composition liability buries the lede for the treatment arm.** The abstract states gucy2c_bb2 "carries high alanine content" and should be "treated as provisional." But it still leads with ipTM 0.915 before this qualification. Given that the composition analysis is a genuine new finding that reranks the leads, the abstract could foreground dkk1_bb1 as the cleaner lead first.

**9. Author list and affiliation remain placeholders.** This is a preprint and the placeholder is flagged, but for a manuscript that includes a specific AI-disclosure section crediting human supervisors, the identity of those supervisors is relevant to the disclosure's completeness.

---

### Feasible Revisions

**R1. Address the DKK1 GTEx fibroblast artifact explicitly.** Add one paragraph in Results or a footnote to Table S4 explaining that DKK1's top GTEx "tissue" is cell culture (fibroblasts, serum-inducible Wnt signaling artifact), report the top in-vivo tissue by median TPM, and reassess whether the tumor_selectivity score of 3 is defensible given the actual in-vivo normal-tissue profile. If DKK1 is broadly expressed in normal stromal/connective tissue, state the implications for a neutralizing trap in a surveillance population. This is a public-data reanalysis with no wet-lab component.

**R2. Query GUCY2C mRNA in the TCGA EAC cohort.** Retrieve mRNA expression values for GUCY2C from cBioPortal's mRNA expression data for *esca_tcga_pan_can_atlas_2018* (RNA-seq RSEM or z-score). Report median expression and the fraction of EAC tumors with detectable expression (>1 TPM or equivalent). If expression is low or absent, state that explicitly and lower the treatment-arm confidence accordingly. This closes the ER-1 gap the authors themselves identified without requiring any new data generation.

**R3. Provide a numbering reconciliation table for interface residues.** Add a short table (or footnote to S4) mapping isolated-domain numbering → complex-file numbering → UniProt canonical positions for both the GUCY2C and DKK1 hotspot/interface residue sets. Confirm explicitly that the DKK1 interface residues [10,11,12,13,28,29,30,31] in complex numbering correspond to CRD2 residues within 178–256 UniProt O94907.

**R4. Define the binder-diversity metric.** Report pairwise sequence identity (or BLOSUM-based similarity) across all 16 designs, state the threshold used to call "no high-similarity pairs," and report the two most similar pairs numerically. If high-alanine designs are near-identical, state this.

**R5. Move the trial-count non-summability note to the Figure 1b legend.** One sentence added to the figure legend is sufficient; no figure redrawing required.

**R6. Add a one-sentence note in the Methods triage section** flagging that the DKK1 Monte Carlo result is structurally trivial (single-candidate lane), cross-referencing the Table S0 footnote.

---

### Assessment of Each Essential Revision

| ER | Verdict | Notes |
|---|---|---|
| **ER-1** (GUCY2C selectivity/GTEx) | **PARTIALLY ADEQUATE** | GTEx data added and GUCY2C selectivity quantified correctly. The DKK1 fibroblast artifact is a new problem introduced by this analysis and is not addressed. The TCGA EAC tumor-mRNA query — an in-silico step the authors themselves flag as the remaining gap — was not performed. |
| **ER-2** (Orthogonal structural validation) | **ADEQUATE for a preprint, with appropriate scoping** | The authors honestly declined to fabricate GPU-dependent Chai-1/AF-Multimer fold-back numbers they did not run. The lowering of claim language, the explicit O-1 designation, and the addition of the composition screen as the check they *could* run is an acceptable and transparent scoping decision for a preprint. Reject-contingent on orthogonal fold-back would be an out-of-scope demand. |
| **ER-3** (Negative/specificity control) | **ADEQUATE** | The composition liability screen is a genuine contribution. It produced a real finding (gucy2c_bb2 reassigned from headline to provisional), and dkk1_bb1 is correctly identified as the cleaner lead. The remaining monomer/decoy fold-back controls are appropriately named as O-1. The binder-diversity claim lacks a defined metric (Minor Concern 4 / R4), which should be fixed. |
| **ER-4** (Triage rubric disclosure and sensitivity test) | **ADEQUATE** | The disclosed formula, component values, and Monte Carlo analysis are exemplary. The DKK1 singleton acknowledgment is appropriately candid. Minor: the Methods paragraph should cross-reference the DKK1 triviality note (R6). |
| **ER-5** (Pilot-scale reframing) | **ADEQUATE** | "Pilot-scale" language appears in Abstract, Results, and Figure 3 legend. The ipTM > 0.5 permissiveness and single-seed generation are now explicitly stated. |
| **ER-6** (Trial-count methodology and query date) | **PARTIALLY ADEQUATE** | The authors honestly acknowledge they could not recover query strings/snapshot dates from upstream artifacts. The non-summability caveat is appropriate. However, this caveat appears in Methods only, not in the figure legend where readers encounter the counts first (R5). The reproducibility gap acknowledgment is honest; the fix is incomplete. |
| **ER-7** (Non-placeholder AI disclosure) | **ADEQUATE** | The three-category framework is substantive and operationally specific. The explicit labeling of ordinal scores as analyst judgment is the right disclosure. This revision exceeds what most preprints provide. |
| **ER-8** (DKK1 precedent segregation by population) | **ADEQUATE** | The DisTinGuish/DKN-01 precedent is now clearly scoped to advanced GEJ disease. The language "not evidence for DKK1 interception in a largely pre-malignant Barrett's surveillance population" is direct and appropriately limits the inference. |

---

### Recommendation

**Minor revision.**

The revision is substantive, honest, and represents a meaningful improvement over the round-1 submission. The authors correctly declined to fabricate orthogonal fold-back numbers (ER-2), produced a real composition-screen finding that changed the scientific conclusion (ER-3), and delivered a disclosed, sensitivity-tested triage rubric (ER-4). Claim strength is now better calibrated to evidence throughout. Three issues require attention before acceptance: (1) the DKK1 fibroblast artifact in GTEx is a new unaddressed problem introduced by the ER-1 analysis that undermines the interception-arm selectivity argument as currently framed; (2) the GUCY2C EAC tumor-mRNA query from cBioPortal — the authors' own identified gap — is an achievable in-silico step that should be taken; (3) the interface residue numbering is ambiguous in a way that prevents independent verification of whether the binders engage the intended functional sites. All three are addressable by public-data queries and text/table edits with no GPU compute or wet-lab work required.

---

## Cross-Cutting Comment

The revision demonstrates that honest, incremental tightening of an in-silico preprint is both achievable and scientifically valuable: the composition liability screen alone is a genuine finding that changed which design is the lead, and the triage transparency sets a standard that most computational drug-discovery papers do not meet. The one systemic risk that carries through both rounds is that analyses added to address one concern can introduce new problems if not fully interpreted — the DKK1 GTEx fibroblast result being the clearest example here. Future revisions should treat every new data exhibit as requiring the same level of critical engagement as the original claims. The trial-count reproducibility gap (ER-6) remains the most straightforward outstanding item and is the lowest-effort fix on the list; it should be resolved before any submission to a journal that requires methodological reproducibility as a condition of publication.


---

# AUTHORS' RESPONSE TO ROUND 2

# Point-by-Point Response to Round-2 Reviews (v2 → v3)

**Decision:** Minor revision (unanimous; all three reviewers moved up from Major). The editor named four blocking items and six recommended items. Below, each is answered. All changes were made in silico or by editing; no wet-lab work was required.

## Blocking items

**B-1 — GUCY2C EAC tumor-mRNA expression from the esca_tcga_pan_can_atlas_2018 cohort.**
Attempted; honestly scoped. The cBioPortal access available to us exposes mutation, copy-number, and clinical endpoints but **no per-sample mRNA-expression retrieval**, so we could not pull GUCY2C tumor mRNA from that cohort without fabricating values. We did not invent them. Instead we (a) sharpened the flag in Results R3 to name this as the single highest-priority follow-up query and (b) state explicitly that the treatment-arm rationale is provisional until GUCY2C EAC tumor expression is reported. This preserves candor over a manufactured result; the query should be run against a portal build with an mRNA endpoint (or the GDC) before external submission.

**B-2 — DKK1 systemic safety in a surveillance population.**
Done. A new Discussion paragraph addresses DKK1's canonical Wnt-dependent bone/skeletal biology, interprets the fibroblast-dominant GTEx pattern (a cell-culture artifact; low solid-tissue levels consistent with a stromal/paracrine source), and states why systemic neutralization in otherwise-well Barrett's patients faces a higher safety bar than late-stage DKN-01 dosing. Gate G5 now names skeletal/systemic toxicity explicitly.

**B-3 — Interface residue → UniProt mapping.**
Done. Computed directly from the complex PDBs: GUCY2C interface residues map to UniProt P25092 positions 153–402 (within ECD 24–430); DKK1 interface residues map to O94907 positions 187–208 (within CRD2 178–256) — confirming domain engagement. Added to Results R5, Methods, and new **Table S6** (full offset mapping).

**B-4 — Figure 4 in-figure confidence notation.**
Done. Each Figure 4 panel now carries the in-panel annotation "single-model Boltz-2 · not consensus-validated," and the legend states the depicted structure is the highest-confidence of five diffusion seeds, not an ensemble or orthogonal consensus.

## Recommended items (addressed now; non-blocking)

**R-6 — GUCY2C tumor_selectivity 5 → 4 inconsistency.** Resolved, and the recomputation revealed something we now report honestly rather than assert away: recalibrating GUCY2C selectivity from 5 to 4 (window pending, O-2) lowers its composite 13.00→12.00 and *does* change the lane ranking — CDH17 (12.38) overtakes it and its Monte-Carlo rank-1 frequency falls from 80.7% to ~0% (top-3 ~48%). We state in Results and Table S0 that GUCY2C's top-target status is contingent on the optimistic selectivity score (so O-2 is decisive for lead choice as well as safety) and now carry CDH17 as a credible backup treatment-lane target.

**R-7 — Figure 1b non-dedup caveat on the figure.** Done — added to the Figure 1b legend.

**R-9 — Binder-diversity metric defined.** Done. Length-penalized pairwise identity: max 21.4% (GUCY2C set), 32.9% (DKK1 set; top pair dkk1_bb2/bb3, both alanine-rich), means ≈15%. Reported in Results R5 and Table S5, with the most-similar pairs named numerically.

**R-10 — DKK1 fibroblast artifact contextualized.** Done — Table S4 footnote flags the cultured-fibroblast value as non-physiological and frames the likely stromal/paracrine DKK1 source as compatible with a neutralizing-trap strategy.

**R-5 (association-weight grid) and R-8 (benchmarked %Ala calibration):** acknowledged as recommended-before-journal-submission. For R-8 we downgraded "clean" language to an internal-threshold framing and cite the calibration references to add. For R-5 we note the ±50% Monte Carlo already spans association-weight parity-to-premium within its U(2.5,7.5) draw and that a formal grid is a documented next step. Neither is blocking for preprint posting per the editor's letter.

## Net effect
All four blocking items are resolved except B-1, which is honestly scoped as a tooling limitation rather than papered over. The manuscript's claim strengths remain matched to (or below) its evidence, and the one item we could not compute is flagged as the top priority follow-up rather than fabricated.


---

# Revision Tracking — every change between versions

## v1 → v2 (in response to Round 1)
| Round-1 essential revision | Change made | Where |
|---|---|---|
| ER-1 GUCY2C selectivity/expression | Added GTEx v8 normal-tissue analysis (Fig S1, Table S4); 85% GI-restricted; added explicit "normal ≠ tumor expression" and "~28 TPM in normal GI → window is polarity not absence (O-2)" caveats | Results R3, Discussion, Fig S1, Table S4 |
| ER-2 Orthogonal structural validation | Removed all multi-model-validation language; named AF-Multimer/Chai-1 fold-back as not-yet-run next step (O-1); did NOT fabricate consensus numbers | Abstract, Results R5, Discussion |
| ER-3 Specificity/negative control | Added sequence-composition liability screen (Table S5); confirmed the predicted poly-alanine failure mode; re-ranked leads | Results R5, Table S5 |
| ER-4 Disclose + sensitivity-test triage rubric | Disclosed exact additive rubric (reproduces scores to 0.005); ran 20,000-draw Monte-Carlo robustness (GUCY2C #1 in 80.7%); disclosed DKK1 sole-addressable-target caveat | Results R3, Methods, Table S0 |
| ER-5 Reframe "15/16 pass" | Reframed as pilot-scale (n=8/target, single seed, permissive bar) throughout | Abstract, Results R5, Fig 3 legend |
| ER-6 Trial-count methodology | Labeled counts as non-de-duplicated relative crowding indicators; flagged missing query strings/dates as reproducibility gap; removed unsupported Fig 1b point | Methods, Fig 1b, Limitations |
| ER-7 AI-disclosure | Replaced placeholder with 3-category disclosure (deterministic bioinformatics / ML methods / AI-assisted synthesis); labeled ordinal scores as analyst judgment | Author contributions |
| ER-8 DKK1 precedent by population | Segregated DKN-01 (advanced disease) from interception (Barrett's); stated precursor-setting evidence absent; carried DKK1 on mechanistic grounds with mismatch as risk O-4 | Results R3 |

## v2 → v3 (in response to Round 2)
| Round-2 item | Change made | Where |
|---|---|---|
| B-1 GUCY2C EAC tumor mRNA (blocking) | Honestly scoped: available cBioPortal access has no per-sample mRNA endpoint; flagged as #1 follow-up rather than fabricated; treatment rationale marked provisional | Results R3 |
| B-2 DKK1 systemic safety (blocking) | Added Discussion paragraph on Wnt/bone biology + fibroblast-artifact interpretation + higher surveillance-population bar; G5 names skeletal/systemic tox | Discussion, Table S6/gates |
| B-3 Residue→UniProt mapping (blocking) | Mapped interface residues to P25092 153–402 and O94907 187–208 (within CRD2); added Table S6 | Results R5, Methods, Table S6 |
| B-4 Fig 4 in-panel confidence notation (blocking) | Added "single-model Boltz-2 · not consensus-validated" in each panel + legend note on best-of-5-seed | Figure 4, legend |
| R-6 selectivity 5→4 inconsistency | Recalibrated AND recomputed: composite 13.00→12.00, CDH17 overtakes, GUCY2C rank-1 freq 80.7%→~0%; reported as a material contingency (O-2 decides lead choice); CDH17 carried as backup | Results R3, Table S0 |
| R-7 Fig 1b caveat on figure | Added non-dedup caveat to legend | Fig 1b legend |
| R-9 diversity metric | Reported max pairwise identity 21.4%/32.9%, means ~15%, named top pairs | Results R5, Table S5 |
| R-10 fibroblast artifact | Table S4 footnote; framed stromal/paracrine DKK1 as trap-compatible | Table S4 |
| R-5 assoc-weight grid / R-8 %Ala benchmark | Acknowledged as recommended-before-journal-submission (non-blocking for preprint) | response letter |

**Residual open item:** B-1 (GUCY2C EAC tumor-cell mRNA quantification) remains to be run against a portal build exposing an mRNA endpoint; it is flagged as the top follow-up, and the treatment-arm rationale is explicitly provisional pending it.
