# Mock Peer-Review Report — EAC manuscript, Round 2 (revised submission)

**Manuscript assessed:** M1 v2 — "Steering into competitive whitespace…" (revised after round-1 Major Revision), with the authors' point-by-point response letter.

**Review model:** Same 3-persona panel + handling-editor synthesis. Reviewers were given the round-1 essential revisions and the authors' response, and asked to judge the adequacy of each revision and flag any newly-introduced problems. (Reviews run on claude-sonnet-4-6 after transient generation refusals on the primary reasoning model; content unaffected.)

**Recommendations:** Reviewer 1 — Minor revision · Reviewer 2 — Minor revision · Reviewer 3 — Minor revision  *(all three moved up from Major revision in round 1)*

*Advisory only — every specific factual claim was treated as a lead to verify, not ground truth.*

---

## Editor's Decision Letter (Round 2)

## Decision Letter — Round 2

**Journal of Computational Biology & Drug Discovery (Preprint Track)**
**Manuscript:** "Steering into competitive whitespace — a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory"
**Decision Date:** [current date]
**Handling Editor:** [HE]

---

### Opening Assessment

The revised manuscript represents a substantive and honest response to the round-1 Major Revision decision. The authors successfully resolved five of the eight essential revisions to an adequate standard and made creditable progress on two others, most notably by (a) delivering a genuine new scientific finding — the alanine-scaffold liability of the headline GUCY2C binder — that reranks the leads in a direction that reduces rather than inflates confidence in their most dramatic result, and (b) declining, rather than fabricating, orthogonal fold-back validation numbers they could not responsibly compute. The response letter is internally consistent with the manuscript text throughout. That said, three reviewers converge on a common set of residual problems — one introduced by the revision itself — that are materially important for the manuscript's core biological claims and are all addressable by public-data query and text editing.

---

### Decision: **Minor Revision**

All three reviewers independently recommend Minor Revision. The manuscript is not ready to accept as-is, but none of the remaining issues requires new experimental data, GPU-intensive recomputation, or a full-cycle structural overhaul. A focused author response addressing the items below should allow acceptance without a further formal review round, at editor discretion.

---

### Status of the Eight Essential Revisions

| ER | Status | Notes |
|---|---|---|
| **ER-1** GUCY2C selectivity/expression from public data | **PARTIALLY ADEQUATE** | GTEx normal-tissue selectivity quantified correctly and well-contextualized; the normal-tissue vs. tumor-expression distinction is now explicit. Two residual problems: (i) GUCY2C mRNA expression in the TCGA *esca_tcga_pan_can_atlas_2018* cohort — from the same public repository used for Figure 1a — was not queried; (ii) the DKK1 fibroblast artifact introduced by the ER-1 GTEx panel is unaddressed (see Remaining Items). |
| **ER-2** Orthogonal structural validation | **ADEQUATE for preprint scope** — see separate judgment below | |
| **ER-3** Negative/specificity control | **ADEQUATE** | Composition liability screen is a genuine contribution; gucy2c_bb2 correctly reassigned from headline to provisional; dkk1_bb1 elevated on combined grounds. Residual: binder-diversity metric is undefined (see Remaining Items). |
| **ER-4** Triage rubric disclosure and sensitivity test | **PARTIALLY ADEQUATE** | Full disclosure and reproduction are exemplary; Monte Carlo design is creditable. Residual: the ±50% perturbation preserves the ×5 association-weight premium rather than testing it structurally; a narrow grid sensitivity (association weight ∈ {1,3,5,7,10}) would constitute genuine robustness vs. local noise. |
| **ER-5** Reframe "15/16 pass" as pilot-scale | **ADEQUATE** | Consistently applied across Abstract, Results, and legends. |
| **ER-6** Trial-count methodology and query date | **PARTIALLY ADEQUATE** | The non-de-duplication caveat and honest acknowledgment of missing query strings are appropriate. Residual: the caveat appears in Methods but not in the Figure 1b legend, where readers encounter the counts first. A one-sentence label fix closes this. |
| **ER-7** Specific non-placeholder AI disclosure | **ADEQUATE** | Three-category framework is substantive and operationally specific; explicit labeling of ordinal scores as analyst judgment is the correct disclosure. |
| **ER-8** DKK1 precedent segregated by patient population | **ADEQUATE** | Population mismatch now explicitly stated in Results R3 as a risk. Note: this adequate resolution made the DKK1 systemic safety gap more urgent (see Remaining Items). |

---

### Remaining Items — Minor Revision Required

The following items are **blocking** (manuscript cannot be accepted without them) or **strongly recommended before external submission**. All are achievable by in-silico query and/or text and figure editing; none requires wet-lab work.

**Blocking items:**

**1. GUCY2C EAC tumor-mRNA expression — close the gap the authors themselves identified** *(R1 by Reviewers 1, 2, and 3 in consensus)*
Query GUCY2C (and CDH17 as a comparator) mRNA expression in cBioPortal's *esca_tcga_pan_can_atlas_2018* dataset — the same cohort used for Figure 1a alteration landscape. Report median z-score or RSEM, percent of tumors with detectable expression, and range. If expression is low, heterogeneous, or absent, report that result and lower the treatment-arm confidence accordingly. This is not a new data-generation ask; it is a public query on a resource the paper already cites. Suppressing an unflattering result from the TCGA cohort already in use would be a more serious problem than reporting it candidly.

**2. DKK1 systemic safety in a surveillance population — follow through on what the GTEx data now expose** *(Reviewer 1, Major Concern 2; Reviewers 2 and 3 in agreement)*
The revision correctly added GTEx data for DKK1 and correctly segregated the DKN-01/DisTinGuish precedent. The result is that Table S4 now shows DKK1 at 852.8 TPM in cultured fibroblasts and 0.39 TPM at GEJ — data the paper added but did not interpret with respect to the interception-arm patient population. A surveillance Barrett's population is not an advanced-cancer population; the safety bar is fundamentally different. A substantive paragraph in the Discussion addressing (a) the DKK1 canonical bone/skeletal biology and published DKN-01 adverse-event profiles, (b) the implication of the fibroblast-dominant GTEx expression pattern for systemic exposure in otherwise-well patients, and (c) why this safety uncertainty is acknowledged as a threshold risk distinct from GI-tract on-target effects, is both achievable by literature synthesis and necessary to frame the interception arm honestly. Gate G5 should list skeletal/systemic toxicity explicitly alongside "bone-GI safety margin."

**3. Interface residue numbering — provide a mapping to canonical UniProt positions** *(Reviewers 2 and 3 in agreement)*
Methods and S4 report GUCY2C interface residues in complex-file numbering and DKK1 hotspots in isolated-domain numbering; neither is mapped to UniProt canonical coordinates. A reader cannot independently verify that the DKK1 binder engages CRD2 (UniProt O94907 residues 178–256) when the interface residues start at position 10 in complex-file numbering. A one-line offset table or in-text mapping (e.g., "complex-file residue = UniProt residue − N; residues [10–13, 28–31] correspond to UniProt positions [188–191, 206–209]") resolves this. This is a reproducibility requirement, not a stylistic preference.

**4. Figure 4 in-figure confidence notation** *(Reviewer 1, Major Concern 1)*
The figure caption correctly caveats that backbone traces are single-model, single-seed Boltz-2 predictions. However, a reader who sees the rendered complex images and glances at ipTM > 0.85 without reading the caption in full will draw an incorrect inference about structural certainty. Add an in-figure text annotation (e.g., "Single-model Boltz-2 / not consensus-validated") visibly in each panel. The legend should also note that the depicted structure is the highest-confidence of five Boltz-2 seeds rather than an ensemble or consensus. This is a figure-editing task.

**Strongly recommended before external journal submission (non-blocking for preprint acceptance):**

**5. Monte Carlo robustness — add a narrow grid test of the ×5 association premium** *(Reviewer 2, Major Concern 1)*
The current ±50% perturbation tests noise around the chosen point but does not test whether GUCY2C's rank-1 frequency changes if the association weight is reduced to parity with the other components. Running a grid over association weight ∈ {1, 3, 5, 7, 10} (other weights held at 1) and reporting GUCY2C's rank-1 frequency at each point would constitute genuine structural robustness rather than local-neighbourhood sensitivity. If GUCY2C leads at all association weights ≥ 3, that is a strong result; if it leads only when association weight ≥ 5, that should be stated.

**6. GUCY2C tumor_selectivity score — resolve internal inconsistency** *(Reviewer 2, Major Concern 2)*
The rubric anchor for score = 5 is "strong therapeutic window," but the manuscript simultaneously states that the GI safety window for GUCY2C is open question O-2 — i.e., unresolved. Recalibrate the score to 4 (or redefine the anchor for 5 as "tissue-restricted, window pending") and report the resulting composite. GUCY2C will almost certainly still lead the treatment lane; show it.

**7. Figure 1b legend — move the non-de-duplication caveat to the figure itself** *(All three reviewers)*
One sentence added to the figure legend: "x-axis values reflect non-de-duplicated active-EAC trial intervention mentions and should be read as relative crowding indicators; see Methods for caveats." No figure redrawing required.

**8. Composition liability screen — anchor against a published reference distribution** *(Reviewer 2, Major Concern 4)*
Cite a published %Ala distribution from the RFdiffusion/ProteinMPNN benchmark literature (e.g., Bennett et al. 2023 *Nature*, Cao et al. 2022 *Nature*) to ground the 17.5% vs. 43.7% comparison, or explicitly downgrade "clean" to "below the threshold for obvious alanine-scaffold artifacts, pending benchmarked calibration."

**9. Define the binder-diversity metric** *(Reviewer 3, Major Concern 4)*
Report pairwise sequence identity (or BLOSUM-based similarity) across all 16 designs, state the threshold used to call "no high-similarity pairs," and report the two most similar pairs numerically.

**10. DKK1 fibroblast artifact — contextualise in the table and note the in-vivo expression pattern** *(All three reviewers)*
Add a table footnote in S4 noting that the "Cells_Cultured_fibroblasts" value (852.8 TPM) is a non-physiological cell-culture artifact, that DKK1's in-vivo normal-tissue expression pattern is better captured by the remaining tissue rows, and that the biologically relevant DKK1 source in Barrett's tissue is likely stromal/paracrine — a framing that may actually support the neutralizing-trap rationale if made explicit.

---

### Judgment on ER-2: GPU-Dependent Orthogonal Fold-Back

The authors' decision to **scope rather than fabricate** orthogonal structural validation (AlphaFold-Multimer / Chai-1 consensus fold-back) is **acceptable for a preprint at this stage, and the handling editor concurs with all three reviewers on this point.** The authors correctly declined to generate numbers they did not have, lowered claim language throughout the manuscript, explicitly named the missing validation as open question O-1, and substituted the composition liability screen as a check they *could* responsibly run — which produced a genuine finding rather than cosmetic reassurance. That is the scientifically honest behavior this journal's scope rewards. The reviewers' uniform verdict that this scoping is appropriate for a preprint is adopted here. The absence of orthogonal fold-back is not grounds for rejection or for demanding another Major Revision cycle; it should be named prominently in the preprint's cover note as the next experimental (or high-compute) step for any group that wishes to take either lead into a wet-lab program.

---

### Path to Acceptance

The authors should prepare a revised manuscript addressing the four blocking items above, with a brief point-by-point response letter. For items 5–10 (non-blocking), a statement of intent (or completion) in the response letter is sufficient; the editor will not require a further formal review round if the blocking items are clearly resolved. The response letter should, for each blocking item, either show the result (e.g., the GUCY2C TCGA mRNA query output, the DKK1 safety discussion paragraph, the residue-mapping table, the Figure 4 annotation) or, if a result is unflattering, show it anyway — candor has been this manuscript's principal strength across both rounds of revision and should not be abandoned at the finish line.

If the four blocking items are resolved without introducing new unaddressed data exhibits, the manuscript will be accepted for posting on the preprint track. Authors should note that items 5–10 remain open and will be expected to be resolved before submission to any indexing journal.

The editor thanks all three reviewers for the rigor and specificity of their round-2 assessment, and thanks the authors for a revision that demonstrated what responsible, scope-constrained iteration on a computational preprint looks like.

---

*[Handling Editor signature block]*

---

# Full Reviews (Round 2)

## Reviewer 1 (Round 2)

### Reviewer 1: Clinician-scientist (GI oncology / translational EAC biology; 15+ years of Barrett's surveillance and GEJ-cancer clinical trial experience)

---

## Manuscript: "Steering into competitive whitespace" — EAC dual-arm in-silico discovery-to-design campaign (v2, revised)

---

### Brief Summary

This revised preprint presents a fully computational, public-data-driven campaign to nominate and pursue two differentiated therapeutic targets in EAC — GUCY2C (T-cell-engager treatment arm) and DKK1 (neutralizing-trap interception arm) — spanning the Barrett's-to-carcinoma disease trajectory. The authors mined TCGA genomics, Open Targets associations, ClinicalTrials.gov landscape data, and GTEx expression profiles to motivate target selection, then applied an explicit, now-disclosed druggability triage and generated de novo protein binders via RFdiffusion → ProteinMPNN → Boltz-2 fold-back. The revision meaningfully strengthens the first submission by adding a GTEx-based selectivity analysis with appropriately caveated interpretation, a Monte Carlo weight-perturbation robustness test of the triage rubric, a sequence-composition liability screen that reverses the naïve ipTM-ranking of leads, and a clearly segregated DKK1 precedent acknowledgement. The honest, unvarnished admission that the top GUCY2C binder has a high-alanine developability liability — and that the DKK1 lead is the more credible starting point — is the kind of author candor that a competitive whitespace preprint needs. Several problems from round 1 remain incompletely resolved, one new issue is introduced, and a few biological framing points still need tightening, but the manuscript is substantially improved.

---

### Strengths

- **Honest, internally consistent claim calibration throughout.** The abstract now reads like a methods-section summary rather than a promotional document: "ipTM is a single-model docking-confidence proxy, not a measured affinity; these are pilot-scale, in-silico leads, not experimentally tested therapeutics." The Discussion's limitations section is comprehensive and candid.
- **Composition liability screen (ER-3) is the standout new contribution.** The finding that gucy2c_bb2's headline ipTM 0.915 co-occurs with 44% alanine and an 8-residue homopolymer run — and that this reverses the single-metric ranking in favor of dkk1_bb1 — is both scientifically important and a demonstration that the authors are willing to undercut their own headlineable result when the data warrant it.
- **Monte Carlo rubric robustness (ER-4)** is well-executed: 20,000 draws, ±50% weight perturbation, results segregated by lane, and the DKK1 trivially-rank-1 result is flagged as reflecting a "thin addressable space" rather than a contested victory. That disclosure is appropriate and rare.
- **GTEx selectivity panel (ER-1)** is appropriately contextualized: the split between normal-tissue selectivity (what GTEx supports) and EAC tumor-cell expression (what it does not) is now clearly stated in Results, Discussion, and the figure legend.
- **DKK1 precedent segregation (ER-8)** is handled well. The population mismatch — advanced GEJ disease vs. pre-malignant Barrett's — is now stated explicitly in Results R3 as a risk rather than collapsed into a single favorable citation.
- **AI-disclosure (ER-7)** is now specific, structured, and honest about which outputs are analyst judgment vs. deterministic computation. The three-category taxonomy is a useful model.
- **Pilot-scale reframing (ER-5)** is applied consistently through Abstract, Results, and the legend. The phrase "n=8/target is too small to estimate a rate" is exactly the right language.

---

### Major Concerns

**1. ER-2 (orthogonal structural validation) — the scoping is appropriate but the ipTM values are still interpreted without sufficient hedge in the main-text visual narrative.**
The authors correctly decline to fabricate AlphaFold-Multimer/Chai-1 numbers they have not run, and I find that stance acceptable for a preprint. The problem is that Figure 4 is captioned with "Backbone traces depict topology and interface localization, not experimentally-determined secondary structure; ipTM is a docking-confidence proxy, not a measured affinity" — but the figure panels themselves show what look like cleanly folded, precisely docked complexes rendered in full-detail backbone trace. A reader who does not read every caption caveat will see two polished molecular structures with ipTM > 0.85 and draw exactly the wrong inference. The authors should add a visible, in-figure notation (e.g., a superimposed text box or a distinctive border) that marks these as single-model, non-consensus predictions, or rephrase the legend to state explicitly: "These are single-seed Boltz-2 predictions without consensus validation; visual precision should not be read as structural accuracy." This is achievable by editing the figure legend and/or the figure itself.

*Severity:* Moderate — the text is correct, but the figure communicates something the text walks back. In a preprint that will be read by molecular biologists who skim figures, the mismatch matters.

**2. DKK1 normal-tissue expression profile is alarming and underinterpreted.**
Table S4 shows DKK1's top GTEx tissue as "Cells_Cultured_fibroblasts" at 852.8 TPM with a GI fraction of exactly 0.00 — meaning DKK1 has essentially no GI-tissue restriction in normal tissue profiling. Esophagus mucosa is 1.19 TPM and GEJ is 0.39 TPM. The manuscript is intended to place DKK1 in an *interception arm for a largely pre-malignant Barrett's population* — a setting that sets, as the authors correctly acknowledge, a very high safety bar. Yet the normal-tissue data show that DKK1 neutralization would affect a protein whose highest normal expression is in fibroblasts, with potentially relevant expression in bone (DKK1 is a canonical bone-metabolism regulator), skin, and kidney. The DKK1 preclinical safety literature documents skeletal toxicity (heterotopic ossification, bone density effects) in DKK1-pathway interference models. The manuscript mentions "bone-GI safety margin" briefly in the Gate G5 criterion table and lists O-4 as the regulatory/surrogate-endpoint question — but there is no substantive discussion of the systemic (non-GI) safety risk of DKK1 neutralization in a surveillance population of patients who are otherwise well. This is not a request for wet-lab toxicology; it is a request for a paragraph in the Discussion that honestly characterizes the DKK1 safety landscape for the interception-arm patient population, drawing on the existing clinical and preclinical literature for DKK1 inhibitors. The DKN-01 trial data and the anti-DKK1 preclinical literature have published adverse-event profiles that the authors can cite.

*Severity:* Major — the central clinical framing of the DKK1 interception arm rests on the assumption that DKK1 neutralization is tolerable in an essentially well, surveillance population. The manuscript does not address this anywhere substantively, and the GTEx data the authors added make it more urgent, not less.

**3. GUCY2C EAC tumor-cell expression gap is flagged but its severity is understated.**
The authors commendably note that GTEx establishes *normal-tissue* selectivity but not EAC *tumor-cell* expression. However, the GUCY2C EAC tumor-expression question is not merely a gap to fill in future studies — it is a threshold criterion for the entire treatment-arm rationale. The EAC literature is not silent on this: there are IHC studies in esophageal/GEJ cancer (some negative, some reporting low or heterogeneous expression) and RNA data in TCGA ESCA that the authors could have examined but apparently have not. cBioPortal provides TCGA mRNA expression data for *esca_tcga_pan_can_atlas_2018* (the same cohort used for the alteration frequency analysis in Figure 1a). Checking GUCY2C mRNA expression in that dataset — even a simple median-expression-and-percent-detectable query — would directly address whether the target is expressed in EAC tumors at levels comparable to the CRC setting where clinical precedent exists. This is achievable in-silico (it requires a cBioPortal query, not a wet-lab experiment). The current manuscript uses the same cohort for alteration frequency but apparently did not retrieve expression data from it. I consider this a feasible, mandatory revision: query GUCY2C (and CDH17 as a comparator) mRNA expression in *esca_tcga_pan_can_atlas_2018* and report the result, even if it is unflattering.

*Severity:* Major — the selectivity argument is structurally sound, but the treatment-arm rationale requires tumor-expressed antigen. The data exist in the already-used public source and the authors chose not to retrieve them.

---

### Minor Concerns

**4. The DKK1 composite score disease_association (Barrett's) is 0.074, while GPX7 is 0.372 and REG4 is 0.094.** DKK1 has among the weakest Open Targets Barrett's association scores in the interception lane. The Monte Carlo robustness check cannot rescue this when DKK1 is literally the only biologic-addressable interception candidate after the filter — the filter is doing all the work, not the score. The text acknowledges this ("trivial"), but the Figure 2 visualization and the Abstract still imply that DKK1 "emerged" from a competitive field. The Abstract should say explicitly that DKK1 was selected because it is the only biologic-addressable interception candidate, not because it out-scored others.

**5. The interface patch numbering in S4 uses "complex-file numbering" for DKK1 (residues [10,11,12,13,28,29,30,31]) — these are clearly the isolated-domain residue numbers, not the full-protein UniProt numbering.** The DKK1 CRD2 domain is described as residues 178–256 of the full protein; residue 10 in the isolated domain is approximately residue 188 in the full protein. The manuscript should either provide a mapping table or report numbers in full-protein coordinates to allow independent structural interpretation.

**6. Figure S1 legend notes DKK1's top tissue as "Cells_Cultured_fibroblasts" at 852.8 TPM** — a cell-culture artifact rather than a tissue measurement. This value likely reflects culture-condition upregulation and is not directly comparable to in-vivo tissue TPM values. The authors should note in the S4/S1 context that fibroblast culture values are not an in-vivo tissue profile and that DKK1's in-vivo normal-tissue expression pattern is better captured by the remaining tissues.

**7. The Methods statement "The per-axis 'active EAC trial intervention mentions' … were not de-duplicated across combination trials"** is appropriately caveated, but Figure 1b remains a log-scale scatter plot with specific numerical x-axis positions that will be read as exact counts. The figure legend should note the non-de-duplicated nature of the counts and the missing query date.

**8. The CEACAM5 esophagus_mucosa_TPM is 293.4 in Table S4** — by far the highest esophageal mucosa expression of any comparator — but the manuscript is silent on what this means for CEACAM5 as a potential comparator target in EAC. This is a minor point that need not be expanded, but the figure legend should note that CEACAM5's high esophageal mucosa expression reflects normal esophageal expression (CEA is a known normal GI-tract antigen), not a favorable therapeutic window.

---

### Assessment of Each Essential Revision

| ER | Verdict | Notes |
|---|---|---|
| **ER-1** (GUCY2C selectivity/expression from public data) | **PARTIALLY ADEQUATE** | GTEx normal-tissue data added and appropriately caveated. However, TCGA *esca_tcga_pan_can_atlas_2018* mRNA expression data — from the same cohort already used — were not retrieved. This is a consequential omission that this round-2 review escalates to Major Concern 3. |
| **ER-2** (orthogonal structural validation) | **ADEQUATE (for preprint scope)** | Authors correctly decline to fabricate consensus numbers unavailable without GPU re-run; claim language is lowered throughout. The figure-text mismatch (Major Concern 1) is a presentation problem introduced by the revision, not a scientific fabrication. |
| **ER-3** (negative/specificity control for binder designs) | **ADEQUATE** | Composition liability screen is rigorous, directly addresses the predicted alanine-scaffold failure mode, and the reversal of the gucy2c_bb2 ranking is the intellectually honest result. Monomer/decoy controls named as O-1 remaining work. |
| **ER-4** (disclose + sensitivity-test triage rubric) | **ADEQUATE** | Exact formula disclosed, all 16 scores reproduced within 0.005, Monte Carlo robustness reported with appropriate lane-segregation. DKK1 trivial-rank disclosure is laudable. |
| **ER-5** (reframe "15/16 pass" as pilot-scale) | **ADEQUATE** | Consistently applied; language in Abstract, Results, and legends is appropriately hedged. |
| **ER-6** (trial-count methodology + query date) | **PARTIALLY ADEQUATE** | The non-de-duplication caveat and crowding-indicator framing are helpful. The missing query strings and snapshot date are acknowledged as a reproducibility gap. This is acceptable for a preprint but must be resolved before journal submission, as stated. Minor Concern 7 persists regarding Figure 1b axis labeling. |
| **ER-7** (specific non-placeholder AI disclosure) | **ADEQUATE** | The three-category taxonomy is clear and specific; the identification of ordinal scores and "whitespace verdict" as analyst judgment is exactly right. |
| **ER-8** (segregate DKK1 precedent by patient population) | **ADEQUATE** | The population mismatch is now stated explicitly in Results R3. However, this segregation makes Major Concern 2 (DKK1 normal-tissue safety in a surveillance population) more urgent, not less — the revision correctly exposed an issue that requires follow-through discussion. |

**New problems introduced by the revision:**
- Table S4 DKK1 fibroblast-culture artifact (Minor Concern 6) — not present in v1 because the table did not exist.
- Figure 4 visual-confidence mismatch (Major Concern 1) — the added render makes the single-model caveat more important to surface in the figure itself.
- The DKK1 GI-fraction = 0.00 finding now visible in Table S4 makes the interception-arm safety framing materially weaker than stated (Major Concern 2), an issue the revision created by honestly reporting the GTEx data without following through on their implication.

---

### Feasible Revisions

1. **[Major Concern 3 — achievable in-silico]** Query GUCY2C (and CDH17 as a comparator) mRNA expression in the TCGA *esca_tcga_pan_can_atlas_2018* cohort via cBioPortal. Report median z-score or RSEM, percent of tumors with detectable expression (e.g., z-score > 0), and range. Add a single panel or supplementary table. If expression is low or variable, say so and adjust the treatment-arm confidence accordingly — do not suppress an unflattering result.

2. **[Major Concern 2 — achievable by manuscript editing]** Add a paragraph to the Discussion on DKK1 systemic safety in the interception-arm population, citing published DKN-01 and anti-DKK1 preclinical adverse-event data. Address specifically: (a) DKK1's canonical role in bone homeostasis and the skeletal safety risk; (b) the implication of the GTEx fibroblast-dominant expression pattern for systemic (non-GI) exposure in otherwise-well Barrett's patients; (c) why the safety bar for Barrett's interception differs fundamentally from the advanced-cancer bar of the DisTinGuish trial. Gate G5 should list "skeletal/systemic toxicity" alongside "bone-GI safety margin."

3. **[Major Concern 1 — achievable by figure editing]** Add an in-figure text annotation to Figure 4 panels making explicit that these are single-model, single-seed Boltz-2 predictions without cross-model consensus validation. A bold-text inset ("Single-model Boltz-2; not consensus-validated") is sufficient. Update the legend to match.

4. **[Minor Concern 4 — achievable by editing one sentence]** Revise the Abstract to state explicitly that DKK1 was selected as the only biologic-addressable interception candidate after the druggability filter, not as the highest composite scorer in a competitive field.

5. **[Minor Concern 5 — achievable by editing S4]** Add a coordinate-mapping footnote to Section S4 converting the DKK1 complex-file interface residue numbers [10–13, 28–31] to full-protein UniProt O94907 coordinates (~188–191, ~206–209).

6. **[Minor Concern 6 — achievable by editing S1/S4]** Add a note in Table S4 and the Figure S1 legend that the DKK1 "Cells_Cultured_fibroblasts" value (852.8 TPM) is a cell-culture artifact not directly comparable to in-vivo tissue expression, and that DKK1 in-vivo normal-tissue expression is best assessed from the remaining tissue rows.

7. **[Minor Concern 7 — achievable by editing the figure legend]** Add to the Figure 1b legend: "x-axis positions reflect non-de-duplicated axis-level active-trial mentions; query date not recorded; values should be read as relative crowding indicators, not exact disjoint counts."

---

### Recommendation

**Minor revision.**

The revision substantially addressed five of the eight essential revisions to an adequate or better standard, and the honest composition-liability finding that demotes the headline GUCY2C binder is a genuine scientific contribution. Three problems remain — a consequential omission (TCGA EAC mRNA expression for GUCY2C), a clinical-framing gap opened by the authors' own GTEx data (DKK1 systemic safety in a surveillance population), and a figure-text mismatch in Figure 4 — but all three are achievable by in-silico query and manuscript editing. None requires wet-lab work. The manuscript should not be accepted without these three revisions; it does not require another full major-revision cycle if the authors respond directly and completely to them.

---

### Cross-Cutting Comment

The revision demonstrates that the authors are capable of honest self-correction under critical pressure — the alanine-scaffold finding in particular is the kind of uncomfortable result that many authors bury, and its elevation to a key finding reflects well on the transparency commitment. The two remaining biological framing gaps — GUCY2C EAC tumor-cell expression and DKK1 systemic safety in a surveillance population — are not minor housekeeping items: they bear directly on whether each arm's lead rationale is credible enough to motivate a wet-lab follow-on program, which is the preprint's stated purpose. The campaign is most useful to the field if its known weaknesses are as visible as its strengths. Closing these two gaps with the data that are already publicly available (TCGA expression, published DKN-01 safety data) would make the revised manuscript a genuinely rigorous and reproducible template, rather than a well-caveated aspiration.

---

## Reviewer 2 (Round 2)

### Reviewer 2: Computational biologist & biostatistician — statistical rigor, methods transparency, prediction-tool validity, and reproducibility

---

## Manuscript: "Steering into competitive whitespace — a public-data, computation-driven campaign nominates GUCY2C and DKK1 and delivers de novo binder leads across the esophageal adenocarcinoma trajectory"

---

### Brief Summary

The revised manuscript presents a human-supervised, AI-assisted computational campaign that (i) characterizes EAC as a copy-number-driven disease using TCGA data, (ii) maps competitive whitespace among 351 EAC trials and an Open Targets known-drug set, (iii) applies a disclosed and now robustness-tested additive triage rubric to nominate GUCY2C and DKK1 across two trajectory positions, and (iv) delivers eight de novo binder designs per target via RFdiffusion→ProteinMPNN→Boltz-2, evaluated by ipTM/pLDDT and a new sequence-composition liability screen. The revision adds a GTEx normal-tissue selectivity analysis for GUCY2C, fully discloses the triage rubric with Monte-Carlo sensitivity testing, introduces a composition-based specificity proxy, reframes the pass-rate claim as pilot-scale, partially addresses the trial-count methodology gap, delivers a three-category AI-assistance disclosure, and segregates DKK1 clinical precedent by patient population. Overall claim strength has been deliberately lowered to match the evidence.

---

### Strengths

- **Honest, layered self-assessment.** The revision does not paper over what was not done: single-model Boltz-2 confidence, missing orthogonal fold-back, unrecorded query dates, thin interception-lane addressable space. These admissions are integrated into the text rather than buried in footnotes, and the response letter is consistent with the manuscript.
- **Triage rubric disclosure and robustness (ER-4).** The exact formula is now published with sufficient detail for independent reproduction; the Monte-Carlo perturbation (20,000 draws, ±50% on all weights) is a creditable sensitivity test, and the candid disclosure that DKK1's 100% rank-1 rate is trivially forced by being the sole addressable interception candidate is scientifically exemplary.
- **Composition liability screen (ER-3).** Adding Table S5 was the correct move. It immediately identified that the headline GUCY2C binder (gucy2c_bb2, ipTM 0.915) is 43.7% alanine with an 8-residue homopolymer run, directly confirming the low-complexity scaffold failure mode reviewers anticipated. The paper now correctly demotes this to "provisional pending redesign" and elevates dkk1_bb1 (17.5% Ala, max run 2) as the most robust lead. This is a genuine new finding, not a cosmetic fix.
- **DKK1 population-mismatch segregation (ER-8).** The DKN-01 precedent is now cleanly separated by population — advanced GEJ/gastric treatment vs. pre-malignant Barrett's interception — and the gap is explicitly stated as risk O-4. This is exactly the correction that was requested.
- **GUCY2C selectivity framing (ER-1).** The GTEx S4/S1 addition successfully distinguishes tissue selectivity (supportable: 85% GI-restricted, top non-GI 1.7 TPM) from tumor expression (explicitly not established from these data, correctly flagged as a gap). The polarity-not-absence framing for luminal restriction is appropriately cautious.
- **Pilot-scale reframing (ER-5).** "Pilot-scale," "permissive bar," and "n=8 cannot support a generalizable success rate" now appear in the Abstract, Results, and legend — adequate.
- **Three-category AI disclosure (ER-7).** Separating deterministic bioinformatics, published ML methods as objects of study, and LLM-assisted synthesis/drafting — with explicit identification of the ordinal scores as analyst judgment — is a meaningful and usable disclosure, not a placeholder.

---

### Major Concerns

**1. The Monte-Carlo weight perturbation is methodologically misspecified in a way that potentially overstates GUCY2C's robustness.**

Table S0 specifies that all four weights are perturbed: tumor_selectivity, novelty, tractability, and disease_association. However, the association weight is varied over U(2.5, 7.5) while the other three are varied over U(0.5, 1.5) — i.e., a ×5 mean multiplied by a ±50% factor. This preserves the *ratio* of the association weight to the other three weights within a narrow band, rather than testing whether the dominance of the ×5 association term is itself driving the result. A multiplicatively-symmetric perturbation over, say, U(1, 10) for the association coefficient (or equivalently, asking how often GUCY2C leads if association is down-weighted to equal the other components) would constitute a genuine robustness check. The current design is a robustness check around the chosen point, not across the structural choice to weight association at 5×. The manuscript should either (a) run a second sensitivity scenario with association weight ∈ {1, 3, 5, 7, 10} on a grid and report how GUCY2C's rank-1 frequency varies, or (b) explicitly acknowledge that the U(2.5, 7.5) perturbation preserves the ×5 premium and that the robustness result applies only in the neighbourhood of the published rubric. Currently the text implies more sweeping robustness than the analysis demonstrates.

**2. The GUCY2C tumor_selectivity score of 5 is retained despite the GTEx data now demonstrating substantial normal-gut expression (~28 TPM in small intestine) — an internal inconsistency with the rubric anchor.**

Table S0 defines tumor_selectivity = 5 as "tissue-restricted with a strong therapeutic window." Yet the manuscript simultaneously states that GUCY2C is "genuinely expressed in normal small intestine and colon (~28 TPM median)" and that the normal-GI safety window is "precisely open question O-2" — i.e., the therapeutic window is *not established*. A score of 5 on a scale where 5 = strong therapeutic window is therefore inconsistent with the paper's own caveat that the window is unresolved. A score of 3–4 ("moderate window / some normal expression") would be more defensible given the GTEx data. The authors should either (a) recalibrate the score with a brief justification and re-report the affected composite (GUCY2C composite would drop to ~11–12, which likely still leads the treatment lane but should be shown), or (b) define selectivity = 5 as "tissue-restricted" (not "strong window") and state the window assessment is deferred — and make sure the rubric anchor matches the score for all 16 candidates under the same logic.

**3. The trial-count methodology gap (ER-6) remains inadequately resolved for a submission claiming quantitative competitive-landscape analysis.**

The Methods now state that per-axis trial mentions "were not de-duplicated across combination trials and therefore do not sum to the 106 active-trial total" and that query strings and snapshot dates were not recorded. This is honest, but it means the x-axis of Figure 1b — the primary evidence for competitive crowding, which drives the central "whitespace" argument — is neither reproducible nor de-duplicated. The paper flags this as "a reproducibility gap to close before external submission." For a preprint that will presumably circulate and be cited, "gap to close before external submission" is insufficient. The minimum requirement is: provide the raw trial-ID list (even approximately reconstructed) or a ClinicalTrials.gov API query string with the approximate date, and label the axis unambiguously as "approximate, non-deduplicated axis-level mentions" both in the legend and in the figure caption. The current legend says "active EAC trial intervention mentions (x, log scale)" without qualification, so a reader who does not read the Methods text carefully will not know the values are approximate and non-deduplicated. This is a presentation/reproducibility defect in the primary evidence figure.

**4. The composition liability screen (Table S5) does not include a calibrated threshold or a reference distribution, so it cannot be used to rule designs in or out — it only flags the worst cases visually.**

The screen computes %Ala, %hydrophobic, max homopolymer run, and frac_charged, and correctly identifies gucy2c_bb2 as a likely generic scaffold. But no threshold is applied (e.g., %Ala > X% → fails), no reference distribution from the literature (e.g., known binders designed with RFdiffusion/ProteinMPNN from published benchmarks) is provided, and the "clean" verdict for dkk1_bb1 (17.5% Ala) is asserted but not formally justified against a baseline. Without a calibrated cutoff, the screen is qualitatively useful but does not constitute a specificity filter. The paper should either (a) cite a published threshold or distribution (%Ala in published RFdiffusion/ProteinMPNN binder sets, e.g. from Bennett et al. 2023 *Nature* or Cao et al. 2022 *Nature*) and anchor the 17.5% vs. 43.7% comparison against it, or (b) explicitly state that the screen is a qualitative flag, not a quantitative filter, and that its interpretation is provisional pending orthogonal fold-back.

---

### Minor Concerns

**1. DKK1 esophagus_gej_TPM = 0.39 in Table S4 is neither discussed nor contextualized.** The paper uses GTEx to argue for GUCY2C GI restriction, but DKK1's top tissue in GTEx is "Cells_Cultured_fibroblasts" at 852.8 TPM — a completely non-physiological ex-vivo cell culture context — while its esophageal/GEJ expression is 0.39–1.19 TPM. The main text describes DKK1 as a "secreted Wnt-pathway modulator" but does not note that its dominant GTEx signal is in fibroblasts, which raises a question about whether the biologically relevant source of serum/lumenal DKK1 in Barrett's tissue is stromal rather than epithelial — a detail that actually strengthens the neutralizing-trap rationale (serum-accessible target) but should be stated rather than passed over silently in Table S4.

**2. hotspot residue numbering is not cross-walked to UniProt canonical positions.** Methods and S4 report GUCY2C interface residues as [130,132,146,192,325,326,327,330,347,348,349,378,379] and DKK1 as [10,11,12,13,28,29,30,31] "in complex-file numbering," while the design hotspots use "isolated-domain numbering." A one-line offset table (e.g., complex-file residue = UniProt residue − N for GUCY2C ECD, given chain starts at residue 24) would allow any reader to map these to canonical coordinates for comparison with published functional sites — otherwise the epitope localization claim in Figure 4 is not independently checkable.

**3. The MPNN score is reported in Table S1 but its units and directionality are not defined in the table or the legend.** ProteinMPNN outputs a log-likelihood / negative log-likelihood score depending on implementation; "1.232" for dkk1_bb1 and "0.95" for gucy2c_bb2 should be labeled (lower = better or higher = better; sign convention) so readers can interpret Table S1 without reading the ProteinMPNN paper.

**4. Figure 4 caption says "Backbone traces depict topology and interface localization, not experimentally-determined secondary structure" — this caveat is good, but the legend should also note that the depicted structure is the highest-confidence of five Boltz-2 samples rather than an ensemble-averaged or consensus model.** A single-sample structure depicted without that label may give a false sense of structural certainty.

**5. The Methods state binder length ranges of 70–100 aa (GUCY2C) and 60–90 aa (DKK1), yet dkk1_bb1 is 63 aa.** The text should either correct the range (≤60 aa designs were accepted) or note that the lower bound was relaxed, and confirm this was by design rather than a reporting error.

---

### Feasible Revisions

1. **[ER-4 / Major Concern 1]** Run a grid sensitivity analysis varying the association weight coefficient across {1, 3, 5, 7, 10} (all other weights held at 1, then separately varied), and add a 2–3 sentence paragraph in the Methods/Table S0 footnote reporting how GUCY2C's rank-1 frequency changes as the association premium varies. If GUCY2C leads at coefficient ≥ 3, this is genuine robustness; if it leads only when association weight ≥ 5, that should be stated.

2. **[Major Concern 2]** Recalibrate the tumor_selectivity score for GUCY2C from 5 to 4 (or redefine the rubric anchor for 5 as "tissue-restricted" rather than "strong therapeutic window"), report the resulting composite (likely ~12 rather than 13), and confirm that GUCY2C still leads the treatment lane. Add one sentence noting the internal consistency between the score and the therapeutic-window caveats.

3. **[ER-6 / Major Concern 3]** Add an in-text qualifier to Figure 1b's axis label and caption (e.g., "approximate, non-deduplicated active-EAC trial intervention mentions, x-axis; see Methods for caveats") and provide, as a supplement, either a reconstructed list of the trial IDs queried or the ClinicalTrials.gov API query parameters with best-available date, so the figure is at minimum partially reproducible.

4. **[Major Concern 4]** Anchor the composition liability screen to a published reference distribution or threshold from the RFdiffusion/ProteinMPNN literature, or explicitly downgrade the "dkk1_bb1 is clean" language to "dkk1_bb1 is below the threshold for obvious alanine-scaffold artifacts, pending benchmarked calibration."

5. **[Minor Concern 1]** Add 1–2 sentences noting DKK1's dominant GTEx signal in cultured fibroblasts, clarify that this is a non-physiological artifact of the tissue panel, and note that the relevant DKK1 source in Barrett's tissue is likely stromal/paracrine — a framing that supports rather than undermines the neutralizing-trap rationale.

6. **[Minor Concern 2]** Add a one-line offset table or in-text mapping from complex-file residue numbers to UniProt canonical positions for both targets, so the interface patches in Figure 4 can be compared with published structure–function data.

7. **[Minor Concern 3]** Define the MPNN score in Table S1's legend (e.g., "MPNN log-likelihood; lower = better design confidence" or equivalent), including sign convention.

8. **[Minor Concern 5]** Reconcile the stated DKK1 length range (60–90 aa) with dkk1_bb1's 63 aa length — either correct the range to ≥60 aa or add a note that the lower bound was a flexible constraint.

---

### Assessment of Each Essential Revision

**ER-1 (GUCY2C EAC selectivity/expression from public data):** **ADEQUATE with one residual inconsistency (see Major Concern 2).** The GTEx analysis is the right data source, the selectivity quantification is clear, and the normal-tissue vs. tumor-expression distinction is correctly drawn. The only problem introduced is an internal inconsistency: the tumor_selectivity score of 5 (= "strong therapeutic window") is now in tension with the paper's own statement that the window is an open question.

**ER-2 (Orthogonal structural validation):** **PARTIALLY ADEQUATE — honestly scoped.** The authors correctly declined to fabricate consensus numbers, lowered claim strength throughout, and named orthogonal fold-back + monomer/decoy controls as the not-yet-run O-1 step. For a preprint at this stage, this scoping is acceptable; the paper does not over-claim. No new problem introduced.

**ER-3 (Negative/specificity control):** **ADEQUATE as a first-order proxy, with the caveat in Major Concern 4.** Table S5 and its integration into the Results and Discussion are a genuine improvement. The alanine-content finding for gucy2c_bb2 is important and correctly reported. The remaining gap is the absence of a calibrated threshold, but the paper's treatment of this as qualitative is defensible.

**ER-4 (Triage rubric disclosure + sensitivity test):** **PARTIALLY ADEQUATE.** The rubric is now fully disclosed and reproducible; this was the most important part of the ask and is done well. The Monte-Carlo test is creditable but the perturbation design preserves the relative advantage of the ×5 association weight rather than testing it structurally (see Major Concern 1). A minor additional analysis would make this fully adequate.

**ER-5 ("15/16 pass" reframed as pilot-scale):** **ADEQUATE.** The reframing is complete and consistently applied across Abstract, Results, and legend.

**ER-6 (Trial-count methodology + query date):** **PARTIALLY ADEQUATE — honest about the gap, but the gap is too large for the primary evidence figure.** The Methods acknowledgment is correct. The deficiency is that the Figure 1b axis label and caption are still presented without qualification, so the non-deduplicated, unanchored nature of the counts is not visible at the figure level. This is a presentation fix, not a data-generation problem (see Feasible Revision 3).

**ER-7 (Non-placeholder AI disclosure):** **ADEQUATE.** The three-category structure is specific, useful, and appropriately distinguishes deterministic computation from model judgment.

**ER-8 (DKK1 precedent segregated by patient population):** **ADEQUATE.** The population mismatch is now explicitly stated in Results R3, named as risk O-4, and the mechanistic/druggability rationale is correctly offered as the only legitimate justification for carrying DKK1 into the interception arm.

---

### Recommendation

**Minor revision.**

The revision is substantive and honest: it lowered claim strength to match evidence, produced a genuinely new finding (alanine-scaffold liability in the top GUCY2C binder), fully disclosed the triage rubric, and correctly segregated the DKK1 precedent. Two issues prevent immediate acceptance: the tumor_selectivity = 5 score for GUCY2C is now internally inconsistent with the paper's own caveats about the unresolved therapeutic window (Major Concern 2, resolvable with a one-line score adjustment), and the Figure 1b axis — the central whitespace evidence — lacks in-figure qualification of its non-deduplicated, date-unanchored nature (Major Concern 3, resolvable with a caption edit and a supplementary query log). Major Concern 1 (Monte-Carlo design) and Major Concern 4 (composition threshold calibration) are methodological improvements that would materially strengthen the paper but can be addressed with small additional analyses or explicit scope-limitation statements. None of the required revisions involves wet-lab work or new data generation.

---

### Cross-Cutting Comment

This revision demonstrates what honest iteration under scope constraints looks like: the authors declined to fabricate orthogonal fold-back numbers, instead lowering claim strength and making the absent validation conspicuous and named, and they used the revision cycle to surface a new finding (the gucy2c_bb2 alanine liability) that actually *reduces* confidence in their most dramatic result rather than inflating it. That is the right scientific behavior and should be recognized. The manuscript's principal remaining statistical vulnerability is the triage rubric's sensitivity analysis, which tests noise around the ×5 association premium without questioning whether the premium itself is justified — a relatively minor fix. Taken as a package, the work is a legitimate, reproducible template for competition-aware computational target nomination and early binder generation in a disease where the field genuinely needs differentiated starting points, and the dual-lead output (one arm now stronger than the other on developability grounds, for honest reasons) is more credible after revision than before it.

---

## Reviewer 3 (Round 2)

### Reviewer 3: Methodology, research ethics, and responsible-AI lens — computational drug discovery

---

## Manuscript 1 (single manuscript): "Steering into competitive whitespace"

---

### Brief Summary

This revised preprint reports a human-supervised, AI-assisted computational campaign to nominate differentiated therapeutic targets for esophageal adenocarcinoma (EAC) and design de novo protein binders against them, working entirely from public data (TCGA/cBioPortal genomics, Open Targets, ClinicalTrials.gov, GTEx, AlphaFold DB). Two leads — GUCY2C for a T-cell-engager treatment arm and DKK1 for a neutralizing-trap interception arm — were selected under an explicit novelty steer and druggability filter; eight binder backbones per target were generated via RFdiffusion→ProteinMPNN→Boltz-2, with 15/16 designs clearing ipTM > 0.5. The revision adds GTEx selectivity data, a sequence-composition liability screen, a disclosed triage rubric with Monte Carlo robustness analysis, and a substantially more honest framing of both confidence levels and remaining unknowns.

---

### Strengths

- **Epistemic honesty materially improved.** The revision genuinely lowered claim strength to match evidence rather than patching cosmetically. The abstract now leads with the composition liability of the top-ipTM GUCY2C design; the pilot-scale framing is explicit; the DKK1 precedent population mismatch is stated as a risk rather than buried.
- **Triage rubric fully disclosed and reproducible.** Table S0 provides the exact formula, component values, ordinal anchors, and a 20,000-draw Monte Carlo sensitivity analysis. The candid acknowledgment that DKK1's rank-1 status is trivial (it is the only biologic-addressable interception candidate) is methodologically exemplary and rare in this kind of work.
- **Composition liability screen (ER-3).** The alanine-content and homopolymer analysis of all 16 ProteinMPNN sequences is a real finding — confirming the failure mode reviewers predicted — and it produced a genuine reranking: dkk1_bb1 is now correctly identified as the most robust lead on combined ipTM + developability grounds, while gucy2c_bb2's headline ipTM is appropriately qualified.
- **GTEx selectivity (ER-1).** Quantified normal-tissue expression with a proper comparator panel (CDH17 as a like-minded GI-restricted control; TROP2/TACSTD2 as a cautionary contrast). The 85% GI fraction and ~1.7 TPM top non-GI tissue are specific and traceable.
- **AI-disclosure structure (ER-7).** The three-category framework (deterministic bioinformatics / published ML tools / AI-assisted synthesis with human review) is the most substantive AI disclosure I have seen in a preprint of this type. Explicitly labeling the Table S2 ordinal scores and "whitespace verdict" column as analyst judgment, not measurement, is appropriate.
- **Open questions and gates remain prominently flagged.** O-1 through O-5 and G1–G5 are retained and specific; the Discussion limitation list is honest and reasonably complete.
- **DKK1 precedent segregation (ER-8).** The Klempner/DisTinGuish precedent is now explicitly scoped to advanced gastric/GEJ disease with a statement that it is not evidence for precursor-setting interception.

---

### Major Concerns

**1. The DKK1 GTEx profile is never explained and constitutes an unresolved embarrassment for the selectivity argument.**
Table S4 shows DKK1's top tissue is "Cells_Cultured_fibroblasts" at 852.8 TPM, with GI_fraction = 0.0. This is the opposite of selective GI expression; fibroblast culture artifact (serum-induced Wnt signaling in vitro) almost certainly explains the number, but it is never addressed in the text. The tumor_selectivity score of 3 for DKK1 is stated without justification anywhere in the manuscript. A secreted protein with 852.8 TPM in fibroblasts and effectively no GI enrichment requires explicit engagement: what tissues express DKK1 at physiologically relevant levels in vivo, and does that expression pattern create a therapeutic index problem for a neutralizing trap in a surveillance population? The GTEx data as presented actually deepens concern rather than alleviating it. This is a new problem introduced by ER-1 — the very analysis done to address selectivity created a data exhibit that undermines it for DKK1.

**2. The GUCY2C EAC tumor-expression gap remains unaddressed beyond acknowledgment.**
ER-1 is credited as "done," and the GTEx analysis is genuine. However, the authors' own ER-1 response distinguishes "establishes tissue selectivity" from "establishes EAC tumor-cell expression." For the treatment-arm rationale to be complete, tumor expression must be established, and cBioPortal itself contains RNA expression data (TCGA EAC cohort) that could be queried in silico — mRNA expression from the same n=182 cohort used for the alteration landscape. Not querying GUCY2C mRNA in the TCGA EAC RNA-seq data when that data is in the same public repository used for Figure 1a is an unexplained gap. This is not a wet-lab ask; it is a public-data query. The authors flag the gap; they should close it, or explicitly explain why cBioPortal/UCSC Xena GUCY2C mRNA data in EAC is not informative (e.g., low purity, relevant only for protein, etc.).

**3. Interface residue numbering is ambiguous and unreproducible.**
Methods/S4 reports GUCY2C interface residues as [130,132,146,192,325,326,327,330,347,348,349,378,379] and DKK1 residues as [10,11,12,13,28,29,30,31] in "complex-file numbering." The hotspot definition used as RFdiffusion input uses "isolated-domain numbering." These two numbering systems are never reconciled, and neither is mapped back to canonical UniProt residue positions. A reader cannot verify whether the binder engages the intended CRD2/LRP6-binding region of DKK1 (residues 178–256 UniProt O94907) when the reported interface residues start at position 10 in complex-file numbering. This is a reproducibility and interpretability failure that can be fixed by providing a numbering-offset table or by reporting all residues in UniProt coordinates throughout.

**4. The "no high-similarity binder pairs" diversity claim lacks a defined metric.**
Table S5/Section S5 states "no high-similarity binder pairs detected" with reference to `epitope_cross_reactivity.csv`. No threshold, alignment method, or similarity metric is stated — not sequence identity cutoff, not structural RMSD, not anything. The claim cannot be evaluated or reproduced. Given the high alanine content of several designs, pairwise sequence identity across the alanine-rich designs (dkk1_bb2/bb3/bb5, gucy2c_bb2/bb5) may be artifactually high or may be masked by the generic composition. This should be reported as a defined number with a defined method.

---

### Minor Concerns

**5. Table S4 column redundancy creates confusion.** The `top_nonGI_tissue` and `top_nonGI_TPM` columns are redundant with `top_tissue`/`top_median_TPM` for DKK1 (both point to fibroblasts). The table would be cleaner and more honest if it reported the top *in-vivo* tissue separately from cell-culture artifacts, with a footnote.

**6. The trial-count x-axis (Figure 1b) still lacks a defined unit.** The Methods now state these are "active-EAC-trial intervention mentions" that "were not de-duplicated across combination trials." However, Figure 1b's x-axis label in the legend reads "active-EAC-trial intervention mentions (x, log scale)" — if log scale is used, the statement that individual axis counts "do not sum to 106" should appear in the figure legend, not just Methods, to prevent misreading at the figure level.

**7. Monte Carlo robustness framing for DKK1 is technically correct but potentially misleading.** Table S0 correctly notes DKK1 ranks #1 in 100% of draws because it is the only addressable interception target. This is right to disclose, but the "Rank robustness" table header and column headers (% draws rank #1, nearest competitor) applied to a singleton give a false air of statistical validation. The footnote explanation is adequate but should also appear in the Methods paragraph on triage, not only in the table.

**8. The one-sentence abstract framing of composition liability buries the lede for the treatment arm.** The abstract states gucy2c_bb2 "carries high alanine content" and should be "treated as provisional." But it still leads with ipTM 0.915 before this qualification. Given that the composition analysis is a genuine new finding that reranks the leads, the abstract could foreground dkk1_bb1 as the cleaner lead first.

**9. Author list and affiliation remain placeholders.** This is a preprint and the placeholder is flagged, but for a manuscript that includes a specific AI-disclosure section crediting human supervisors, the identity of those supervisors is relevant to the disclosure's completeness.

---

### Feasible Revisions

**R1. Address the DKK1 GTEx fibroblast artifact explicitly.** Add one paragraph in Results or a footnote to Table S4 explaining that DKK1's top GTEx "tissue" is cell culture (fibroblasts, serum-inducible Wnt signaling artifact), report the top in-vivo tissue by median TPM, and reassess whether the tumor_selectivity score of 3 is defensible given the actual in-vivo normal-tissue profile. If DKK1 is broadly expressed in normal stromal/connective tissue, state the implications for a neutralizing trap in a surveillance population. This is a public-data reanalysis with no wet-lab component.

**R2. Query GUCY2C mRNA in the TCGA EAC cohort.** Retrieve mRNA expression values for GUCY2C from cBioPortal's mRNA expression data for *esca_tcga_pan_can_atlas_2018* (RNA-seq RSEM or z-score). Report median expression and the fraction of EAC tumors with detectable expression (>1 TPM or equivalent). If expression is low or absent, state that explicitly and lower the treatment-arm confidence accordingly. This closes the ER-1 gap the authors themselves identified without requiring any new data generation.

**R3. Provide a numbering reconciliation table for interface residues.** Add a short table (or footnote to S4) mapping isolated-domain numbering → complex-file numbering → UniProt canonical positions for both the GUCY2C and DKK1 hotspot/interface residue sets. Confirm explicitly that the DKK1 interface residues [10,11,12,13,28,29,30,31] in complex numbering correspond to CRD2 residues within 178–256 UniProt O94907.

**R4. Define the binder-diversity metric.** Report pairwise sequence identity (or BLOSUM-based similarity) across all 16 designs, state the threshold used to call "no high-similarity pairs," and report the two most similar pairs numerically. If high-alanine designs are near-identical, state this.

**R5. Move the trial-count non-summability note to the Figure 1b legend.** One sentence added to the figure legend is sufficient; no figure redrawing required.

**R6. Add a one-sentence note in the Methods triage section** flagging that the DKK1 Monte Carlo result is structurally trivial (single-candidate lane), cross-referencing the Table S0 footnote.

---

### Assessment of Each Essential Revision

| ER | Verdict | Notes |
|---|---|---|
| **ER-1** (GUCY2C selectivity/GTEx) | **PARTIALLY ADEQUATE** | GTEx data added and GUCY2C selectivity quantified correctly. The DKK1 fibroblast artifact is a new problem introduced by this analysis and is not addressed. The TCGA EAC tumor-mRNA query — an in-silico step the authors themselves flag as the remaining gap — was not performed. |
| **ER-2** (Orthogonal structural validation) | **ADEQUATE for a preprint, with appropriate scoping** | The authors honestly declined to fabricate GPU-dependent Chai-1/AF-Multimer fold-back numbers they did not run. The lowering of claim language, the explicit O-1 designation, and the addition of the composition screen as the check they *could* run is an acceptable and transparent scoping decision for a preprint. Reject-contingent on orthogonal fold-back would be an out-of-scope demand. |
| **ER-3** (Negative/specificity control) | **ADEQUATE** | The composition liability screen is a genuine contribution. It produced a real finding (gucy2c_bb2 reassigned from headline to provisional), and dkk1_bb1 is correctly identified as the cleaner lead. The remaining monomer/decoy fold-back controls are appropriately named as O-1. The binder-diversity claim lacks a defined metric (Minor Concern 4 / R4), which should be fixed. |
| **ER-4** (Triage rubric disclosure and sensitivity test) | **ADEQUATE** | The disclosed formula, component values, and Monte Carlo analysis are exemplary. The DKK1 singleton acknowledgment is appropriately candid. Minor: the Methods paragraph should cross-reference the DKK1 triviality note (R6). |
| **ER-5** (Pilot-scale reframing) | **ADEQUATE** | "Pilot-scale" language appears in Abstract, Results, and Figure 3 legend. The ipTM > 0.5 permissiveness and single-seed generation are now explicitly stated. |
| **ER-6** (Trial-count methodology and query date) | **PARTIALLY ADEQUATE** | The authors honestly acknowledge they could not recover query strings/snapshot dates from upstream artifacts. The non-summability caveat is appropriate. However, this caveat appears in Methods only, not in the figure legend where readers encounter the counts first (R5). The reproducibility gap acknowledgment is honest; the fix is incomplete. |
| **ER-7** (Non-placeholder AI disclosure) | **ADEQUATE** | The three-category framework is substantive and operationally specific. The explicit labeling of ordinal scores as analyst judgment is the right disclosure. This revision exceeds what most preprints provide. |
| **ER-8** (DKK1 precedent segregation by population) | **ADEQUATE** | The DisTinGuish/DKN-01 precedent is now clearly scoped to advanced GEJ disease. The language "not evidence for DKK1 interception in a largely pre-malignant Barrett's surveillance population" is direct and appropriately limits the inference. |

---

### Recommendation

**Minor revision.**

The revision is substantive, honest, and represents a meaningful improvement over the round-1 submission. The authors correctly declined to fabricate orthogonal fold-back numbers (ER-2), produced a real composition-screen finding that changed the scientific conclusion (ER-3), and delivered a disclosed, sensitivity-tested triage rubric (ER-4). Claim strength is now better calibrated to evidence throughout. Three issues require attention before acceptance: (1) the DKK1 fibroblast artifact in GTEx is a new unaddressed problem introduced by the ER-1 analysis that undermines the interception-arm selectivity argument as currently framed; (2) the GUCY2C EAC tumor-mRNA query from cBioPortal — the authors' own identified gap — is an achievable in-silico step that should be taken; (3) the interface residue numbering is ambiguous in a way that prevents independent verification of whether the binders engage the intended functional sites. All three are addressable by public-data queries and text/table edits with no GPU compute or wet-lab work required.

---

## Cross-Cutting Comment

The revision demonstrates that honest, incremental tightening of an in-silico preprint is both achievable and scientifically valuable: the composition liability screen alone is a genuine finding that changed which design is the lead, and the triage transparency sets a standard that most computational drug-discovery papers do not meet. The one systemic risk that carries through both rounds is that analyses added to address one concern can introduce new problems if not fully interpreted — the DKK1 GTEx fibroblast result being the clearest example here. Future revisions should treat every new data exhibit as requiring the same level of critical engagement as the original claims. The trial-count reproducibility gap (ER-6) remains the most straightforward outstanding item and is the lowest-effort fix on the list; it should be resolved before any submission to a journal that requires methodological reproducibility as a condition of publication.
