# From disease name to design specification in a single working session: a human-supervised, self-correcting AI-agent campaign for eosinophilic esophagitis therapeutic discovery

**Authors:** [to be completed] · **Affiliations:** [to be completed]
**Manuscript type:** Methods / Perspective with a worked case study
**Corresponding author:** [to be completed]

---

## Abstract

General-purpose AI agents are increasingly capable of executing multi-step scientific analyses, yet published accounts rarely document an agent completing an entire discovery-to-design arc on real data — and almost never report where the agent was wrong. We describe a human-supervised campaign in which an AI agent, working from public data alone, carried a therapeutic-discovery programme for eosinophilic esophagitis (EoE) from a disease name through target discovery, mechanistic validation, computational protein design, and a prospective antigen-directed framework. The agent surveyed 57 public omics datasets, derived a 567-gene cross-study meta-signature from 9 harmonized bulk cohorts (235 samples) and validated it in 166,420 single cells, nominated three protein-tractable leads (SIGLEC6, IL1RL1/ST2, CCL26), and specified an MHC-II/CD4 antigen-directed strategy across 8 food allergens — producing 147 lineage-tracked artifacts in approximately 11.4 hours of active compute. The scientific interest of the campaign lies less in any single output than in its *conduct*: of 12 findings classified against the literature, 8 were grounded in prior work and only 3 claimed as novel, one finding was reversed by the agent's own orthogonal check (an S100A8/9 alarmin axis corrected to S100A4-driven mast remodeling), and one conceptual error was caught by the human supervisor (an antigen-presentation analysis mis-scoped to MHC-I, re-scoped to the disease-relevant MHC-II arm). We present the campaign as a template for *trustworthy* agentic science: require orthogonal confirmation for novel claims, keep an auditable correction ledger, retain the human as a supervisor rather than a spectator, and track provenance for every result. All findings are computational and require experimental validation.

---

## 1. Introduction

Therapeutic-target discovery is rate-limited less by any single analysis than by the number of steps between a disease and an actionable design specification. A team must inventory the relevant public data, harmonize heterogeneous assays, run reproducible statistics, validate at single-cell resolution, cross-reference genetics and treatment-response data, annotate druggability, and only then reason about modality and design. Each hand-off is a place where effort is lost and where analyses drift out of sync with one another.

Eosinophilic esophagitis (EoE) is a useful proving ground for this problem. It is a chronic, food-antigen-driven type-2 inflammatory disease of the esophagus with a large treatment-refractory population and, at present, a narrow targeted-therapy landscape. It has accumulated a substantial body of public omics data across bulk transcriptomics, single-cell RNA-seq, miRNA, methylation, and proteomics — enough to support an integrative discovery effort, but distributed across dozens of studies with incompatible platforms and annotations.

AI agents promise to compress the discovery-to-design arc by executing all of these steps in one continuous session. What the literature largely lacks is a candid, end-to-end account of an agent doing so — and, critically, an honest treatment of where the agent erred and how those errors were caught. An agent that reports only its successes is indistinguishable from one that fabricates; the trustworthiness of agentic science rests on the visibility of its failure modes.

This paper contributes: (i) a documented end-to-end agentic campaign spanning discovery, validation, design, and a prospective framework; (ii) an explicit model of the human–agent division of labor; (iii) empirical evidence of calibration and self-correction, including a machine-initiated reversal and a human-initiated conceptual correction; and (iv) a fully artifact-tracked output set in which every result is reproducible from its inputs. EoE is the worked example; the transferable object is the *practice*.

---

## 2. The campaign at a glance

The campaign proceeded in four phases — Discovery, Validation, Design, and a Prospective antigen-directed framework (**Figure 1**). Discovery mined public omics to a reproducible molecular signature; Validation subjected the resulting hypotheses to mechanistic, genetic, and treatment-response scrutiny; Design converted the surviving targets into structure-grounded specifications; and the Prospective phase laid out an antigen-directed strategy and preclinical roadmap.

The quantitative envelope of the campaign is summarized in **Table 1**. From an inventory of 57 public datasets, 9 case-control bulk transcriptomic cohorts (235 samples: 152 EoE, 83 control) were harmonized to a common gene space, from which a 567-gene high-confidence meta-signature (383 up, 184 down) was derived and then validated in 166,420 esophageal single cells. The programme nominated three protein-tractable leads and mapped eight food allergens for the antigen-directed arm. It produced 147 lineage-tracked deliverable artifacts — figures, data tables, reports, and structures — in approximately 11.4 hours of active compute, packaged alongside one reusable analysis pipeline. The 11.4-hour figure is the active-compute span measured from artifact timestamps; the calendar window was the one-week hackathon in which the work was situated.

![Figure 1. Campaign schematic: the four-phase arc from disease name to design specification, with the two documented intervention points marked (A, agent self-correction; H, human domain catch), and the verified campaign metrics.]({{artifact:art_e6ba6c6e-66f0-4493-a968-2fdc33797cf7}})

*Figure 1. The campaign at a glance. Four phases carry the programme from a disease name to protein-design specifications and a prospective antigen-directed framework. The narrowing data funnel (57 datasets → 235 samples / 166,420 cells → 3 leads → 8 allergens) tracks the convergence from broad mining to specific design. Two process-intervention points are marked at the phase where each occurred: an agent self-correction (A) at the discovery→validation boundary, and a human domain catch (H) in the prospective phase. Metrics are read from the tracked-artifact store; the active-compute span is marked approximate.*

---

## 3. The human–AI operating model

The campaign was not autonomous and was not intended to be. It ran under a division of labor in which the agent carried the execution volume and the human contributed low-frequency, high-leverage judgment (**Figure 2**, **Table 3**).

The agent executed the mechanical and statistical work: dataset retrieval and quality control, per-study differential expression (variance-moderated t-tests), cross-study random-effects meta-analysis, single-cell processing and cell-type differential expression, pathway and regulator and ligand-receptor enrichment, druggability annotation, protein-structure retrieval, protein-language-model epitope analysis, developability screening, and manuscript drafting.

The human set the frame and exercised judgment: defining the disease goal and the scope of the data survey, sanity-checking biological plausibility, judging whether a signal was disease-specific or an artifact, and — most consequentially — catching a conceptual error in the antigen-presentation analysis (Section 5). The human's contributions were sparse in volume but decisive in effect.

A third category of decisions was neither purely executed nor purely directed but negotiated between the two: the criteria for target prioritization, the operational definition of "novel," the choice of therapeutic modality per target, and the decision to lead the antigen-directed strategy with the MHC-II/CD4 arm. These are the decisions where domain values and analytical evidence meet, and they are the ones least suited to full automation.

![Figure 2. Human–AI operating model: task ownership across the four phases, arranged as swimlanes for the human supervisor, the executing agent, and the negotiated middle ground.]({{artifact:art_5bc59e84-7439-4b59-bad4-97be36757ca6}})

*Figure 2. The operating model across the campaign. The agent lane carries the execution volume; the human lane contributes sparse, decisive judgment; a negotiated band holds the decisions where values meet evidence. The two intervention badges sit on the specific tasks they correspond to, reinforcing that self-correction occurred inside the agent's own workflow while the domain catch came from the human.*

---

## 4. Case-study results

The EoE findings are reported here in compressed form; they are the subject of companion manuscripts (in preparation) and appear here as evidence that the campaign produced valid science, not as the primary contribution.

**Discovery.** Random-effects meta-analysis across the harmonized cohorts yielded a 567-gene high-confidence signature reproducible in at least six studies each. It recovered every canonical EoE marker with the correct direction and was dominated by an interferon, IL-6/JAK-STAT, and NF-κB inflammatory program superimposed on a loss of epithelial keratinization and barrier function. Single-cell analysis revealed a roughly 13-fold expansion of mast cells (0.2% to 2.6% of cells) and basal-epithelial depletion. Druggability annotation surfaced a shortlist dominated by secreted and surface-accessible proteins — the substrate class best suited to biologic therapeutics.

**Validation.** Diffusion-pseudotime analysis showed that EoE epithelium is developmentally arrested, with far fewer cells reaching terminal differentiation than in health. Deconvolution resolved the interferon dominance as a type-II/IFN-γ axis sourced from T cells with the epithelium as the primary responder. Patients stratified into a reproducible mild-to-severe inflammatory gradient across all six cohorts tested. Using recovered proton-pump-inhibitor response labels, the lead targets normalized toward control in responders but persisted in non-responders — locating the targets in precisely the refractory population that most needs a novel agent. Seven signature genes carried EoE genetic risk, with support concentrated in susceptibility genes rather than downstream effectors.

**Design.** Three leads were carried into structure-grounded design. SIGLEC6, a mast-cell-restricted receptor expressed on 84.5% of EoE mast cells versus 34.2% in health, was specified for a depleting antibody or antibody-drug conjugate. IL1RL1/ST2, the IL-33 alarmin receptor, was mapped from an experimental IL-33/ST2 complex to a large buried interface suitable for a blocking antibody. CCL26/eotaxin-3, the master eosinophil chemoattractant and most EoE-specific gene, was specified for neutralization. Each lead received interface mapping, protein-language-model epitope-conservation analysis, developability screening, and a paired companion-diagnostic assay.

**Prospective framework.** Because EoE is a food-antigen-driven CD4/Th2 disease, the antigen-directed arm was scoped to MHC class II. The agent mapped MHC-II/CD4 epitopes across eight food allergens, specified an HLA-based patient-stratification scheme, and drafted a T-cell-assay and preclinical roadmap.

---

## 5. Self-correction and calibration

The distinguishing evidence of this campaign is not that the agent produced plausible outputs but that it produced *calibrated* ones — and corrected itself, and was corrected, in ways that left an auditable trail (**Figure 3**).

**Calibration against the literature.** The agent classified twelve principal findings against retrieved literature. Eight were grounded in prior work (seven confirmatory, one confirmatory with a novel refinement), three were claimed as genuinely novel, and one was flagged as contradictory (Figure 3a). The agent did not label everything novel; two-thirds of its findings were explicitly anchored to existing work. This is the behavior one wants from a discovery agent — novelty asserted sparingly and only where the literature does not already cover the ground.

**A machine-initiated reversal.** In Discovery, the agent nominated an S100A8/9 alarmin signaling axis from cell-cell communication analysis. On subjecting the claim to an orthogonal check — the direction of the ligand genes across nine bulk cohorts — the agent found the ligands were down-regulated, tracking the loss of epithelial differentiation rather than an active alarmin program, and that the true mast-remodeling signal was S100A4. The agent reversed its own claim (Figure 3b, row 1). No human prompted this; the reversal was forced by evidence the agent itself gathered.

**A human-initiated conceptual correction.** The agent's first antigen-presentation analysis scored MHC class I (CD8) peptides. The human supervisor noted that EoE is a CD4/Th2 food-antigen disease, for which the disease-relevant antigen-presentation arm is MHC class II. The agent integrated the correction and re-scoped the antigen-directed strategy to lead with class II (Figure 3b, row 2). This is the class of error an agent is least likely to catch unaided — not a statistical slip but a mis-framing that is internally consistent and only visibly wrong from domain knowledge.

**Contradiction contained, not hidden.** Filaggrin (FLG) is canonically down-regulated in active EoE. The agent observed up-regulation in one recurrence cohort, and rather than suppress the discrepancy, recorded it: canonical down-regulation held in eight of nine cohorts, and the reversal was isolated to a single recurrence dataset (Figure 3b, row 3).

**Uncertainty carried forward.** The proton-pump-inhibitor-refractory persistence finding rested on the only paired cohort available. The agent confirmed the result within that cohort and flagged explicitly that it could not be replicated for lack of a second (Figure 3b, row 4) — a caveat carried into every downstream use of the finding.

**Novel claim upheld under test.** The novel SIGLEC6 mast-state finding, derived from single-cell data, was tested independently in eight bulk cohorts and held (up in all eight, significant in seven), confirming it was not an artifact of the single-cell derivation (Figure 3b, row 5).

![Figure 3. Calibration and correction. (a) Classification of twelve findings against the literature. (b) The verification ledger: each novel or surprising claim, the orthogonal test applied, who caught the issue, and the resolution.]({{artifact:art_fb30a287-1576-4330-bfae-9c296e20c022}})

*Figure 3. The campaign's calibration profile and correction ledger. (a) Eight of twelve findings were grounded in prior work; three were claimed novel and one flagged contradictory — no over-claiming of novelty. (b) Five episodes in which a novel or surprising claim was subjected to an orthogonal test, showing the trigger, whether the agent or the human caught the issue, and the effect on the conclusion. Two of the five corrections changed a conclusion; all five are auditable.*

---

## 6. Reproducibility and provenance

Every result in the campaign is an artifact with tracked lineage — the code that produced it, its input artifacts, and its computational environment. Provenance is therefore not a claim but a graph that can be walked.

**Figure 4** shows that graph for the SIGLEC6 design brief: 36 upstream artifacts connected by 83 dependency edges, laid out by dependency depth from raw public inputs on the left to the design brief at the terminus. The brief is not an assertion; it is the endpoint of a traceable chain that runs from GEO downloads and per-study differential expression, through the meta-signature and its single-cell validation, into the SIGLEC6-specific structure, epitope, and interface analyses and the cross-referencing ledgers, and finally into the revised target dossier and the brief itself. Any node can be regenerated from its parents.

Beyond individual artifacts, the campaign produced a reusable object: the discovery pipeline was packaged as a self-contained analysis skill (documentation plus a dependency-light statistics module), so the method transfers to other diseases rather than remaining locked to this one instance.

![Figure 4. Provenance of the SIGLEC6 design brief: the real dependency graph of 36 artifacts and 83 edges, from raw public inputs to the terminal design brief.]({{artifact:art_fd8d199b-1b1c-49da-b200-c924c1b9e21b}})

*Figure 4. The lineage graph behind one design brief. Nodes are colored by campaign phase; the nine per-study differential-expression tables are collapsed to a single node for legibility. The left-to-right arrangement follows dependency depth, terminating at the highlighted SIGLEC6 design brief. Every result in the campaign carries a graph of this kind.*

---

## 7. Discussion

**What worked.** The agent integrated a large, heterogeneous body of public data at a speed no manual pipeline would match, applied appropriate statistics when the methods were specified, maintained continuity across four phases without losing track of earlier decisions, and — most importantly — expressed calibrated uncertainty rather than uniform confidence. The correction ledger is the clearest evidence: an agent that grounds two-thirds of its findings in prior work and reverses itself when its own check contradicts it is behaving as a careful analyst, not a confident confabulator.

**What required a human.** The conceptual mis-framing of the antigen-presentation arm (MHC-I versus MHC-II) is the instructive case. It was not a statistical error and would not have been caught by any orthogonal quantitative check, because the analysis was internally consistent — it was simply answering the wrong question. Catching it required knowing that EoE is a CD4-driven disease. Biological plausibility judgments and disease-specificity calls fell similarly to the human. The lesson is not that the agent is unreliable but that its failure modes are specific and predictable, and that a supervisor who knows the domain closes exactly the gaps the agent cannot see.

**Hard limits.** Every finding reported here is computational. No result has been validated at the bench. The GPU-dependent design steps — de novo binder-backbone generation and designed-complex refolding — were deferred and delivered as executable specifications rather than run. The proton-pump-inhibitor-refractory finding rests on a single cohort. Eosinophils, central to the disease, are under-captured by droplet single-cell methods and their biology was inferred from bulk data. These are limits of the data and the compute available, not of the workflow, but they bound what the campaign can claim.

**Generalizable practice.** The campaign suggests four practices for trustworthy agentic science, none EoE-specific: require an orthogonal confirmation for any claim asserted as novel; maintain an auditable correction ledger so that reversals and caveats are visible rather than silently overwritten; keep the human as a supervisor with defined judgment responsibilities rather than a passive observer; and track provenance for every artifact so that reproducibility is a property of the record, not a promise. An agent that can be confidently wrong in ways no internal check catches is safe to deploy only inside a process built to catch exactly that.

---

## 8. Conclusion

A domain-supervised AI agent compressed a discovery-to-design arc — from a disease name to structure-grounded therapeutic specifications and a prospective antigen-directed framework — into a single working session on public data, producing 147 reproducible, lineage-tracked artifacts. The scientific outputs for EoE are the subject of companion work; the contribution here is the demonstration that such a campaign can be conducted in a *calibrated and self-correcting* way, with its errors and their corrections on the record. That combination — end-to-end reach, honest uncertainty, and auditable provenance — is what makes agentic science trustworthy enough to build on.

---

## Methods (summary)

Dataset inventory spanned GEO, ArrayExpress, PRIDE, and CELLxGENE. Bulk transcriptomic cohorts were harmonized to a common gene space and normalized; per-study differential expression used variance-moderated Welch t-tests with Benjamini-Hochberg correction, and cross-study effects were pooled by DerSimonian-Laird random-effects meta-analysis with heterogeneity statistics. Single-cell data were processed with standard quality control, highly-variable-gene selection, batch integration, and Leiden clustering, followed by composition-shift testing and per-cell-type differential expression. Enrichment used gene-set and transcription-factor libraries and a ligand-receptor inference method. Druggability was annotated from Open Targets. Protein structures were retrieved from experimental sources and AlphaFold; epitope conservation was analyzed with a protein-language model. Full parameters and code are captured in each artifact's lineage record. The discovery pipeline is available as a reusable analysis skill.

## Data and code availability

All input datasets are public (accessions listed in the dataset-inventory artifact). Every result is a lineage-tracked artifact; the discovery pipeline is packaged for reuse. Companion manuscripts (in preparation) report the EoE biology, the pipeline as a method, and the antigen-directed framework in full.

## Author contributions and AI-disclosure statement

[To be completed per target-journal policy.] The analytical work described here was executed by an AI agent under human scientific supervision; the division of labor is documented in Table 3 and Figure 2. This disclosure is provided in the body of the paper, consistent with the paper's own subject matter.

## Competing interests

[To be completed.]

## Limitations

All findings are computational and require experimental validation. The work does not constitute clinical or therapeutic advice. Human oversight was integral to the campaign and to the validity of its outputs.
