---
title: "From disease name to design specification: a self-correcting AI-agent campaign for eosinophilic esophagitis"
subtitle: "Perspective"
author: "[Authors to be completed]"
---

# From disease name to design specification: a self-correcting AI-agent campaign for eosinophilic esophagitis

**Perspective**

*Standfirst —* General-purpose AI agents can now execute long chains of scientific analysis, but their trustworthiness is rarely tested in the open. We describe a human-supervised agent that carried a therapeutic-discovery programme for eosinophilic esophagitis from a disease name to protein-design specifications on public data alone — and we use its documented errors, and their corrections, to argue for how such campaigns should be conducted. The picture that emerges is deliberately unromantic: what we describe is a human-supervised system with a functioning correction loop, not an agent that reliably catches its own errors — of the corrections we log, one was an unprompted machine reversal while the decisive conceptual, evidential and ethical catches were made by the human supervisor.

---

## The trust problem in agentic science

The promise of an AI research agent is continuity: a single system that inventories the relevant public data, harmonises incompatible assays, runs reproducible statistics, validates at single-cell resolution, annotates druggability, and reasons about therapeutic modality — without the losses that accumulate at each hand-off between specialists. The obstacle is not capability but trust. An agent that reports only its successes is indistinguishable from one that fabricates them, and the failure modes of a confident language model are precisely the ones a casual reader will not notice.

We think the way to earn that trust is to show the whole campaign, errors included. Here we report one such campaign, conducted for eosinophilic esophagitis (EoE), and we foreground not the disease findings — which are the subject of companion work — but the *conduct*: where the agent calibrated its claims against the literature, where it reversed itself, and where a human supervisor caught what the agent could not.

EoE is a fitting proving ground. It is a chronic, food-antigen-driven type-2 inflammatory disease of the esophagus with a large treatment-refractory population and, despite the recent approval of dupilumab, no therapy that selectively depletes the expanded esophageal mast compartment^1,2^, diagnosed and managed against well-established clinical criteria^3^, and it has accumulated a large but fragmented public-omics literature — enough to support integrative discovery, but scattered across dozens of studies on incompatible platforms.

## Anatomy of the campaign

The agent worked in four phases — Discovery, Validation, Design, and a prospective antigen-directed framework (**Fig. 1**). Discovery mined public omics to a reproducible molecular signature; Validation subjected the resulting hypotheses to mechanistic, genetic and treatment-response scrutiny; Design converted the surviving targets into structure-grounded specifications; and the prospective phase set out an antigen-directed strategy.

The scale is summarised in **Table 1**. From 57 public datasets, nine case-control bulk transcriptomic cohorts (235 samples: 152 EoE, 83 control) were harmonised to a common gene space and pooled by random-effects meta-analysis into a 567-gene high-confidence signature (383 up, 184 down) — an approach concordant with prior multi-cohort syntheses of the EoE tissue transcriptome^15^ — which was then validated in 166,420 esophageal single cells. The programme nominated three protein-tractable leads, mapped ten food allergens for the antigen-directed arm, and — as the campaign continued — carried the antigen-directed strategy into a worked personalised example calibrated against a real patient's clinical ground truth. In total it produced 214 lineage-tracked artifacts in approximately 23 hours of active compute — the timestamped compute span within a one-week campaign — packaged with four reusable analysis pipelines.

## A supervisor, not a spectator

The campaign was never autonomous by design (**Fig. 2**, **Table 3**). The agent carried the execution volume: dataset retrieval and quality control, per-study differential expression, cross-study meta-analysis, single-cell processing, enrichment and ligand–receptor inference, druggability annotation, structure retrieval, protein-language-model epitope analysis, and drafting. The human set the frame and exercised judgement — defining the disease goal and survey scope, testing biological plausibility, deciding whether a signal was disease-specific, introducing a key piece of prior art the agent had missed, rejecting an ethically unacceptable trial design, and catching conceptual errors that no quantitative check would have surfaced. A third class of decisions was negotiated between the two: the target-prioritisation criteria, the operational definition of "novel", modality choice, and the decision to lead the antigen-directed strategy with the MHC class II arm.

The asymmetry is the point. The agent's contribution was high-volume and mechanical; the human's was sparse and decisive. Neither could have run this campaign alone, and the interesting decisions lived in the negotiated middle.

## What the campaign found

We report the EoE findings compactly; they appear here as evidence that the campaign produced valid science.

In Discovery, the meta-signature recovered every canonical EoE marker with the correct direction — including the eosinophil chemoattractant CCL26/eotaxin-3, the most disease-specific transcript^4^ — over a dominant interferon, IL-6/JAK-STAT and NF-κB inflammatory program superimposed on loss of epithelial barrier and keratinisation. Single-cell analysis showed a roughly 14-fold mast-cell expansion (2.63% versus 0.19% of cells; Mann–Whitney *P* = 7.2×10⁻⁵), consistent with recent single-cell characterisation of the EoE mast compartment^5^, and depletion of basal epithelium. Druggability annotation returned a shortlist dominated by secreted and surface-accessible proteins.

In Validation, diffusion-pseudotime analysis placed the epithelial barrier defect in a developmental frame: EoE epithelium is arrested short of terminal differentiation. Deconvolution resolved the interferon dominance as a type-II/IFN-γ axis sourced from T cells with the epithelium as responder. Patients stratified into a reproducible mild-to-severe inflammatory gradient across all six cohorts tested, consistent with established EoE endotypes^6^. Seven signature genes carried EoE genetic risk, with support concentrated in susceptibility rather than effector genes^7^ — and, using recovered proton-pump-inhibitor response labels, the lead targets persisted in non-responders while normalising in responders, locating them in the refractory population that most needs a novel agent.

In Design, three leads were carried into structure-grounded specifications: SIGLEC6, a mast-restricted receptor, for a depleting antibody or antibody–drug conjugate; IL1RL1/ST2, the IL-33 alarmin receptor^9^, for a blocking antibody mapped to an experimental IL-33/ST2 interface; and CCL26 for neutralisation. Each received epitope-conservation analysis, developability screening and a paired companion-diagnostic assay.

The prospective phase followed the disease biology: because EoE is a CD4/Th2 food-antigen disease — the causal role of dietary antigen is established by the efficacy of elimination diets^10^ and is central to current diagnosis and management^16^ — the antigen-directed arm was scoped to MHC class II, with epitopes mapped across ten food allergens. Two supervisor interventions then reshaped this phase, and both are recorded in the correction ledger. First, the supervisor flagged a 2025 report that identified and functionally validated the first food-antigen-specific T-cell receptor in EoE — a β-casein epitope presented on HLA-DRB1\*07:01, confirmed by receptor transduction and MHC-II tetramer staining^17^. This redirected the arm from binding prediction alone toward a functional, T-cell-assay-anchored readout as the specificity-defining step, and exposed a gap in the agent's allergen panel (β-casein had been omitted); notably, once the panel was completed the agent's own predictor independently recovered the same DR7 β-casein window the experimental study had found. Second, the supervisor rejected the agent's initial recruitment design, which stratified enrolment by HLA-predicted epitope burden: because HLA allele frequencies vary by ancestry, burden-based eligibility would have systematically under-enrolled non-European patients. The scheme was rebuilt as an HLA-agnostic personalised-epitope platform in which every patient is eligible and predicted burden is retained only as a mechanistic covariate.

The phase then culminated in a worked personalised example: an index EoE patient's HLA typing and genetic screen were used to build a personalised epitope panel and a food-trigger diagnostic concept, and the predictions were tested against the patient's blinded, clinically-established food triggers (revealed only after the computational work). Both true triggers — milk and soy — ranked in the top two of eight candidate foods once the casein family was complete (chance probability 1/28 = 0.036), and the five tolerated non-wheat foods ranked low. The single discordance, wheat, was the most instructive result: wheat gliadin is heavily presented by the patient's celiac-associated DQ2.2 heterodimer — a mechanism drawn from the celiac literature, since the prediction pipeline modelled HLA-DRB1 only, not DQ — yet the patient tolerated wheat (celiac having been clinically excluded), forcing an explicit correction of a latent assumption — that presentation capacity predicts pathology. Presentation is a pre-test prior; a driving food additionally requires an expanded pathogenic-effector-Th2 clone, which is what the functional assay measures.

## Calibration and self-correction: the core evidence

An agent's outputs are only as trustworthy as its handling of uncertainty. The campaign's clearest evidence is not that its outputs were plausible but that they were *calibrated*, and that its errors were caught and recorded (**Fig. 3**).

The agent classified twelve principal findings against retrieved literature. Eight were grounded in prior work — seven confirmatory, one a novel refinement of a known observation — three were claimed as genuinely novel, and one was flagged as contradictory (Fig. 3a). The agent did not label everything novel; two-thirds of its findings were anchored to existing work^4,6,7,13^. Sparingly asserted novelty is the behaviour one wants from a discovery agent.

Before describing the episodes, we state plainly what they are drawn from, because a ledger without a denominator is a highlight reel. The campaign produced 214 lineage-tracked deliverables. From these we defined a set of twelve principal findings by a stated rule: a disease-level claim substantive enough to be asserted as a stand-alone result — a named target, a mechanism, an endotype structure, a differentiation phenotype — as opposed to the hundreds of intermediate quantities (per-gene statistics, per-cohort tables, quality-control outputs) that supported them. Each of the twelve is enumerated in the accompanying novelty ledger, and it is these twelve that were classified against the literature in Fig. 3a. We are explicit that the resulting 8-grounded / 3-novel / 1-contradictory split describes this curated set of headline claims — it is not a measured base rate over all analytical outputs, and it should not be read as one; a different granularity of "finding" would give a different ratio. The eight-episode correction ledger (Fig. 3b) is a *separate* selection: it records every instance we identified in which a claim or decision was reversed, contained, caveated, or put through an orthogonal test — the selection rule was "a course-correction occurred," not "an interesting result was produced." We are explicit that this ledger was reconstructed from the campaign's artifacts and transcripts *after* the fact, not logged prospectively; it is therefore a lower bound on the true number of corrections (silent fixes that left no artifact are not counted), and it is curated in the sense that we cannot prove completeness. What it is not is cherry-picked for flattery: it includes the campaign's outright reversal, its contained contradiction, and its unresolved uncertainties alongside its successes.

Within that frame, the eight episodes span machine self-correction, human domain catches, contained contradictions, calibrated uncertainty and validated novelty (Fig. 3b). They fall into two classes: errors the agent caught on its own, and errors only a human supervisor caught.

We use "self-correcting" in a deliberately bounded sense, and state its limits up front. Of the eight, five were machine-initiated and three were human catches. Exactly one machine entry is an unprompted, agent-initiated reversal of a substantive claim; the other four are a contained contradiction, two calibrated caveats (the PPI-refractory persistence and the blinded index-case panel), and one novel claim upheld under orthogonal test (SIGLEC6). The three decisive conceptual, evidential and ethical catches were made by the human supervisor, not the agent (below). Just as important, none of the eight episodes reversed a *core methodological* decision — the cohort harmonisation, the meta-analysis, the single-cell annotation, the statistical thresholds. The corrections we document are conceptual, scope, and ethical; the agent did not audit its own statistical scaffolding, and we do not claim it did. "Self-correcting" here therefore denotes a human-supervised system with a functioning correction loop, not an agent that reliably catches its own errors.

The one clean agent-initiated reversal is instructive. The agent nominated an S100A8/9 alarmin signalling axis from cell–cell communication analysis, then subjected it to an orthogonal check — the direction of the ligand genes across nine bulk cohorts — and found the ligands down-regulated, tracking epithelial de-differentiation rather than an active alarmin program, with the true mast-remodelling signal being S100A4 rather than S100A8/9^14^. The agent reversed its own claim, unprompted.

Three episodes were human domain catches — the class of error no orthogonal quantitative check surfaces, because the analysis is internally consistent and simply answers the wrong question. The first was conceptual: the agent's initial antigen-presentation analysis scored MHC class I (CD8) peptides, and the supervisor noted that a CD4/Th2 food-antigen disease demands the MHC class II arm — a mechanism for which esophageal epithelium itself can act as a non-professional antigen-presenting cell^11,12^. The second was evidential: the supervisor introduced the 2025 functional-TCR report^17^, which reoriented the antigen-directed arm toward a T-cell-assay ground truth and revealed the β-casein panel gap. The third was ethical, and it is the episode we most want to foreground: the agent proposed to enrol the prospective trial by HLA-predicted epitope burden, an efficient-looking design that would have encoded ancestry inequity into both the trial and the resulting product, since HLA allele frequencies differ across ancestral populations. The supervisor rejected it and the design was rebuilt HLA-agnostic. No internal check flags an unfair design as wrong; this is precisely the judgement that must remain human.

Read together, the episodes trace a reproducible boundary — the paper's most generalizable finding, and the reason we log them at all. The agent corrected itself precisely when, and only when, an error was reducible to an orthogonal *quantitative* re-check it could run against data already in hand: the S100 reversal fell out of re-examining ligand direction across cohorts. The three errors it did not catch each required something no internal recomputation supplies — the breadth to know which immunological arm a food-antigen disease implicates, the awareness that a decisive 2025 experiment existed in the literature, and the normative judgement that an efficient recruitment design was unjust. The actionable claim is therefore narrow and, we think, durable: on this evidence, agentic self-correction extends to claims with a computable contradiction and stops at claims requiring external knowledge or values. A campaign that relies on the agent to police the second category is mis-designed; the human supervisor is not a formality but the mechanism that covers it.

The remaining episodes show calibrated handling rather than reversal. A contradiction was contained (filaggrin, canonically down-regulated^8^, appeared up in one recurrence cohort — recorded as cohort-specific, not suppressed). An uncertainty was carried forward (the PPI-refractory finding rested on the only paired cohort available and was flagged as unreplicated). A novel claim was upheld under test (the SIGLEC6 mast-state finding, derived from single-cell data, held when tested independently in eight bulk cohorts). And a latent assumption was calibrated against clinical ground truth: the personalised food-trigger panel, which ranked foods by MHC-II presentation, was tested blind against an index patient's established triggers — both true triggers ranked top-two, but the tolerated-yet-highly-presented wheat exposed that presentation is a prior, not pathology (Fig. 3a, *Blinded clinical calibration*). This episode is a single illustrative case (n = 1), not a validation study; with eight candidate foods, two true triggers landing in the top two would occur by chance roughly 1 time in 28 (*P* = 0.036). This figure assumes the eight candidate foods are exchangeable under the null; it does not model the correlation between population-level allergen prevalence and epitope burden, and so is a rough calibration of surprise rather than a formal test. We weight it as a concrete, hypothesis-generating demonstration — valuable because the ground truth was withheld until after the prediction — rather than as statistical proof, and we do not present it as the campaign's decisive evidence.

## Provenance as a graph, not a promise

Every result in the campaign is an artifact with tracked lineage — the code that produced it, its input artifacts, and its computational environment — so provenance can be walked rather than asserted. **Figure 4** shows the dependency graph behind the SIGLEC6 design brief: 36 upstream artifacts and 83 edges, from raw public downloads and per-study differential expression, through the meta-signature and its single-cell validation, into the SIGLEC6-specific structure, epitope and interface analyses and the cross-referencing ledgers, and finally into the design brief. Any node regenerates from its parents. Beyond individual artifacts, the discovery pipeline was packaged as a reusable analysis skill, so the method transfers to other diseases rather than remaining locked to this instance.

## Lessons for trustworthy agentic science

What worked was breadth at speed with calibrated uncertainty: the agent integrated a large, heterogeneous public-data corpus, applied standard statistical methods competently when it selected them — though, as the ledger makes plain, it never audited its own statistical scaffolding, and that self-audit remained a human responsibility — held continuity across four phases, and — decisively — grounded most of its findings in prior work and reversed itself when its own check demanded it.

What required a human was judgement, not arithmetic. The three human domain catches map the territory. The MHC-I-versus-II mis-framing is the classic case: internally consistent, quantitatively unremarkable, and wrong only from the standpoint of disease biology. The missed functional-TCR paper is a second kind — the agent cannot know what it has not retrieved, and a supervisor who reads the field closes that gap. The HLA-burden recruitment design is the third and most important: an *ethical* failure that every quantitative check would have passed, because an inequitable trial is not a numerically wrong one. Biological-plausibility, prior-art and fairness calls fell to the human throughout. The agent's failure modes are specific and predictable, which is what makes a domain supervisor effective rather than merely reassuring — and it is why ethical oversight cannot be delegated to the agent's own checks.

The hard limits are real. Every finding is computational; none has been validated at the bench. The GPU-dependent design steps — de novo binder-backbone generation and designed-complex refolding — were deferred and delivered as executable specifications. The refractory finding rests on one cohort, and eosinophils, central to the disease, are under-captured by droplet single-cell methods.

From this we draw five practices, none EoE-specific. Require an orthogonal confirmation for any claim asserted as novel. Keep an auditable correction ledger, so reversals and caveats are visible rather than silently overwritten. Keep the human as a supervisor with defined judgement responsibilities — biological plausibility, prior art, and *ethics* — not a passive observer, and treat design choices with equity implications (who a trial enrols, who a product reaches) as requiring explicit human sign-off rather than automated optimisation. Where possible, calibrate against withheld ground truth, as the blinded index-case test did. And track provenance for every artifact, so reproducibility is a property of the record. An agent that can be confidently wrong in ways no internal check catches is safe to deploy only inside a process built to catch exactly that — and the campaign we describe is one attempt to build it.

## Outlook

A domain-supervised agent compressed a discovery-to-design arc — from a disease name to structure-grounded therapeutic specifications, a prospective antigen-directed framework, and a blinded, clinically-calibrated personalised example — into a single working campaign on public data, producing 214 reproducible, lineage-tracked artifacts. The transferable contribution is not the EoE targets but the demonstration that such a campaign can be run in a calibrated, self-correcting, auditable way. End-to-end reach, honest uncertainty and walkable provenance are, together, what make agentic science trustworthy enough to build on.

---

## Methods

Dataset inventory spanned GEO, ArrayExpress, PRIDE and CELLxGENE. Bulk transcriptomic cohorts were harmonised to a common gene space and normalised; per-study differential expression used variance-moderated Welch *t*-tests with Benjamini–Hochberg correction, and cross-study effects were pooled by DerSimonian–Laird random-effects meta-analysis with heterogeneity statistics. Single-cell data were processed with standard quality control, highly-variable-gene selection, batch integration and Leiden clustering, followed by composition-shift testing and per-cell-type differential expression. Enrichment used gene-set and transcription-factor libraries with a ligand–receptor inference method; druggability was annotated from Open Targets; protein structures were drawn from experimental sources and AlphaFold, with epitope conservation analysed by a protein-language model. Full parameters and code are captured in each artifact's lineage record, and the discovery pipeline is available as a reusable analysis skill.

## Data and code availability

All input datasets are public; accessions are listed in the dataset-inventory artifact. Every result is a lineage-tracked artifact, and the discovery pipeline is packaged for reuse. Companion manuscripts (in preparation) report the EoE biology, the pipeline as a method, and the antigen-directed framework in full.

## Author contributions and AI use

This work was produced by an AI research agent operating under continuous human scientific supervision; the division of labour between the two is itself the subject of the paper and is documented in Table 3 and Figure 2. In summary: the AI agent performed data inventory and retrieval, cross-cohort harmonisation, differential-expression and single-cell analysis, statistical computation, literature retrieval and classification, epitope and interface prediction, figure and table generation, and manuscript drafting. The human supervisor set objectives and scope, made the domain and ethical judgements recorded in the correction ledger (the MHC class I→II correction, the rejection of HLA-burden-based recruitment on equity grounds, and the introduction of the Hill/Dilollo functional-TCR evidence), verified claims against source data, and approved all outputs; the human author(s) take final responsibility for the manuscript and its conclusions. No patient-identifiable data were used; all inputs were public datasets. The specific AI system and version, and the named human authors and their individual contributions, will be declared in full in accordance with the journal's AI-use and authorship policies. **[Human author names and per-author CRediT contributions to be completed on submission.]**

## Competing interests

The authors are developing an eosinophilic esophagitis therapeutics program based on the targets and antigen-directed strategy described in the companion research article; this constitutes a financial and intellectual interest in the outcomes reported here. The antigen-directed T-cell-assay concept intersects disclosed third-party prior art^17^, for which freedom-to-operate has not been formally assessed. The work used a proprietary AI research agent under human supervision. Final author-specific competing-interest declarations will be completed to the journal's policy on submission. [Author-specific entries to be completed.]

---

## Display items

- **Figure 1** — Campaign schematic ({{artifact:art_e6ba6c6e-66f0-4493-a968-2fdc33797cf7}})
- **Figure 2** — Human–AI operating model ({{artifact:art_5bc59e84-7439-4b59-bad4-97be36757ca6}})
- **Figure 3** — Calibration and correction ledger: 8-episode ledger plus literature calibration and the blinded index-case clinical calibration ({{artifact:art_fb30a287-1576-4330-bfae-9c296e20c022}})
- **Figure 4** — SIGLEC6 provenance graph ({{artifact:art_fd8d199b-1b1c-49da-b200-c924c1b9e21b}})
- **Table 1** — Campaign metrics ({{artifact:art_fe4f45bd-a5b6-458f-a4dc-7675b33cde51}})
- **Table 3** — Human–agent task allocation ({{artifact:art_4f62152a-2dd2-4c46-b47f-29fe9f57525e}})

---

## References

1. Muir, A. & Falk, G. W. et al. Eosinophilic Esophagitis: A Review. *JAMA* **326**, 1310–1318 (2021). https://doi.org/10.1001/jama.2021.14920
2. Spergel, J. M. & Aceves, S. S. Allergic components of eosinophilic esophagitis. *J. Allergy Clin. Immunol.* **142**, 1–8 (2018). https://doi.org/10.1016/j.jaci.2018.05.001
3. Furuta, G. T. et al. Eosinophilic esophagitis in children and adults: a systematic review and consensus recommendations for diagnosis and treatment. *Gastroenterology* **133**, 1342–1363 (2007). https://doi.org/10.1053/j.gastro.2007.08.017
4. Blanchard, C. et al. Eotaxin-3 and a uniquely conserved gene-expression profile in eosinophilic esophagitis. *J. Clin. Invest.* **116**, 536–547 (2006). https://doi.org/10.1172/jci26679
5. Ben-Baruch Morgenstern, N. et al. Single-cell RNA sequencing of mast cells in eosinophilic esophagitis reveals heterogeneity, local proliferation, and activation that persists in remission. *J. Allergy Clin. Immunol.* **149**, 2062–2077 (2022). https://doi.org/10.1016/j.jaci.2022.02.025
6. Ruffner, M. A. et al. Phenotypes and endotypes in eosinophilic esophagitis. *Ann. Allergy Asthma Immunol.* **124**, 233–239 (2020). https://doi.org/10.1016/j.anai.2019.12.011
7. O'Shea, K. M. et al. Pathophysiology of eosinophilic esophagitis. *Gastroenterology* **154**, 333–345 (2018). https://doi.org/10.1053/j.gastro.2017.06.065
8. Politi, E. et al. Filaggrin and periostin expression is altered in eosinophilic esophagitis and normalized with treatment. *J. Pediatr. Gastroenterol. Nutr.* **65**, 47–52 (2017). https://doi.org/10.1097/mpg.0000000000001419
9. Venturelli, N. et al. Allergic skin sensitization promotes eosinophilic esophagitis through the IL-33–basophil axis in mice. *J. Allergy Clin. Immunol.* **138**, 1367–1380.e5 (2016). https://doi.org/10.1016/j.jaci.2016.02.034
10. Chehade, M. & Sampson, H. A. Food allergy and eosinophilic esophagitis. *Curr. Opin. Allergy Clin. Immunol.* **10**, 231–237 (2010). https://doi.org/10.1097/ACI.0b013e328338cbab
11. Mulder, D. J. et al. Antigen presentation and MHC class II expression by human esophageal epithelial cells. *Am. J. Pathol.* **178**, 744–753 (2011). https://doi.org/10.1016/j.ajpath.2010.10.027
12. Azouz, N. P. & Rothenberg, M. E. Mechanisms of gastrointestinal allergic disorders. *J. Clin. Invest.* **129**, 1419–1430 (2019). https://doi.org/10.1172/JCI124604
13. Rothenberg, M. E. Eosinophilic gastrointestinal disorders (EGID). *J. Allergy Clin. Immunol.* **113**, 11–28 (2004). https://doi.org/10.1016/j.jaci.2003.10.047
14. Laky, K. et al. Epithelial-intrinsic defects in TGFβR signaling drive local allergic inflammation manifesting as eosinophilic esophagitis. *Sci. Immunol.* **8**, eabp9940 (2023). https://doi.org/10.1126/sciimmunol.abp9940
15. Jacobse, J. et al. A synthesis and subgroup analysis of the eosinophilic esophagitis tissue transcriptome. *J. Allergy Clin. Immunol.* **153**, 759–771 (2024). https://doi.org/10.1016/j.jaci.2023.10.002
16. Straumann, A. & Katzka, D. A. Diagnosis and treatment of eosinophilic esophagitis. *Gastroenterology* **154**, 346–359 (2018). https://doi.org/10.1053/j.gastro.2017.05.066
17. Dilollo, J. et al. A molecular basis for milk allergen immune recognition in eosinophilic esophagitis. *J. Allergy Clin. Immunol.* (2025). https://doi.org/10.1016/j.jaci.2025.01.008

*All 17 references were retrieved during the campaign's literature mining and verified against CrossRef. Author lists are abbreviated; complete author strings and DOIs resolve via the links above.*
