Open therapeutic science, from target to a fundable plan.
I'm Ruth-Anne Pai, PhD — an immunologist, protein designer, and person living with eosinophilic esophagitis (EoE). In one week I built an end-to-end body of work: mining public omics to nominate targets, designing antigen-specific protein therapeutics with ESM-family and folding models, and packaging it all into reusable software, three manuscripts, and a development plan — then proved the same engine generalizes by pointing it at a second disease, esophageal cancer. Everything here is open.
How to read this work. This is a one-week, one-person citizen-science project, produced with substantial AI assistance (Anthropic Claude Science) under my direction. Every result here is computational (in-silico) and has not been experimentally or clinically validated; the manuscripts are preprints that have not been peer reviewed. It is shared openly as a method and a set of hypotheses — not medical advice, and not established results.
Where would you like to start?
The work spans data mining, protein design, software, and a full biotech plan. Pick the entry point that fits you — each card points to the pages built for you.
Start with the Patient Organization Navigator
A discovery-first AI specialist built for advocates and org leaders: it maps where your disease area stands, what has been done, where the gaps are, and where your funding and capacity sit — resources to review, never a prescription.
- → Patient Organization Navigator (built for you)
- → Therapeutic Program Architect (disease-agnostic)
- → Patient-centered differentiation
De novo design across three programs
Read the methods and structures behind the pMHC-II tolerance work, the CCL26/POSTN effector binders, and the GUCY2C/DKK1 cancer-interception binders — three de novo campaigns, one pipeline.
- → Manuscripts (pMHC-II, CCL26/POSTN & EAC)
- → Epitope & design pipelines
Project Tolera — a roadmap to patients
The full package: platform thesis, market, competitive whitespace, regulatory path, budget, and milestones.
- → Project Tolera (business + scientific plan)
- → Scientific evidence
Reusable skills & pipelines
Twelve open repositories — skills for target mining, epitope mapping, protein design, manuscript drafting, patient-org navigation, market analysis, peer review, media, and project archiving.
Omics-mined targets, reproducibly
A ten-stage pipeline from public GEO/PRIDE/CELLxGENE data to a ranked, druggability-annotated target shortlist.
The one-week story
How a citizen-scientist took EoE from "here is a disease" to IND-ready design and a plan — all in public.
A drug-development engine for any disease
The throughline isn't a single molecule — it's a reusable pipeline and a pair of AI specialists that take any disease from public data to a full program and a written manuscript. Two agents do the work end to end: the Therapeutic Program Architect (disease → program) and the Manuscript Architect (program → preprint). See the engine that works for any disease →
To prove it, I ran it end to end on three programs across two diseases:
Personalized pMHC-II tolerance
An antigen-selection engine that turns a person's HLA type into a ranked, manufacturable epitope set — 4,640 peptide×allele evaluations → 832 strong binders — with a folded, groove-validated milk lead on a tolerogenic nanoparticle.
CCL26 / POSTN binders
A 9-cohort transcriptomic meta-analysis surfaced two high-evidence, no-trial targets; 80→1,920→60 de novo binder designs followed — reported with an honest negative-control calibration.
EAC / Barrett's dual-arm program
The same engine pointed at esophageal cancer: TCGA genomics → a whitespace map → two nominated targets (GUCY2C, DKK1) → 16 de novo binders, 15/16 clearing the interface line — proof the method generalizes.
Browse the complete project archive
Every working session behind these programs — target mining, protein design, the business and scientific plan — is captured in a browsable Work Archive: a per-session summary linked to the figures, tables, structures, manuscripts, and code it produced, with a filterable artifact browser and a provenance page for the large datasets. It's built for others to check the work and reuse it for their own tools and programs.