Stealth · AI Health Platform
AI companions for patients the medical system never understood.
An AI health agentfor rare & genetic disease — a knowledge graph of the domain, with memory. I was asked to design the experience: how a patient meets and uses that intelligence. The patient spends five to eight years just reaching a diagnosis, then is largely on their own — researching at 2am, holding records no one explains, already burned by confident systems that were wrong.
One agent, not five. Specialized skills it routes to, each scoped to what the model can actually do — and allowed to say it doesn't know. Every patient gap has something real answering it:
And under all of it, two systems: the Gen UI design system and the eval platform.





When the model authors its own UI, patterns drift and nothing is reviewable. Designers wrote the vocabulary; the agent assembles from a library it didn't author.





One place to review and evaluate every agent conversation, designed with the ML engineer who uses it. I designed the judges and the rubric myself.
Score each conversation on a Likert rubric beside the full trace; comment, filter, sort. Judgment turned into structured labels.
Configurable judges auto-score at scale, every verdict with its reasoning.
Scripted scenario tests catch regressions before they ship.
Reviewers tag issues per conversation; the queue that drives failure analysis.
Eight properties every behavior is checked against. Published separately — it outlives this project.
Four accounts side by side: what we intended, what the patient sees, what they understand, what engineering can ship. The gaps drive the next move.
When engineering says the agent does X and design isn't sure, we run the probe. The output is the next brief.
In active use today, with patients and caregivers, every day. We meet with a few patients every week to improve it.
“I've been navigating a personal health crisis for my daughter for 13+ years, and nothing has ever helped me as much as this.”
An active platform user, a mother
It's a feedback loop: find a problem the AI creates → design the fix → it reveals the next. (Next: Memory.)