A prototype that re-plans a dialysis unit's chair schedule and the paratransit rides home when the day goes wrong: a van breaks down, a patient runs late, a rider sends a message.
Synthetic data only. Every patient, unit, van, note and message in this repository is generated. Nothing comes from a real person, clinic or transport agency, and nothing here has been tested with a real clinic.
Built at Claude Build Day for Healthcare (17 September 2026) and carried on afterwards as a set of pre-registered tests.
A dialysis session lasts about four hours. The unit books chairs back to back, and a paratransit broker books the rides home against the planned end time. When a session starts late or a van breaks down, the patient waits for a ride that no longer fits the day, and a request such as "I need to ride alone" can sit in a note that the re-plan never reads.
| Who | Does what |
|---|---|
| Software (grey) | Builds the re-plan options, does every calculation and checks every rule (src/c2r/verify.py) |
| Claude (purple) | Picks among the options, reads nurses' notes and riders' messages for needs, writes each rider and staff member a short note |
| A person (amber) | Takes any case the rules cannot settle: a nurse or a dispatcher |
Claude chooses; it never computes. Every constraint lives in the verifier, never in a prompt, and no plan is applied unless the verifier passes it.
All figures are from synthetic days. The full evidence is in
docs/evidence/report.md, generated from committed results.
- The software's gain is ride planning. Planned the day before, the joint chair-and-ride solver picks up 85.6% of rides home within 60 minutes of the patient being ready, against 78.9% for a fair dispatcher (+6.6 points, 95% CI +5.5 to +7.8). Read it with the no-shows: 6.37 a day against 3.44. A plan that books rides only and moves no chair does slightly better (by 0.8 points). Nothing here claims that re-timing chairs improves outcomes.
- Claude did not re-plan better than the software alone. Under the rule registered before the test, the live model "does not add value": it met 3 of 4 conditions and made one needless change in 150 decoy runs (CP10).
- Claude helped with reading. Two claims registered after that both held
(CP19, CP20):
- An AI note reader, with a person confirming, found 88.1% of the needs written in nurses' notes; a keyword search found 67.5%. Many test notes were written to defeat keyword search (it found 6 of 127 of those). On plainly worded notes the keyword search found all 199 and the AI reader 74.3%.
- On a day-of surprise with a rider's message, the live model broke fewer patient needs than the software alone (0.42 fewer a day). Much of that gain is rides handed to a person, not needs met.
- The demo-day headline is retired. "Mean wait from 70 minutes to under 2" was measured against historical practice (one return van per shift). A fair dispatcher planning the day before has a mean wait of 39.1 minutes after the patient is ready (CP6, report).
The API spend summed from the archived run ledgers is $133.09 (cost_report.md).
Needs uv and Python 3.12. The commands below run offline and cost nothing.
make setup # create the venv
make synth # generate the seed-42 synthetic day into data/synthetic/42/
make baseline # print the "before" metrics for that day
make showcase # build the showcase page, then open showcase/dist/index.html
make test # full test suite
make gate # eval gate, offline, under 60 sThe showcase page plays a van breakdown in about 95 seconds (Watch) and lets you replay recorded
runs (Try it). See docs/showcase_runbook.md.
Live runs call the Claude API and bill. Put ANTHROPIC_API_KEY=... in a .env file at the repo
root (it is git-ignored), then use make demo or make showcase-live. Without .env, use the
offline stand-ins: make mediate-fake and make perturb-fake.
To check the evidence from a clean clone with no API key, follow section 14 of
docs/evidence/report.md.
| Path | What it holds |
|---|---|
src/c2r/ |
The engine: synthetic data, solver, verifier, unit and broker parties, mediator, explainer, judge |
src/c2r_kit/ |
Pilot kit: c2r-kit export and c2r-kit intake for a day's data |
prompts/ |
Versioned prompts; a prompt is never edited once it carries an eval result |
evals/ |
Tests, the eval gate and the experiment harness |
specs/ |
One spec per checkpoint; specs/INDEX.md lists them in run order |
docs/evidence/ |
Pre-registration, results pages and the evidence report |
runs/ |
Archived runs the evidence and the showcase are computed from |
showcase/ |
The replay player and try-it page |
.claude/, CLAUDE.md |
How it was built with Claude Code: rules, hooks, skills and subagents |
With Claude Code, one checkpoint at a time. Each checkpoint
starts from a spec, is planned before any code is written, and is checked by a reviewer subagent
before it is committed. Python hooks in .claude/hooks/ hold the rules: a code freeze on the
engine, a guard on privacy-sensitive text, and an eval gate that runs in under a minute. Every
test that makes a claim was registered in
docs/evidence/preregistration.md before it ran, and every
change after a run is listed there as a dated deviation.
- Synthetic data only; not tested with a real clinic, broker or patient.
- No health-outcome claim of any kind.
- The judge that scores the notes is not validated against clinicians.
- This is a research prototype, not a medical device or a dispatch system.
0 comments
log in to comment.