Clinical-AI safety engineering by a board-certified emergency physician who codes. The through-line across these repositories: language models that touch medicine must fail closed, claims must trace to committed artifacts, and evaluation must survive adversarial scrutiny.
| Repository | What it is |
|---|---|
| scribegoat2 | Medical LLM safety evaluation pipeline with reproducible benchmarks and high-risk failure analysis |
| lostbench | Multi-turn safety persistence benchmark: does the model hold the line under sustained patient pressure? |
| healthcraft | Emergency-medicine RL environment (Corecraft architecture) with an MCP tool surface |
| openem-corpus | 370-condition emergency medicine knowledge base; agent-compiled, partially physician-reviewed, grep-friendly |
| abridge | Fail-closed supervising layer for clinical agents (triage, communication, coverage surfaces) |
| radslice | Multimodal radiology benchmark across CT, MRI, X-ray, and ultrasound |
Everything here is research software. Nothing is a medical device; nothing is for clinical use.
Contact: b@thegoatnote.com