Skip to content
@GOATnote-Inc

GOATnote

Clinical-AI safety evaluation and fail-closed supervision, built by an emergency physician. Benchmarks, corpora, and attestation tooling for medical LLMs.

GOATnote

Clinical-AI safety engineering by a board-certified emergency physician who codes. The through-line across these repositories: language models that touch medicine must fail closed, claims must trace to committed artifacts, and evaluation must survive adversarial scrutiny.

Start here

Repository What it is
scribegoat2 Medical LLM safety evaluation pipeline with reproducible benchmarks and high-risk failure analysis
lostbench Multi-turn safety persistence benchmark: does the model hold the line under sustained patient pressure?
healthcraft Emergency-medicine RL environment (Corecraft architecture) with an MCP tool surface
openem-corpus 370-condition emergency medicine knowledge base; agent-compiled, partially physician-reviewed, grep-friendly
abridge Fail-closed supervising layer for clinical agents (triage, communication, coverage surfaces)
radslice Multimodal radiology benchmark across CT, MRI, X-ray, and ultrasound

Everything here is research software. Nothing is a medical device; nothing is for clinical use.

Contact: b@thegoatnote.com

Popular repositories Loading

  1. scribegoat2 scribegoat2 Public

    Open-source medical LLM safety evaluation pipeline with reproducible benchmarks and high-risk clinical failure analysis.

    Python 4 1

  2. robogoat robogoat Public archive

    GPU data-loading/preprocessing kernels for robot learning research. Archived 2026-08; last GPU validation 2025-11-08 (H100 PCIe). Measured speedups documented in-repo.

    Python

  3. lostbench lostbench Public

    Standalone benchmark for multi-turn safety persistence in medical LLM conversations. Measures recommendation monotonicity under sustained patient pressure.

    Python

  4. openem-corpus openem-corpus Public

    The AI-native emergency medicine knowledge base. Agent-compiled, physician-verified, grep-friendly.

    Python 1

  5. safeshift safeshift Public

    Does making the model faster make it less safe? Safety degradation benchmarking under inference optimization.

    Python

  6. radslice radslice Public

    Multimodal radiology LLM benchmark across CT, MRI, X-ray, and Ultrasound

    Python

Repositories

Showing 10 of 20 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…