One command to benchmark AI guardrails and coding agents across safety, security, jailbreak, prompt-injection, and secure-code tasks.
-
Updated
Jun 26, 2026 - Python
One command to benchmark AI guardrails and coding agents across safety, security, jailbreak, prompt-injection, and secure-code tasks.
BYOA Agent orchestration that ensures alignment end-to-end – with built-in eval system.
The runnable companion to the ebook Cracking AI & ML Evaluation System Design Interviews.
Local Codex MCP harness: contracts, persistent RAG memory, raw traces, verification records, governance policy and PASS/FLAG/BLOCK audits, observability reports, harness profiles, eval runs, Meta-Harness-lite promotion records, natural-language harness specs, MCP resources/prompts, multi-client installer, and completion gates.
Evaluate AI agents with Unix-style pipeline commands. Schema-driven adapters for any CLI agent, trajectory capture, pass@k metrics, and multi-run comparison.
Deterministic synthetic two-party conversation corpus generator for testing AI scoring systems.
Locale-aware eval harness for AI agents. Test whether your agent behaves correctly when the same intent is expressed across different languages.
LLM-powered clinical extraction + structured evals. Prompt strategies, hallucination detection, and per-field F1 scoring.
Document -> structured data pipeline (invoices/receipts) with confidence-based review routing and a real eval harness. Python/FastAPI, runs on free GLM flash models.
Production-minded LLM eval harness for safety, reliability, cost, and latency analysis.
Cross-model LLM-as-judge eval harness: validate AI judges with Fleiss' kappa / Krippendorff's alpha, not accuracy. Ships a real 7-model panel (Claude, GPT, Gemini, Grok, Qwen, DeepSeek, GLM) you can replay in ~30s, no API key. MIT.
Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, trigger quality, functional quality, regression protection, baseline value, model variance, rollout safety. Never gradients.
Brainfuck-based generator and exact Rust harness for controlled machine-state tasks; v0.4.0-alpha engineering candidate.
Production-oriented template for building AI agent skills as verifiable software components — offline eval harness, source-grounding validators, structured logs/traces/metrics, replay artifacts, and a CI quality gate. Runs fully offline with a deterministic mock model.
Clinical AI platform: LangGraph RAG over PubMed, ICD-10/CPT/HCPCS billing intelligence, 3 role-based portals. FastAPI · PostgreSQL · Pinecone · React · AWS EC2 · Faithfulness 0.93
Document processing pipeline engine — adapters, contracts, domain packs, eval harnesses
AI tutor agent — architecture & code showcase (student PII and course material excluded)
An evaluation harness that scores LLM reasoning on manufacturing supply chain disruption analysis against hand-derived golden answers.
Agent skills for decisions under uncertainty, plus the evaluation harness that measures them: pre-registered predictions enforced by git ancestry, blind LLM-as-a-judge relabeling with chance-corrected agreement, and every run published with raw transcripts.
AI-powered patient-facing chatbot for a community healthcare provider in Uganda — Claude on Bedrock with intent routing, formulary search, Bedrock Guardrails, and clinical safety evals
Add a description, image, and links to the eval-harness topic page so that developers can more easily learn about it.
To associate your repository with the eval-harness topic, visit your repo's landing page and select "manage topics."