AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
-
Updated
Aug 12, 2026 - Python
AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
The open-source MultiAgentOps evaluation and verification harness for any industry business workflow.
An end-to-end framework for running, sandboxing, and scoring agentic LLMs on complex data-science and econometric replication tasks.
Detecting Relational Boundary Erosion in AI systems. A framework for testing whether models maintain honest, calibrated, and appropriate boundaries.
Field-level accuracy evaluation of LLM document extraction — POs, packing slips, bills of lading to schema-validated JSON, with committed scoring reports
VLA ≠ VLM. Side-by-side viewer running NVIDIA Alpamayo R1 (vision-language-action) alongside Qwen2.5-VL (vision-language) on the same 44-sec SF dashcam clip at 5 Hz. 220 paired traces. Surfaces what an action-trained model sees that a scene-trained model doesn't, and vice versa.
RAG service that treats abstention as a feature: cited answers, a faithfulness gate, and a measured coverage-vs-false-answer curve. Chunking x retriever evaluation grid vs planted ground truth shows why retrieval metrics alone mislead. From-scratch BM25, LSA + RRF hybrid, FastAPI, MLflow, 31 tests, fully offline CI.
Single-file Python library for scanned-document extraction: measured page quality drives adaptive preprocessing, pluggable OCR/VLM backends, schema-driven extraction with bbox provenance, calibrated confidence, and an eval harness. Zero required dependencies.
Continuous Evaluation Infrastructure for Production AI
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
Wellness verification harness for companion AI. Multi-turn adversarial suites grounded in six decades of mental-health research and current clinical standards (988, VERA-MH) and law (SB 243). Point it at any chat endpoint — get an evidence-backed, reproducible report.
An LLM agent over crash telemetry that must cite the tool result behind every number it states - with an automatic citation checker and eval harness.
A contract-first PDF extraction and document intelligence evaluation harness.
Autonomous financial research agent combining live market data, financial news, sentiment analysis, and private RAG with transparent execution.
Eval harness for coding-agent skills and instructions: clean A/B runs, evidence-backed grading, trigger checks, and reproducible iteration workspaces
Independent, cross-tool, register-split benchmark of anti-AI-slop rulesets. It measures rules, not writers.
Does a CLAUDE.md actually change how Claude behaves? An ablation harness: run adversarial traps with the rules and without them, grade blind, and test whether the difference is real.
Open-source, local-first financial-document Q&A with verified citations, hybrid RAG, GraphRAG experiments, and a manifest-backed evaluation harness.
Fail-closed evaluation harness for government-facing chat systems: reproducible, provenance-stamped audit verdicts, byte-identical across Python 3.11 to 3.14, with no third-party dependencies. A silent or unreadable target scores zero rather than passing by absence. Two public projects pin it by exact commit.
Closed-loop LLM factory in one monorepo: data pipeline, trainer, eval-gated checkpoint promotion, serving, and an agent whose single tool is a self-extending CLI.
Add a description, image, and links to the evaluation-harness topic page so that developers can more easily learn about it.
To associate your repository with the evaluation-harness topic, visit your repo's landing page and select "manage topics."