AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
-
Updated
Aug 12, 2026 - Python
AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
An evaluation harness for medical/health AI agents — reproduce and cover multiple benchmarks under one scoring discipline
The open-source MultiAgentOps evaluation and verification harness for any industry business workflow.
Coding Agent Runtime & Evaluation Harness for controlled repository-level software repair
An end-to-end framework for running, sandboxing, and scoring agentic LLMs on complex data-science and econometric replication tasks.
Teach your agent to work with evals: WHEN you actually need an eval or benchmark, HOW to build one that holds up, and how to read what it tells you. Deterministic-first, tool-agnostic.
Run 23 speech language models on your own audio task through one interface: vLLM, transformers, and API backends behind a single JSONL contract. Ships HEAR, the speaker-attribution benchmark from our EMNLP 2026 paper.
An sdk and framework for evaluating and comparing multiple model outputs using configurable LLM-based jurors
Detecting Relational Boundary Erosion in AI systems. A framework for testing whether models maintain honest, calibrated, and appropriate boundaries.
A contract-first PDF extraction and document intelligence evaluation harness.
VLA ≠ VLM. Side-by-side viewer running NVIDIA Alpamayo R1 (vision-language-action) alongside Qwen2.5-VL (vision-language) on the same 44-sec SF dashcam clip at 5 Hz. 220 paired traces. Surfaces what an action-trained model sees that a scene-trained model doesn't, and vice versa.
Continuous Evaluation Infrastructure for Production AI
QwerySmith: open text-to-SQL models that answer from retrieved evidence with verifiable citations - plus the harness that trains, evaluates, and releases them (Qwen3-8B QLoRA, 3 seeds, measured-evaluation gate)
Field-level accuracy evaluation of LLM document extraction — POs, packing slips, bills of lading to schema-validated JSON, with committed scoring reports
MODA_NER: open fashion attribute extraction — three-track benchmark suite and models by Hopit AI. Companion to hopit-ai/Moda.
Autonomous financial research agent combining live market data, financial news, sentiment analysis, and private RAG with transparent execution.
Closed-loop LLM factory in one monorepo: data pipeline, trainer, eval-gated checkpoint promotion, serving, and an agent whose single tool is a self-extending CLI.
HAEnv — synthetic longitudinal patients with code-derived gold, for evaluating health agents over time
Single-file Python library for scanned-document extraction: measured page quality drives adaptive preprocessing, pluggable OCR/VLM backends, schema-driven extraction with bbox provenance, calibrated confidence, and an eval harness. Zero required dependencies.
Measure your agent harness, find where it wastes the model, and prove the fix worked. Harness-agnostic, agent-agnostic, zero dependencies. Reference implementation of HTP-1.
To associate your repository with the evaluation-harness topic, visit your repo's landing page and select "manage topics."