One command to benchmark AI guardrails and coding agents across safety, security, jailbreak, prompt-injection, and secure-code tasks.
-
Updated
Jun 26, 2026 - Python
One command to benchmark AI guardrails and coding agents across safety, security, jailbreak, prompt-injection, and secure-code tasks.
BYOA Agent orchestration that ensures alignment end-to-end – with built-in eval system.
The runnable companion to the ebook Cracking AI & ML Evaluation System Design Interviews.
Local Codex MCP harness: contracts, persistent RAG memory, raw traces, verification records, governance policy and PASS/FLAG/BLOCK audits, observability reports, harness profiles, eval runs, Meta-Harness-lite promotion records, natural-language harness specs, MCP resources/prompts, multi-client installer, and completion gates.
Evaluate AI agents with Unix-style pipeline commands. Schema-driven adapters for any CLI agent, trajectory capture, pass@k metrics, and multi-run comparison.
面向开发者的中文 Eval Harness 源码教材:解析 lm-evaluation-harness、Inspect AI、OpenAI Evals、Promptfoo、DeepEval 与 Harbor,覆盖任务、运行、评分、统计与发布门禁。
Deterministic synthetic two-party conversation corpus generator for testing AI scoring systems.
LLM-powered clinical extraction + structured evals. Prompt strategies, hallucination detection, and per-field F1 scoring.
Locale-aware eval harness for AI agents. Test whether your agent behaves correctly when the same intent is expressed across different languages.
Document -> structured data pipeline (invoices/receipts) with confidence-based review routing and a real eval harness. Python/FastAPI, runs on free GLM flash models.
Production-oriented template for building AI agent skills as verifiable software components — offline eval harness, source-grounding validators, structured logs/traces/metrics, replay artifacts, and a CI quality gate. Runs fully offline with a deterministic mock model.
Agentic reconciliation-break triage system — bounded LangGraph state machine, MCP trust-boundary tool split, human-approval interrupt, and a real eval harness.
Production-minded LLM eval harness for safety, reliability, cost, and latency analysis.
Cross-model LLM-as-judge eval harness: validate AI judges with Fleiss' kappa / Krippendorff's alpha, not accuracy. Ships a real 7-model panel (Claude, GPT, Gemini, Grok, Qwen, DeepSeek, GLM) you can replay in ~30s, no API key. MIT.
AI-powered patient-facing chatbot for a community healthcare provider in Uganda — Claude on Bedrock with intent routing, formulary search, Bedrock Guardrails, and clinical safety evals
AI tutor agent — architecture & code showcase (student PII and course material excluded)
An evaluation harness that scores LLM reasoning on manufacturing supply chain disruption analysis against hand-derived golden answers.
Document processing pipeline engine — adapters, contracts, domain packs, eval harnesses
Clinical AI platform: LangGraph RAG over PubMed, ICD-10/CPT/HCPCS billing intelligence, 3 role-based portals. FastAPI · PostgreSQL · Pinecone · React · AWS EC2 · Faithfulness 0.93
Read-only BigQuery MCP server with mandatory dry-run cost caps, agent-friendly structured errors, and a Spider/BIRD-style NL-to-SQL eval harness.
To associate your repository with the eval-harness topic, visit your repo's landing page and select "manage topics."