Crawler Summary

Agent-Evaluator answer-first brief

LLM agent evaluation framework with 7 Harness Gates (A–G) and 58 metrics. Supports LangChain, CrewAI, AutoGen, DSPy, PydanticAI. Native LLM-as-Judge, OTEL tracing, FastAPI dashboard. Agent Evaluator $1 $1 $1 $1 **Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.** It asks not just "does the agent work well?" but **"is the agent ready for production?"** One decorator line auto-recognizes **24 frameworks** (LangChain, CrewAI, AutoGen, …) and measures **58 metrics (25 Native Trackers + 33 Harness Config)** without touching your agent code — then aggregates Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 10/9/2026.

Freshness

Last checked 10/9/2026

Best For

Agent-Evaluator is best for crewai, multi-agent workflows where OpenClaw compatibility matters.

Not Ideal For

Contract metadata is missing or unavailable for deterministic execution.

Evidence Sources Checked

editorial-content, GITHUB REPOS, runtime-metrics, public facts pack

Agent DossierGITHUB REPOSSafety: 66/100

Agent-Evaluator

LLM agent evaluation framework with 7 Harness Gates (A–G) and 58 metrics. Supports LangChain, CrewAI, AutoGen, DSPy, PydanticAI. Native LLM-as-Judge, OTEL tracing, FastAPI dashboard. Agent Evaluator $1 $1 $1 $1 **Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.** It asks not just "does the agent work well?" but **"is the agent ready for production?"** One decorator line auto-recognizes **24 frameworks** (LangChain, CrewAI, AutoGen, …) and measures **58 metrics (25 Native Trackers + 33 Harness Config)** without touching your agent code — then aggregates

OpenClawself-declared

Public facts

5

Change events

1

Artifacts

0

Freshness

Oct 9, 2026

Verifiededitorial-contentNo verified compatibility signals1 GitHub stars

Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 10/9/2026.

1 GitHub starsTrust evidence available

Trust score

Unknown

Compatibility

OpenClaw

Freshness

Oct 9, 2026

Vendor

Bullpeng72

Artifacts

0

Benchmarks

0

Last release

Unpublished

Executive Summary

Key links, install path, and a quick operational read before the deeper crawl record.

Verifiededitorial-content

Summary

Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 10/9/2026.

Setup snapshot

  1. 1

    Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.

  2. 2

    Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Evidence Ledger

Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.

Verifiededitorial-content
Vendor (1)

Vendor

Bullpeng72

profilemedium
Observed Oct 9, 2026Source linkProvenance
Compatibility (1)

Protocol compatibility

OpenClaw

contractmedium
Observed Oct 9, 2026Source linkProvenance
Adoption (1)

Adoption signal

1 GitHub stars

profilemedium
Observed Oct 9, 2026Source linkProvenance
Security (1)

Handshake status

UNKNOWN

trustmedium
Observed unknownSource linkProvenance
Integration (1)

Crawlable docs

6 indexed pages on the official domain

search_documentmedium
Observed Apr 15, 2026Source linkProvenance

Release & Crawl Timeline

Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.

Self-declaredagent-index

Artifacts Archive

Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.

Self-declaredGITHUB REPOS

Extracted files

0

Examples

6

Snippets

0

Languages

python

Executable Examples

bash

pip install agent-evaluator

python

from agent_evaluator import QuickEval

eval = QuickEval("results/")

@eval.qa
def my_agent(question: str, ground_truth: str = "") -> str:
    return llm.invoke(question)          # your agent code — unchanged

my_agent("What is the capital of South Korea?", ground_truth="Seoul")

eval.save()                                        # results/quickeval.json + .html
eval.gate(tcr=85, accuracy=70, hallucination=5)    # CI/CD gate — sys.exit(1) if unmet

python

@agent_eval(monitor, task_type="qa",
    instructions=InstructionConfig(required_keywords=["Seoul"], fail_on_violation=True),   # Gate A
    loop_detection=LoopDetectionConfig(consecutive_repeat_threshold=6),                    # Gate B
    sla=SLAConfig(p95_ms=3000),                                                            # Gate D
)
def my_agent(question: str, ground_truth: str = "") -> str: ...

bash

pip install "agent-evaluator[examples]"
cd Evaluator_Examples && python ch01_first_eval.py   # ... through ch32_ollama_realtime.py

text

agent_evaluator/
├── decorators.py     # agent_eval · batch_eval · conversation_eval · QuickEval facade (quick_eval.py)
├── gates/            # Gate A–G scoring (gate_a_goal/ … gate_g_observability/) + LiveGuardrail
├── core/trackers/    # 25 Native Trackers (Layer 1 foundation · Layer 2 agentic/security) + monitor.py
├── rca/              # diagnose() — Gate-regression root-cause diagnosis + improvement/experiment logs
├── ontology/         # GATE_GUIDANCE · NATIVE_METRIC_RULES · MAST + single-agent failure taxonomies
├── reporting/        # insights.py (build_insights) + comprehensive_report.py (self-contained HTML)
├── integrations/     # LLMJudge · DeepEval/Ragas adapters · MCP servers · live-guardrail bridges
├── serve/            # FastAPI dashboard ([sdk] extra)
└── cli/              # agent-eval CLI (init, check, gate, decisions, diagnose, abtest, trend, dataset,
                      #   feedback, experiment, target, benchmark, improve, claims, autopilot, monitor,
                      #   opencode, claude)

Evaluator_Examples/   # 32 example files (ch01–ch32)
tests/                # 5,609+ test functions

bash

git clone https://github.com/bullpeng72/Agent-Evaluator.git
cd Agent-Evaluator
pip install -e ".[dev]"

pytest                          # run tests
ruff check agent_evaluator/    # lint
mypy agent_evaluator/          # type check

Docs & README

Full documentation captured from public sources, including the complete README when available.

Self-declaredGITHUB REPOS

Docs source

GITHUB REPOS

Editorial quality

ready

LLM agent evaluation framework with 7 Harness Gates (A–G) and 58 metrics. Supports LangChain, CrewAI, AutoGen, DSPy, PydanticAI. Native LLM-as-Judge, OTEL tracing, FastAPI dashboard. Agent Evaluator $1 $1 $1 $1 **Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.** It asks not just "does the agent work well?" but **"is the agent ready for production?"** One decorator line auto-recognizes **24 frameworks** (LangChain, CrewAI, AutoGen, …) and measures **58 metrics (25 Native Trackers + 33 Harness Config)** without touching your agent code — then aggregates

Full README

Agent Evaluator

PyPI version Python Version License: MIT Version

Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.

It asks not just "does the agent work well?" but "is the agent ready for production?" One decorator line auto-recognizes 24 frameworks (LangChain, CrewAI, AutoGen, …) and measures 58 metrics (25 Native Trackers + 33 Harness Config) without touching your agent code — then aggregates them into 7 Gate pass/warn/fail judgments, a root-cause diagnosis engine for regressions, and statistically valid A/B testing.

pip install agent-evaluator
from agent_evaluator import QuickEval

eval = QuickEval("results/")

@eval.qa
def my_agent(question: str, ground_truth: str = "") -> str:
    return llm.invoke(question)          # your agent code — unchanged

my_agent("What is the capital of South Korea?", ground_truth="Seoul")

eval.save()                                        # results/quickeval.json + .html
eval.gate(tcr=85, accuracy=70, hallucination=5)    # CI/CD gate — sys.exit(1) if unmet

The 7 Harness Gates

| Gate | Area | Judgment Criteria | Harness Config (count) | |------|------|-------------------|----------------------| | A 🟢 | Goal Achievement | Instruction compliance · goal alignment · plan consistency · context retention | InstructionConfig · GoalAlignmentConfig · PlanConfig · SubtaskConfig · ContextRetentionConfig · KnowledgeRetentionConfig (6) | | B 🔵 | Behavioral Integrity | Loop detection · scope deviation · tool safety · state consistency · deadlock detection | LoopDetectionConfig · ScopeConfig · ToolParameterSafetyConfig · ContextWindowConfig · StateConsistencyConfig · DeadlockConfig (6) | | C 🟡 | Reliability | Reproducibility · error recovery rate · hallucination faithfulness · quality floor · idempotency | ReproducibilityConfig · FaultToleranceConfig · GracefulDegradationConfig · RetryConsistencyConfig · IdempotencyConfig (5) | | D 🔵 | Performance Contract | SLA compliance · token efficiency · TTFT variability · cost predictability | SLAConfig · EfficiencyConfig · ResourceBudgetConfig · TTFTVariabilityConfig · CostPredictabilityConfig (5) | | E 🔴 | Security Boundary | Threat severity · compliance · threat response behavior | ThreatSeverityConfig · ComplianceConfig · ThreatResponseConfig (3) | | F 🟣 | Multi-Agent Coordination | Inter-agent consensus · information propagation accuracy · role compliance · conflict resolution | ConsensusConfig · PropagationConfig · AgentRoleConfig · ConflictResolutionConfig (4) | | G 🩵 | Observability | Reasoning explainability · internal state tracking · error diagnosis · latency attribution | ExplainabilityConfig · ObservabilityConfig · ErrorDiagnosisConfig · LatencyAttributionConfig (4) |

Pass any of the 33 Configs above as @agent_eval/@batch_eval/@conversation_eval parameters and PerformanceMonitor auto-aggregates each Gate's pass/warn/fail from the underlying trackers — no separate scoring pass needed.

@agent_eval(monitor, task_type="qa",
    instructions=InstructionConfig(required_keywords=["Seoul"], fail_on_violation=True),   # Gate A
    loop_detection=LoopDetectionConfig(consecutive_repeat_threshold=6),                    # Gate B
    sla=SLAConfig(p95_ms=3000),                                                            # Gate D
)
def my_agent(question: str, ground_truth: str = "") -> str: ...

Full Gate reference: Docs/05_QUALITY_GATE.md · Runnable walkthrough: Evaluator_Examples/ch03_harness_basics.py


What's Inside

  • 3 decorator types — @agent_eval (1 call → 1 result), @batch_eval (1 call → N results), @conversation_eval (N calls → 1 multi-turn result). All non-invasive: your function's signature, return value, and exceptions are untouched. → Docs/01_GETTING_STARTED.md
  • 24 framework adapters — framework="langchain"/"crewai"/"anthropic"/"openai"/… auto-extracts tool_calls/chain_steps/tokens_used from the framework's native response object (duck typing — works without agent-evaluator importing the framework itself). → Docs/03_INTEGRATION_GUIDE.md
  • 58 metrics — 25 Native Trackers (accuracy, hallucination, latency, tool efficiency, 5 security trackers, …) + the 33 Harness Configs above. → Docs/02_METRICS_GUIDE.md or the in-app SDK Reference (agent-eval dashboard → /sdk-docs)
  • Self-contained HTML report — save_to_file() (and agent-eval gate --html-out) write one .html next to the JSON — no server, no external JS. It is laid out in 3 tiers: judgment (a one-line deployment-readiness verdict + a HIGH/MEDIUM/LOW confidence badge + Next actions 1·2·3 + a Path to Green — quantified gap to each failing gate, impact-ordered), iteration (a lifecycle-phase strip — analysis → operations — and a proof panel), and evidence (collapsible groups: per-gate Score Breakdown, worst failure cases each with a tool-call trajectory waterfall, the eval-set design contract, stats, versioning, governance). Pass a baseline and it adds the regressed/new/fixed failure-set diff plus prompt/config change attribution. agent-eval gate --html-summary prints a short Markdown block (verdict + path + the one next command) for a PR body. → Docs/13_OUTPUTS.md
  • CI/CD quality gating — agent-eval gate result.json --tcr 85 --accuracy 70, plus baseline regression detection, per-version baselines, golden-set regression gating, --require-spec-coverage (exit 4 if a declared requirement has no golden case), and --hold-on-undecided (exit 75 — "hold for human" when the verdict is statistically borderline; never overrides a real fail). → Docs/05_QUALITY_GATE.md
  • Root-cause diagnosis (RCA) — agent-eval diagnose / agent_evaluator.rca.diagnose() automates detect → attribute → cross-reference for a Gate regression, and links Gate F findings to the MAST failure-mode taxonomy (Cemri et al., 2025). Candidates and evidence only — HOTL, never a verdict. → Evaluator_Examples/ch28_rca_diagnosis.py
  • Statistically valid A/B testing — agent-eval abtest auto-selects Welch's t-test (2 files), mSPRT always-valid inference (--sequential, safe under repeated peeking), or N-way + FDR correction (3+ files). → Evaluator_Examples/ch29_sequential_ab_test.py
  • Machine-readable insight layer + closed improvement loop — every result JSON carries an extra_metrics.insights object (~65 keys: deployment-readiness verdict + decision_ready, Path to Green, failure clustering, paste-ready fix snippets, per-(gate, change) track record, lifecycle_phase, nondeterminism_repeat, eval_set_delta, spec_coverage, deploy_decision). agent-eval target pins your project SLOs, agent-eval benchmark an external reference distribution, and agent-eval experiment / agent-eval improve register a hypothesis → apply → re-verify loop (improve apply-verify runs the apply in a throw-away git worktree and never merges). Schema-validated, never raises. → Docs/13_OUTPUTS.md
  • Development-support layer (all opt-in, no default behavior change) — create_taskresult(covers=[req_id]) ties golden cases to requirements; a deploy-decision ledger (gate --decision-log + agent-eval decisions record) records who accepted / held / overrode a gate run and why; FaultInjectionConfig injects seeded tool failures / latency into Gate C/D scoring without ever touching the blocking path; run_repeated() folds a K-run verdict-stability summary into insights.nondeterminism_repeat; a StreamingEvaluator(golden_candidate_sink=) queue routes production errors / low-confidence answers to agent-eval dataset review-candidates for human promotion. → Docs/05_QUALITY_GATE.md
  • Real-time guardrail — two reference stacks, plus a host-less mode — the same LiveGuardrail engine blocks a single tool call before it runs (Gate B/E), wired into either AOO (Agent-Evaluator + Ollama + OpenCode — fully local, no cloud model) via agent-eval opencode install, or AC (Agent-Evaluator + Claude Code — native CLI hooks) via agent-eval claude install; identical verdict logic, the difference is the process model (a resident subprocess vs. per-call replay). No host at all? tool_guard() + live_guardrail_session() wrap any Python agent loop directly, and (1.1.0) audit_blocked=True is now the default, audit_log_path= flushes a durable record on exit even if the caller's own error handling is silent, and on_block=webhook_on_block(url) fires an out-of-band alert the instant a call is blocked. A blocked call keeps a redacted excerpt of its command, so agent-eval {claude,opencode} violations (with no query — 1.1.0 — browses recent history; a keyword narrows it) / blocked-detail <task_id> (and the list_violations / show_violation MCP tools) surface what was blocked; doctor now also reports the audit DB's row count proactively, and the HTML report shows blocked attempts directly (insights.blocked_attempts_audit) even though they never move Gate B/E scores. → Docs/06_LIVEGUARDRAIL.md · Docs/07_CLAUDE_CODE_HOOKS.md · Docs/08_AOO_STACK.md · Docs/09_OPENCODE_VS_CLAUDE_CODE.md
  • Dashboard — agent-eval dashboard (FastAPI): Harness Gate breakdown, File Compare with pairwise LLM Judge, anomaly/cost tracking, and a 🔧 Improve tab surfacing the RCA engine.
  • Harness Autopilot — agent-eval autopilot is an optional governance layer on top of this SDK's own Gate/decision data: a multi-task/team registry, a human-in-the-loop approval queue (with dual sign-off for high-stakes kinds), and a phase gate that can require a human approval, a passing Harness Gate verdict, or both, before a task advances — plus its own local dashboard (port 8766). → Docs/15_AUTOPILOT.md

Positioning & Direction

Where it fits

Agent-Evaluator is a deployment-readiness judge and iteration harness — not a tracing platform, and not a metric library:

| Layer | Representative tools | Question they answer | Relation | |-------|--------------------|--------------------|----------| | Tracing / observability | LangSmith · Arize Phoenix · Langfuse | What happened in this run? | Consumes it — ships an OTEL exporter + a Phoenix monitor; not a competitor | | Metric libraries | DeepEval · Ragas · promptfoo | What's the score on metric X? | Plugs them in as adapters; the 25 native trackers mean a base install needs none | | Readiness + iteration | Agent-Evaluator | Is this ready to ship — and if not, what is the smallest set of fixes, in what order? | This layer |

Its center of gravity is where the ship / no-ship decision is made: the CI/CD gate and the analysis → design → development → verification loop before it. Streaming + OTEL cover production, but that is not the focus.

What makes it different

  • Seven non-redundant Gates, not one score. Goal, behavior, reliability, performance, security, multi-agent coordination, observability — each graded pass / warn / fail. "83% accuracy" cannot hide a security regression, because they are scored separately.
  • A verdict and a path, not just numbers. Every report and insights object leads with ship / hold / not-ready + a HIGH/MEDIUM/LOW confidence badge, then an impact-ordered Path to Green with paste-ready @agent_eval snippets — the smallest next step, not a dashboard to interpret.
  • Two closed loops, deliberately separated. A real-time guardrail blocks a dangerous tool call before it runs (Gate B/E); a batch pass grades after. Different code, different timing — the guardrail never scores and the scorer never blocks (FaultInjectionConfig, --hold-on-undecided, everything respects that line).
  • Statistical honesty over a confident number. Wilson CI on the pass-rate; decision_ready → exit 75 ("hold for human") when the verdict is borderline; always-valid A/B inference (safe under repeated peeking); K-run verdict-stability; Benjamini–Hochberg correction across slices. It declines to over-claim on n = 12.
  • Human-on-the-loop by construction. RCA emits candidates + evidence, never a verdict. improve apply-verify runs the apply in a throw-away git worktree and never merges or commits. A deploy-decision ledger records who accepted / held / overrode each gate run, and why.
  • Non-invasive and framework-agnostic. One decorator; your function's signature, return value, and exceptions are untouched. 24 framework adapters by pure duck typing — agent-evaluator never imports the framework itself.
  • A local-first option. The AOO stack (Agent-Evaluator + Ollama + OpenCode) runs the entire loop with no cloud model and no token bill; AC (Agent-Evaluator + Claude Code) targets cloud models. Same verdict engine, different process model.
  • Backed by a written methodology. Harness Methodology (a companion book) states the discipline — decision before code · prove, don't pass · block ≠ score · don't rebuild what already exists · size the model to the task · the human stays accountable — and this SDK is one reference implementation of it.

Direction

The trajectory from 1.0.0 has moved from "score the agent" toward "drive the whole build-and-improve loop, with the human's judgment on record" — and, most recently, toward making the real-time guardrail trustworthy even with no host process watching it:

  • 1.0.0 — the machine-readable insights layer + the target / benchmark / experiment / improve CLI loop.
  • 1.0.5 — Harness Methodology alignment (--hold-on-undecided, human_only_patterns, verdict-stability), a development-support framework (requirement → golden-case coverage, a deploy-decision ledger, thin fault injection, a production → golden candidate queue, improve apply-verify), and the HTML report re-cast as a methodology instrument — three tiers (judgment / iteration / evidence), a lifecycle-phase read, and --html-summary for a PR body.
  • 1.1.0 — LiveGuardrail discovery/durability hardening: a blocked call is durable and discoverable by default now, even for a host-less custom agent loop with no Claude Code/OpenCode bridge and no visible message (tool_guard(audit_blocked=True) default, audit_log_path= crash-safe flush, on_block=webhook_on_block(...) out-of-band alert) — plus keyword-free violations/ list_violations browsing and an HTML report section for everyone, since a blocked attempt never moves Gate B/E scores and a clean scorecard alone would hide it.

Explicit non-goals — the boundary is a deliberate design choice, not a missing feature: it does not author or parse specs (EARS), does not provide a sandbox (E2B / Firecracker — that is the team's infrastructure), does not enforce which model tier a task runs on (that is the runtime's), and never auto-merges or auto-deploys.


Installation

Extras are organized into 5 categories by intent — pick the one(s) that match what you're trying to do. Every category is additive and independent; combine as needed.

| # | Category | Install | What it adds | |---|----------|---------|---------------| | 1 | Base measurement + diagnosis | pip install agent-evaluator | 25 trackers · 33 Harness Config · 7 Gates · LLMJudge · RCA diagnosis engine (agent_evaluator.rca/ontology, no extra deps needed) · full CLI (gate/decisions/diagnose/abtest/trend/dataset/feedback/experiment/target/benchmark/improve/claims) | | 2 | SDK — dashboard + monitoring | pip install "agent-evaluator[sdk]" | FastAPI dashboard (serve), Phoenix/OTEL (otel), Korean RAG PDF processing (pdf+korean) — recommended for most users | | 3 | Real-time guardrail — OpenCode/Claude Code + MCP | pip install "agent-evaluator[mcp]" | search_violations + show_violation, recommend_fix, and ask_insights stdio MCP servers so OpenCode, Claude Code (or another MCP client) can call them as tools during a live session — the underlying functions already work without this (recommend_fix's knowledge is used directly by agent-eval diagnose); this only wires up the MCP protocol layer | | 4 | Your agent's framework | pip install "agent-evaluator[langchain]" (or [crewai]/[autogen]/[dspy]/[pydanticai]/[eval]) | Packages your agent code imports directly — agent-evaluator itself works without them via duck typing; install only what you actually use | | 5 | Examples / full / dev | pip install "agent-evaluator[examples]" | Everything needed to run Evaluator_Examples/ with real (non-mock) DeepEval/Ragas/dashboard/Phoenix output. [full] = category 4's frameworks all at once (⚠️ 10+ min install); [dev] = contributor tooling |

Single-feature extras that don't fit the 5 categories above: [export] (dashboard Parquet/Excel), [wandb], [mlflow]. Full package-by-package breakdown: pyproject.toml.


CLI Commands

| Command | Description | |---------|-------------| | agent-eval init / check | Interactive API key setup / configuration status | | agent-eval dashboard [dir] | FastAPI dashboard web server | | agent-eval gate <result.json> | CI/CD quality gating (--require-spec-coverage exit 4 · --hold-on-undecided exit 75 · --decision-log · --html-out / --html-summary) | | agent-eval decisions list\|record | Deploy-decision ledger — record who accepted / held / overrode / rejected a gate run, and why | | agent-eval diagnose <result.json> | Root-cause diagnosis for a Gate regression | | agent-eval abtest <files...> | Statistical A/B / N-way comparison | | agent-eval trend <dir> | Regression detection across sequential results | | agent-eval dataset build\|promote\|health\|review-candidates | Golden-dataset extraction / HITL promotion / coverage health / production-candidate review queue | | agent-eval feedback export-preferences | Export A/B preference rows (pairwise-judge / annotation / contrast-pair) to JSONL — export only | | agent-eval target set\|show\|clear | Pin project SLOs (.aoo/targets.json) — used by gate and the report's "below target" lines | | agent-eval benchmark set\|show\|clear | Pin an external reference distribution (.aoo/reference.json) for percentile + gap-to-frontier | | agent-eval experiment register\|list\|score | Register a Gate/field hypothesis, score predicted vs actual | | agent-eval improve plan\|start\|verify\|patch\|apply-verify | Closed loop: proposal → experiment → re-verify → outcome log (apply-verify applies in an isolated git worktree, never merges) | | agent-eval monitor | Arize Phoenix + OTEL real-time monitoring | | agent-eval opencode / claude install\|upgrade\|doctor\|test-config\|uninstall | Install & manage the LiveGuardrail OpenCode plugin / Claude Code CLI hooks (test-config asserts the resolved guardrail config against a case file) | | agent-eval opencode / claude violations\|blocked-detail | Search (or, with no query, browse most-recent-first) past Gate B/E blocks; show the exact blocked command for a session | | agent-eval claims add\|list\|release\|audit | Team scope-claim management (.aoo/claims.jsonl) | | agent-eval autopilot install\|dashboard\|new-task\|... | Harness Autopilot — an HITL approval queue, task/team registry, and phase-gate governance layer on top of this SDK's own Gate/decision data, with a local dashboard (port 8766) |


Examples

32 standalone, book-chapter-based files in Evaluator_Examples/ (ch01–ch32), covering everything from a first evaluation to the full RCA/A/B-testing improvement loop:

pip install "agent-evaluator[examples]"
cd Evaluator_Examples && python ch01_first_eval.py   # ... through ch32_ollama_realtime.py

Project Structure

agent_evaluator/
├── decorators.py     # agent_eval · batch_eval · conversation_eval · QuickEval facade (quick_eval.py)
├── gates/            # Gate A–G scoring (gate_a_goal/ … gate_g_observability/) + LiveGuardrail
├── core/trackers/    # 25 Native Trackers (Layer 1 foundation · Layer 2 agentic/security) + monitor.py
├── rca/              # diagnose() — Gate-regression root-cause diagnosis + improvement/experiment logs
├── ontology/         # GATE_GUIDANCE · NATIVE_METRIC_RULES · MAST + single-agent failure taxonomies
├── reporting/        # insights.py (build_insights) + comprehensive_report.py (self-contained HTML)
├── integrations/     # LLMJudge · DeepEval/Ragas adapters · MCP servers · live-guardrail bridges
├── serve/            # FastAPI dashboard ([sdk] extra)
└── cli/              # agent-eval CLI (init, check, gate, decisions, diagnose, abtest, trend, dataset,
                      #   feedback, experiment, target, benchmark, improve, claims, autopilot, monitor,
                      #   opencode, claude)

Evaluator_Examples/   # 32 example files (ch01–ch32)
tests/                # 5,609+ test functions

Changelog

  • v1.1.1 (2026-09-23) — Harness Autopilot: new agent-eval autopilot subcommand — HITL approval queue, multi-task/team registry, local dashboard with full CLI↔dashboard write parity (task/team/approval CRUD, phase transitions gated on both HITL approval and the SDK's own Harness Gate A–G verdict via --require-gate-ready), actor attribution + optimistic-concurrency on task/team edits, ops-page Gate scoreboard and open-experiment visibility, claims audit/filtering. CLI plugin architecture (entry-points, not a hardcoded import). Skills/ grows from 5 to 15 (10 backported) and now ships as real package data, fixing agent-eval autopilot install silently finding nothing on a real (non-editable) pip install; new agent-eval autopilot skills install [NAME|--all]. Additive/opt-in, no Gate/schema change. See CHANGELOG.md for the full list — this release absorbed a large round of hardening found via real multi-week use in the AOO Stack workbook.
  • v1.1.0 (2026-09-11) — LiveGuardrail discovery/durability hardening. tool_guard(audit_blocked=True) now default; on_block=/webhook_on_block() out-of-band block alerts; keyword-free list_violations() / violations browsing; blocked_attempt_capture.max_chars 240→500; insights.blocked_attempts_audit HTML section. Additive/opt-in beyond the default change.
  • v1.0.6 (2026-09-10) — Maintenance: decorators.py split into framework_adapters.py + _eval_shared.py (re-exported, no API change); --help now lists all 18 subcommands.
  • v1.0.5 (2026-09-09) — Harness Methodology alignment + dev-support framework + 3-tier HTML report. --hold-on-undecided, --require-spec-coverage, deploy-decision ledger, run_repeated(), FaultInjectionConfig, dataset review-candidates, improve apply-verify. All opt-in.
  • v1.0.4 (2026-09-08) — Blocked calls keep a redacted command excerpt; new show_violation MCP tool + {claude,opencode} violations/blocked-detail CLI.
  • v1.0.3 (2026-09-08) — Fixed search_violations MCP to point at the Claude Code DB (was defaulting to OpenCode's).
  • v1.0.2 (2026-09-04) — Phoenix/OTEL: version-scoped pin, span/trace/session annotations, ASCII console output.
  • v1.0.1 (2026-09-03) — Report-generation hardening (malformed JSON no longer crashes), dashboard/report value parity, English-only runtime output.
  • v1.0.0 (2026-08-31) — General Availability: machine-readable insights layer (~62 keys) + target/benchmark/experiment/improve CLI loop.

Full history (incl. the 1.0.0-rc.1–rc4 series): CHANGELOG.md.


Documentation

| | | |---|---| | Docs/01_GETTING_STARTED.md | Decorators, QuickEval, first evaluation | | Docs/02_METRICS_GUIDE.md | All 58 metrics — formulas, activation conditions | | Docs/03_INTEGRATION_GUIDE.md | 24 framework adapters, auto-detection | | Docs/04_DATA_GUIDE.md | Golden datasets, evaluation data design | | Docs/05_QUALITY_GATE.md | Harness Gates, CI/CD gating, RCA diagnosis | | Docs/06_LIVEGUARDRAIL.md | LiveGuardrail subsystem reference — all usage modes + v1.1.0 discovery/durability hardening | | Docs/07_CLAUDE_CODE_HOOKS.md | AC stack (Agent-Evaluator + Claude Code) — the same guardrail via native Claude Code CLI hooks | | Docs/08_AOO_STACK.md | AOO stack (Agent-Evaluator + Ollama + OpenCode) — the fully-local real-time-guardrail reference integration | | Docs/09_OPENCODE_VS_CLAUDE_CODE.md | AOO vs AC — detailed side-by-side comparison | | Docs/10_OBSERVABILITY.md | Dashboard, alerts, anomaly detection | | Docs/11_OTEL_DATA_REFERENCE.md | Every span, attribute, metric & Phoenix annotation sent over OpenTelemetry | | Docs/12_OPERATIONS.md | Install variants, Docker, per-environment config, performance tuning, troubleshooting | | Docs/13_OUTPUTS.md | Result JSON · HTML reports · CLI · dashboard · AI-runtime output system | | Docs/14_API_REFERENCE.md | Full public API reference | | Docs/15_AUTOPILOT.md | Harness Autopilot — HITL approval queue, task/team registry, phase gates, dashboard | | CHANGELOG.md | Version history |

Also available in-app once the dashboard is running: agent-eval dashboard → SDK Reference (/sdk-docs) and REST API (/api/docs).


Development

git clone https://github.com/bullpeng72/Agent-Evaluator.git
cd Agent-Evaluator
pip install -e ".[dev]"

pytest                          # run tests
ruff check agent_evaluator/    # lint
mypy agent_evaluator/          # type check

License

MIT — see LICENSE.

Contract & API

Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.

MissingGITHUB REPOS

Contract coverage

Status

missing

Auth

None

Streaming

No

Data region

Unspecified

Protocol support

OpenClaw: self-declared

Requires: none

Forbidden: none

Guardrails

Operational confidence: low

No positive guardrails captured.
Invocation examples
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust"

Reliability & Benchmarks

Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.

Missingruntime-metrics

Trust signals

Handshake

UNKNOWN

Confidence

unknown

Attempts 30d

unknown

Fallback rate

unknown

Runtime metrics

Observed P50

unknown

Observed P95

unknown

Rate limit

unknown

Estimated cost

unknown

Do not use if

Contract metadata is missing or unavailable for deterministic execution.
No benchmark suites or observed failure patterns are available.

Media & Demo

Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.

Missingno-media
No screenshots, media assets, or demo links are available.

Related Agents

Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.

Self-declaredprotocol-neighbors
Github ReposUpdated 2h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW
Machine Appendix

Contract JSON

{
  "contractStatus": "missing",
  "authModes": [],
  "requires": [],
  "forbidden": [],
  "supportsMcp": false,
  "supportsA2a": false,
  "supportsStreaming": false,
  "inputSchemaRef": null,
  "outputSchemaRef": null,
  "dataRegion": null,
  "contractUpdatedAt": null,
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Invocation Guide

{
  "preferredApi": {
    "snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/snapshot",
    "contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract",
    "trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust"
  },
  "curlExamples": [
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/snapshot\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust\""
  ],
  "jsonRequestTemplate": {
    "query": "summarize this repo",
    "constraints": {
      "maxLatencyMs": 2000,
      "protocolPreference": [
        "OPENCLEW"
      ]
    }
  },
  "jsonResponseTemplate": {
    "ok": true,
    "result": {
      "summary": "...",
      "confidence": 0.9
    },
    "meta": {
      "source": "GITHUB_REPOS",
      "generatedAt": "2026-10-09T20:53:14.278Z"
    }
  },
  "retryPolicy": {
    "maxAttempts": 3,
    "backoffMs": [
      500,
      1500,
      3500
    ],
    "retryableConditions": [
      "HTTP_429",
      "HTTP_503",
      "NETWORK_TIMEOUT"
    ]
  }
}

Trust JSON

{
  "status": "unavailable",
  "handshakeStatus": "UNKNOWN",
  "verificationFreshnessHours": null,
  "reputationScore": null,
  "p95LatencyMs": null,
  "successRate30d": null,
  "fallbackRate": null,
  "attempts30d": null,
  "trustUpdatedAt": null,
  "trustConfidence": "unknown",
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Capability Matrix

{
  "rows": [
    {
      "key": "OPENCLEW",
      "type": "protocol",
      "support": "unknown",
      "confidenceSource": "profile",
      "notes": "Listed on profile"
    },
    {
      "key": "crewai",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    },
    {
      "key": "multi-agent",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    }
  ],
  "flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}

Facts JSON

[
  {
    "factKey": "vendor",
    "category": "vendor",
    "label": "Vendor",
    "value": "Bullpeng72",
    "href": "https://github.com/bullpeng72/Agent-Evaluator",
    "sourceUrl": "https://github.com/bullpeng72/Agent-Evaluator",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-10-09T12:48:04.956Z",
    "isPublic": true
  },
  {
    "factKey": "protocols",
    "category": "compatibility",
    "label": "Protocol compatibility",
    "value": "OpenClaw",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract",
    "sourceType": "contract",
    "confidence": "medium",
    "observedAt": "2026-10-09T12:48:04.956Z",
    "isPublic": true
  },
  {
    "factKey": "traction",
    "category": "adoption",
    "label": "Adoption signal",
    "value": "1 GitHub stars",
    "href": "https://github.com/bullpeng72/Agent-Evaluator",
    "sourceUrl": "https://github.com/bullpeng72/Agent-Evaluator",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-10-09T12:48:04.956Z",
    "isPublic": true
  },
  {
    "factKey": "docs_crawl",
    "category": "integration",
    "label": "Crawlable docs",
    "value": "6 indexed pages on the official domain",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  },
  {
    "factKey": "handshake_status",
    "category": "security",
    "label": "Handshake status",
    "value": "UNKNOWN",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust",
    "sourceType": "trust",
    "confidence": "medium",
    "observedAt": null,
    "isPublic": true
  }
]

Change Events JSON

[
  {
    "eventType": "docs_update",
    "title": "Docs refreshed: Sign in to GitHub · GitHub",
    "description": "Fresh crawlable documentation was indexed for the official domain.",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  }
]

Sponsored

Ads related to Agent-Evaluator and adjacent AI workflows.