AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
Crawler Summary
LLM agent evaluation framework with 7 Harness Gates (A–G) and 58 metrics. Supports LangChain, CrewAI, AutoGen, DSPy, PydanticAI. Native LLM-as-Judge, OTEL tracing, FastAPI dashboard. Agent Evaluator $1 $1 $1 $1 **Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.** It asks not just "does the agent work well?" but **"is the agent ready for production?"** One decorator line auto-recognizes **24 frameworks** (LangChain, CrewAI, AutoGen, …) and measures **58 metrics (25 Native Trackers + 33 Harness Config)** without touching your agent code — then aggregates Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 10/9/2026.
Freshness
Last checked 10/9/2026
Best For
Agent-Evaluator is best for crewai, multi-agent workflows where OpenClaw compatibility matters.
Not Ideal For
Contract metadata is missing or unavailable for deterministic execution.
Evidence Sources Checked
editorial-content, GITHUB REPOS, runtime-metrics, public facts pack
LLM agent evaluation framework with 7 Harness Gates (A–G) and 58 metrics. Supports LangChain, CrewAI, AutoGen, DSPy, PydanticAI. Native LLM-as-Judge, OTEL tracing, FastAPI dashboard. Agent Evaluator $1 $1 $1 $1 **Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.** It asks not just "does the agent work well?" but **"is the agent ready for production?"** One decorator line auto-recognizes **24 frameworks** (LangChain, CrewAI, AutoGen, …) and measures **58 metrics (25 Native Trackers + 33 Harness Config)** without touching your agent code — then aggregates
Public facts
5
Change events
1
Artifacts
0
Freshness
Oct 9, 2026
Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 10/9/2026.
Trust score
Unknown
Compatibility
OpenClaw
Freshness
Oct 9, 2026
Vendor
Bullpeng72
Artifacts
0
Benchmarks
0
Last release
Unpublished
Key links, install path, and a quick operational read before the deeper crawl record.
Summary
Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 10/9/2026.
Setup snapshot
Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.
Vendor
Bullpeng72
Protocol compatibility
OpenClaw
Adoption signal
1 GitHub stars
Handshake status
UNKNOWN
Crawlable docs
6 indexed pages on the official domain
Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.
Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.
Extracted files
0
Examples
6
Snippets
0
Languages
python
bash
pip install agent-evaluator
python
from agent_evaluator import QuickEval
eval = QuickEval("results/")
@eval.qa
def my_agent(question: str, ground_truth: str = "") -> str:
return llm.invoke(question) # your agent code — unchanged
my_agent("What is the capital of South Korea?", ground_truth="Seoul")
eval.save() # results/quickeval.json + .html
eval.gate(tcr=85, accuracy=70, hallucination=5) # CI/CD gate — sys.exit(1) if unmetpython
@agent_eval(monitor, task_type="qa",
instructions=InstructionConfig(required_keywords=["Seoul"], fail_on_violation=True), # Gate A
loop_detection=LoopDetectionConfig(consecutive_repeat_threshold=6), # Gate B
sla=SLAConfig(p95_ms=3000), # Gate D
)
def my_agent(question: str, ground_truth: str = "") -> str: ...bash
pip install "agent-evaluator[examples]" cd Evaluator_Examples && python ch01_first_eval.py # ... through ch32_ollama_realtime.py
text
agent_evaluator/
├── decorators.py # agent_eval · batch_eval · conversation_eval · QuickEval facade (quick_eval.py)
├── gates/ # Gate A–G scoring (gate_a_goal/ … gate_g_observability/) + LiveGuardrail
├── core/trackers/ # 25 Native Trackers (Layer 1 foundation · Layer 2 agentic/security) + monitor.py
├── rca/ # diagnose() — Gate-regression root-cause diagnosis + improvement/experiment logs
├── ontology/ # GATE_GUIDANCE · NATIVE_METRIC_RULES · MAST + single-agent failure taxonomies
├── reporting/ # insights.py (build_insights) + comprehensive_report.py (self-contained HTML)
├── integrations/ # LLMJudge · DeepEval/Ragas adapters · MCP servers · live-guardrail bridges
├── serve/ # FastAPI dashboard ([sdk] extra)
└── cli/ # agent-eval CLI (init, check, gate, decisions, diagnose, abtest, trend, dataset,
# feedback, experiment, target, benchmark, improve, claims, autopilot, monitor,
# opencode, claude)
Evaluator_Examples/ # 32 example files (ch01–ch32)
tests/ # 5,609+ test functionsbash
git clone https://github.com/bullpeng72/Agent-Evaluator.git cd Agent-Evaluator pip install -e ".[dev]" pytest # run tests ruff check agent_evaluator/ # lint mypy agent_evaluator/ # type check
Full documentation captured from public sources, including the complete README when available.
Docs source
GITHUB REPOS
Editorial quality
ready
LLM agent evaluation framework with 7 Harness Gates (A–G) and 58 metrics. Supports LangChain, CrewAI, AutoGen, DSPy, PydanticAI. Native LLM-as-Judge, OTEL tracing, FastAPI dashboard. Agent Evaluator $1 $1 $1 $1 **Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.** It asks not just "does the agent work well?" but **"is the agent ready for production?"** One decorator line auto-recognizes **24 frameworks** (LangChain, CrewAI, AutoGen, …) and measures **58 metrics (25 Native Trackers + 33 Harness Config)** without touching your agent code — then aggregates
Harness Engineering evaluation SDK that judges AI agent deployment readiness through 7 Gates.
It asks not just "does the agent work well?" but "is the agent ready for production?" One decorator line auto-recognizes 24 frameworks (LangChain, CrewAI, AutoGen, …) and measures 58 metrics (25 Native Trackers + 33 Harness Config) without touching your agent code — then aggregates them into 7 Gate pass/warn/fail judgments, a root-cause diagnosis engine for regressions, and statistically valid A/B testing.
pip install agent-evaluator
from agent_evaluator import QuickEval
eval = QuickEval("results/")
@eval.qa
def my_agent(question: str, ground_truth: str = "") -> str:
return llm.invoke(question) # your agent code — unchanged
my_agent("What is the capital of South Korea?", ground_truth="Seoul")
eval.save() # results/quickeval.json + .html
eval.gate(tcr=85, accuracy=70, hallucination=5) # CI/CD gate — sys.exit(1) if unmet
| Gate | Area | Judgment Criteria | Harness Config (count) | |------|------|-------------------|----------------------| | A 🟢 | Goal Achievement | Instruction compliance · goal alignment · plan consistency · context retention | InstructionConfig · GoalAlignmentConfig · PlanConfig · SubtaskConfig · ContextRetentionConfig · KnowledgeRetentionConfig (6) | | B 🔵 | Behavioral Integrity | Loop detection · scope deviation · tool safety · state consistency · deadlock detection | LoopDetectionConfig · ScopeConfig · ToolParameterSafetyConfig · ContextWindowConfig · StateConsistencyConfig · DeadlockConfig (6) | | C 🟡 | Reliability | Reproducibility · error recovery rate · hallucination faithfulness · quality floor · idempotency | ReproducibilityConfig · FaultToleranceConfig · GracefulDegradationConfig · RetryConsistencyConfig · IdempotencyConfig (5) | | D 🔵 | Performance Contract | SLA compliance · token efficiency · TTFT variability · cost predictability | SLAConfig · EfficiencyConfig · ResourceBudgetConfig · TTFTVariabilityConfig · CostPredictabilityConfig (5) | | E 🔴 | Security Boundary | Threat severity · compliance · threat response behavior | ThreatSeverityConfig · ComplianceConfig · ThreatResponseConfig (3) | | F 🟣 | Multi-Agent Coordination | Inter-agent consensus · information propagation accuracy · role compliance · conflict resolution | ConsensusConfig · PropagationConfig · AgentRoleConfig · ConflictResolutionConfig (4) | | G 🩵 | Observability | Reasoning explainability · internal state tracking · error diagnosis · latency attribution | ExplainabilityConfig · ObservabilityConfig · ErrorDiagnosisConfig · LatencyAttributionConfig (4) |
Pass any of the 33 Configs above as @agent_eval/@batch_eval/@conversation_eval parameters and
PerformanceMonitor auto-aggregates each Gate's pass/warn/fail from the underlying trackers — no
separate scoring pass needed.
@agent_eval(monitor, task_type="qa",
instructions=InstructionConfig(required_keywords=["Seoul"], fail_on_violation=True), # Gate A
loop_detection=LoopDetectionConfig(consecutive_repeat_threshold=6), # Gate B
sla=SLAConfig(p95_ms=3000), # Gate D
)
def my_agent(question: str, ground_truth: str = "") -> str: ...
Full Gate reference: Docs/05_QUALITY_GATE.md · Runnable walkthrough:
Evaluator_Examples/ch03_harness_basics.py
@agent_eval (1 call → 1 result), @batch_eval (1 call → N results),
@conversation_eval (N calls → 1 multi-turn result). All non-invasive: your function's signature,
return value, and exceptions are untouched. → Docs/01_GETTING_STARTED.mdframework="langchain"/"crewai"/"anthropic"/"openai"/… auto-extracts
tool_calls/chain_steps/tokens_used from the framework's native response object (duck typing —
works without agent-evaluator importing the framework itself). → Docs/03_INTEGRATION_GUIDE.mdDocs/02_METRICS_GUIDE.md
or the in-app SDK Reference (agent-eval dashboard → /sdk-docs)save_to_file() (and agent-eval gate --html-out) write one .html
next to the JSON — no server, no external JS. It is laid out in 3 tiers: judgment (a one-line
deployment-readiness verdict + a HIGH/MEDIUM/LOW confidence badge + Next actions 1·2·3 + a Path to
Green — quantified gap to each failing gate, impact-ordered), iteration (a lifecycle-phase strip —
analysis → operations — and a proof panel), and evidence (collapsible groups: per-gate Score
Breakdown, worst failure cases each with a tool-call trajectory waterfall, the eval-set design
contract, stats, versioning, governance). Pass a baseline and it adds the regressed/new/fixed
failure-set diff plus prompt/config change attribution. agent-eval gate --html-summary prints a
short Markdown block (verdict + path + the one next command) for a PR body.
→ Docs/13_OUTPUTS.mdagent-eval gate result.json --tcr 85 --accuracy 70, plus baseline
regression detection, per-version baselines, golden-set regression gating, --require-spec-coverage
(exit 4 if a declared requirement has no golden case), and --hold-on-undecided (exit 75 — "hold for
human" when the verdict is statistically borderline; never overrides a real fail).
→ Docs/05_QUALITY_GATE.mdagent-eval diagnose / agent_evaluator.rca.diagnose() automates
detect → attribute → cross-reference for a Gate regression, and links Gate F findings to the MAST
failure-mode taxonomy (Cemri et al., 2025). Candidates and evidence only — HOTL, never a
verdict. → Evaluator_Examples/ch28_rca_diagnosis.pyagent-eval abtest auto-selects Welch's t-test (2 files),
mSPRT always-valid inference (--sequential, safe under repeated peeking), or N-way + FDR correction
(3+ files). → Evaluator_Examples/ch29_sequential_ab_test.pyextra_metrics.insights object (~65 keys: deployment-readiness verdict + decision_ready, Path to
Green, failure clustering, paste-ready fix snippets, per-(gate, change) track record,
lifecycle_phase, nondeterminism_repeat, eval_set_delta, spec_coverage, deploy_decision).
agent-eval target pins your project SLOs, agent-eval benchmark an external reference distribution,
and agent-eval experiment / agent-eval improve register a hypothesis → apply → re-verify loop
(improve apply-verify runs the apply in a throw-away git worktree and never merges).
Schema-validated, never raises.
→ Docs/13_OUTPUTS.mdcreate_taskresult(covers=[req_id]) ties golden cases to requirements; a deploy-decision ledger
(gate --decision-log + agent-eval decisions record) records who accepted / held / overrode a gate
run and why; FaultInjectionConfig injects seeded tool failures / latency into Gate C/D scoring
without ever touching the blocking path; run_repeated() folds a K-run verdict-stability summary
into insights.nondeterminism_repeat; a StreamingEvaluator(golden_candidate_sink=) queue routes
production errors / low-confidence answers to agent-eval dataset review-candidates for human
promotion. → Docs/05_QUALITY_GATE.mdLiveGuardrail
engine blocks a single tool call before it runs (Gate B/E), wired into either AOO
(Agent-Evaluator + Ollama + OpenCode — fully local, no cloud model) via
agent-eval opencode install, or AC (Agent-Evaluator + Claude Code
— native CLI hooks) via agent-eval claude install; identical verdict logic, the difference is the
process model (a resident subprocess vs. per-call replay). No host at all? tool_guard() +
live_guardrail_session() wrap any Python agent loop directly, and (1.1.0) audit_blocked=True is
now the default, audit_log_path= flushes a durable record on exit even if the caller's own error
handling is silent, and on_block=webhook_on_block(url) fires an out-of-band alert the instant a call
is blocked. A blocked call keeps a redacted excerpt of its command, so agent-eval {claude,opencode} violations (with no query — 1.1.0 — browses recent history; a keyword narrows it) /
blocked-detail <task_id> (and the list_violations / show_violation MCP tools) surface what was
blocked; doctor now also reports the audit DB's row count proactively, and the HTML report shows
blocked attempts directly (insights.blocked_attempts_audit) even though they never move Gate B/E
scores. →
Docs/06_LIVEGUARDRAIL.md ·
Docs/07_CLAUDE_CODE_HOOKS.md ·
Docs/08_AOO_STACK.md ·
Docs/09_OPENCODE_VS_CLAUDE_CODE.mdagent-eval dashboard (FastAPI): Harness Gate breakdown, File Compare with pairwise
LLM Judge, anomaly/cost tracking, and a 🔧 Improve tab surfacing the RCA engine.agent-eval autopilot is an optional governance layer on top of this SDK's
own Gate/decision data: a multi-task/team registry, a human-in-the-loop approval queue (with dual
sign-off for high-stakes kinds), and a phase gate that can require a human approval, a passing
Harness Gate verdict, or both, before a task advances — plus its own local dashboard (port 8766).
→ Docs/15_AUTOPILOT.mdAgent-Evaluator is a deployment-readiness judge and iteration harness — not a tracing platform, and not a metric library:
| Layer | Representative tools | Question they answer | Relation | |-------|--------------------|--------------------|----------| | Tracing / observability | LangSmith · Arize Phoenix · Langfuse | What happened in this run? | Consumes it — ships an OTEL exporter + a Phoenix monitor; not a competitor | | Metric libraries | DeepEval · Ragas · promptfoo | What's the score on metric X? | Plugs them in as adapters; the 25 native trackers mean a base install needs none | | Readiness + iteration | Agent-Evaluator | Is this ready to ship — and if not, what is the smallest set of fixes, in what order? | This layer |
Its center of gravity is where the ship / no-ship decision is made: the CI/CD gate and the analysis → design → development → verification loop before it. Streaming + OTEL cover production, but that is not the focus.
insights object leads with
ship / hold / not-ready + a HIGH/MEDIUM/LOW confidence badge, then an impact-ordered Path to Green
with paste-ready @agent_eval snippets — the smallest next step, not a dashboard to interpret.FaultInjectionConfig, --hold-on-undecided,
everything respects that line).decision_ready →
exit 75 ("hold for human") when the verdict is borderline; always-valid A/B inference (safe under
repeated peeking); K-run verdict-stability; Benjamini–Hochberg correction across slices. It declines
to over-claim on n = 12.improve apply-verify runs the apply in a throw-away git worktree and never merges or commits.
A deploy-decision ledger records who accepted / held / overrode each gate run, and why.The trajectory from 1.0.0 has moved from "score the agent" toward "drive the whole build-and-improve loop, with the human's judgment on record" — and, most recently, toward making the real-time guardrail trustworthy even with no host process watching it:
insights layer + the target / benchmark / experiment /
improve CLI loop.--hold-on-undecided, human_only_patterns,
verdict-stability), a development-support framework (requirement → golden-case coverage, a
deploy-decision ledger, thin fault injection, a production → golden candidate queue,
improve apply-verify), and the HTML report re-cast as a methodology instrument — three tiers
(judgment / iteration / evidence), a lifecycle-phase read, and --html-summary for a PR body.LiveGuardrail discovery/durability hardening: a blocked call is durable and
discoverable by default now, even for a host-less custom agent loop with no Claude Code/OpenCode
bridge and no visible message (tool_guard(audit_blocked=True) default, audit_log_path= crash-safe
flush, on_block=webhook_on_block(...) out-of-band alert) — plus keyword-free violations/
list_violations browsing and an HTML report section for everyone, since a blocked attempt never
moves Gate B/E scores and a clean scorecard alone would hide it.Explicit non-goals — the boundary is a deliberate design choice, not a missing feature: it does not author or parse specs (EARS), does not provide a sandbox (E2B / Firecracker — that is the team's infrastructure), does not enforce which model tier a task runs on (that is the runtime's), and never auto-merges or auto-deploys.
Extras are organized into 5 categories by intent — pick the one(s) that match what you're trying to do. Every category is additive and independent; combine as needed.
| # | Category | Install | What it adds |
|---|----------|---------|---------------|
| 1 | Base measurement + diagnosis | pip install agent-evaluator | 25 trackers · 33 Harness Config · 7 Gates · LLMJudge · RCA diagnosis engine (agent_evaluator.rca/ontology, no extra deps needed) · full CLI (gate/decisions/diagnose/abtest/trend/dataset/feedback/experiment/target/benchmark/improve/claims) |
| 2 | SDK — dashboard + monitoring | pip install "agent-evaluator[sdk]" | FastAPI dashboard (serve), Phoenix/OTEL (otel), Korean RAG PDF processing (pdf+korean) — recommended for most users |
| 3 | Real-time guardrail — OpenCode/Claude Code + MCP | pip install "agent-evaluator[mcp]" | search_violations + show_violation, recommend_fix, and ask_insights stdio MCP servers so OpenCode, Claude Code (or another MCP client) can call them as tools during a live session — the underlying functions already work without this (recommend_fix's knowledge is used directly by agent-eval diagnose); this only wires up the MCP protocol layer |
| 4 | Your agent's framework | pip install "agent-evaluator[langchain]" (or [crewai]/[autogen]/[dspy]/[pydanticai]/[eval]) | Packages your agent code imports directly — agent-evaluator itself works without them via duck typing; install only what you actually use |
| 5 | Examples / full / dev | pip install "agent-evaluator[examples]" | Everything needed to run Evaluator_Examples/ with real (non-mock) DeepEval/Ragas/dashboard/Phoenix output. [full] = category 4's frameworks all at once (⚠️ 10+ min install); [dev] = contributor tooling |
Single-feature extras that don't fit the 5 categories above: [export] (dashboard Parquet/Excel),
[wandb], [mlflow]. Full package-by-package breakdown: pyproject.toml.
| Command | Description |
|---------|-------------|
| agent-eval init / check | Interactive API key setup / configuration status |
| agent-eval dashboard [dir] | FastAPI dashboard web server |
| agent-eval gate <result.json> | CI/CD quality gating (--require-spec-coverage exit 4 · --hold-on-undecided exit 75 · --decision-log · --html-out / --html-summary) |
| agent-eval decisions list\|record | Deploy-decision ledger — record who accepted / held / overrode / rejected a gate run, and why |
| agent-eval diagnose <result.json> | Root-cause diagnosis for a Gate regression |
| agent-eval abtest <files...> | Statistical A/B / N-way comparison |
| agent-eval trend <dir> | Regression detection across sequential results |
| agent-eval dataset build\|promote\|health\|review-candidates | Golden-dataset extraction / HITL promotion / coverage health / production-candidate review queue |
| agent-eval feedback export-preferences | Export A/B preference rows (pairwise-judge / annotation / contrast-pair) to JSONL — export only |
| agent-eval target set\|show\|clear | Pin project SLOs (.aoo/targets.json) — used by gate and the report's "below target" lines |
| agent-eval benchmark set\|show\|clear | Pin an external reference distribution (.aoo/reference.json) for percentile + gap-to-frontier |
| agent-eval experiment register\|list\|score | Register a Gate/field hypothesis, score predicted vs actual |
| agent-eval improve plan\|start\|verify\|patch\|apply-verify | Closed loop: proposal → experiment → re-verify → outcome log (apply-verify applies in an isolated git worktree, never merges) |
| agent-eval monitor | Arize Phoenix + OTEL real-time monitoring |
| agent-eval opencode / claude install\|upgrade\|doctor\|test-config\|uninstall | Install & manage the LiveGuardrail OpenCode plugin / Claude Code CLI hooks (test-config asserts the resolved guardrail config against a case file) |
| agent-eval opencode / claude violations\|blocked-detail | Search (or, with no query, browse most-recent-first) past Gate B/E blocks; show the exact blocked command for a session |
| agent-eval claims add\|list\|release\|audit | Team scope-claim management (.aoo/claims.jsonl) |
| agent-eval autopilot install\|dashboard\|new-task\|... | Harness Autopilot — an HITL approval queue, task/team registry, and phase-gate governance layer on top of this SDK's own Gate/decision data, with a local dashboard (port 8766) |
32 standalone, book-chapter-based files in Evaluator_Examples/ (ch01–ch32),
covering everything from a first evaluation to the full RCA/A/B-testing improvement loop:
pip install "agent-evaluator[examples]"
cd Evaluator_Examples && python ch01_first_eval.py # ... through ch32_ollama_realtime.py
agent_evaluator/
├── decorators.py # agent_eval · batch_eval · conversation_eval · QuickEval facade (quick_eval.py)
├── gates/ # Gate A–G scoring (gate_a_goal/ … gate_g_observability/) + LiveGuardrail
├── core/trackers/ # 25 Native Trackers (Layer 1 foundation · Layer 2 agentic/security) + monitor.py
├── rca/ # diagnose() — Gate-regression root-cause diagnosis + improvement/experiment logs
├── ontology/ # GATE_GUIDANCE · NATIVE_METRIC_RULES · MAST + single-agent failure taxonomies
├── reporting/ # insights.py (build_insights) + comprehensive_report.py (self-contained HTML)
├── integrations/ # LLMJudge · DeepEval/Ragas adapters · MCP servers · live-guardrail bridges
├── serve/ # FastAPI dashboard ([sdk] extra)
└── cli/ # agent-eval CLI (init, check, gate, decisions, diagnose, abtest, trend, dataset,
# feedback, experiment, target, benchmark, improve, claims, autopilot, monitor,
# opencode, claude)
Evaluator_Examples/ # 32 example files (ch01–ch32)
tests/ # 5,609+ test functions
agent-eval autopilot subcommand — HITL approval queue, multi-task/team registry, local dashboard with full CLI↔dashboard write parity (task/team/approval CRUD, phase transitions gated on both HITL approval and the SDK's own Harness Gate A–G verdict via --require-gate-ready), actor attribution + optimistic-concurrency on task/team edits, ops-page Gate scoreboard and open-experiment visibility, claims audit/filtering. CLI plugin architecture (entry-points, not a hardcoded import). Skills/ grows from 5 to 15 (10 backported) and now ships as real package data, fixing agent-eval autopilot install silently finding nothing on a real (non-editable) pip install; new agent-eval autopilot skills install [NAME|--all]. Additive/opt-in, no Gate/schema change. See CHANGELOG.md for the full list — this release absorbed a large round of hardening found via real multi-week use in the AOO Stack workbook.tool_guard(audit_blocked=True) now default; on_block=/webhook_on_block() out-of-band block alerts; keyword-free list_violations() / violations browsing; blocked_attempt_capture.max_chars 240→500; insights.blocked_attempts_audit HTML section. Additive/opt-in beyond the default change.decorators.py split into framework_adapters.py + _eval_shared.py (re-exported, no API change); --help now lists all 18 subcommands.--hold-on-undecided, --require-spec-coverage, deploy-decision ledger, run_repeated(), FaultInjectionConfig, dataset review-candidates, improve apply-verify. All opt-in.show_violation MCP tool + {claude,opencode} violations/blocked-detail CLI.search_violations MCP to point at the Claude Code DB (was defaulting to OpenCode's).insights layer (~62 keys) + target/benchmark/experiment/improve CLI loop.Full history (incl. the 1.0.0-rc.1–rc4 series): CHANGELOG.md.
| | |
|---|---|
| Docs/01_GETTING_STARTED.md | Decorators, QuickEval, first evaluation |
| Docs/02_METRICS_GUIDE.md | All 58 metrics — formulas, activation conditions |
| Docs/03_INTEGRATION_GUIDE.md | 24 framework adapters, auto-detection |
| Docs/04_DATA_GUIDE.md | Golden datasets, evaluation data design |
| Docs/05_QUALITY_GATE.md | Harness Gates, CI/CD gating, RCA diagnosis |
| Docs/06_LIVEGUARDRAIL.md | LiveGuardrail subsystem reference — all usage modes + v1.1.0 discovery/durability hardening |
| Docs/07_CLAUDE_CODE_HOOKS.md | AC stack (Agent-Evaluator + Claude Code) — the same guardrail via native Claude Code CLI hooks |
| Docs/08_AOO_STACK.md | AOO stack (Agent-Evaluator + Ollama + OpenCode) — the fully-local real-time-guardrail reference integration |
| Docs/09_OPENCODE_VS_CLAUDE_CODE.md | AOO vs AC — detailed side-by-side comparison |
| Docs/10_OBSERVABILITY.md | Dashboard, alerts, anomaly detection |
| Docs/11_OTEL_DATA_REFERENCE.md | Every span, attribute, metric & Phoenix annotation sent over OpenTelemetry |
| Docs/12_OPERATIONS.md | Install variants, Docker, per-environment config, performance tuning, troubleshooting |
| Docs/13_OUTPUTS.md | Result JSON · HTML reports · CLI · dashboard · AI-runtime output system |
| Docs/14_API_REFERENCE.md | Full public API reference |
| Docs/15_AUTOPILOT.md | Harness Autopilot — HITL approval queue, task/team registry, phase gates, dashboard |
| CHANGELOG.md | Version history |
Also available in-app once the dashboard is running: agent-eval dashboard → SDK Reference
(/sdk-docs) and REST API (/api/docs).
git clone https://github.com/bullpeng72/Agent-Evaluator.git
cd Agent-Evaluator
pip install -e ".[dev]"
pytest # run tests
ruff check agent_evaluator/ # lint
mypy agent_evaluator/ # type check
MIT — see LICENSE.
Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.
Contract coverage
Status
missing
Auth
None
Streaming
No
Data region
Unspecified
Protocol support
Requires: none
Forbidden: none
Guardrails
Operational confidence: low
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust"
Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.
Trust signals
Handshake
UNKNOWN
Confidence
unknown
Attempts 30d
unknown
Fallback rate
unknown
Runtime metrics
Observed P50
unknown
Observed P95
unknown
Rate limit
unknown
Estimated cost
unknown
Do not use if
Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.
Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
The Frontend for Agents & Generative UI. React + Angular
Contract JSON
{
"contractStatus": "missing",
"authModes": [],
"requires": [],
"forbidden": [],
"supportsMcp": false,
"supportsA2a": false,
"supportsStreaming": false,
"inputSchemaRef": null,
"outputSchemaRef": null,
"dataRegion": null,
"contractUpdatedAt": null,
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Invocation Guide
{
"preferredApi": {
"snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/snapshot",
"contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract",
"trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust"
},
"curlExamples": [
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/snapshot\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust\""
],
"jsonRequestTemplate": {
"query": "summarize this repo",
"constraints": {
"maxLatencyMs": 2000,
"protocolPreference": [
"OPENCLEW"
]
}
},
"jsonResponseTemplate": {
"ok": true,
"result": {
"summary": "...",
"confidence": 0.9
},
"meta": {
"source": "GITHUB_REPOS",
"generatedAt": "2026-10-09T20:53:14.278Z"
}
},
"retryPolicy": {
"maxAttempts": 3,
"backoffMs": [
500,
1500,
3500
],
"retryableConditions": [
"HTTP_429",
"HTTP_503",
"NETWORK_TIMEOUT"
]
}
}Trust JSON
{
"status": "unavailable",
"handshakeStatus": "UNKNOWN",
"verificationFreshnessHours": null,
"reputationScore": null,
"p95LatencyMs": null,
"successRate30d": null,
"fallbackRate": null,
"attempts30d": null,
"trustUpdatedAt": null,
"trustConfidence": "unknown",
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Capability Matrix
{
"rows": [
{
"key": "OPENCLEW",
"type": "protocol",
"support": "unknown",
"confidenceSource": "profile",
"notes": "Listed on profile"
},
{
"key": "crewai",
"type": "capability",
"support": "supported",
"confidenceSource": "profile",
"notes": "Declared in agent profile metadata"
},
{
"key": "multi-agent",
"type": "capability",
"support": "supported",
"confidenceSource": "profile",
"notes": "Declared in agent profile metadata"
}
],
"flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}Facts JSON
[
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Bullpeng72",
"href": "https://github.com/bullpeng72/Agent-Evaluator",
"sourceUrl": "https://github.com/bullpeng72/Agent-Evaluator",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T12:48:04.956Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T12:48:04.956Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1 GitHub stars",
"href": "https://github.com/bullpeng72/Agent-Evaluator",
"sourceUrl": "https://github.com/bullpeng72/Agent-Evaluator",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T12:48:04.956Z",
"isPublic": true
},
{
"factKey": "docs_crawl",
"category": "integration",
"label": "Crawlable docs",
"value": "6 indexed pages on the official domain",
"href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceType": "search_document",
"confidence": "medium",
"observedAt": "2026-04-15T05:03:46.393Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-bullpeng72-agent-evaluator/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
]Change Events JSON
[
{
"eventType": "docs_update",
"title": "Docs refreshed: Sign in to GitHub · GitHub",
"description": "Fresh crawlable documentation was indexed for the official domain.",
"href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceType": "search_document",
"confidence": "medium",
"observedAt": "2026-04-15T05:03:46.393Z",
"isPublic": true
}
]Sponsored
Ads related to Agent-Evaluator and adjacent AI workflows.