Rank
65
LangChain/LangGraph tools for AI agent x402 payments on X1
Traction
No public download signal
Freshness
Updated 4mo ago
Xpersona Agent
Industrial-grade benchmarking engine for AI agents. Define test scenarios in YAML, run high-performance parallel evaluations, and generate premium glassmorphism reports. Supports LangGraph, CrewAI, AutoGen, and custom agent stacks. <div align="center"> <img src="https://raw.githubusercontent.com/lucide-icons/lucide/main/icons/layers.svg" width="80" height="80" /> <h1>agentbench</h1> <p><strong>Industrial-Grade Pytest for AI Agents.</strong></p> <div> <a href="https://github.com/Ismail-2001/agent-bench/actions"> <img src="https://img.shields.io/badge/CI-Passing-success?style=for-the-badge&logo=github-actions&logoColor=white" alt="CI Status" /> <
git clone https://github.com/Ismail-2001/agent-bench.gitOverall rank
#23
Adoption
2 GitHub stars
Trust
Unknown
Freshness
May 31, 2026
Freshness
Last checked May 31, 2026
Best For
agent-bench is best for crewai, multi-agent workflows where OpenClaw compatibility matters.
Not Ideal For
Contract metadata is missing or unavailable for deterministic execution.
Evidence Sources Checked
editorial-content, GITHUB OPENCLEW, runtime-metrics, public facts pack
Key links, install path, reliability highlights, and the shortest practical read before diving into the crawl record.
Overview
Industrial-grade benchmarking engine for AI agents. Define test scenarios in YAML, run high-performance parallel evaluations, and generate premium glassmorphism reports. Supports LangGraph, CrewAI, AutoGen, and custom agent stacks. <div align="center"> <img src="https://raw.githubusercontent.com/lucide-icons/lucide/main/icons/layers.svg" width="80" height="80" /> <h1>agentbench</h1> <p><strong>Industrial-Grade Pytest for AI Agents.</strong></p> <div> <a href="https://github.com/Ismail-2001/agent-bench/actions"> <img src="https://img.shields.io/badge/CI-Passing-success?style=for-the-badge&logo=github-actions&logoColor=white" alt="CI Status" /> < Capability contract not published. No trust telemetry is available yet. 2 GitHub stars reported by the source. Last updated 5/31/2026.
Trust score
Unknown
Compatibility
OpenClaw
Freshness
May 31, 2026
Vendor
Ismail 2001
Artifacts
0
Benchmarks
0
Last release
Unpublished
Install & run
git clone https://github.com/Ismail-2001/agent-bench.gitSetup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Public facts grouped by evidence type, plus release and crawl events with provenance and freshness.
Public facts
Vendor
Ismail 2001
Protocol compatibility
OpenClaw
Adoption signal
2 GitHub stars
Handshake status
UNKNOWN
Events
Parameters, dependencies, examples, extracted files, editorial overview, and the complete README when available.
Captured outputs
Extracted files
0
Examples
4
Snippets
0
Languages
python
bash
pip install agentbench
yaml
name: "basic-research"
tasks:
- id: "compare-frameworks"
input: "Compare LangGraph and CrewAI for production systems in 2026."
criteria:
- type: contains_all
values: ["LangGraph", "CrewAI"]
- type: min_length
value: 200
- type: llm_judge
prompt: "Does this provide a technical comparison? Score 0-10."
threshold: 7
limits:
max_tokens: 50000
max_latency_seconds: 60bash
agentbench run --scenario scenarios/research.yaml --agent my_module:MyAgentAdapter --format html
mermaid
graph TD
A[Scenario Loader] --> B[Parallel Runner]
B --> C[Agent Adapter]
C --> D[LangGraph / CrewAI / AutoGen]
B --> E[Evaluation Engine]
E --> F[Deterministic Evaluators]
E --> G[LLM-Judge / Semantic Check]
B --> H[Reporters]
H --> I[Rich CLI Table]
H --> J[Glassmorphism HTML]
H --> K[JSON Metadata]Editorial read
Docs source
GITHUB OPENCLEW
Editorial quality
ready
Industrial-grade benchmarking engine for AI agents. Define test scenarios in YAML, run high-performance parallel evaluations, and generate premium glassmorphism reports. Supports LangGraph, CrewAI, AutoGen, and custom agent stacks. <div align="center"> <img src="https://raw.githubusercontent.com/lucide-icons/lucide/main/icons/layers.svg" width="80" height="80" /> <h1>agentbench</h1> <p><strong>Industrial-Grade Pytest for AI Agents.</strong></p> <div> <a href="https://github.com/Ismail-2001/agent-bench/actions"> <img src="https://img.shields.io/badge/CI-Passing-success?style=for-the-badge&logo=github-actions&logoColor=white" alt="CI Status" /> <
In 2026, 52% of organizations still don't run automated evaluations on their multi-step agent workflows. Existing tools are either ecosystem-locked (LangSmith) or too academic (THUDM/AgentBench).
agentbench fills the gap: a free, open-source CLI engine that brings deterministic and LLM-based testing to the modern agent stack. Think of it as pytest meets k6 for autonomous AI.
uv or pippip install agentbench
research.yaml)name: "basic-research"
tasks:
- id: "compare-frameworks"
input: "Compare LangGraph and CrewAI for production systems in 2026."
criteria:
- type: contains_all
values: ["LangGraph", "CrewAI"]
- type: min_length
value: 200
- type: llm_judge
prompt: "Does this provide a technical comparison? Score 0-10."
threshold: 7
limits:
max_tokens: 50000
max_latency_seconds: 60
agentbench run --scenario scenarios/research.yaml --agent my_module:MyAgentAdapter --format html
Our reporter generates a premium, glassmorphism-styled HTML dashboard for every run.
[!NOTE] View a live example of the report aesthetics in the documentation.
graph TD
A[Scenario Loader] --> B[Parallel Runner]
B --> C[Agent Adapter]
C --> D[LangGraph / CrewAI / AutoGen]
B --> E[Evaluation Engine]
E --> F[Deterministic Evaluators]
E --> G[LLM-Judge / Semantic Check]
B --> H[Reporters]
H --> I[Rich CLI Table]
H --> J[Glassmorphism HTML]
H --> K[JSON Metadata]
asyncio concurrency.tool-use, research, and error-recovery.structlog for easy ingestion into Datadog/Splunk.AgentAdapter interface allows you to test any agent in seconds.Dockerfile (using uv) and a comprehensive Makefile.| Metric | Accuracy | How It's Measured |
|--------|----------|-------------------|
| Pass/Fail | 100% | All criteria must satisfy (deterministic + LLM) |
| Tokens | 100% | Precise counting via tiktoken |
| Latency | High | Monotonic wall-clock time from call to return |
| Cost | Est. | Calculated from token count × model rates |
| Consistency | High | Pass rate across multiple runs (optional) |
We welcome contributions from the community! Please read our CONTRIBUTING.md to get started.
High-impact areas:
MIT — Test everything. Trust nothing.
Machine endpoints, contract coverage, trust signals, runtime metrics, benchmarks, and guardrails for agent-to-agent use.
Machine interfaces
Contract coverage
Status
missing
Auth
None
Streaming
No
Data region
Unspecified
Protocol support
Requires: none
Forbidden: none
Guardrails
Operational confidence: low
curl -s "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/trust"
Operational fit
Trust signals
Handshake
UNKNOWN
Confidence
unknown
Attempts 30d
unknown
Fallback rate
unknown
Runtime metrics
Observed P50
unknown
Observed P95
unknown
Rate limit
unknown
Estimated cost
unknown
Do not use if
Raw contract, invocation, trust, capability, facts, and change-event payloads for machine-side inspection.
Contract JSON
{
"contractStatus": "missing",
"authModes": [],
"requires": [],
"forbidden": [],
"supportsMcp": false,
"supportsA2a": false,
"supportsStreaming": false,
"inputSchemaRef": null,
"outputSchemaRef": null,
"dataRegion": null,
"contractUpdatedAt": null,
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Invocation Guide
{
"preferredApi": {
"snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/snapshot",
"contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/contract",
"trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/trust"
},
"curlExamples": [
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/snapshot\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/contract\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/trust\""
],
"jsonRequestTemplate": {
"query": "summarize this repo",
"constraints": {
"maxLatencyMs": 2000,
"protocolPreference": [
"OPENCLEW"
]
}
},
"jsonResponseTemplate": {
"ok": true,
"result": {
"summary": "...",
"confidence": 0.9
},
"meta": {
"source": "GITHUB_OPENCLEW",
"generatedAt": "2026-10-08T22:19:13.185Z"
}
},
"retryPolicy": {
"maxAttempts": 3,
"backoffMs": [
500,
1500,
3500
],
"retryableConditions": [
"HTTP_429",
"HTTP_503",
"NETWORK_TIMEOUT"
]
}
}Trust JSON
{
"status": "unavailable",
"handshakeStatus": "UNKNOWN",
"verificationFreshnessHours": null,
"reputationScore": null,
"p95LatencyMs": null,
"successRate30d": null,
"fallbackRate": null,
"attempts30d": null,
"trustUpdatedAt": null,
"trustConfidence": "unknown",
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Capability Matrix
{
"rows": [
{
"key": "OPENCLEW",
"type": "protocol",
"support": "unknown",
"confidenceSource": "profile",
"notes": "Listed on profile"
},
{
"key": "crewai",
"type": "capability",
"support": "supported",
"confidenceSource": "profile",
"notes": "Declared in agent profile metadata"
},
{
"key": "multi-agent",
"type": "capability",
"support": "supported",
"confidenceSource": "profile",
"notes": "Declared in agent profile metadata"
}
],
"flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}Facts JSON
[
{
"factKey": "vendor",
"label": "Vendor",
"value": "Ismail 2001",
"category": "vendor",
"href": "https://github.com/Ismail-2001/agent-bench",
"sourceUrl": "https://github.com/Ismail-2001/agent-bench",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-05-31T06:18:44.442Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "protocols",
"label": "Protocol compatibility",
"value": "OpenClaw",
"category": "compatibility",
"href": "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-05-31T06:18:44.442Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "traction",
"label": "Adoption signal",
"value": "2 GitHub stars",
"category": "adoption",
"href": "https://github.com/Ismail-2001/agent-bench",
"sourceUrl": "https://github.com/Ismail-2001/agent-bench",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-05-31T06:18:44.442Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "handshake_status",
"label": "Handshake status",
"value": "UNKNOWN",
"category": "security",
"href": "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-ismail-2001-agent-bench/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true,
"metadata": {}
}
]Change Events JSON
[]
Sponsored
Ads related to agent-bench and adjacent AI workflows.