Crawler Summary

crewai-evaluation answer-first brief

A CrewAI support-triage crew tested three ways: crewai test CLI baseline scoring, DeepEval span-level assertions, and LangWatch Scenario delegation testing. Testing a CrewAI Support-Triage Crew: Baseline, Span-Level, and Delegation A CrewAI support-triage crew (triage agent + billing resolver) tested three ways, each layer catching what the one before it can't: 1. **Baseline quality** with the crewai test CLI, scored by an evaluator model over several iterations. 2. **Span-level tool-call assertions** with DeepEval, grading the billing resolver's tool call (did it use th Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Freshness

Last checked 10/9/2026

Best For

crewai-evaluation is best for crewai, multi-agent workflows where OpenClaw compatibility matters.

Not Ideal For

Contract metadata is missing or unavailable for deterministic execution.

Evidence Sources Checked

editorial-content, GITHUB REPOS, runtime-metrics, public facts pack

Claim this agent
Agent DossierGITHUB REPOSSafety: 66/100

crewai-evaluation

A CrewAI support-triage crew tested three ways: crewai test CLI baseline scoring, DeepEval span-level assertions, and LangWatch Scenario delegation testing. Testing a CrewAI Support-Triage Crew: Baseline, Span-Level, and Delegation A CrewAI support-triage crew (triage agent + billing resolver) tested three ways, each layer catching what the one before it can't: 1. **Baseline quality** with the crewai test CLI, scored by an evaluator model over several iterations. 2. **Span-level tool-call assertions** with DeepEval, grading the billing resolver's tool call (did it use th

OpenClawself-declared

Public facts

4

Change events

1

Artifacts

0

Freshness

Oct 9, 2026

Verifiededitorial-contentNo verified compatibility signals

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Trust evidence available

Trust score

Unknown

Compatibility

OpenClaw

Freshness

Oct 9, 2026

Vendor

Autonoma Tools

Artifacts

0

Benchmarks

0

Last release

Unpublished

Executive Summary

Key links, install path, and a quick operational read before the deeper crawl record.

Verifiededitorial-content

Summary

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Setup snapshot

  1. 1

    Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.

  2. 2

    Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Evidence Ledger

Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.

Verifiededitorial-content
Vendor (1)

Vendor

Autonoma Tools

profilemedium
Observed Oct 9, 2026Source linkProvenance
Compatibility (1)

Protocol compatibility

OpenClaw

contractmedium
Observed Oct 9, 2026Source linkProvenance
Security (1)

Handshake status

UNKNOWN

trustmedium
Observed unknownSource linkProvenance
Integration (1)

Crawlable docs

6 indexed pages on the official domain

search_documentmedium
Observed Apr 15, 2026Source linkProvenance

Release & Crawl Timeline

Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.

Self-declaredagent-index

Artifacts Archive

Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.

Self-declaredGITHUB REPOS

Extracted files

0

Examples

2

Snippets

0

Languages

python

Executable Examples

bash

git clone https://github.com/Autonoma-Tools/crewai-evaluation.git
cd crewai-evaluation
pip install -r requirements.txt
export OPENAI_API_KEY=sk-...

# Layer 1 - baseline quality score via the crewai test CLI
bash scripts/baseline_eval.sh

# Layer 2 - span-level tool-call assertions with DeepEval
pytest tests/test_deepeval_spans.py

# Layer 3 - multi-agent delegation testing with LangWatch Scenario
pytest tests/test_delegation_scenario.py

text

crewai-evaluation/
├── src/
│   └── crew.py                        # the two-agent crew every layer tests against
├── scripts/
│   └── baseline_eval.sh               # Layer 1: crewai test CLI baseline scoring
├── tests/
│   ├── test_deepeval_spans.py         # Layer 2: DeepEval span-level tool-call assertions
│   └── test_delegation_scenario.py    # Layer 3: LangWatch Scenario delegation test
├── examples/
│   ├── autogen_groupchat_test.py      # AutoGen equivalent of the delegation assertion
│   └── openai_agents_sdk_test.py      # OpenAI Agents SDK equivalent (HandoffOutputItem)
├── .github/workflows/agent-tests.yml  # PR smoke checks + nightly delegation suite
└── requirements.txt

Docs & README

Full documentation captured from public sources, including the complete README when available.

Self-declaredGITHUB REPOS

Docs source

GITHUB REPOS

Editorial quality

ready

A CrewAI support-triage crew tested three ways: crewai test CLI baseline scoring, DeepEval span-level assertions, and LangWatch Scenario delegation testing. Testing a CrewAI Support-Triage Crew: Baseline, Span-Level, and Delegation A CrewAI support-triage crew (triage agent + billing resolver) tested three ways, each layer catching what the one before it can't: 1. **Baseline quality** with the crewai test CLI, scored by an evaluator model over several iterations. 2. **Span-level tool-call assertions** with DeepEval, grading the billing resolver's tool call (did it use th

Full README

Testing a CrewAI Support-Triage Crew: Baseline, Span-Level, and Delegation

A CrewAI support-triage crew (triage agent + billing resolver) tested three ways, each layer catching what the one before it can't:

  1. Baseline quality with the crewai test CLI, scored by an evaluator model over several iterations.
  2. Span-level tool-call assertions with DeepEval, grading the billing resolver's tool call (did it use the exact order ID?) separately from the crew's overall answer.
  3. Multi-agent delegation testing with LangWatch Scenario, asserting the triage agent actually hands the question off to the billing resolver with the order ID intact, rather than answering it itself.

The repo also ships a CI workflow that runs the baseline + span checks on every pull request and the full delegation suite nightly, plus AutoGen and OpenAI Agents SDK equivalents of the same delegation assertion.

Companion code for the Autonoma blog post: Testing a CrewAI Support-Triage Crew: Baseline, Span-Level, and Delegation

Requirements

Python 3.11+ and an OPENAI_API_KEY environment variable (used by the crew, the evaluator, and the Scenario judge/simulator).

Quickstart

git clone https://github.com/Autonoma-Tools/crewai-evaluation.git
cd crewai-evaluation
pip install -r requirements.txt
export OPENAI_API_KEY=sk-...

# Layer 1 - baseline quality score via the crewai test CLI
bash scripts/baseline_eval.sh

# Layer 2 - span-level tool-call assertions with DeepEval
pytest tests/test_deepeval_spans.py

# Layer 3 - multi-agent delegation testing with LangWatch Scenario
pytest tests/test_delegation_scenario.py

Project structure

crewai-evaluation/
├── src/
│   └── crew.py                        # the two-agent crew every layer tests against
├── scripts/
│   └── baseline_eval.sh               # Layer 1: crewai test CLI baseline scoring
├── tests/
│   ├── test_deepeval_spans.py         # Layer 2: DeepEval span-level tool-call assertions
│   └── test_delegation_scenario.py    # Layer 3: LangWatch Scenario delegation test
├── examples/
│   ├── autogen_groupchat_test.py      # AutoGen equivalent of the delegation assertion
│   └── openai_agents_sdk_test.py      # OpenAI Agents SDK equivalent (HandoffOutputItem)
├── .github/workflows/agent-tests.yml  # PR smoke checks + nightly delegation suite
└── requirements.txt
  • src/ — primary source files for the snippets referenced in the blog post.
  • examples/ — runnable examples you can execute as-is.
  • docs/ — extended notes, diagrams, or supporting material (when present).

About

This repository is maintained by Autonoma as reference material for the linked blog post. Autonoma builds autonomous AI agents that plan, execute, and maintain end-to-end tests directly from your codebase.

If something here is wrong, out of date, or unclear, please open an issue.

License

Released under the MIT License © 2026 Autonoma Labs.

Contract & API

Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.

MissingGITHUB REPOS

Contract coverage

Status

missing

Auth

None

Streaming

No

Data region

Unspecified

Protocol support

OpenClaw: self-declared

Requires: none

Forbidden: none

Guardrails

Operational confidence: low

No positive guardrails captured.
Invocation examples
curl -s "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/trust"

Reliability & Benchmarks

Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.

Missingruntime-metrics

Trust signals

Handshake

UNKNOWN

Confidence

unknown

Attempts 30d

unknown

Fallback rate

unknown

Runtime metrics

Observed P50

unknown

Observed P95

unknown

Rate limit

unknown

Estimated cost

unknown

Do not use if

Contract metadata is missing or unavailable for deterministic execution.
No benchmark suites or observed failure patterns are available.

Media & Demo

Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.

Missingno-media
No screenshots, media assets, or demo links are available.

Related Agents

Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.

Self-declaredprotocol-neighbors
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW
Machine Appendix

Contract JSON

{
  "contractStatus": "missing",
  "authModes": [],
  "requires": [],
  "forbidden": [],
  "supportsMcp": false,
  "supportsA2a": false,
  "supportsStreaming": false,
  "inputSchemaRef": null,
  "outputSchemaRef": null,
  "dataRegion": null,
  "contractUpdatedAt": null,
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Invocation Guide

{
  "preferredApi": {
    "snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/snapshot",
    "contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/contract",
    "trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/trust"
  },
  "curlExamples": [
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/snapshot\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/contract\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/trust\""
  ],
  "jsonRequestTemplate": {
    "query": "summarize this repo",
    "constraints": {
      "maxLatencyMs": 2000,
      "protocolPreference": [
        "OPENCLEW"
      ]
    }
  },
  "jsonResponseTemplate": {
    "ok": true,
    "result": {
      "summary": "...",
      "confidence": 0.9
    },
    "meta": {
      "source": "GITHUB_REPOS",
      "generatedAt": "2026-10-09T18:50:23.075Z"
    }
  },
  "retryPolicy": {
    "maxAttempts": 3,
    "backoffMs": [
      500,
      1500,
      3500
    ],
    "retryableConditions": [
      "HTTP_429",
      "HTTP_503",
      "NETWORK_TIMEOUT"
    ]
  }
}

Trust JSON

{
  "status": "unavailable",
  "handshakeStatus": "UNKNOWN",
  "verificationFreshnessHours": null,
  "reputationScore": null,
  "p95LatencyMs": null,
  "successRate30d": null,
  "fallbackRate": null,
  "attempts30d": null,
  "trustUpdatedAt": null,
  "trustConfidence": "unknown",
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Capability Matrix

{
  "rows": [
    {
      "key": "OPENCLEW",
      "type": "protocol",
      "support": "unknown",
      "confidenceSource": "profile",
      "notes": "Listed on profile"
    },
    {
      "key": "crewai",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    },
    {
      "key": "multi-agent",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    }
  ],
  "flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}

Facts JSON

[
  {
    "factKey": "vendor",
    "category": "vendor",
    "label": "Vendor",
    "value": "Autonoma Tools",
    "href": "https://github.com/Autonoma-Tools/crewai-evaluation",
    "sourceUrl": "https://github.com/Autonoma-Tools/crewai-evaluation",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-10-09T16:16:47.043Z",
    "isPublic": true
  },
  {
    "factKey": "protocols",
    "category": "compatibility",
    "label": "Protocol compatibility",
    "value": "OpenClaw",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/contract",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/contract",
    "sourceType": "contract",
    "confidence": "medium",
    "observedAt": "2026-10-09T16:16:47.043Z",
    "isPublic": true
  },
  {
    "factKey": "docs_crawl",
    "category": "integration",
    "label": "Crawlable docs",
    "value": "6 indexed pages on the official domain",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  },
  {
    "factKey": "handshake_status",
    "category": "security",
    "label": "Handshake status",
    "value": "UNKNOWN",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/trust",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-autonoma-tools-crewai-evaluation/trust",
    "sourceType": "trust",
    "confidence": "medium",
    "observedAt": null,
    "isPublic": true
  }
]

Change Events JSON

[
  {
    "eventType": "docs_update",
    "title": "Docs refreshed: Sign in to GitHub · GitHub",
    "description": "Fresh crawlable documentation was indexed for the official domain.",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  }
]

Sponsored

Ads related to crewai-evaluation and adjacent AI workflows.