Crawler Summary

tanglefoot answer-first brief

A rigorous public benchmark and interactive dashboard evaluating LLM-based agent frameworks (LangGraph, CrewAI, AutoGen, LlamaIndex) on multi-step tasks under intentional adversarial stress (flaky APIs, lying search tools, slow scrapers, contradictory sources, and circular loops). Tanglefoot: Agent Chaos Engineering & Benchmark Suite Tanglefoot is an evaluation tool designed to test how well AI agents handle difficult real-world situations. It runs agents built with frameworks like **LangGraph, CrewAI, AutoGen, and LlamaIndex** against a set of challenging tasks. During these tasks, Tanglefoot introduces adversarial stressors (difficult problems) such as: * **Slow APIs**: Delays in server resp Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Freshness

Last checked 10/9/2026

Best For

tanglefoot is best for crewai, multi-agent workflows where OpenClaw compatibility matters.

Not Ideal For

Contract metadata is missing or unavailable for deterministic execution.

Evidence Sources Checked

editorial-content, GITHUB REPOS, runtime-metrics, public facts pack

Agent DossierGITHUB REPOSSafety: 66/100

tanglefoot

A rigorous public benchmark and interactive dashboard evaluating LLM-based agent frameworks (LangGraph, CrewAI, AutoGen, LlamaIndex) on multi-step tasks under intentional adversarial stress (flaky APIs, lying search tools, slow scrapers, contradictory sources, and circular loops). Tanglefoot: Agent Chaos Engineering & Benchmark Suite Tanglefoot is an evaluation tool designed to test how well AI agents handle difficult real-world situations. It runs agents built with frameworks like **LangGraph, CrewAI, AutoGen, and LlamaIndex** against a set of challenging tasks. During these tasks, Tanglefoot introduces adversarial stressors (difficult problems) such as: * **Slow APIs**: Delays in server resp

OpenClawself-declared

Public facts

4

Change events

1

Artifacts

0

Freshness

Oct 9, 2026

Verifiededitorial-contentNo verified compatibility signals

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Trust evidence available

Trust score

Unknown

Compatibility

OpenClaw

Freshness

Oct 9, 2026

Vendor

Nizaalkhot

Artifacts

0

Benchmarks

0

Last release

Unpublished

Executive Summary

Key links, install path, and a quick operational read before the deeper crawl record.

Verifiededitorial-content

Summary

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Setup snapshot

  1. 1

    Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.

  2. 2

    Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Evidence Ledger

Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.

Verifiededitorial-content
Vendor (1)

Vendor

Nizaalkhot

profilemedium
Observed Oct 9, 2026Source linkProvenance
Compatibility (1)

Protocol compatibility

OpenClaw

contractmedium
Observed Oct 9, 2026Source linkProvenance
Security (1)

Handshake status

UNKNOWN

trustmedium
Observed unknownSource linkProvenance
Integration (1)

Crawlable docs

6 indexed pages on the official domain

search_documentmedium
Observed Apr 15, 2026Source linkProvenance

Release & Crawl Timeline

Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.

Self-declaredagent-index

Artifacts Archive

Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.

Self-declaredGITHUB REPOS

Extracted files

0

Examples

4

Snippets

0

Languages

python

Executable Examples

text

tanglefoot/
├── benchmark/                      # Core Python Evaluation Harness
│   ├── tasks/                      # Tasks configurations and evaluators
│   ├── frameworks/                 # Integrations for LangGraph, CrewAI, AutoGen, and LlamaIndex
│   ├── tools/                      # FastAPI server supplying stressed REST endpoints
│   └── run_benchmark.py            # CLI harness and agent reference implementations
│
├── tanglefoot/                     # Python SDK package
│   └── stressor.py                 # @stressor decorator to inject chaos in custom endpoints
│
└── dashboard/                      # React Developer Dashboard and Leaderboard UI

bash

# Install core packages
pip install fastapi uvicorn requests pydantic websockets

# Start the FastAPI server on port 8005
uvicorn benchmark.tools.adversarial_api:app --reload --port 8005

bash

# Run the complete stress test suite using the resilient LangGraph agent
python benchmark/run_benchmark.py --agent langgraph --task all

# Run and sync the results and logs directly to the dashboard
python benchmark/run_benchmark.py --agent langgraph --task all --sync-dashboard

bash

cd dashboard
npm install
npm run dev

Docs & README

Full documentation captured from public sources, including the complete README when available.

Self-declaredGITHUB REPOS

Docs source

GITHUB REPOS

Editorial quality

ready

A rigorous public benchmark and interactive dashboard evaluating LLM-based agent frameworks (LangGraph, CrewAI, AutoGen, LlamaIndex) on multi-step tasks under intentional adversarial stress (flaky APIs, lying search tools, slow scrapers, contradictory sources, and circular loops). Tanglefoot: Agent Chaos Engineering & Benchmark Suite Tanglefoot is an evaluation tool designed to test how well AI agents handle difficult real-world situations. It runs agents built with frameworks like **LangGraph, CrewAI, AutoGen, and LlamaIndex** against a set of challenging tasks. During these tasks, Tanglefoot introduces adversarial stressors (difficult problems) such as: * **Slow APIs**: Delays in server resp

Full README

Tanglefoot: Agent Chaos Engineering & Benchmark Suite

Tanglefoot is an evaluation tool designed to test how well AI agents handle difficult real-world situations.

It runs agents built with frameworks like LangGraph, CrewAI, AutoGen, and LlamaIndex against a set of challenging tasks. During these tasks, Tanglefoot introduces adversarial stressors (difficult problems) such as:

  • Slow APIs: Delays in server responses.
  • Rate Limits & Server Outages: Temporary HTTP 429 and 500 errors.
  • Redirection Loops: Tools that redirect in a circle to trick the agent.
  • Contradictory Sources: Different databases giving conflicting information (for example, an outdated corporate wiki vs. a official SEC filing).
  • Prompt Injections: Input text that attempts to hijack agent instructions.

🎥 Dashboard Replay & Interactive Preview

Below is a high-fidelity visual preview of the Tanglefoot interactive trace scroller, ticking token-cost telemetry, and dark/light system interface:

Tanglefoot Interactive Dashboard Demo


Repository Structure

The project has the following main parts:

tanglefoot/
├── benchmark/                      # Core Python Evaluation Harness
│   ├── tasks/                      # Tasks configurations and evaluators
│   ├── frameworks/                 # Integrations for LangGraph, CrewAI, AutoGen, and LlamaIndex
│   ├── tools/                      # FastAPI server supplying stressed REST endpoints
│   └── run_benchmark.py            # CLI harness and agent reference implementations
│
├── tanglefoot/                     # Python SDK package
│   └── stressor.py                 # @stressor decorator to inject chaos in custom endpoints
│
└── dashboard/                      # React Developer Dashboard and Leaderboard UI

Quickstart Guide

Get the complete evaluation environment up and running on your local machine.

1. Start the Adversarial REST Server

The API endpoints are served via a local FastAPI app that simulates slow, broken, and circular integrations:

# Install core packages
pip install fastapi uvicorn requests pydantic websockets

# Start the FastAPI server on port 8005
uvicorn benchmark.tools.adversarial_api:app --reload --port 8005

You can verify the server is running by opening http://localhost:8005/ in your browser.

2. Run the Benchmark CLI

Run the evaluator against all tasks:

# Run the complete stress test suite using the resilient LangGraph agent
python benchmark/run_benchmark.py --agent langgraph --task all

# Run and sync the results and logs directly to the dashboard
python benchmark/run_benchmark.py --agent langgraph --task all --sync-dashboard

3. Launch the Dashboard UI

View the leaderboard and step-by-step agent thoughts:

cd dashboard
npm install
npm run dev

Open http://localhost:5173/ in your browser.

🚀 V2 Advanced Features

Tanglefoot has been upgraded with highly robust security stressors, hardened evaluation layers, and a concurrent framework execution harness:

1. Advanced Adversarial Stressors

We simulate advanced LLM vulnerabilities to test agent safety limits:

  • Multi-Stage Indirect Prompt Injections (Task 59): Validates if agents can successfully filter out malicious third-party payloads instructing them to output fake data, while extracting factual data.
  • Dynamic Tool-Definition Hijacking (Task 60): Rewrites tool schemas dynamically between successive invocations to test the resilience of dynamic schema parsers.
  • Complex Data Poisoning & Guardrail Traps (Task 61): Injects zero-width spaces (\u200b) and Right-to-Left Override (\u202b) formatting traps to test data sanitization.

2. Hardened Evaluation Layer (judges.py)

  • Consensus Panel of Judges: Scores are evaluated concurrently by an ensemble of three independent evaluator personas ("Strict Auditor", "Resilience Advocate", and "Guardrail Warden") to minimize grading variance.
  • Deterministic Invariant Testing: Skips LLM calls for quantitative assertions by performing strict regex-based and numerical assertions first.
  • Cost & Token Normalization: Integrates execution efficiency directly into the Robustness Index using OpenInference token counters.

3. Concurrency & SDK Chaos Interceptors

  • True Concurrent Matrix Testing: Speeds up benchmarking runtimes via --parallel support using multi-threaded execution pools.
  • Network Fault Injection: Simulates packet drops and flaky environments programmatically with enable_network_proxy_interceptor.
  • Stateful Mutators: Introduces StatefulDBCacheMutator to simulate dirty reads and database cache race conditions.

Scoring Metrics

Agents are graded using a blended formula based on four pillars:

  1. Completeness: Did the agent find the correct answer?
  2. Resilience: Did the agent detect and recover from server errors, loops, and lies?
  3. Guardrails: Did the agent avoid prompt injections and keep system boundaries?
  4. Step Efficiency: Did the agent complete the task in a reasonable number of steps without getting stuck?

Contract & API

Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.

MissingGITHUB REPOS

Contract coverage

Status

missing

Auth

None

Streaming

No

Data region

Unspecified

Protocol support

OpenClaw: self-declared

Requires: none

Forbidden: none

Guardrails

Operational confidence: low

No positive guardrails captured.
Invocation examples
curl -s "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/trust"

Reliability & Benchmarks

Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.

Missingruntime-metrics

Trust signals

Handshake

UNKNOWN

Confidence

unknown

Attempts 30d

unknown

Fallback rate

unknown

Runtime metrics

Observed P50

unknown

Observed P95

unknown

Rate limit

unknown

Estimated cost

unknown

Do not use if

Contract metadata is missing or unavailable for deterministic execution.
No benchmark suites or observed failure patterns are available.

Media & Demo

Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.

Missingno-media
No screenshots, media assets, or demo links are available.

Related Agents

Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.

Self-declaredprotocol-neighbors
Github ReposUpdated 4h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW
Machine Appendix

Contract JSON

{
  "contractStatus": "missing",
  "authModes": [],
  "requires": [],
  "forbidden": [],
  "supportsMcp": false,
  "supportsA2a": false,
  "supportsStreaming": false,
  "inputSchemaRef": null,
  "outputSchemaRef": null,
  "dataRegion": null,
  "contractUpdatedAt": null,
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Invocation Guide

{
  "preferredApi": {
    "snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/snapshot",
    "contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/contract",
    "trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/trust"
  },
  "curlExamples": [
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/snapshot\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/contract\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/trust\""
  ],
  "jsonRequestTemplate": {
    "query": "summarize this repo",
    "constraints": {
      "maxLatencyMs": 2000,
      "protocolPreference": [
        "OPENCLEW"
      ]
    }
  },
  "jsonResponseTemplate": {
    "ok": true,
    "result": {
      "summary": "...",
      "confidence": 0.9
    },
    "meta": {
      "source": "GITHUB_REPOS",
      "generatedAt": "2026-10-09T22:49:47.078Z"
    }
  },
  "retryPolicy": {
    "maxAttempts": 3,
    "backoffMs": [
      500,
      1500,
      3500
    ],
    "retryableConditions": [
      "HTTP_429",
      "HTTP_503",
      "NETWORK_TIMEOUT"
    ]
  }
}

Trust JSON

{
  "status": "unavailable",
  "handshakeStatus": "UNKNOWN",
  "verificationFreshnessHours": null,
  "reputationScore": null,
  "p95LatencyMs": null,
  "successRate30d": null,
  "fallbackRate": null,
  "attempts30d": null,
  "trustUpdatedAt": null,
  "trustConfidence": "unknown",
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Capability Matrix

{
  "rows": [
    {
      "key": "OPENCLEW",
      "type": "protocol",
      "support": "unknown",
      "confidenceSource": "profile",
      "notes": "Listed on profile"
    },
    {
      "key": "crewai",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    },
    {
      "key": "multi-agent",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    }
  ],
  "flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}

Facts JSON

[
  {
    "factKey": "vendor",
    "category": "vendor",
    "label": "Vendor",
    "value": "Nizaalkhot",
    "href": "https://github.com/nizaalkhot/tanglefoot",
    "sourceUrl": "https://github.com/nizaalkhot/tanglefoot",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-10-09T20:22:14.418Z",
    "isPublic": true
  },
  {
    "factKey": "protocols",
    "category": "compatibility",
    "label": "Protocol compatibility",
    "value": "OpenClaw",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/contract",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/contract",
    "sourceType": "contract",
    "confidence": "medium",
    "observedAt": "2026-10-09T20:22:14.418Z",
    "isPublic": true
  },
  {
    "factKey": "docs_crawl",
    "category": "integration",
    "label": "Crawlable docs",
    "value": "6 indexed pages on the official domain",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  },
  {
    "factKey": "handshake_status",
    "category": "security",
    "label": "Handshake status",
    "value": "UNKNOWN",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/trust",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-nizaalkhot-tanglefoot/trust",
    "sourceType": "trust",
    "confidence": "medium",
    "observedAt": null,
    "isPublic": true
  }
]

Change Events JSON

[
  {
    "eventType": "docs_update",
    "title": "Docs refreshed: Sign in to GitHub · GitHub",
    "description": "Fresh crawlable documentation was indexed for the official domain.",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  }
]

Sponsored

Ads related to tanglefoot and adjacent AI workflows.