Crawler Summary

deep_research_engine_v1_crewai-project answer-first brief

multi-agent AI research system built on [crewAI](https://crewai.com), designed for deep technical research, fact-checking, and content synthesis. DeepResearchEngine Crew A production-grade multi-agent AI research system built on $1, designed for deep technical research, fact-checking, and content synthesis. This system leverages a sophisticated content ingestion pipeline, vector embeddings, and hierarchical reasoning to deliver accurate, grounded research outputs. Features Core Capabilities - **Deep Technical Research** - Systematic analysis bypassing marketin Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Freshness

Last checked 10/9/2026

Best For

deep_research_engine_v1_crewai-project is best for crewai, multi-agent workflows where OpenClaw compatibility matters.

Not Ideal For

Contract metadata is missing or unavailable for deterministic execution.

Evidence Sources Checked

editorial-content, GITHUB REPOS, runtime-metrics, public facts pack

Agent DossierGITHUB REPOSSafety: 66/100

deep_research_engine_v1_crewai-project

multi-agent AI research system built on [crewAI](https://crewai.com), designed for deep technical research, fact-checking, and content synthesis. DeepResearchEngine Crew A production-grade multi-agent AI research system built on $1, designed for deep technical research, fact-checking, and content synthesis. This system leverages a sophisticated content ingestion pipeline, vector embeddings, and hierarchical reasoning to deliver accurate, grounded research outputs. Features Core Capabilities - **Deep Technical Research** - Systematic analysis bypassing marketin

OpenClawself-declared

Public facts

4

Change events

1

Artifacts

0

Freshness

Oct 9, 2026

Verifiededitorial-contentNo verified compatibility signals

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Trust evidence available

Trust score

Unknown

Compatibility

OpenClaw

Freshness

Oct 9, 2026

Vendor

B08x

Artifacts

0

Benchmarks

0

Last release

Unpublished

Executive Summary

Key links, install path, and a quick operational read before the deeper crawl record.

Verifiededitorial-content

Summary

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Setup snapshot

  1. 1

    Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.

  2. 2

    Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Evidence Ledger

Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.

Verifiededitorial-content
Vendor (1)

Vendor

B08x

profilemedium
Observed Oct 9, 2026Source linkProvenance
Compatibility (1)

Protocol compatibility

OpenClaw

contractmedium
Observed Oct 9, 2026Source linkProvenance
Security (1)

Handshake status

UNKNOWN

trustmedium
Observed unknownSource linkProvenance
Integration (1)

Crawlable docs

6 indexed pages on the official domain

search_documentmedium
Observed Apr 15, 2026Source linkProvenance

Release & Crawl Timeline

Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.

Self-declaredagent-index

Artifacts Archive

Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.

Self-declaredGITHUB REPOS

Extracted files

0

Examples

6

Snippets

0

Languages

python

Executable Examples

bash

# Install UV if not already present
pip install uv

# Install dependencies
uv sync

# Or use crewai CLI
crewai install

bash

# Create virtual environment
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

# Install dependencies
pip install -e .

# Install spaCy model (required for content ingestion)
python -m spacy download en_core_web_sm

bash

OPENAI_API_KEY=your_openai_key
ANTHROPIC_API_KEY=your_anthropic_key
GEMINI_API_KEY=your_gemini_key
GROQ_API_KEY=your_groq_key
MISTRAL_API_KEY=your_mistral_key
OPENROUTER_API_KEY=your_openrouter_key

# Optional: Redis configuration for caching
REDIS_HOST=localhost
REDIS_PORT=6379
REDIS_DB=0

# Optional: Embedding provider (default: ollama)
EMBEDDING_PROVIDER=ollama

# Optional: Summary model for ingestion pipeline
SUMMARY_MODEL=openrouter/google/gemini-2.0-flash-001

bash

uv run configure_crew

bash

# Run with default inputs
uv run deep_research_engine run

# Or using crewai CLI
crewai run

bash

# Run with custom document ingestion
uv run deep_research_engine run --doc-path ./research_paper.pdf --embedding-provider ollama

# Resume from checkpoint
uv run deep_research_engine run --resume

# Train the crew
uv run deep_research_engine train <iterations> <filename>

# Replay from specific task
uv run deep_research_engine replay <task_id>

# Test execution
uv run deep_research_engine test <iterations> <openai_model_name>

Docs & README

Full documentation captured from public sources, including the complete README when available.

Self-declaredGITHUB REPOS

Docs source

GITHUB REPOS

Editorial quality

ready

multi-agent AI research system built on [crewAI](https://crewai.com), designed for deep technical research, fact-checking, and content synthesis. DeepResearchEngine Crew A production-grade multi-agent AI research system built on $1, designed for deep technical research, fact-checking, and content synthesis. This system leverages a sophisticated content ingestion pipeline, vector embeddings, and hierarchical reasoning to deliver accurate, grounded research outputs. Features Core Capabilities - **Deep Technical Research** - Systematic analysis bypassing marketin

Full README

DeepResearchEngine Crew

A production-grade multi-agent AI research system built on crewAI, designed for deep technical research, fact-checking, and content synthesis. This system leverages a sophisticated content ingestion pipeline, vector embeddings, and hierarchical reasoning to deliver accurate, grounded research outputs.

Features

Core Capabilities

  • Deep Technical Research - Systematic analysis bypassing marketing rhetoric to extract technical reality
  • SIFT Fact-Checking - Rigorous evidence verification using SIFT evaluation protocols
  • Hierarchical Content Ingestion - Multi-stage processing with parent-child chunking and two-pass LLM contextualization
  • Vector Embeddings - Semantic search with txtai, supporting Ollama and Sentence-Transformers
  • Checkpoint & Resume - Save and restore execution state for long-running research tasks
  • Fallback LLM Support - Automatic failover to backup models when primary fails

Content Ingestion Pipeline

The ContentIngestionTool implements a robust 6-stage processing pipeline:

  1. Cache Check - Redis-based caching for previously processed content
  2. Content Extraction - Web (trafilatura) and local file (Kreuzberg) extraction
  3. Noise Removal - spaCy-based text cleaning with custom pipeline components
  4. Hierarchical Recursive Chunking - Parent-child separator strategy (2000 token parents → 500 token children)
  5. Two-Pass Contextualization - Parent summary generation followed by context injection into children
  6. Vector Indexing - Automatic embedding and indexing with dimensionality guards

Vector Store

  • Multi-Provider Support - Ollama (nomic-embed-text), Sentence-Transformers (all-MiniLM-L6-v2), Mistral
  • Dimensionality Guards - Automatic backup of existing index when model dimensions don't match
  • Graphify Integration - Exports processed chunks to graphify-out/ for downstream visualization
  • Metadata Preservation - Stores source hashes and provenance with each chunk

Installation

Prerequisites

  • Python >=3.10, <3.14
  • UV (recommended) or pip

Quick Install

# Install UV if not already present
pip install uv

# Install dependencies
uv sync

# Or use crewai CLI
crewai install

Manual Setup

# Create virtual environment
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

# Install dependencies
pip install -e .

# Install spaCy model (required for content ingestion)
python -m spacy download en_core_web_sm

Configuration

API Keys

Add your provider API keys to .env:

OPENAI_API_KEY=your_openai_key
ANTHROPIC_API_KEY=your_anthropic_key
GEMINI_API_KEY=your_gemini_key
GROQ_API_KEY=your_groq_key
MISTRAL_API_KEY=your_mistral_key
OPENROUTER_API_KEY=your_openrouter_key

# Optional: Redis configuration for caching
REDIS_HOST=localhost
REDIS_PORT=6379
REDIS_DB=0

# Optional: Embedding provider (default: ollama)
EMBEDDING_PROVIDER=ollama

# Optional: Summary model for ingestion pipeline
SUMMARY_MODEL=openrouter/google/gemini-2.0-flash-001

Interactive Configuration (Recommended)

Launch the visual configuration dashboard to manage agents, tasks, and run parameters:

uv run configure_crew

The TUI provides:

  • Agent Editor - Configure role, goal, backstory, LLM provider/model, and capability requirements
  • Task Editor - Define task descriptions, expected outputs, and agent assignments
  • Input Parameters - Set grounding context, primary topic, sub-nodes, and target audience
  • API Key Validation - Real-time checking of provider API keys
  • Model Capability Matching - Validates that selected models support required capabilities (vision, reasoning, tool calling, structured output)
  • Fallback LLM Configuration - Set primary and fallback models with automatic failover

Manual Configuration Files

  • src/deep_research_engine/config/agents.yaml - Agent definitions
  • src/deep_research_engine/config/tasks.yaml - Task definitions
  • src/deep_research_engine/config/inputs.yaml - Run parameters

Usage

Basic Research Run

# Run with default inputs
uv run deep_research_engine run

# Or using crewai CLI
crewai run

Advanced Usage

# Run with custom document ingestion
uv run deep_research_engine run --doc-path ./research_paper.pdf --embedding-provider ollama

# Resume from checkpoint
uv run deep_research_engine run --resume

# Train the crew
uv run deep_research_engine train <iterations> <filename>

# Replay from specific task
uv run deep_research_engine replay <task_id>

# Test execution
uv run deep_research_engine test <iterations> <openai_model_name>

Input Parameters

Create or modify src/deep_research_engine/config/inputs.yaml:

grounding_context: ""  # Optional: Pre-loaded context for grounding
primary_topic: "AI Research Methodologies"
sub_nodes: 
  - "Large Language Models"
  - "Evaluation Metrics"
  - "Ethical Considerations"
target_audience: "Senior Engineers"

Architecture

Agents

The system includes five specialized agents, each with distinct capabilities:

| Agent | Role | Default Model | Tools | Specialization | |-------|------|---------------|-------|----------------| | senior_research_strategist | Strategy Formulation | glm-4.7-flash | SerperDev, EXA Search, Content Ingestion, FileWriter | Identifies research gaps, audience analysis | | other_steve_deep_research_mode | Deep Technical Research | mistral-medium-latest | SerperDev, EXA Search, Content Ingestion, Arxiv, FileWriter | Architecture analysis, trade-off evaluation | | other_steve_sift_fact_checker | SIFT Fact-Checker | glm-4.7 | SerperDev, EXA Search, Content Ingestion, FileWriter | Evidence verification, source reliability | | other_steve_pragmatic_editor | Pragmatic Editor | magistral-medium-latest | FileWriter | Precision filtering, conciseness | | other_steve_tree_of_thoughts_evaluator | Tree of Thoughts Evaluator | glm-4.7 | FileWriter | Multi-path analysis, failure mode identification |

All agents support:

  • Custom LLM configuration (temperature, top_p, top_k, max_tokens)
  • Fallback LLM models
  • Capability requirement validation (vision, reasoning, tool calling, structured output)

Workflow

Strategy Formulation → Deep Technical Research → SIFT Fact-Checking → 
Research Methodology Evaluation → Technical Content Synthesis

Tools

  • SerperDevTool - Web search
  • EXASearchTool - Enhanced search
  • ArxivPaperTool - Academic paper retrieval
  • FileWriterTool - Output generation
  • ContentIngestionTool - Multi-stage content processing

Development

Running Tests

# Run vector store tests
uv run pytest tests/test_vector_store.py

# Or with pytest directly
pytest tests/test_vector_store.py

Project Structure

deep_research_engine_v1_crewai-project/
├── src/
│   └── deep_research_engine/
│       ├── __init__.py
│       ├── crew.py              # Agent and task definitions
│       ├── main.py              # CLI entry points
│       ├── tui_config.py        # Textual-based configuration TUI
│       ├── utils.py             # Redis cache, config helpers
│       ├── models_fetcher.py    # Dynamic model loading
│       ├── vector_store.py      # txtai-based vector embeddings
│       └── tools/
│           ├── __init__.py
│           └── ingestion_tools.py # Content ingestion pipeline
│       └── config/
│           ├── agents.yaml      # Agent configurations
│           ├── tasks.yaml       # Task definitions
│           └── inputs.yaml      # Run parameters
├── tests/
│   └── test_vector_store.py
├── pyproject.toml
├── uv.lock
└── README.md

Advanced Configuration

Custom Embedding Providers

Set EMBEDDING_PROVIDER in .env:

  • ollama - Uses Ollama's nomic-embed-text (default)
  • mistral - Uses Mistral embeddings
  • local - Uses Sentence-Transformers (all-MiniLM-L6-v2)

Redis Configuration

For distributed caching:

REDIS_HOST=your_redis_host
REDIS_PORT=6379
REDIS_DB=0

Model Fallback Configuration

Each agent can have a primary and fallback LLM:

senior_research_strategist:
  llm: openrouter/z-ai/glm-4.7-flash
  fallback_llm: openai/gpt-4o-mini
  llm_config:
    temperature: 0.2
    top_p: 0.9
    top_k: null
    max_tokens: 8192

Checkpoint Management

Checkpoints are automatically saved to .checkpoints/ directory. To resume:

uv run deep_research_engine run --resume

The system will automatically find the latest checkpoint and restore execution state.

Dependencies

Core Dependencies

  • crewai[file-processing,litellm,tools]==1.14.4
  • exa-py
  • pydantic
  • python-dotenv

Content Ingestion

  • trafilatura - Web content extraction
  • spacy - NLP processing
  • spacy-llm - LLM integration for spaCy
  • kreuzberg - Local file extraction

Vector Store

  • txtai[ann,pipeline,similarity] - Embeddings and semantic search
  • redis - Distributed caching

Development

  • pytest>=9.0.3

Support

For support, questions, or feedback regarding DeepResearchEngine:


Built with crewAI - Let's create wonders together with the power and simplicity of multi-agent AI systems.

Contract & API

Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.

MissingGITHUB REPOS

Contract coverage

Status

missing

Auth

None

Streaming

No

Data region

Unspecified

Protocol support

OpenClaw: self-declared

Requires: none

Forbidden: none

Guardrails

Operational confidence: low

No positive guardrails captured.
Invocation examples
curl -s "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/trust"

Reliability & Benchmarks

Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.

Missingruntime-metrics

Trust signals

Handshake

UNKNOWN

Confidence

unknown

Attempts 30d

unknown

Fallback rate

unknown

Runtime metrics

Observed P50

unknown

Observed P95

unknown

Rate limit

unknown

Estimated cost

unknown

Do not use if

Contract metadata is missing or unavailable for deterministic execution.
No benchmark suites or observed failure patterns are available.

Media & Demo

Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.

Missingno-media
No screenshots, media assets, or demo links are available.

Related Agents

Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.

Self-declaredprotocol-neighbors
Github ReposUpdated 7h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW
Machine Appendix

Contract JSON

{
  "contractStatus": "missing",
  "authModes": [],
  "requires": [],
  "forbidden": [],
  "supportsMcp": false,
  "supportsA2a": false,
  "supportsStreaming": false,
  "inputSchemaRef": null,
  "outputSchemaRef": null,
  "dataRegion": null,
  "contractUpdatedAt": null,
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Invocation Guide

{
  "preferredApi": {
    "snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/snapshot",
    "contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/contract",
    "trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/trust"
  },
  "curlExamples": [
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/snapshot\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/contract\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/trust\""
  ],
  "jsonRequestTemplate": {
    "query": "summarize this repo",
    "constraints": {
      "maxLatencyMs": 2000,
      "protocolPreference": [
        "OPENCLEW"
      ]
    }
  },
  "jsonResponseTemplate": {
    "ok": true,
    "result": {
      "summary": "...",
      "confidence": 0.9
    },
    "meta": {
      "source": "GITHUB_REPOS",
      "generatedAt": "2026-10-10T01:53:08.893Z"
    }
  },
  "retryPolicy": {
    "maxAttempts": 3,
    "backoffMs": [
      500,
      1500,
      3500
    ],
    "retryableConditions": [
      "HTTP_429",
      "HTTP_503",
      "NETWORK_TIMEOUT"
    ]
  }
}

Trust JSON

{
  "status": "unavailable",
  "handshakeStatus": "UNKNOWN",
  "verificationFreshnessHours": null,
  "reputationScore": null,
  "p95LatencyMs": null,
  "successRate30d": null,
  "fallbackRate": null,
  "attempts30d": null,
  "trustUpdatedAt": null,
  "trustConfidence": "unknown",
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Capability Matrix

{
  "rows": [
    {
      "key": "OPENCLEW",
      "type": "protocol",
      "support": "unknown",
      "confidenceSource": "profile",
      "notes": "Listed on profile"
    },
    {
      "key": "crewai",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    },
    {
      "key": "multi-agent",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    }
  ],
  "flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}

Facts JSON

[
  {
    "factKey": "vendor",
    "category": "vendor",
    "label": "Vendor",
    "value": "B08x",
    "href": "https://github.com/b08x/deep_research_engine_v1_crewai-project",
    "sourceUrl": "https://github.com/b08x/deep_research_engine_v1_crewai-project",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-10-09T22:15:07.427Z",
    "isPublic": true
  },
  {
    "factKey": "protocols",
    "category": "compatibility",
    "label": "Protocol compatibility",
    "value": "OpenClaw",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/contract",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/contract",
    "sourceType": "contract",
    "confidence": "medium",
    "observedAt": "2026-10-09T22:15:07.427Z",
    "isPublic": true
  },
  {
    "factKey": "docs_crawl",
    "category": "integration",
    "label": "Crawlable docs",
    "value": "6 indexed pages on the official domain",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  },
  {
    "factKey": "handshake_status",
    "category": "security",
    "label": "Handshake status",
    "value": "UNKNOWN",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/trust",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-b08x-deep-research-engine-v1-crewai-project/trust",
    "sourceType": "trust",
    "confidence": "medium",
    "observedAt": null,
    "isPublic": true
  }
]

Change Events JSON

[
  {
    "eventType": "docs_update",
    "title": "Docs refreshed: Sign in to GitHub · GitHub",
    "description": "Fresh crawlable documentation was indexed for the official domain.",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  }
]

Sponsored

Ads related to deep_research_engine_v1_crewai-project and adjacent AI workflows.