Tavily Best Practices
Build production-ready Tavily integrations with best practices baked in. Reference documentation for developers using coding assistants (Claude Code, Cursor, etc.) to implement web search, content extraction, crawling, and research in agentic workflows, RAG systems, or autonomous agents. Skill: Tavily Best Practices Owner: barneyjm Summary: Build production-ready Tavily integrations with best practices baked in. Reference documentation for developers using coding assistants (Claude Code, Cursor, etc.) to implement web search, content extraction, crawling, and research in agentic workflows, RAG systems, or autonomous agents. Tags: latest:0.1.0 Version history: v0.1.0 | 2026-02-03T15:27:32.248Z | auto
Rank
62
Safety
84
Downloads
1.7k
Updated
Oct 10, 2026
Version
0.1.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.7K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.7K downloadsadoption · observed Oct 10, 2026
- Latest release
- 0.1.0release · observed Feb 3, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s175g00ed7n2ej0r13jbm8m3k18842ww:tavily-best-practices- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-tavily-best-practices/snapshot"
Documentation
CLAWHUB
64,393 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: tavily-best-practices
description: "Build production-ready Tavily integrations with best practices baked in. Reference documentation for developers using coding assistants (Claude Code, Cursor, etc.) to implement web search, content extraction, crawling, and research in agentic workflows, RAG systems, or autonomous agents."
---
# Tavily
Tavily is a search API designed for LLMs, enabling AI applications to access real-time web data.
## Prerequisites
**Tavily API Key Required** - Get your key at https://app.tavily.com (1,000 free API credits/month, no credit card required)
Add to `~/.claude/settings.json`:
```json
{
"env": {
"TAVILY_API_KEY": "tvly-YOUR_API_KEY"
}
}
```
Restart Claude Code after adding your API key.
## Installation
**Python:**
```bash
pip install tavily-python
```
**JavaScript:**
```bash
npm install @tavily/core
```
See **[references/sdk.md](references/sdk.md)** for complete SDK reference.
## Client Initialization
```python
from tavily import TavilyClient
# Option 1: Uses TAVILY_API_KEY env var (recommended)
client = TavilyClient()
# Option 2: Explicit API key
client = TavilyClient(api_key="tvly-YOUR_API_KEY")
# Option 3: With project tracking (for usage organization)
client = TavilyClient(api_key="tvly-YOUR_API_KEY", project_id="your-project-id")
# Async client for parallel queries
from tavily import AsyncTavilyClient
async_client = AsyncTavilyClient()
```
## Choosing the Right Method
**For custom agents/workflows:**
| Need | Method |
|------|--------|
| Web search results | `search()` |
| Content from specific URLs | `extract()` |
| Content from entire site | `crawl()` |
| URL discovery from site | `map()` |
**For out-of-the-box research:**
| Need | Method |
|------|--------|
| End-to-end research with AI synthesis | `research()` |
## Quick Reference
### search() - Web Search
```python
response = client.search(
query="quantum computing breakthroughs", # Keep under 400 chars
max_results=10,
search_depth="advanced", # 2 credits, highest relevance
topic="general" # or "news", "finance"
)
for result in response["results"]:
print(f"{result['title']}: {result['score']}")
```
Key parameters: `query`, `max_results`, `search_depth` (ultra-fast/fast/basic/advanced), `topic`, `include_domains`, `exclude_domains`, `time_range`
### extract() - URL Content Extraction
```python
# Two-step pattern (recommended for control)
search_results = client.search(query="Python async best practices")
urls = [r["url"] for r in search_results["results"] if r["score"] > 0.5]
extracted = client.extract(
urls=urls[:20],
query="async patterns", # Reranks chunks by relevance
chunks_per_source=3 # Prevents context explosion
)
```
Key parameters: `urls` (max 20), `extract_depth`, `query`, `chunks_per_source` (1-5)
### crawl() - Site-Wide Extraction
```python
response = client.crawl(
url="https://docs.example.com",
max_depth=2,
instructions="Find API documentation pages_meta.json
{
"ownerId": "kn7e3j3w1x5et0yppacy4m90y18084ba",
"slug": "tavily-best-practices",
"version": "0.1.0",
"publishedAt": 1770132452248
}references/crawl.md
# Crawl & Map API Reference
## Table of Contents
- [Crawl vs Map](#crawl-vs-map)
- [Key Parameters](#key-parameters)
- [Instructions and Chunks](#instructions-and-chunks)
- [Path and Domain Filtering](#path-and-domain-filtering)
- [Use Cases](#use-cases)
- [Map then Extract Pattern](#map-then-extract-pattern)
- [Performance Optimization](#performance-optimization)
- [Common Pitfalls](#common-pitfalls)
- [Response Fields](#response-fields)
- [Summary](#summary)
---
## Crawl vs Map
| Feature | Crawl | Map |
|---------|-------|-----|
| **Returns** | Full content | URLs only |
| **Speed** | Slower | Faster |
| **Best for** | RAG, deep analysis, documentation | Site structure discovery, URL collection |
**Use Crawl when:**
- Full content extraction needed
- Building RAG systems
- Processing paginated/nested content
- Integration with knowledge bases
**Use Map when:**
- Quick site structure discovery
- URL collection without content
- Planning before crawling
- Sitemap generation
---
## Key Parameters
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `url` | string | Required | Root URL to begin |
| `max_depth` | integer | 1 | Levels deep to crawl (1-5). **Start with 1-2** |
| `max_breadth` | integer | 20 | Links per page. 50-100 for focused crawls |
| `limit` | integer | 50 | Total pages cap |
| `instructions` | string | null | Natural language guidance (2 credits/10 pages) |
| `chunks_per_source` | integer | 3 | Chunks per page (1-5). Only with `instructions` |
| `extract_depth` | enum | `"basic"` | `"basic"` (1 credit/5 URLs) or `"advanced"` (2 credits/5 URLs) |
| `format` | enum | `"markdown"` | `"markdown"` or `"text"` |
| `select_paths` | array | null | Regex patterns to include |
| `exclude_paths` | array | null | Regex patterns to exclude |
| `select_domains` | array | null | Regex for domains to include |
| `exclude_domains` | array | null | Regex for domains to exclude |
| `allow_external` | boolean | true (crawl) / false (map) | Include external domain links |
| `include_images` | boolean | false | Include images (crawl only) |
| `include_favicon` | boolean | false | Include favicon URL (crawl only) |
| `include_usage` | boolean | false | Include credit usage info |
| `timeout` | float | 150 | Max wait (10-150 seconds) |
---
## Instructions and Chunks
Use `instructions` and `chunks_per_source` for semantic focus and token optimization:
```python
response = client.crawl(
url="https://docs.example.com",
max_depth=2,
instructions="Find all documentation about authentication and security",
chunks_per_source=3 # Only top 3 relevant chunks per page
)
```
**Key benefits:**
- `instructions` guides crawler semantically, focusing on relevant content
- `chunks_per_source` returns only relevant snippets (max 500 chars each)
- Prevents context window explosion in agentic use cases
- Chunks appear in `raw_content` as: `<chunk 1> [...] <chunk 2> [...] <chunk 3>`
**Note:** `chunks_pereferences/extract.md
# Extract API Reference
## Table of Contents
- [Extraction Approaches](#extraction-approaches)
- [Key Parameters](#key-parameters)
- [Query and Chunks](#query-and-chunks)
- [Extract Depth](#extract-depth)
- [Advanced Filtering Strategies](#advanced-filtering-strategies)
- [Response Fields](#response-fields)
- [Summary](#summary)
---
## Extraction Approaches
### Search with include_raw_content
Get search results and content in one call:
```python
response = client.search(
query="AI healthcare applications",
include_raw_content=True,
max_results=5
)
```
**When to use:**
- Quick prototyping
- Simple queries where search results are likely relevant
- Single API call convenience
### Direct Extract API (Recommended)
Two-step pattern for more control:
```python
# Step 1: Search
search_results = client.search(
query="Python async best practices",
max_results=10
)
# Step 2: Filter by relevance score
relevant_urls = [
r["url"] for r in search_results["results"]
if r["score"] > 0.5
]
# Step 3: Extract with targeting
extracted = client.extract(
urls=relevant_urls[:20],
query="async patterns and concurrency", # Reranks chunks
chunks_per_source=3 # Prevents context explosion
)
for item in extracted["results"]:
print(f"URL: {item['url']}")
print(f"Content: {item['raw_content'][:500]}...")
```
**When to use:**
- You want control over which URLs to extract
- You need to filter/curate URLs before extraction
- You want targeted extraction with query and chunks_per_source
---
## Key Parameters
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `urls` | string/array | Required | Single URL or list (max 20) |
| `extract_depth` | enum | `"basic"` | `"basic"` or `"advanced"` (for complex/JS pages) |
| `query` | string | null | Reranks chunks by relevance to this query |
| `chunks_per_source` | integer | 3 | Chunks per source (1-5, max 500 chars each). Only with `query` |
| `format` | enum | `"markdown"` | Output: `"markdown"` or `"text"` |
| `include_images` | boolean | false | Include image URLs |
| `include_favicon` | boolean | false | Include favicon URL |
| `include_usage` | boolean | false | Include credit consumption data in response |
| `timeout` | float | varies | Max wait time (1.0-60.0 seconds) |
---
## Query and Chunks
Use `query` and `chunks_per_source` to get only relevant content and prevent context window explosion:
```python
extracted = client.extract(
urls=[
"https://example.com/ml-healthcare",
"https://example.com/ai-diagnostics",
"https://example.com/medical-ai"
],
query="AI diagnostic tools accuracy",
chunks_per_source=2 # 2 most relevant chunks per URL
)
```
**When to use query:**
- To extract only relevant portions of long documents
- When you need focused content instead of full page extraction
- For targeted information retrieval from specific URLs
**Key benefits of chunks_per_source:**
- Retureferences/integrations.md
# Framework Integrations
## Table of Contents
- [LangChain](#langchain)
- [LlamaIndex](#llamaindex)
- [OpenAI Function Calling](#openai-function-calling)
- [Anthropic Tool Use](#anthropic-tool-use)
- [Vercel AI SDK](#vercel-ai-sdk)
- [CrewAI](#crewai)
- [No-Code Platforms](#no-code-platforms)
---
## LangChain
The `langchain-tavily` package is the official LangChain integration supporting Search, Extract, Map, Crawl, and Research.
### Installation
```bash
pip install -U langchain-tavily
```
### Search
```python
from langchain_tavily import TavilySearch
tool = TavilySearch(
max_results=5,
topic="general", # or "news", "finance"
# search_depth="basic",
# include_answer=False,
# include_raw_content=False,
)
# Direct invocation
result = tool.invoke({"query": "What happened at Wimbledon?"})
# With agent
from langchain.agents import create_agent
from langchain_openai import ChatOpenAI
agent = create_agent(
model=ChatOpenAI(model="gpt-4"),
tools=[tool],
system_prompt="You are a helpful research assistant."
)
response = agent.invoke({
"messages": [{"role": "user", "content": "What are the latest AI trends?"}]
})
```
**Dynamic parameters at invocation:**
- `include_images`, `search_depth`, `time_range`, `include_domains`, `exclude_domains`, `start_date`, `end_date`
### Extract
```python
from langchain_tavily import TavilyExtract
tool = TavilyExtract(
extract_depth="basic", # or "advanced"
# include_images=False
)
result = tool.invoke({
"urls": ["https://en.wikipedia.org/wiki/Lionel_Messi"]
})
```
### Map
```python
from langchain_tavily import TavilyMap
tool = TavilyMap()
result = tool.invoke({
"url": "https://docs.example.com",
"instructions": "Find all documentation and tutorial pages"
})
# Returns: {"base_url": ..., "results": [urls...], "response_time": ...}
```
### Crawl
```python
from langchain_tavily import TavilyCrawl
tool = TavilyCrawl()
result = tool.invoke({
"url": "https://docs.example.com",
"instructions": "Extract API documentation and code examples"
})
# Returns: {"base_url": ..., "results": [{url, raw_content}...], "response_time": ...}
```
### Research
```python
from langchain_tavily import TavilyResearch, TavilyGetResearch
# Start research
research_tool = TavilyResearch(model="mini")
result = research_tool.invoke({
"input": "Research the latest developments in AI",
"citation_format": "apa"
})
# Get results
get_tool = TavilyGetResearch()
final = get_tool.invoke({"request_id": result["request_id"]})
```
---
## LlamaIndex
```python
from llama_index.tools.tavily_research import TavilyToolSpec
# Initialize tools
tavily_tool = TavilyToolSpec(api_key="tvly-YOUR_API_KEY")
tools = tavily_tool.to_tool_list()
# Use with agent
from llama_index.agent.openai import OpenAIAgent
agent = OpenAIAgent.from_tools(tools)
response = agent.chat("What are the latest AI developments?")
```
---
## OpenAI Function Calling
Define Tavily as an OpenAI functiAionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/barneyjm/skills/tavily-best-practices",
"sourceUrl": "https://clawhub.ai/barneyjm/skills/tavily-best-practices",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T03:29:59.974Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-tavily-best-practices/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-tavily-best-practices/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T03:29:59.974Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.7K downloads",
"href": "https://clawhub.ai/barneyjm/tavily-best-practices",
"sourceUrl": "https://clawhub.ai/barneyjm/tavily-best-practices",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T03:29:59.974Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.1.0",
"href": "https://clawhub.ai/barneyjm/tavily-best-practices",
"sourceUrl": "https://clawhub.ai/barneyjm/tavily-best-practices",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-02-03T15:27:32.248Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-tavily-best-practices/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-tavily-best-practices/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.1.0",
"description": "- Initial release of tavily-best-practices skill. - Provides reference documentation for production-ready Tavily integrations, including best practices. - Covers usage for web search, content extraction, crawling, research, and agentic workflows. - Includes SDK quickstart and parameter guides for Python and JavaScript. - Links to detailed guides for search, extraction, crawling, research, and integrations.",
"href": "https://clawhub.ai/barneyjm/tavily-best-practices",
"sourceUrl": "https://clawhub.ai/barneyjm/tavily-best-practices",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-02-03T15:27:32.248Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
