Document-Analyzer-CrewAI
Multi-agent system that reads financial PDFs and returns structured investment insights via REST API Financial Document Analyzer — Multi-Agent AI Pipeline Upload any financial PDF (annual report, investor deck, earnings statement) and get structured investment insights returned via a REST API. Built with CrewAI multi-agent orchestration — a financial analyst agent reads the document, reasons over it, and optionally searches the web for market context before producing a structured report. --- Tech stack | Layer | Too
Rank
27
Safety
66
Updated
May 31, 2026
Source
GITHUB OPENCLEW
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. Last updated 5/31/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Mrunmayeenaikvendor · observed May 31, 2026
- Protocol compatibility
- OpenClawcompatibility · observed May 31, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
git clone https://github.com/MrunmayeeNaik/Document-Analyzer-CrewAI.git- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/crewai-mrunmayeenaik-document-analyzer-crewai/snapshot"
Documentation
GITHUB OPENCLEW
Read the full documentation
Financial Document Analyzer — Multi-Agent AI Pipeline
Upload any financial PDF (annual report, investor deck, earnings statement) and get structured investment insights returned via a REST API.
Built with CrewAI multi-agent orchestration — a financial analyst agent reads the document, reasons over it, and optionally searches the web for market context before producing a structured report.
Tech stack
| Layer | Tool | |---|---| | Agent orchestration | CrewAI | | LLM backend | OpenAI-compatible (Groq) | | PDF extraction | Custom tool (PyMuPDF) | | Web search | Serper API | | API layer | FastAPI | | Storage | SQLite via SQLAlchemy | | Testing | Postman |
What it produces
For any uploaded financial PDF, the system returns:
- Document summary
- Key financial metrics and trends
- Analysis relevant to your specific query
- Risks and uncertainties
- High-level recommendations (with disclaimer)
Previous analyses are stored and retrievable via GET /history.
1. Bugs Found and How They Were Fixed
-
LLM ignored the uploaded PDF and returned a generic answer
- Symptom:
analysisfield said things like “without access to the specific financial document…” even though a PDF was uploaded. - Cause: The
analyze_financial_documenttask intask.pyonly referenced{query}and never mentioned{file_path}, so thefinancial_analystagent was not clearly instructed to use the PDF reader tool with the actual path. - Fix: Updated the task description to explicitly include
{file_path}and to instruct the agent to call thefinancial_document_readertool with that path before analyzing.
- Symptom:
-
HTTP 500 with LLM error
Request too large / rate_limit_exceeded- Symptom:
POST /analyzereturned a 500 with a nested 413‑style error from Groq: “Request too large for model … Limit 12000, Requested 13632”. - Cause:
FinancialDocumentToolintools.pyalways returned the full text of the PDF, which made some prompts exceed the model’s token‑per‑minute limits for large documents. - Fix: Added truncation in
FinancialDocumentTool._runso only the first N characters (default 20,000, configurable viaMAX_PDF_CHARS) are passed to the model, with a note appended when truncation occurs.
- Symptom:
-
No records visible in
/history- Symptom:
GET /historyalways returned an empty list. - Causes:
- Earlier
/analyzecalls failed (see above), so the DB write inmain.pynever executed. - The SQLite URL (
sqlite:///./analysis.db) is relative, so starting the server from the wrong working directory can create or query the wrong DB file.
- Earlier
- Fix: After fixing the LLM errors and ensuring the server is started from the project root, successful
/analyzecalls now insert rows intoAnalysis, and/historyreturns data.
- Symptom:
-
Incorrect install command in README
- Symptom: README instructed
pip install -r requirement.txt. - Cause: The actual dependency file in the repo is
requirements.txt. - Fix: Updated the installation instructions to use
requirements.txt.
- Symptom: README instructed
2. Setup Instructions
-
Requirements
- Python 3.10+ recommended
pipfor installing dependencies
-
Install dependencies
pip install -r requirements.txt
- Environment variables
Create a .env file in the project root with at least:
OPENAI_API_KEY: Groq/OpenAI‑compatible key used by CrewAI’sLLM(backed by the Groq endpoint).OPENAI_BASE_URL: Base URL for the Groq OpenAI‑compatible API (already set tohttps://api.groq.com/openai/v1in this project).SERPER_API_KEY: API key for the Serper search tool used by the agents.- Optional:
MAX_PDF_CHARS– maximum number of characters of PDF text passed to the LLM (defaults to20000).
Security note: Keep
.envout of version control; do not share your API keys.
- Run the API
From the project root (financial-document-analyzer-debug):
uvicorn main:app --host 0.0.0.0 --port 8000 --reload
The API will be available at http://localhost:8000.
3. Usage Instructions
-
1) Start the server
- Run the
uvicorncommand above from the project root so thatanalysis.dbis created and used in the correct directory.
- Run the
-
2) Analyze a financial PDF
- Endpoint:
POST /analyze - Use a tool like Postman, curl, or a frontend form with:
- A file field named
file(PDF). - An optional form field
query(defaults to: “Analyze this financial document for investment insights”).
- A file field named
- The backend will:
- Save the uploaded PDF to
data/financial_document_<uuid>.pdf. - Run the CrewAI pipeline (financial analyst agent + tools).
- Store the result in SQLite (
analysis.db, tableanalyses). - Return the structured analysis in the response.
- Save the uploaded PDF to
- Endpoint:
-
3) View analysis history
- Endpoint:
GET /history - Returns basic metadata for all stored analyses (ID, file name, query, timestamp).
- Endpoint:
-
4) Inspect the database (optional)
- A SQLite DB file named
analysis.dbis created in the project root. - You can open it with any SQLite browser to inspect the
analysestable.
- A SQLite DB file named
4. API Documentation
GET /
- Description: Health check.
- Response:
200 OK– JSON:message:"Financial Document Analyzer API is running"
POST /analyze
-
Description: Analyze an uploaded financial PDF and persist the result.
-
Request
- Content-Type:
multipart/form-data - Fields:
file(required): The PDF to analyze (UploadFile).query(optional,Formstring):- Default:
"Analyze this financial document for investment insights". - Used as the user’s prompt for the analysis.
- Default:
- Content-Type:
-
Processing steps
- Save the uploaded file to
data/financial_document_<uuid>.pdf. - Call
run_crew(query, file_path):- Creates a
Crewwith thefinancial_analystagent andanalyze_financial_documenttask. - The task:
- Uses
financial_document_readerto read and clean the PDF text (truncated to avoid token limits). - Optionally uses the Serper search tool for market context.
- Produces a structured analysis with:
- Document Summary
- Key Financial Metrics and Trends
- Analysis Relevant to the User's Query
- Risks and Uncertainties
- High-level, non-personalized recommendations and a disclaimer
- Uses
- Creates a
- Store the result in SQLite via the
Analysismodel. - Delete the temporary PDF from
data/in thefinallyblock.
- Save the uploaded file to
-
Successful Response
- Status:
200 OK - Body (JSON):
status:"success"query: The final query string used.analysis: The full textual analysis produced by the Crew.file_processed: Original filename of the uploaded PDF.
- Status:
-
Error Responses
400/422: Validation errors (e.g., missing file field) handled by FastAPI automatically.500: Internal errors, wrapped as:{"detail": "Error processing financial document: <message>"}- Examples:
- File not found / invalid PDF path.
- LLM API issues (e.g., network or credential problems).
- Unexpected exceptions during crew execution.
GET /history
-
Description: Return a list of previously stored analyses (metadata only).
-
Response
- Status:
200 OK - Body (JSON array):
- Status:
[
{
"id": "uuid-string",
"file_name": "example.pdf",
"query": "Analyze this financial document for investment insights",
"created_at": "2026-02-26T12:34:56.789000"
}
]
The full analysis text is stored in the database but not returned by
/historyto keep the payload small. You can extend the API with a/history/{id}endpoint if you need to retrieve full analyses by ID.
5. High-Level Architecture
-
FastAPI (
main.py)- Defines the HTTP endpoints.
- Orchestrates file upload, crew execution, and DB persistence.
-
CrewAI Agents and Tasks (
agents.py,task.py)financial_analyst: Main agent that reads the PDF and generates the structured report.- Tasks describe what to do with the uploaded document and the user query.
-
Tools (
tools.py)financial_document_reader: Reads and cleans PDF content, now with length limiting to avoid LLM token issues.search_tool: Serper‑based web search for financial and market context.
-
Database (
database.py)- SQLite via SQLAlchemy.
Analysismodel stores file name, query, analysis text, and creation timestamp.
@x1pay/langchain
LangChain/LangGraph tools for AI agent x402 payments on X1
@langchain/langgraph-swarm
An implementation of a multi-agent swarm using LangGraph
@langchain/langgraph-supervisor
LangGraph Multi-Agent Supervisor
oceanbus-langchain
LangChain tools for OceanBus — give your LangChain and CrewAI agents a global identity, encrypted messaging, and Yellow Pages service discovery with a single import.
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"label": "Vendor",
"value": "Mrunmayeenaik",
"category": "vendor",
"href": "https://github.com/MrunmayeeNaik/Document-Analyzer-CrewAI",
"sourceUrl": "https://github.com/MrunmayeeNaik/Document-Analyzer-CrewAI",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-05-31T06:18:24.024Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "protocols",
"label": "Protocol compatibility",
"value": "OpenClaw",
"category": "compatibility",
"href": "https://www.xpersona.co/api/v1/agents/crewai-mrunmayeenaik-document-analyzer-crewai/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-mrunmayeenaik-document-analyzer-crewai/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-05-31T06:18:24.024Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "handshake_status",
"label": "Handshake status",
"value": "UNKNOWN",
"category": "security",
"href": "https://www.xpersona.co/api/v1/agents/crewai-mrunmayeenaik-document-analyzer-crewai/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-mrunmayeenaik-document-analyzer-crewai/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true,
"metadata": {}
}
],
"events": []
}Record generated Oct 8, 2026.
