Crawler Summary

heftra answer-first brief

Autonomous invoice processing & spend categorization (CrewAI Flow + RAG over a chart of accounts, local vLLM) <picture> <source media="(prefers-color-scheme: dark)" srcset="brand/png/heftra-lockup-white-600.png"> <img alt="Heftra" src="brand/png/heftra-lockup-black-600.png" width="300"> </picture> Heftra **Know the true price of everything you buy.** Heftra ingests a company's spend from its ERP, categorizes every invoice line against the company's own spend tree, and surfaces redundant suppliers and product-level savings. T Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Freshness

Last checked 10/9/2026

Best For

heftra is best for crewai, multi-agent workflows where OpenClaw compatibility matters.

Not Ideal For

Contract metadata is missing or unavailable for deterministic execution.

Evidence Sources Checked

editorial-content, GITHUB REPOS, runtime-metrics, public facts pack

Agent DossierGITHUB REPOSSafety: 66/100

heftra

Autonomous invoice processing & spend categorization (CrewAI Flow + RAG over a chart of accounts, local vLLM) <picture> <source media="(prefers-color-scheme: dark)" srcset="brand/png/heftra-lockup-white-600.png"> <img alt="Heftra" src="brand/png/heftra-lockup-black-600.png" width="300"> </picture> Heftra **Know the true price of everything you buy.** Heftra ingests a company's spend from its ERP, categorizes every invoice line against the company's own spend tree, and surfaces redundant suppliers and product-level savings. T

OpenClawself-declared

Public facts

4

Change events

1

Artifacts

0

Freshness

Oct 9, 2026

Verifiededitorial-contentNo verified compatibility signals

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Trust evidence available

Trust score

Unknown

Compatibility

OpenClaw

Freshness

Oct 9, 2026

Vendor

Hadisdev

Artifacts

0

Benchmarks

0

Last release

Unpublished

Executive Summary

Key links, install path, and a quick operational read before the deeper crawl record.

Verifiededitorial-content

Summary

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Setup snapshot

  1. 1

    Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.

  2. 2

    Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Evidence Ledger

Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.

Verifiededitorial-content
Vendor (1)

Vendor

Hadisdev

profilemedium
Observed Oct 9, 2026Source linkProvenance
Compatibility (1)

Protocol compatibility

OpenClaw

contractmedium
Observed Oct 9, 2026Source linkProvenance
Security (1)

Handshake status

UNKNOWN

trustmedium
Observed unknownSource linkProvenance
Integration (1)

Crawlable docs

6 indexed pages on the official domain

search_documentmedium
Observed Apr 15, 2026Source linkProvenance

Release & Crawl Timeline

Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.

Self-declaredagent-index

Artifacts Archive

Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.

Self-declaredGITHUB REPOS

Extracted files

0

Examples

6

Snippets

0

Languages

python

Executable Examples

text

PDF -> extract -> verify -> research products -> categorize (leaf + Direct/Indirect) -> output/ledger.csv

bash

# 1. Install dependencies (creates the .venv on Python 3.12)
uv sync

# 2. Create your env file and set the BUYER (this drives Direct/Indirect)
cp .env.example .env
#    then edit .env and set at least:
#      BUYER_NAME=Your Company, Inc.
#      BUYER_WEBSITE=https://yourcompany.example
#    (VLLM_* defaults already point at http://localhost:8000/v1)

bash

docker volume create heftra_rustfs_data
docker run --rm -v steelyard_rustfs_data:/from:ro -v heftra_rustfs_data:/to \
  alpine cp -a /from/. /to/
# .env: S3_BUCKET=steelyard  S3_ACCESS_KEY=steelyard  S3_SECRET_KEY=steelyard-dev-secret
#       DATABASE_URL=postgresql://heftra:heftra@localhost:5432/heftra

bash

# 1. Drop PDF invoices into data/invoices/ (a sample is included).

# 2. Make sure your vLLM server is running (see Prerequisites).

# 3. Run the pipeline:
uv run apps/ai-api/main.py

# 4. Inspect the results:
cat output/ledger.csv

bash

rm -rf chroma_db

bash

# Web API (http://localhost:8100), from apps/web-api
cd apps/web-api && uv run src/web_api/app.py

# Pipeline worker: executes the runs a system admin requests from
# Settings -> Companies, and reads pending documents on its own
uv run python -m ai_api.worker

# One pass only (handy for scripts and checks)
uv run python -m ai_api.worker --once

Docs & README

Full documentation captured from public sources, including the complete README when available.

Self-declaredGITHUB REPOS

Docs source

GITHUB REPOS

Editorial quality

ready

Autonomous invoice processing & spend categorization (CrewAI Flow + RAG over a chart of accounts, local vLLM) <picture> <source media="(prefers-color-scheme: dark)" srcset="brand/png/heftra-lockup-white-600.png"> <img alt="Heftra" src="brand/png/heftra-lockup-black-600.png" width="300"> </picture> Heftra **Know the true price of everything you buy.** Heftra ingests a company's spend from its ERP, categorizes every invoice line against the company's own spend tree, and surfaces redundant suppliers and product-level savings. T

Full README
<picture> <source media="(prefers-color-scheme: dark)" srcset="brand/png/heftra-lockup-white-600.png"> <img alt="Heftra" src="brand/png/heftra-lockup-black-600.png" width="300"> </picture>

Heftra

Know the true price of everything you buy. Heftra ingests a company's spend from its ERP, categorizes every invoice line against the company's own spend tree, and surfaces redundant suppliers and product-level savings. The brand pack — logo, favicons, usage rules — lives in brand/.

Autonomous invoice processing (CrewAI + RAG)

The original PDF pipeline: a multi-agent pipeline that reads PDF invoices and codes each to a corporate chart of accounts, writing results to a CSV ledger. Built on CrewAI (Flow + Agent.kickoff) with RAG-backed categorization over a persisted ChromaDB index.

Pipeline

PDF -> extract -> verify -> research products -> categorize (leaf + Direct/Indirect) -> output/ledger.csv
  1. Extract structured fields from the invoice text.
  2. Verify the arithmetic (line items -> subtotal -> total).
  3. Categorize picks the leaf account (L2/L3/leaf from the chart of accounts) and classifies Direct/Indirect from buyer and product context using RAG retrieval.

Prerequisites

  • Astral uv and Python 3.12 (uv installs it).
  • A local vLLM server running and serving the model on an OpenAI-compatible API at http://localhost:8000/v1 (model google/gemma-4-E4B-it, referenced by CrewAI as hosted_vllm/google/gemma-4-E4B-it). Confirm it's up with curl -s http://localhost:8000/v1/models.
  • Internet access at run time — the buyer's website is scraped and each line item is web-searched (both degrade gracefully if offline).
  • Embeddings use a local sentence-transformers model (Snowflake/snowflake-arctic-embed-l-v2.0, multilingual, about 2 GB, downloaded automatically on first run). On a GPU it runs in half precision and needs about 1.5 GB of VRAM beside the LLM.

Setup

# 1. Install dependencies (creates the .venv on Python 3.12)
uv sync

# 2. Create your env file and set the BUYER (this drives Direct/Indirect)
cp .env.example .env
#    then edit .env and set at least:
#      BUYER_NAME=Your Company, Inc.
#      BUYER_WEBSITE=https://yourcompany.example
#    (VLLM_* defaults already point at http://localhost:8000/v1)

The compose project, database, role and default bucket are named heftra. The POSTGRES_* values only apply when pgdata/ is first initialised. To carry over a stack created under an older name, stop it with docker compose -p steelyard down, start postgres and run scripts/rename-dev-db.sh (pass --from spend_predictor for the oldest stacks). Then copy the rustfs_data volume across and keep the old S3_* values in .env:

docker volume create heftra_rustfs_data
docker run --rm -v steelyard_rustfs_data:/from:ro -v heftra_rustfs_data:/to \
  alpine cp -a /from/. /to/
# .env: S3_BUCKET=steelyard  S3_ACCESS_KEY=steelyard  S3_SECRET_KEY=steelyard-dev-secret
#       DATABASE_URL=postgresql://heftra:heftra@localhost:5432/heftra

Run

# 1. Drop PDF invoices into data/invoices/ (a sample is included).

# 2. Make sure your vLLM server is running (see Prerequisites).

# 3. Run the pipeline:
uv run apps/ai-api/main.py

# 4. Inspect the results:
cat output/ledger.csv

Each invoice produces exactly one ledger row. Columns include vendor_name, supplier_country_code, supplier_vat_number, buyer_name, buyer_country_code, buyer_vat_number, level1 (Direct/Indirect), level2, level3, account_code, account_name, total, arithmetic_ok, confidence, and notes. Status is processed, skipped (unreadable/empty PDF), or error (a stage failed, e.g. an LLM timeout) — skipped/error rows record the reason in notes and the batch continues.

First run is slower: it scrapes the buyer site and searches each line item; both are cached under data/web_cache/ (gitignored), so re-runs are faster.

Invoices are independent and processed concurrently — set INVOICE_CONCURRENCY in .env (default 4) to match what your vLLM server handles; 1 forces strictly sequential processing (deterministic ledger order). The shared index build and buyer-website scrape happen once, before the batch.

Changing the chart of accounts

data/chart_of_accounts.csv provides level2/level3 and the leaf account (account_code, account_name). Direct/Indirect (level1) is not stored here — it's judged per invoice from the buyer context. Replace the sample with your real chart (same columns). The ChromaDB index rebuilds automatically when the row count changes; if you edit rows without changing the count, force a rebuild:

rm -rf chroma_db

Run the app

The app needs two long-running processes next to the database: the web API and the pipeline worker. Start each in its own terminal:

# Web API (http://localhost:8100), from apps/web-api
cd apps/web-api && uv run src/web_api/app.py

# Pipeline worker: executes the runs a system admin requests from
# Settings -> Companies, and reads pending documents on its own
uv run python -m ai_api.worker

# One pass only (handy for scripts and checks)
uv run python -m ai_api.worker --once

Without the worker, requested runs stay queued. It polls every WORKER_POLL_SECONDS (default 5) when idle and reads at most WORKER_DOCUMENT_BATCH (default 5) pending documents per pass, so a requested run never waits behind a long backlog. The stage CLIs (python -m ai_api.sync.runner, python -m ai_api.documents.runner) keep working on their own.

Landing site

The public marketing site at heftra.com lives in apps/landing, a static Astro site with its own bun project. It runs on http://localhost:3200 and posts demo requests to the web API, so add http://localhost:3200 to WEB_API_CORS_ORIGINS and set DEMO_BOOKING_URL to hand out the booking link.

cd apps/landing && cp .env.example .env && ./node_modules/.bin/astro dev --port 3200

Investor demo

apps/demo deploys a read-only copy of the web app at demo.heftra.com with Dokploy and Cloudflare Tunnel. It serves only the fictional Nordlys Byg A/S, from a database dump built by scripts/build-demo-db.sh, behind one shared Clerk login.

Test

uv run pytest

Unit tests run fully offline (no LLM, no network, no model download — web lookups and embeddings are faked via dependency injection).

Synthetic data & benchmarking

The ai_api.synthdata subpackage generates labeled synthetic invoice fixtures (PDF + structured fields + ERP journal entries + category labels) and benchmarks the extraction/categorization pipeline against them using ANLS. Labels are chosen programmatically from the chart of accounts and buyer profiles (Direct/Indirect is derived from the buyer's business), so every fixture's labels are ground-truth by construction.

By default, the generator produces richly varied data with no LLM required: industry-flavored vendor names, per-account realistic line-item catalogs, 9 distinct invoice templates (modern, classic, minimal, corporate, eu_vat, us_net30, freelancer, saas_receipt, utility) chosen at random, plus per-invoice randomized accent color, font, logo/monogram, and realistic extra fields (addresses, PO number, payment terms, due date, bank/IBAN, notes). The --live flag (which requires uv sync --group live

  • a running vLLM) is optional — it only swaps in LLM-written line-item descriptions for extra realism; all other variation and quality work without it. Templates are auto-discovered from apps/ai-api/src/ai_api/synthdata/render/templates/*.html, so you can drop in your own .html template and it joins the rotation automatically.

Installation

Scoring and the unit tests do NOT need the live group. PDF rendering uses WeasyPrint (a regular project dependency); it needs system libraries that are usually already present on desktop Linux, but on a bare system install them with:

sudo apt-get install -y libpango-1.0-0 libpangocairo-1.0-0 libgdk-pixbuf-2.0-0 libffi-dev libcairo2

To use LLM-generated line-item descriptions, install the optional dependency group:

uv sync --group live

Generate synthetic invoices

uv run python -m ai_api.synthdata.generate --n 100 --seed 7 --out data/synthetic

Requires: by default, NO LLM or vLLM — the generator produces richly varied invoices deterministically. To use LLM-written line-item descriptions, pass --live (which requires uv sync --group live + a running vLLM server). The --cryptic flag (terse, harder-to-categorize descriptions) only has an effect with --live.

Each fixture is written to its own directory:

  • data/synthetic/<id>/invoice.pdf — the rendered invoice
  • data/synthetic/<id>/labels.json — ground-truth: the extracted fields, the category (account_code, level1/level2/level3), the buyer, and the double-entry journal
  • data/synthetic/manifest.jsonl — an index of all generated fixtures

Author new templates from web references (optional)

Grow the template library by drafting new templates from real invoice designs found online. This is an offline developer tool — separate from the generator — and is human-gated: it stages drafts for you to review, and never writes into render/templates/ itself.

uv run python -m ai_api.synthdata.templategen --n 5
# or drive the search yourself:
uv run python -m ai_api.synthdata.templategen --query "eu vat invoice template" --n 8

It searches DuckDuckGo images (no key), drafts a Jinja2 template per image via the local vision LLM (requires your vLLM server to serve the model with vision enabled), then validates each draft — it must render cleanly, contain the required placeholders, and pass a no-real-data lint (no emails, contiguous runs of 4+ digits — so spaced or hyphenated numbers may slip through — or embedded image URLs; human review is the real guarantee). Results land in data/template_drafts/ (gitignored): passing drafts as <name>.html + <name>.pdf preview, failures under _rejected/ with a reason, plus a report.md. Review them, then move the good .html files into apps/ai-api/src/ai_api/synthdata/render/templates/ — the generator auto-discovers them.

No real data ever enters a template: the vision model is instructed to copy only layout/styling and use Jinja placeholders for all data; the lint and your manual review are the backstops.

Score extraction & categorization accuracy

uv run python -m ai_api.synthdata.score --fixtures data/synthetic

Requires: the local vLLM server running — scoring runs the real extract → categorize pipeline (which calls the model) over each fixture PDF and compares the result to labels.json. It reports per-field ANLS for extraction plus exact-match accuracy for the leaf account code and the Direct/Indirect (L1), L2, and L3 labels.

Output

The data/synthetic/ directory is gitignored — it is regenerated per run and not committed.

Contract & API

Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.

MissingGITHUB REPOS

Contract coverage

Status

missing

Auth

None

Streaming

No

Data region

Unspecified

Protocol support

OpenClaw: self-declared

Requires: none

Forbidden: none

Guardrails

Operational confidence: low

No positive guardrails captured.
Invocation examples
curl -s "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/trust"

Reliability & Benchmarks

Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.

Missingruntime-metrics

Trust signals

Handshake

UNKNOWN

Confidence

unknown

Attempts 30d

unknown

Fallback rate

unknown

Runtime metrics

Observed P50

unknown

Observed P95

unknown

Rate limit

unknown

Estimated cost

unknown

Do not use if

Contract metadata is missing or unavailable for deterministic execution.
No benchmark suites or observed failure patterns are available.

Media & Demo

Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.

Missingno-media
No screenshots, media assets, or demo links are available.

Related Agents

Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.

Self-declaredprotocol-neighbors
Github ReposUpdated 4h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW
Machine Appendix

Contract JSON

{
  "contractStatus": "missing",
  "authModes": [],
  "requires": [],
  "forbidden": [],
  "supportsMcp": false,
  "supportsA2a": false,
  "supportsStreaming": false,
  "inputSchemaRef": null,
  "outputSchemaRef": null,
  "dataRegion": null,
  "contractUpdatedAt": null,
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Invocation Guide

{
  "preferredApi": {
    "snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/snapshot",
    "contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/contract",
    "trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/trust"
  },
  "curlExamples": [
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/snapshot\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/contract\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/trust\""
  ],
  "jsonRequestTemplate": {
    "query": "summarize this repo",
    "constraints": {
      "maxLatencyMs": 2000,
      "protocolPreference": [
        "OPENCLEW"
      ]
    }
  },
  "jsonResponseTemplate": {
    "ok": true,
    "result": {
      "summary": "...",
      "confidence": 0.9
    },
    "meta": {
      "source": "GITHUB_REPOS",
      "generatedAt": "2026-10-09T23:08:36.584Z"
    }
  },
  "retryPolicy": {
    "maxAttempts": 3,
    "backoffMs": [
      500,
      1500,
      3500
    ],
    "retryableConditions": [
      "HTTP_429",
      "HTTP_503",
      "NETWORK_TIMEOUT"
    ]
  }
}

Trust JSON

{
  "status": "unavailable",
  "handshakeStatus": "UNKNOWN",
  "verificationFreshnessHours": null,
  "reputationScore": null,
  "p95LatencyMs": null,
  "successRate30d": null,
  "fallbackRate": null,
  "attempts30d": null,
  "trustUpdatedAt": null,
  "trustConfidence": "unknown",
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Capability Matrix

{
  "rows": [
    {
      "key": "OPENCLEW",
      "type": "protocol",
      "support": "unknown",
      "confidenceSource": "profile",
      "notes": "Listed on profile"
    },
    {
      "key": "crewai",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    },
    {
      "key": "multi-agent",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    }
  ],
  "flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}

Facts JSON

[
  {
    "factKey": "vendor",
    "category": "vendor",
    "label": "Vendor",
    "value": "Hadisdev",
    "href": "https://github.com/HadiSDev/heftra",
    "sourceUrl": "https://github.com/HadiSDev/heftra",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-10-09T10:43:30.741Z",
    "isPublic": true
  },
  {
    "factKey": "protocols",
    "category": "compatibility",
    "label": "Protocol compatibility",
    "value": "OpenClaw",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/contract",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/contract",
    "sourceType": "contract",
    "confidence": "medium",
    "observedAt": "2026-10-09T10:43:30.741Z",
    "isPublic": true
  },
  {
    "factKey": "docs_crawl",
    "category": "integration",
    "label": "Crawlable docs",
    "value": "6 indexed pages on the official domain",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  },
  {
    "factKey": "handshake_status",
    "category": "security",
    "label": "Handshake status",
    "value": "UNKNOWN",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/trust",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-hadisdev-heftra/trust",
    "sourceType": "trust",
    "confidence": "medium",
    "observedAt": null,
    "isPublic": true
  }
]

Change Events JSON

[
  {
    "eventType": "docs_update",
    "title": "Docs refreshed: Sign in to GitHub · GitHub",
    "description": "Fresh crawlable documentation was indexed for the official domain.",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  }
]

Sponsored

Ads related to heftra and adjacent AI workflows.