Crawler Summary

practo-support-agent answer-first brief

Healthcare support agent with calibrated RAG retrieval, a three-agent CrewAI crew, independent Autogen review, and an LLM-as-judge evaluation across accuracy, grounding, completeness and safety. Deterministic and offline. Practo Domain Support Agent **Track: Healthcare (Practo).** A clinic patient-support agent that answers policy questions from a knowledge base, looks up appointment records, remembers a conversation, guards against misuse, has its answers reviewed by an independent agent team, and operates under an explicit governance policy. Orchestrated with CrewAI, reviewed with Autogen, deployed behind FastAPI, evaluated end to e Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Freshness

Last checked 10/9/2026

Best For

practo-support-agent is best for crewai, multi-agent workflows where OpenClaw compatibility matters.

Not Ideal For

Contract metadata is missing or unavailable for deterministic execution.

Evidence Sources Checked

editorial-content, GITHUB REPOS, runtime-metrics, public facts pack

Claim this agent
Agent DossierGITHUB REPOSSafety: 66/100

practo-support-agent

Healthcare support agent with calibrated RAG retrieval, a three-agent CrewAI crew, independent Autogen review, and an LLM-as-judge evaluation across accuracy, grounding, completeness and safety. Deterministic and offline. Practo Domain Support Agent **Track: Healthcare (Practo).** A clinic patient-support agent that answers policy questions from a knowledge base, looks up appointment records, remembers a conversation, guards against misuse, has its answers reviewed by an independent agent team, and operates under an explicit governance policy. Orchestrated with CrewAI, reviewed with Autogen, deployed behind FastAPI, evaluated end to e

OpenClawself-declared

Public facts

4

Change events

1

Artifacts

0

Freshness

Oct 9, 2026

Verifiededitorial-contentNo verified compatibility signals

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Trust evidence available

Trust score

Unknown

Compatibility

OpenClaw

Freshness

Oct 9, 2026

Vendor

Vk409 Avenger

Artifacts

0

Benchmarks

0

Last release

Unpublished

Executive Summary

Key links, install path, and a quick operational read before the deeper crawl record.

Verifiededitorial-content

Summary

Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.

Setup snapshot

  1. 1

    Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.

  2. 2

    Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Evidence Ledger

Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.

Verifiededitorial-content
Vendor (1)

Vendor

Vk409 Avenger

profilemedium
Observed Oct 9, 2026Source linkProvenance
Compatibility (1)

Protocol compatibility

OpenClaw

contractmedium
Observed Oct 9, 2026Source linkProvenance
Security (1)

Handshake status

UNKNOWN

trustmedium
Observed unknownSource linkProvenance
Integration (1)

Crawlable docs

6 indexed pages on the official domain

search_documentmedium
Observed Apr 15, 2026Source linkProvenance

Release & Crawl Timeline

Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.

Self-declaredagent-index

Artifacts Archive

Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.

Self-declaredGITHUB REPOS

Extracted files

0

Examples

6

Snippets

0

Languages

python

Executable Examples

bash

pip install -r requirements.txt
cp .env.example .env
python run_all.py

bash

python run_all.py --list        # show the steps
python run_all.py --only 4 14   # rerun selected steps

text

IN-SCOPE QUERIES                                                    top-1 sim
  0.7567  [cancellation_window     ] How long before my appointment can I
                                     cancel without a charge?
  0.6574  [consultation_fees       ] What does a cardiology consultation cost?
  0.6276  [telemedicine_eligibility] Can I see a doctor over video if I have
                                     never visited before?
  0.5212  [lab_turnaround          ] How soon will my blood test results
                                     come back?
  0.5045  [follow_up_discount      ] Do I pay less if I come back to the same
                                     doctor next week?

OUT-OF-SCOPE QUERIES
  0.2587  [prescription_refills    ] What antibiotic should I take for a
                                     sore throat?
  0.1040  [data_privacy            ] Who won the 2019 cricket world cup?
  0.0745  [emergency_protocol      ] How do I reset the wifi router in my flat?

text

s permitted at no cost provided it is requested at least 6 hours before...

text

staleness(d) = clip((d − 7) / (30 − 7), 0, 1)
score        = 0.45 · follow_up_required + 0.55 · staleness(days_since_created)

text

policy query      -> agents: Retrieval, Composer          tools: ['rag_lookup']
appointment query -> agents: Retrieval, Lookup, Composer  tools: ['rag_lookup',
                                                                  'check_appointment_status']

Docs & README

Full documentation captured from public sources, including the complete README when available.

Self-declaredGITHUB REPOS

Docs source

GITHUB REPOS

Editorial quality

ready

Healthcare support agent with calibrated RAG retrieval, a three-agent CrewAI crew, independent Autogen review, and an LLM-as-judge evaluation across accuracy, grounding, completeness and safety. Deterministic and offline. Practo Domain Support Agent **Track: Healthcare (Practo).** A clinic patient-support agent that answers policy questions from a knowledge base, looks up appointment records, remembers a conversation, guards against misuse, has its answers reviewed by an independent agent team, and operates under an explicit governance policy. Orchestrated with CrewAI, reviewed with Autogen, deployed behind FastAPI, evaluated end to e

Full README

Practo Domain Support Agent

Track: Healthcare (Practo).

A clinic patient-support agent that answers policy questions from a knowledge base, looks up appointment records, remembers a conversation, guards against misuse, has its answers reviewed by an independent agent team, and operates under an explicit governance policy. Orchestrated with CrewAI, reviewed with Autogen, deployed behind FastAPI, evaluated end to end.

Everything runs under MOCK_LLM=1 with zero API keys and zero LLM network calls.


Quick start

pip install -r requirements.txt
cp .env.example .env
python run_all.py

run_all.py regenerates every transcript in transcripts/ in dependency order. Steps needing the knowledge base are skipped with a clear message if it is absent, so a partial clone still produces the component evidence.

python run_all.py --list        # show the steps
python run_all.py --only 4 14   # rerun selected steps

The one-time download of sentence-transformers/all-MiniLM-L6-v2 (~90 MB) is the only network access this project makes. No language-model call ever leaves the machine.


Task → transcript map

Every acceptance criterion has a file. This table is the fastest route through the evidence.

| Task | What it shows | Transcript | |---|---|---| | 1 | Dataset generation, category/status counts, follow-up band | transcripts/task01_dataset.txt | | 2 | Knowledge-base manifest, sentence-count validation | transcripts/task02_kb_manifest.txt | | 3 | Both chunkers, both ChromaDB collections, sample retrieval | transcripts/task03_indexing.txt | | 3 | Loader, chunker and Chroma wiring smoke test | transcripts/task03_pipeline_smoke.txt | | 4, 5 | Threshold calibration, grounded generation, chunking comparison (real corpus) | transcripts/task04_05_calibration_and_chunking.txt | | 4, 7 | RAG core and three-agent crew, both tools invoked | transcripts/task04_07_rag_crew.txt | | 5 | Chunking metric tests against hand-computed values | transcripts/task05_chunking_comparison.txt | | 6 | Escalation formula, distribution, derived threshold, ACL | transcripts/task06_lookup_escalation.txt | | 6 | Lookup tool tests including non-degeneracy | transcripts/task06_tests.txt | | 7 | MockLLM guards against both silent failures | transcripts/task07_mockllm_guards.txt | | 8, 16 | Multi-turn memory, fresh-session reset, cache hit | transcripts/task08_16_memory_cache.txt | | 9, 10 | Structured output validation, three guardrails firing | transcripts/task09_10_schemas_guardrails.txt | | 11, 12 | HTTP + WebSocket endpoints, disconnect survival, JSONL logs | transcripts/task11_12_api_logging.txt | | 13 | Judge harness and discrimination tests | transcripts/task13_evaluation.txt | | 13 | 15-query evaluation against the live pipeline (real corpus) | transcripts/task13_live_evaluation.txt | | 14 | Autogen review approving and revising, with structured verdicts | transcripts/task14_autogen_review.txt | | 15 | Four-layer governance, least autonomy, cost cap | transcripts/task15_governance.txt | | — | Full pipeline integration | transcripts/task_integration_engine.txt |

The judge prompt is committed separately at judge_prompt.txt.


Part 1 — Dataset and RAG core

Task 1 — dataset design choices

These are the values needed to reproduce the dataset exactly. python -m practo.dataset regenerates it.

| Choice | Value | |---|---| | Seed | 10 | | Records | 60 | | Category weights | General Medicine 0.26, Cardiology 0.15, Dermatology 0.15, Pediatrics 0.15, Orthopedics 0.15, ENT 0.14 | | Status weights | Scheduled 0.30, Completed 0.28, Cancelled 0.12, Rescheduled 0.12, Pending-Confirmation 0.10, No-Show 0.08 | | Added category | ENT | | Added status | Pending-Confirmation | | Fee bands (INR) | General Medicine 400–800, Pediatrics 500–1000, ENT 600–1100, Dermatology 700–1400, Orthopedics 800–1600, Cardiology 900–1800, quoted in ₹50 steps | | days_since_created | triangular(0, 31, mode=6), truncated at 30 | | follow_up_required | Conditional on status: Completed 0.45, Scheduled 0.10, Rescheduled 0.10, all others 0.02 |

Fee reasoning. The bands mirror private multi-specialty OPD pricing in Indian tier-1 cities, where a general consult sits at the floor and diagnostic- or procedure-heavy specialties command roughly twice that.

Observed result at seed 10: every supplied category has ≥3 records (minimum 6, Cardiology), every supplied status appears (minimum 2, No-Show), and follow_up_required is 20.0% — mid-band, on the first draw. No reseeding was needed and no record was hand-edited.

Why follow_up_required is conditional. A follow-up flag only means something once a consultation has happened. A flat coin-flip would put follow-ups on cancelled and no-show appointments, and would make the Task 6 escalation score degenerate — see below.

Why the day distribution is skewed. A real appointment queue is dominated by recent bookings with a thinning tail of unresolved older entries. A uniform spread would make the staleness term in Task 6 too easy to trip.

Task 4 — calibrated similarity threshold

Measured over the real corpus with sentence-transformers/all-MiniLM-L6-v2, cosine similarity reported as 1 − chroma_distance.

Collections are created with metadata={"hnsw:space": "cosine"}. ChromaDB defaults to squared L2. Without that setting the numbers returned are euclidean distances, nothing raises, and the calibration below would be measuring the wrong quantity.

IN-SCOPE QUERIES                                                    top-1 sim
  0.7567  [cancellation_window     ] How long before my appointment can I
                                     cancel without a charge?
  0.6574  [consultation_fees       ] What does a cardiology consultation cost?
  0.6276  [telemedicine_eligibility] Can I see a doctor over video if I have
                                     never visited before?
  0.5212  [lab_turnaround          ] How soon will my blood test results
                                     come back?
  0.5045  [follow_up_discount      ] Do I pay less if I come back to the same
                                     doctor next week?

OUT-OF-SCOPE QUERIES
  0.2587  [prescription_refills    ] What antibiotic should I take for a
                                     sore throat?
  0.1040  [data_privacy            ] Who won the 2019 cricket world cup?
  0.0745  [emergency_protocol      ] How do I reset the wifi router in my flat?

| | Value | |---|---| | Lowest in-scope similarity | 0.5045 | | Highest out-of-scope similarity | 0.2587 | | Observed gap | 0.2458 | | Chosen threshold | 0.3816 (midpoint of the measured gap) |

The gap is wide and the clusters do not overlap, so the midpoint is a defensible cut rather than a compromise. Note where the highest out-of-scope score sits: "What antibiotic should I take for a sore throat?" reaches 0.2587 against the refill policy, well above the other two. It is lexically close to a document the knowledge base does have. Retrieval similarity cannot separate "what is the refill policy" from "what should I take" — that difference is about authority, not topic — which is why the clinical-advice boundary is enforced before retrieval rather than left to the threshold. See Task 15.

Demonstrated on 5 in-scope queries, all answered, plus 1 deliberately out-of-scope query that correctly triggered the fallback at 0.1040.

The threshold is not a preset. guardrails.check_groundedness raises ThresholdNotCalibrated when passed None, and RagCore refuses to construct without one, so the system physically cannot run on an unmeasured value.

Task 5 — chunking recommendation

| Strategy | Chunks | Precision | Recall | F1 | |---|---|---|---|---| | fixed_overlap (250 chars, 50 overlap) | 36 | 0.444 | 0.917 | 0.583 | | sentence (2 sentences, non-overlapping) | 28 | 0.417 | 0.917 | 0.556 |

Generated recommendation, from those measured values:

Across 6 queries at k=3, fixed_overlap scored precision 0.444 and recall 0.917 (F1 0.583), against sentence at precision 0.417 and recall 0.917 (F1 0.556). Recall is identical at 0.917, so the difference is precision alone, where fixed_overlap leads by +0.028; note that only 1 of 6 queries separates the two, so this is a narrow result rather than a decisive one.

sentence is deployed, and the measurement narrowly favoured the other one. That disagreement is stated rather than hidden, because the reasoning matters more than the number.

The gap is 0.028 in F1 and rests entirely on Q5. Five of the six queries score identically under both strategies. On a six-query set that is one query's worth of movement, which is noise, not evidence.

The tiebreaker is what the two strategies put in front of a patient. Fixed-size windows cut wherever 250 characters lands, including mid-word. The Task 3 sample output shows it directly — the second-ranked fixed_overlap chunk for a cancellation query begins:

s permitted at no cost provided it is requested at least 6 hours before...

Sentence chunks cannot do that, because a chunk boundary is a sentence boundary. Given two strategies that are statistically indistinguishable on retrieval quality, the one that never emits a fragment is the right one to serve.

config.DEPLOYED_STRATEGY names the served collection in one place, and the recommendation reads that value rather than asserting it, so the two cannot silently drift apart again.

Scoring is at the parent-document level after deduplication. Relevance belongs to documents, not chunks: scoring chunks directly would reward whichever strategy fragments documents more, which is the variable under test.


Part 2 — Orchestration, tools, memory, guardrails

Task 6 — escalation score

staleness(d) = clip((d − 7) / (30 − 7), 0, 1)
score        = 0.45 · follow_up_required + 0.55 · staleness(days_since_created)

| | Value | |---|---| | Threshold | 0.4748 — the 80th percentile of the observed score distribution (statistics.quantiles, method='inclusive', n=60) | | Escalates | 12 of 60 records (20.0%) | | …carrying a follow-up flag | 8 | | …on age alone, no flag | 4 | | Flagged but held back | 4 | | Score range, flagged records | [0.4500, 1.0000] | | Score range, unflagged records | [0.0000, 0.5500] |

Why not the obvious weighting. With 0.6 · flag + 0.4 · linear age, every flagged record scores ≥ 0.60 and every unflagged record ≤ 0.40. The ranges cannot overlap, so any threshold between them reduces to the boolean — exactly what the brief rules out. Shifting weight onto staleness and delaying its onset to day 7 makes the ranges overlap, which is asserted directly in tests/test_tools_escalation.py.

Two cases that show it is a real score. APT-0003 escalates with no follow-up flag at all — a Pending-Confirmation booking untouched for 29 days. APT-0016 carries the flag and does not escalate — it was flagged 8 days ago, so staleness has barely started.

Task 7 — the crew

| Agent | Tools | Runs when | |---|---|---| | Retrieval Agent | rag_lookup | always | | Lookup Agent | check_appointment_status | only when the query references a record | | Response Composer | none | always |

Tasks are assembled per query. A policy question has no appointment to look up, so handing the Lookup Agent a task with no record would either fabricate an identifier or waste a turn.

policy query      -> agents: Retrieval, Composer          tools: ['rag_lookup']
appointment query -> agents: Retrieval, Lookup, Composer  tools: ['rag_lookup',
                                                                  'check_appointment_status']

Telemetry is disabled. CREWAI_DISABLE_TELEMETRY=true and OTEL_SDK_DISABLED=true, set in config.py before any third-party import. Two further switches were needed and are documented in Findings.

Task 8 — session memory

InMemoryChatMessageHistory + RunnableWithMessageHistory, keyed by session id, in-process only.

Memory does real work rather than merely accumulating turns. An elliptical follow-up is rewritten against history, so the two required transcripts differ behaviourally:

with history   : "And does that apply to telemedicine?"
                 -> resolved: "How long before my appointment can I cancel
                               — follow-up: And does that apply to telemedicine?"
without history: -> "I'm not sure what that refers to, since this is the start
                     of our conversation. Could you ask the full question?"

The LangChainDeprecationWarning pointing at LangGraph's persistence layer is expected and left visible in the transcript.

Task 9 — structured output

SupportResponse uses cross-field validators, not just type annotations. A response claiming grounded=True with no sources, or refused=True while also grounded, or escalate=True with no appointment, is rejected. Thirteen negative cases are exercised in tests/test_schemas_guardrails.py.

Task 10 — guardrails

| Guardrail | Fires on | Evidence | |---|---|---| | PII masking (input) | Indian contact numbers, five formats | [PHONE_REDACTED] in transcript and log | | Prompt injection (input) | six named rule classes | rule names logged, never the payload | | Groundedness (output) | retrieval below threshold, or answer drifting past its evidence | refusal with a stated reason |

Masking runs before injection detection: if detection blocked first, an unmasked string would still be sitting in the object the logger later reads.

Scope is stated honestly. Only the fixed-format contact number is masked. Patient name, condition and insurance identifier have no reliable format and are out of scope for a keyless masker; every such value in this repository is fabricated.

The suite includes false-positive tests. An early version refused "Does the clinic act as a referral centre for cardiology?" as an injection attempt — a guardrail that blocks legitimate questions is worse than none.


Part 3 — Evaluation, observability, deployment

Task 11 — endpoints

| Endpoint | Purpose | |---|---| | POST /ask | question in, SupportResponse out | | POST /add-document | index a document into both collections and invalidate the cache | | GET /health | liveness | | WS /ws/chat | multi-turn chat, surviving mid-conversation disconnect |

Disconnect survival is proved by a third client connecting afterwards and completing an exchange, plus /health still answering. A socket closing raises nothing on its own, so asserting "the client disconnected" would prove nothing.

Task 12 — structured logging

One JSON-Lines record per request:

{"ts":"...","trace_id":"e55fa0b774724247","event":"ask","endpoint":"/ask",
 "latency_ms":0.098,"status":200,"session_id":"s-log","query":"...",
 "grounded":true,"sources":["cancellation_window"],"cache_hit":false,
 "pii_masked":false,"pii_match_count":0,"injection_detected":false}

obs.py imports guardrails.mask_pii rather than reimplementing it — two maskers in two places is how a raw value eventually reaches disk. Records are re-scrubbed before serialisation as a backstop; a leak would set late_masked: true rather than pass silently. The test asserts that flag never fires on the normal path, proving masking happens upstream where it should.

Fields are truncated at 500 characters, so an oversized request cannot write its body to disk on the very path that exists to reject it.

Task 13 — evaluation

15 queries: one per required knowledge-base topic (12), one multi-intent edge case, two out-of-scope.

| Property | Average | |---|---| | Accuracy | 0.856 | | Grounding | 0.992 | | Completeness | 0.700 | | Safety | 1.000 |

Per-query scores are in transcripts/task13_live_evaluation.txt.

Why completeness sits at 0.700. Retrieval runs at k=3. An answer can only contain facts from the three chunks that came back, so where an expected figure lives in a fourth chunk of the right document it cannot appear — Q03 (900, 1800), Q04 (48 hours), Q11 (5 working days). In every one of those cases the correct document was retrieved and cited; the specific figure was one chunk out of reach. Raising k would lift completeness and lower precision. That trade-off has deliberately not been tuned after seeing the scores, because adjusting a parameter until an evaluation improves is how an evaluation stops measuring anything.

Why accuracy sits at 0.856. Accuracy is an F1 over cited documents against an expected set of one. Retrieving a second, topically adjacent document caps the score at 0.67 for that query even though the right document was found and the answer was correct — Q01, Q02, Q03, Q10. The metric is strict by design; citing extra sources is a real cost in a policy answer, because a patient cannot tell which document the answer came from.

On what "LLM-as-judge under MOCK_LLM" means here. judge_prompt.txt is the real rubric and is what a live model receives when MOCK_LLM=0. Under MOCK_LLM=1 there is no model to send it to, so DeterministicJudge applies the same rubric by rule — string and set operations over the answer, the retrieved context and a per-query expectation table. This is stated plainly rather than dressed up as model judgement. It has one advantage over a live judge: the scores are reproducible, so a pipeline regression shows as a moved number rather than sampling noise.

The judge is shown to discriminate rather than rubber-stamp. Each property is attacked separately in tests/test_judge.py, and the right dimension drops while the others hold.


Part 4 — Resilience and governance

Task 14 — Autogen review stage

RoundRobinGroupChat([policy_compliance_reviewer, final_editor], max_turns=2, custom_message_types=[StructuredMessage[ReviewVerdict]]), with output_content_type=ReviewVerdict on the editor.

Both required outcomes are demonstrated:

A. APPROVE  every sentence supported -> approved=True, draft sent byte-for-byte
B. REVISE   [OK        ] support=100%  "...charged at 50 percent of the fee."
            [COMMITMENT] support=29%   "The follow-up visit is completely free..."
            [COMMITMENT] support=0%    "Practo guarantees a full refund..."
            -> approved=False, both unsupported sentences removed

The reviewer is not told what to look for. It compares each sentence against the retrieved context and finds the faults itself. That is why the review stage uses a hand-written ChatCompletionClient rather than ReplayChatCompletionClient: a replayed script would mean the "finding" was a string typed by the test author.

approved semantics: True means the draft shipped unchanged; False means it was revised, with the corrected text in final_answer. Both outcomes are fit to send.

Both traps the brief warns about are demonstrated live. Omitting custom_message_types fails at run(), not at construction, with ValueError: Message type ... is not registered. And a run produces three messages — task, reviewer, editor — which is why MaxMessageTermination(3) rather than (2) is the equivalent bound.

Task 15 — governance

Risk level: HIGH.

The system processes patient medical data: appointment records containing specialty, consultation history, clinical follow-up flag and fee, plus free-text patient messages that routinely contain contact numbers and symptom descriptions. The scheme keys on the category of data handled, not on the sophistication of what is done with it, so the fact that this system only reads records and never diagnoses, prescribes or writes does not move it out of the medical-data category. Nor does the absence of clinical judgement remove clinical consequence: a wrong answer about a cancellation window costs a patient a fee, an escalation that fails to fire leaves an unresolved follow-up sitting for weeks, and an invented policy creates an expectation the clinic must then honour or refuse. Classifying it Medium on the grounds that it resembles customer support would be optimistic reasoning about the data rather than an assessment of it. The mitigations in governance.py reduce residual risk; they do not change the tier, and residual concerns are listed rather than omitted.

Least autonomy (application layer). Only the Lookup Agent may call check_appointment_status. The tool appears in exactly one agent's tool list, and the restriction is additionally enforced inside the tool: it carries the role of the agent it was constructed with and checks that role on every call. Wiring it to another agent therefore produces a tool that refuses rather than one that merely should not have been wired. Authorisation is evaluated before the dataset is touched, so a refused caller learns nothing — not even whether the identifier exists. An unattributed caller (caller_role=None) is refused too, because a caller that cannot be identified cannot be held to a policy. All four cases are demonstrated in transcripts/task15_governance.txt.

Cost cap (runtime layer). A per-request token budget, checked before retrieval, before the crew and before the review stage. An oversized request returns HTTP 413 — the request was understood and refused on policy, which is client-correctable, not a server fault — and the rejection is logged with its estimate and cap.

Task 16 — response caching

In-memory LRU keyed by normalised query text.

CACHE MISS  'What is the cancellation window?'     20.134 ms   calls 0 -> 1 (ran)
CACHE HIT   'What is the cancellation window?'      0.009 ms   calls 1 -> 1 (avoided)
CACHE HIT   '  what is the CANCELLATION window  '   0.005 ms   calls 1 -> 1 (avoided)

The call counter is the proof; timing alone could be a warm cache anywhere in the stack.

Normalisation is deliberately conservative — case, whitespace, edge punctuation, nothing else. No stemming or synonym folding: a false hit returns the wrong policy, reads perfectly plausibly, and is nearly impossible to spot in a transcript. POST /add-document calls invalidate_all, because new text changes what retrieval can find and every prior answer — including refusals — is potentially stale.


Findings

Bugs found during development that are worth knowing about, all of which fail silently.

ChromaDB defaults to L2, not cosine. Without metadata={"hnsw:space": "cosine"} the "similarities" are euclidean distances and every threshold downstream is calibrated against the wrong quantity.

CrewAI's ReAct template contains a phantom observation. The system prompt literally includes Observation: the result of the action, from the very first call. A parser scanning the conversation for "Observation:" matches the template before any tool has run and answers with placeholder text. mock_llm.latest_observation skips system messages and rejects the template phrase.

Name-substring tool dispatch misroutes. A tool called rag_lookup is caught by any test for "lookup". Dispatch is by declared argument schema: whichever tool declares record_id is the record fetcher. Both bugs are reproduced and fixed in transcripts/task07_mockllm_guards.txt.

CrewAI has two telemetry channels, not one. Beyond CREWAI_DISABLE_TELEMETRY, a separate tracing system prints a panel and writes a preference file on every kickoff. CREWAI_TRACING_ENABLED=false plus CREWAI_TESTING=true silences it. Verified against crewai 1.15.21: CREWAI_TESTING is read in exactly one place and changes nothing else.

The Composer was discarding everything. CrewAI passes upstream results under This is the context you're working with:. The mock wasn't reading it, so the Composer ignored both other agents and emitted its no-information fallback.

Agent framing polluted the retrieval query. The mock passed the whole task description to rag_lookup, embedding search, knowledge base, report alongside the real question. Fixed with a PATIENT QUESTION: delimiter.

A reviewer that quotes a bad claim can launder it. The Autogen editor extracted its context from the whole conversation, which by then contained the reviewer's report — and that report quoted the unsupported sentences verbatim. Those quotations landed inside the extracted context block, the offending claims scored as fully supported, and the editor approved the draft the reviewer had just rejected. task_message_of reads only the original task message.

The groundedness gate discarded successful record lookups. Asking about an appointment fetched the record fine, but with no matching policy text the gate refused and threw the record away — telling a patient nothing was known about their own appointment. The gate now governs only the policy half and returns a partial answer.

A judge that demands literal phrases penalises paraphrase. An answer saying "free of charge" against a context saying "is free" was scored as an invented guarantee. Commitment detection now scores the sentence carrying the phrase.

The composer was deleting facts it had already retrieved. Answers were assembled from the best chunk per document, on the assumption that a second chunk of the same document would restate the first. For four-sentence policy documents that is wrong — chunks are disjoint, so each carries different facts. Asked whether a prescription could be repeated, the system returned the exclusions chunk and omitted the 90-day validity rule sitting one chunk away in the same document. Retrieval had found it; composition threw it away. No retrieval metric can see this, which is a fair argument for why an answer-level evaluation exists alongside Task 5.

Grounding was judged against half the evidence. On a merged answer — appointment record plus policy text — the record half was scored only against retrieved knowledge-base text, which never contained it. Overlap fell below the floor and the system took the partial-answer path, so it was penalised for correctly using the appointments dataset. The record is now part of the evidence set.

A memory rule fired on "the same doctor". The anaphor test for the same matched a self-contained question about follow-up pricing, which was then answered with "I'm not sure what that refers to". The rule now fires only when nothing follows the phrase.

A guardrail that blocks legitimate questions is worse than none. An early injection rule matched "Does the clinic act as a referral centre for cardiology?" as an attempted role reassignment. Every guardrail in this project now has explicit false-positive tests alongside its detection tests.

A diagnosis I got wrong, recorded because it is instructive. The completeness misses were initially attributed to fixed-size chunks cutting mid-sentence, and that reasoning was used to argue for sentence chunking. It was wrong: the cause was the composer bug above, present under either strategy, and switching chunkers made completeness slightly worse. The chunking decision still stands, but on the Task 5 numbers and the fragment problem — not on completeness. A plausible explanation that happens to support the conclusion you already favour is worth checking twice.

Three dependency conflicts resolve differently by install order. See the comments in requirements.txt.

One flaky test, isolated rather than ignored. ChromaDB's persistent backend occasionally fails a query issued shortly after a small upsert (Error creating hnsw segment reader: Nothing found on disk). It is intermittent and specific to collections holding a handful of vectors. The wiring smoke test now runs against an ephemeral store, since it tests the chunker and the metadata round trip rather than persistence; the graded corpus run still uses the persistent backend.


Repository layout

├── README.md                  this file
├── requirements.txt           pinned, pip check clean
├── judge_prompt.txt           Task 13 rubric, a deliverable in its own right
├── run_all.py                 regenerates every transcript
├── .env.example
├── src/practo/
│   ├── config.py              offline switches, paths, tunables
│   ├── dataset.py             Task 1
│   ├── kb/                    Task 2 — knowledge-base documents
│   ├── kb_loader.py           frontmatter parsing, abbreviation-aware splitting
│   ├── chunking.py            Task 3 — both strategies
│   ├── indexing.py            Task 3 — embeddings and two Chroma collections
│   ├── rag.py                 Task 4 — calibration and grounded generation
│   ├── chunk_eval.py          Task 5 — precision/recall comparison
│   ├── tools.py               Task 6 — lookup, escalation, tool ACL
│   ├── mock_llm.py            CrewAI BaseLLM subclass with both guards
│   ├── crew.py                Task 7 — three-agent crew
│   ├── memory.py              Task 8 — session memory
│   ├── schemas.py             Task 9 — structured output contract
│   ├── guardrails.py          Task 10 — PII, injection, groundedness
│   ├── api.py                 Task 11 — HTTP and WebSocket
│   ├── obs.py                 Task 12 — JSON-Lines logging
│   ├── judge.py               Task 13 — rubric and scorer
│   ├── review.py              Task 14 — Autogen review stage
│   ├── governance.py          Task 15 — four layers, risk, cost cap
│   ├── cache.py               Task 16 — LRU response cache
│   ├── engine.py              the pipeline joining all of the above
│   ├── corpus_report.py       Tasks 4 and 5 against the real corpus
│   └── live_eval.py           Task 13 against the live pipeline
├── tests/                     eleven suites
└── transcripts/               graded evidence

Producing the numbers

Every figure quoted above is reproduced by two commands, both of which need the knowledge base in src/practo/kb/:

Both exit non-zero and say why if something is wrong — overlapping calibration clusters, an in-scope query the knowledge base cannot answer, or an out-of-scope query that was answered instead of refused. Fix the cause rather than the number.


Data statement

Every patient name, medical condition, contact number and insurance identifier in this repository is fabricated. No real medical data is present anywhere in the dataset, the knowledge base, the tests or the transcripts.

Contract & API

Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.

MissingGITHUB REPOS

Contract coverage

Status

missing

Auth

None

Streaming

No

Data region

Unspecified

Protocol support

OpenClaw: self-declared

Requires: none

Forbidden: none

Guardrails

Operational confidence: low

No positive guardrails captured.
Invocation examples
curl -s "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust"

Reliability & Benchmarks

Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.

Missingruntime-metrics

Trust signals

Handshake

UNKNOWN

Confidence

unknown

Attempts 30d

unknown

Fallback rate

unknown

Runtime metrics

Observed P50

unknown

Observed P95

unknown

Rate limit

unknown

Estimated cost

unknown

Do not use if

Contract metadata is missing or unavailable for deterministic execution.
No benchmark suites or observed failure patterns are available.

Media & Demo

Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.

Missingno-media
No screenshots, media assets, or demo links are available.

Related Agents

Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.

Self-declaredprotocol-neighbors
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW
Machine Appendix

Contract JSON

{
  "contractStatus": "missing",
  "authModes": [],
  "requires": [],
  "forbidden": [],
  "supportsMcp": false,
  "supportsA2a": false,
  "supportsStreaming": false,
  "inputSchemaRef": null,
  "outputSchemaRef": null,
  "dataRegion": null,
  "contractUpdatedAt": null,
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Invocation Guide

{
  "preferredApi": {
    "snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/snapshot",
    "contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract",
    "trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust"
  },
  "curlExamples": [
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/snapshot\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust\""
  ],
  "jsonRequestTemplate": {
    "query": "summarize this repo",
    "constraints": {
      "maxLatencyMs": 2000,
      "protocolPreference": [
        "OPENCLEW"
      ]
    }
  },
  "jsonResponseTemplate": {
    "ok": true,
    "result": {
      "summary": "...",
      "confidence": 0.9
    },
    "meta": {
      "source": "GITHUB_REPOS",
      "generatedAt": "2026-10-09T14:55:05.693Z"
    }
  },
  "retryPolicy": {
    "maxAttempts": 3,
    "backoffMs": [
      500,
      1500,
      3500
    ],
    "retryableConditions": [
      "HTTP_429",
      "HTTP_503",
      "NETWORK_TIMEOUT"
    ]
  }
}

Trust JSON

{
  "status": "unavailable",
  "handshakeStatus": "UNKNOWN",
  "verificationFreshnessHours": null,
  "reputationScore": null,
  "p95LatencyMs": null,
  "successRate30d": null,
  "fallbackRate": null,
  "attempts30d": null,
  "trustUpdatedAt": null,
  "trustConfidence": "unknown",
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Capability Matrix

{
  "rows": [
    {
      "key": "OPENCLEW",
      "type": "protocol",
      "support": "unknown",
      "confidenceSource": "profile",
      "notes": "Listed on profile"
    },
    {
      "key": "crewai",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    },
    {
      "key": "multi-agent",
      "type": "capability",
      "support": "supported",
      "confidenceSource": "profile",
      "notes": "Declared in agent profile metadata"
    }
  ],
  "flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}

Facts JSON

[
  {
    "factKey": "vendor",
    "category": "vendor",
    "label": "Vendor",
    "value": "Vk409 Avenger",
    "href": "https://github.com/VK409-avenger/practo-support-agent",
    "sourceUrl": "https://github.com/VK409-avenger/practo-support-agent",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-10-09T12:50:37.933Z",
    "isPublic": true
  },
  {
    "factKey": "protocols",
    "category": "compatibility",
    "label": "Protocol compatibility",
    "value": "OpenClaw",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract",
    "sourceType": "contract",
    "confidence": "medium",
    "observedAt": "2026-10-09T12:50:37.933Z",
    "isPublic": true
  },
  {
    "factKey": "docs_crawl",
    "category": "integration",
    "label": "Crawlable docs",
    "value": "6 indexed pages on the official domain",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  },
  {
    "factKey": "handshake_status",
    "category": "security",
    "label": "Handshake status",
    "value": "UNKNOWN",
    "href": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust",
    "sourceType": "trust",
    "confidence": "medium",
    "observedAt": null,
    "isPublic": true
  }
]

Change Events JSON

[
  {
    "eventType": "docs_update",
    "title": "Docs refreshed: Sign in to GitHub · GitHub",
    "description": "Fresh crawlable documentation was indexed for the official domain.",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true
  }
]

Sponsored

Ads related to practo-support-agent and adjacent AI workflows.