activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
Crawler Summary
Healthcare support agent with calibrated RAG retrieval, a three-agent CrewAI crew, independent Autogen review, and an LLM-as-judge evaluation across accuracy, grounding, completeness and safety. Deterministic and offline. Practo Domain Support Agent **Track: Healthcare (Practo).** A clinic patient-support agent that answers policy questions from a knowledge base, looks up appointment records, remembers a conversation, guards against misuse, has its answers reviewed by an independent agent team, and operates under an explicit governance policy. Orchestrated with CrewAI, reviewed with Autogen, deployed behind FastAPI, evaluated end to e Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.
Freshness
Last checked 10/9/2026
Best For
practo-support-agent is best for crewai, multi-agent workflows where OpenClaw compatibility matters.
Not Ideal For
Contract metadata is missing or unavailable for deterministic execution.
Evidence Sources Checked
editorial-content, GITHUB REPOS, runtime-metrics, public facts pack
Healthcare support agent with calibrated RAG retrieval, a three-agent CrewAI crew, independent Autogen review, and an LLM-as-judge evaluation across accuracy, grounding, completeness and safety. Deterministic and offline. Practo Domain Support Agent **Track: Healthcare (Practo).** A clinic patient-support agent that answers policy questions from a knowledge base, looks up appointment records, remembers a conversation, guards against misuse, has its answers reviewed by an independent agent team, and operates under an explicit governance policy. Orchestrated with CrewAI, reviewed with Autogen, deployed behind FastAPI, evaluated end to e
Public facts
4
Change events
1
Artifacts
0
Freshness
Oct 9, 2026
Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.
Trust score
Unknown
Compatibility
OpenClaw
Freshness
Oct 9, 2026
Vendor
Vk409 Avenger
Artifacts
0
Benchmarks
0
Last release
Unpublished
Key links, install path, and a quick operational read before the deeper crawl record.
Summary
Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.
Setup snapshot
Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.
Vendor
Vk409 Avenger
Protocol compatibility
OpenClaw
Handshake status
UNKNOWN
Crawlable docs
6 indexed pages on the official domain
Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.
Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.
Extracted files
0
Examples
6
Snippets
0
Languages
python
bash
pip install -r requirements.txt cp .env.example .env python run_all.py
bash
python run_all.py --list # show the steps python run_all.py --only 4 14 # rerun selected steps
text
IN-SCOPE QUERIES top-1 sim
0.7567 [cancellation_window ] How long before my appointment can I
cancel without a charge?
0.6574 [consultation_fees ] What does a cardiology consultation cost?
0.6276 [telemedicine_eligibility] Can I see a doctor over video if I have
never visited before?
0.5212 [lab_turnaround ] How soon will my blood test results
come back?
0.5045 [follow_up_discount ] Do I pay less if I come back to the same
doctor next week?
OUT-OF-SCOPE QUERIES
0.2587 [prescription_refills ] What antibiotic should I take for a
sore throat?
0.1040 [data_privacy ] Who won the 2019 cricket world cup?
0.0745 [emergency_protocol ] How do I reset the wifi router in my flat?text
s permitted at no cost provided it is requested at least 6 hours before...
text
staleness(d) = clip((d − 7) / (30 − 7), 0, 1) score = 0.45 · follow_up_required + 0.55 · staleness(days_since_created)
text
policy query -> agents: Retrieval, Composer tools: ['rag_lookup']
appointment query -> agents: Retrieval, Lookup, Composer tools: ['rag_lookup',
'check_appointment_status']Full documentation captured from public sources, including the complete README when available.
Docs source
GITHUB REPOS
Editorial quality
ready
Healthcare support agent with calibrated RAG retrieval, a three-agent CrewAI crew, independent Autogen review, and an LLM-as-judge evaluation across accuracy, grounding, completeness and safety. Deterministic and offline. Practo Domain Support Agent **Track: Healthcare (Practo).** A clinic patient-support agent that answers policy questions from a knowledge base, looks up appointment records, remembers a conversation, guards against misuse, has its answers reviewed by an independent agent team, and operates under an explicit governance policy. Orchestrated with CrewAI, reviewed with Autogen, deployed behind FastAPI, evaluated end to e
Track: Healthcare (Practo).
A clinic patient-support agent that answers policy questions from a knowledge base, looks up appointment records, remembers a conversation, guards against misuse, has its answers reviewed by an independent agent team, and operates under an explicit governance policy. Orchestrated with CrewAI, reviewed with Autogen, deployed behind FastAPI, evaluated end to end.
Everything runs under MOCK_LLM=1 with zero API keys and zero LLM network calls.
pip install -r requirements.txt
cp .env.example .env
python run_all.py
run_all.py regenerates every transcript in transcripts/ in dependency order. Steps needing the knowledge base are skipped with a clear message if it is absent, so a partial clone still produces the component evidence.
python run_all.py --list # show the steps
python run_all.py --only 4 14 # rerun selected steps
The one-time download of sentence-transformers/all-MiniLM-L6-v2 (~90 MB) is the only network access this project makes. No language-model call ever leaves the machine.
Every acceptance criterion has a file. This table is the fastest route through the evidence.
| Task | What it shows | Transcript |
|---|---|---|
| 1 | Dataset generation, category/status counts, follow-up band | transcripts/task01_dataset.txt |
| 2 | Knowledge-base manifest, sentence-count validation | transcripts/task02_kb_manifest.txt |
| 3 | Both chunkers, both ChromaDB collections, sample retrieval | transcripts/task03_indexing.txt |
| 3 | Loader, chunker and Chroma wiring smoke test | transcripts/task03_pipeline_smoke.txt |
| 4, 5 | Threshold calibration, grounded generation, chunking comparison (real corpus) | transcripts/task04_05_calibration_and_chunking.txt |
| 4, 7 | RAG core and three-agent crew, both tools invoked | transcripts/task04_07_rag_crew.txt |
| 5 | Chunking metric tests against hand-computed values | transcripts/task05_chunking_comparison.txt |
| 6 | Escalation formula, distribution, derived threshold, ACL | transcripts/task06_lookup_escalation.txt |
| 6 | Lookup tool tests including non-degeneracy | transcripts/task06_tests.txt |
| 7 | MockLLM guards against both silent failures | transcripts/task07_mockllm_guards.txt |
| 8, 16 | Multi-turn memory, fresh-session reset, cache hit | transcripts/task08_16_memory_cache.txt |
| 9, 10 | Structured output validation, three guardrails firing | transcripts/task09_10_schemas_guardrails.txt |
| 11, 12 | HTTP + WebSocket endpoints, disconnect survival, JSONL logs | transcripts/task11_12_api_logging.txt |
| 13 | Judge harness and discrimination tests | transcripts/task13_evaluation.txt |
| 13 | 15-query evaluation against the live pipeline (real corpus) | transcripts/task13_live_evaluation.txt |
| 14 | Autogen review approving and revising, with structured verdicts | transcripts/task14_autogen_review.txt |
| 15 | Four-layer governance, least autonomy, cost cap | transcripts/task15_governance.txt |
| — | Full pipeline integration | transcripts/task_integration_engine.txt |
The judge prompt is committed separately at judge_prompt.txt.
These are the values needed to reproduce the dataset exactly. python -m practo.dataset regenerates it.
| Choice | Value |
|---|---|
| Seed | 10 |
| Records | 60 |
| Category weights | General Medicine 0.26, Cardiology 0.15, Dermatology 0.15, Pediatrics 0.15, Orthopedics 0.15, ENT 0.14 |
| Status weights | Scheduled 0.30, Completed 0.28, Cancelled 0.12, Rescheduled 0.12, Pending-Confirmation 0.10, No-Show 0.08 |
| Added category | ENT |
| Added status | Pending-Confirmation |
| Fee bands (INR) | General Medicine 400–800, Pediatrics 500–1000, ENT 600–1100, Dermatology 700–1400, Orthopedics 800–1600, Cardiology 900–1800, quoted in ₹50 steps |
| days_since_created | triangular(0, 31, mode=6), truncated at 30 |
| follow_up_required | Conditional on status: Completed 0.45, Scheduled 0.10, Rescheduled 0.10, all others 0.02 |
Fee reasoning. The bands mirror private multi-specialty OPD pricing in Indian tier-1 cities, where a general consult sits at the floor and diagnostic- or procedure-heavy specialties command roughly twice that.
Observed result at seed 10: every supplied category has ≥3 records (minimum 6, Cardiology), every supplied status appears (minimum 2, No-Show), and follow_up_required is 20.0% — mid-band, on the first draw. No reseeding was needed and no record was hand-edited.
Why follow_up_required is conditional. A follow-up flag only means something once a consultation has happened. A flat coin-flip would put follow-ups on cancelled and no-show appointments, and would make the Task 6 escalation score degenerate — see below.
Why the day distribution is skewed. A real appointment queue is dominated by recent bookings with a thinning tail of unresolved older entries. A uniform spread would make the staleness term in Task 6 too easy to trip.
Measured over the real corpus with sentence-transformers/all-MiniLM-L6-v2, cosine similarity reported as 1 − chroma_distance.
Collections are created with
metadata={"hnsw:space": "cosine"}. ChromaDB defaults to squared L2. Without that setting the numbers returned are euclidean distances, nothing raises, and the calibration below would be measuring the wrong quantity.
IN-SCOPE QUERIES top-1 sim
0.7567 [cancellation_window ] How long before my appointment can I
cancel without a charge?
0.6574 [consultation_fees ] What does a cardiology consultation cost?
0.6276 [telemedicine_eligibility] Can I see a doctor over video if I have
never visited before?
0.5212 [lab_turnaround ] How soon will my blood test results
come back?
0.5045 [follow_up_discount ] Do I pay less if I come back to the same
doctor next week?
OUT-OF-SCOPE QUERIES
0.2587 [prescription_refills ] What antibiotic should I take for a
sore throat?
0.1040 [data_privacy ] Who won the 2019 cricket world cup?
0.0745 [emergency_protocol ] How do I reset the wifi router in my flat?
| | Value | |---|---| | Lowest in-scope similarity | 0.5045 | | Highest out-of-scope similarity | 0.2587 | | Observed gap | 0.2458 | | Chosen threshold | 0.3816 (midpoint of the measured gap) |
The gap is wide and the clusters do not overlap, so the midpoint is a defensible cut rather than a compromise. Note where the highest out-of-scope score sits: "What antibiotic should I take for a sore throat?" reaches 0.2587 against the refill policy, well above the other two. It is lexically close to a document the knowledge base does have. Retrieval similarity cannot separate "what is the refill policy" from "what should I take" — that difference is about authority, not topic — which is why the clinical-advice boundary is enforced before retrieval rather than left to the threshold. See Task 15.
Demonstrated on 5 in-scope queries, all answered, plus 1 deliberately out-of-scope query that correctly triggered the fallback at 0.1040.
The threshold is not a preset. guardrails.check_groundedness raises ThresholdNotCalibrated when passed None, and RagCore refuses to construct without one, so the system physically cannot run on an unmeasured value.
| Strategy | Chunks | Precision | Recall | F1 |
|---|---|---|---|---|
| fixed_overlap (250 chars, 50 overlap) | 36 | 0.444 | 0.917 | 0.583 |
| sentence (2 sentences, non-overlapping) | 28 | 0.417 | 0.917 | 0.556 |
Generated recommendation, from those measured values:
Across 6 queries at k=3,
fixed_overlapscored precision 0.444 and recall 0.917 (F1 0.583), againstsentenceat precision 0.417 and recall 0.917 (F1 0.556). Recall is identical at 0.917, so the difference is precision alone, wherefixed_overlapleads by +0.028; note that only 1 of 6 queries separates the two, so this is a narrow result rather than a decisive one.
sentence is deployed, and the measurement narrowly favoured the other one.
That disagreement is stated rather than hidden, because the reasoning matters
more than the number.
The gap is 0.028 in F1 and rests entirely on Q5. Five of the six queries score identically under both strategies. On a six-query set that is one query's worth of movement, which is noise, not evidence.
The tiebreaker is what the two strategies put in front of a patient. Fixed-size
windows cut wherever 250 characters lands, including mid-word. The Task 3 sample
output shows it directly — the second-ranked fixed_overlap chunk for a
cancellation query begins:
s permitted at no cost provided it is requested at least 6 hours before...
Sentence chunks cannot do that, because a chunk boundary is a sentence boundary. Given two strategies that are statistically indistinguishable on retrieval quality, the one that never emits a fragment is the right one to serve.
config.DEPLOYED_STRATEGY names the served collection in one place, and the
recommendation reads that value rather than asserting it, so the two cannot
silently drift apart again.
Scoring is at the parent-document level after deduplication. Relevance belongs to documents, not chunks: scoring chunks directly would reward whichever strategy fragments documents more, which is the variable under test.
staleness(d) = clip((d − 7) / (30 − 7), 0, 1)
score = 0.45 · follow_up_required + 0.55 · staleness(days_since_created)
| | Value |
|---|---|
| Threshold | 0.4748 — the 80th percentile of the observed score distribution (statistics.quantiles, method='inclusive', n=60) |
| Escalates | 12 of 60 records (20.0%) |
| …carrying a follow-up flag | 8 |
| …on age alone, no flag | 4 |
| Flagged but held back | 4 |
| Score range, flagged records | [0.4500, 1.0000] |
| Score range, unflagged records | [0.0000, 0.5500] |
Why not the obvious weighting. With 0.6 · flag + 0.4 · linear age, every flagged record scores ≥ 0.60 and every unflagged record ≤ 0.40. The ranges cannot overlap, so any threshold between them reduces to the boolean — exactly what the brief rules out. Shifting weight onto staleness and delaying its onset to day 7 makes the ranges overlap, which is asserted directly in tests/test_tools_escalation.py.
Two cases that show it is a real score. APT-0003 escalates with no follow-up flag at all — a Pending-Confirmation booking untouched for 29 days. APT-0016 carries the flag and does not escalate — it was flagged 8 days ago, so staleness has barely started.
| Agent | Tools | Runs when |
|---|---|---|
| Retrieval Agent | rag_lookup | always |
| Lookup Agent | check_appointment_status | only when the query references a record |
| Response Composer | none | always |
Tasks are assembled per query. A policy question has no appointment to look up, so handing the Lookup Agent a task with no record would either fabricate an identifier or waste a turn.
policy query -> agents: Retrieval, Composer tools: ['rag_lookup']
appointment query -> agents: Retrieval, Lookup, Composer tools: ['rag_lookup',
'check_appointment_status']
Telemetry is disabled. CREWAI_DISABLE_TELEMETRY=true and OTEL_SDK_DISABLED=true, set in config.py before any third-party import. Two further switches were needed and are documented in Findings.
InMemoryChatMessageHistory + RunnableWithMessageHistory, keyed by session id, in-process only.
Memory does real work rather than merely accumulating turns. An elliptical follow-up is rewritten against history, so the two required transcripts differ behaviourally:
with history : "And does that apply to telemedicine?"
-> resolved: "How long before my appointment can I cancel
— follow-up: And does that apply to telemedicine?"
without history: -> "I'm not sure what that refers to, since this is the start
of our conversation. Could you ask the full question?"
The LangChainDeprecationWarning pointing at LangGraph's persistence layer is expected and left visible in the transcript.
SupportResponse uses cross-field validators, not just type annotations. A response claiming grounded=True with no sources, or refused=True while also grounded, or escalate=True with no appointment, is rejected. Thirteen negative cases are exercised in tests/test_schemas_guardrails.py.
| Guardrail | Fires on | Evidence |
|---|---|---|
| PII masking (input) | Indian contact numbers, five formats | [PHONE_REDACTED] in transcript and log |
| Prompt injection (input) | six named rule classes | rule names logged, never the payload |
| Groundedness (output) | retrieval below threshold, or answer drifting past its evidence | refusal with a stated reason |
Masking runs before injection detection: if detection blocked first, an unmasked string would still be sitting in the object the logger later reads.
Scope is stated honestly. Only the fixed-format contact number is masked. Patient name, condition and insurance identifier have no reliable format and are out of scope for a keyless masker; every such value in this repository is fabricated.
The suite includes false-positive tests. An early version refused "Does the clinic act as a referral centre for cardiology?" as an injection attempt — a guardrail that blocks legitimate questions is worse than none.
| Endpoint | Purpose |
|---|---|
| POST /ask | question in, SupportResponse out |
| POST /add-document | index a document into both collections and invalidate the cache |
| GET /health | liveness |
| WS /ws/chat | multi-turn chat, surviving mid-conversation disconnect |
Disconnect survival is proved by a third client connecting afterwards and completing an exchange, plus /health still answering. A socket closing raises nothing on its own, so asserting "the client disconnected" would prove nothing.
One JSON-Lines record per request:
{"ts":"...","trace_id":"e55fa0b774724247","event":"ask","endpoint":"/ask",
"latency_ms":0.098,"status":200,"session_id":"s-log","query":"...",
"grounded":true,"sources":["cancellation_window"],"cache_hit":false,
"pii_masked":false,"pii_match_count":0,"injection_detected":false}
obs.py imports guardrails.mask_pii rather than reimplementing it — two maskers in two places is how a raw value eventually reaches disk. Records are re-scrubbed before serialisation as a backstop; a leak would set late_masked: true rather than pass silently. The test asserts that flag never fires on the normal path, proving masking happens upstream where it should.
Fields are truncated at 500 characters, so an oversized request cannot write its body to disk on the very path that exists to reject it.
15 queries: one per required knowledge-base topic (12), one multi-intent edge case, two out-of-scope.
| Property | Average | |---|---| | Accuracy | 0.856 | | Grounding | 0.992 | | Completeness | 0.700 | | Safety | 1.000 |
Per-query scores are in transcripts/task13_live_evaluation.txt.
Why completeness sits at 0.700. Retrieval runs at k=3. An answer can only
contain facts from the three chunks that came back, so where an expected figure
lives in a fourth chunk of the right document it cannot appear — Q03 (900,
1800), Q04 (48 hours), Q11 (5 working days). In every one of those cases
the correct document was retrieved and cited; the specific figure was one chunk
out of reach. Raising k would lift completeness and lower precision. That
trade-off has deliberately not been tuned after seeing the scores, because
adjusting a parameter until an evaluation improves is how an evaluation stops
measuring anything.
Why accuracy sits at 0.856. Accuracy is an F1 over cited documents against an expected set of one. Retrieving a second, topically adjacent document caps the score at 0.67 for that query even though the right document was found and the answer was correct — Q01, Q02, Q03, Q10. The metric is strict by design; citing extra sources is a real cost in a policy answer, because a patient cannot tell which document the answer came from.
On what "LLM-as-judge under MOCK_LLM" means here. judge_prompt.txt is the real rubric and is what a live model receives when MOCK_LLM=0. Under MOCK_LLM=1 there is no model to send it to, so DeterministicJudge applies the same rubric by rule — string and set operations over the answer, the retrieved context and a per-query expectation table. This is stated plainly rather than dressed up as model judgement. It has one advantage over a live judge: the scores are reproducible, so a pipeline regression shows as a moved number rather than sampling noise.
The judge is shown to discriminate rather than rubber-stamp. Each property is attacked separately in tests/test_judge.py, and the right dimension drops while the others hold.
RoundRobinGroupChat([policy_compliance_reviewer, final_editor], max_turns=2, custom_message_types=[StructuredMessage[ReviewVerdict]]), with output_content_type=ReviewVerdict on the editor.
Both required outcomes are demonstrated:
A. APPROVE every sentence supported -> approved=True, draft sent byte-for-byte
B. REVISE [OK ] support=100% "...charged at 50 percent of the fee."
[COMMITMENT] support=29% "The follow-up visit is completely free..."
[COMMITMENT] support=0% "Practo guarantees a full refund..."
-> approved=False, both unsupported sentences removed
The reviewer is not told what to look for. It compares each sentence against the retrieved context and finds the faults itself. That is why the review stage uses a hand-written ChatCompletionClient rather than ReplayChatCompletionClient: a replayed script would mean the "finding" was a string typed by the test author.
approved semantics: True means the draft shipped unchanged; False means it was revised, with the corrected text in final_answer. Both outcomes are fit to send.
Both traps the brief warns about are demonstrated live. Omitting custom_message_types fails at run(), not at construction, with ValueError: Message type ... is not registered. And a run produces three messages — task, reviewer, editor — which is why MaxMessageTermination(3) rather than (2) is the equivalent bound.
Risk level: HIGH.
The system processes patient medical data: appointment records containing specialty, consultation history, clinical follow-up flag and fee, plus free-text patient messages that routinely contain contact numbers and symptom descriptions. The scheme keys on the category of data handled, not on the sophistication of what is done with it, so the fact that this system only reads records and never diagnoses, prescribes or writes does not move it out of the medical-data category. Nor does the absence of clinical judgement remove clinical consequence: a wrong answer about a cancellation window costs a patient a fee, an escalation that fails to fire leaves an unresolved follow-up sitting for weeks, and an invented policy creates an expectation the clinic must then honour or refuse. Classifying it Medium on the grounds that it resembles customer support would be optimistic reasoning about the data rather than an assessment of it. The mitigations in governance.py reduce residual risk; they do not change the tier, and residual concerns are listed rather than omitted.
Least autonomy (application layer). Only the Lookup Agent may call check_appointment_status. The tool appears in exactly one agent's tool list, and the restriction is additionally enforced inside the tool: it carries the role of the agent it was constructed with and checks that role on every call. Wiring it to another agent therefore produces a tool that refuses rather than one that merely should not have been wired. Authorisation is evaluated before the dataset is touched, so a refused caller learns nothing — not even whether the identifier exists. An unattributed caller (caller_role=None) is refused too, because a caller that cannot be identified cannot be held to a policy. All four cases are demonstrated in transcripts/task15_governance.txt.
Cost cap (runtime layer). A per-request token budget, checked before retrieval, before the crew and before the review stage. An oversized request returns HTTP 413 — the request was understood and refused on policy, which is client-correctable, not a server fault — and the rejection is logged with its estimate and cap.
In-memory LRU keyed by normalised query text.
CACHE MISS 'What is the cancellation window?' 20.134 ms calls 0 -> 1 (ran)
CACHE HIT 'What is the cancellation window?' 0.009 ms calls 1 -> 1 (avoided)
CACHE HIT ' what is the CANCELLATION window ' 0.005 ms calls 1 -> 1 (avoided)
The call counter is the proof; timing alone could be a warm cache anywhere in the stack.
Normalisation is deliberately conservative — case, whitespace, edge punctuation, nothing else. No stemming or synonym folding: a false hit returns the wrong policy, reads perfectly plausibly, and is nearly impossible to spot in a transcript. POST /add-document calls invalidate_all, because new text changes what retrieval can find and every prior answer — including refusals — is potentially stale.
Bugs found during development that are worth knowing about, all of which fail silently.
ChromaDB defaults to L2, not cosine. Without metadata={"hnsw:space": "cosine"} the "similarities" are euclidean distances and every threshold downstream is calibrated against the wrong quantity.
CrewAI's ReAct template contains a phantom observation. The system prompt literally includes Observation: the result of the action, from the very first call. A parser scanning the conversation for "Observation:" matches the template before any tool has run and answers with placeholder text. mock_llm.latest_observation skips system messages and rejects the template phrase.
Name-substring tool dispatch misroutes. A tool called rag_lookup is caught by any test for "lookup". Dispatch is by declared argument schema: whichever tool declares record_id is the record fetcher. Both bugs are reproduced and fixed in transcripts/task07_mockllm_guards.txt.
CrewAI has two telemetry channels, not one. Beyond CREWAI_DISABLE_TELEMETRY, a separate tracing system prints a panel and writes a preference file on every kickoff. CREWAI_TRACING_ENABLED=false plus CREWAI_TESTING=true silences it. Verified against crewai 1.15.21: CREWAI_TESTING is read in exactly one place and changes nothing else.
The Composer was discarding everything. CrewAI passes upstream results under This is the context you're working with:. The mock wasn't reading it, so the Composer ignored both other agents and emitted its no-information fallback.
Agent framing polluted the retrieval query. The mock passed the whole task description to rag_lookup, embedding search, knowledge base, report alongside the real question. Fixed with a PATIENT QUESTION: delimiter.
A reviewer that quotes a bad claim can launder it. The Autogen editor extracted its context from the whole conversation, which by then contained the reviewer's report — and that report quoted the unsupported sentences verbatim. Those quotations landed inside the extracted context block, the offending claims scored as fully supported, and the editor approved the draft the reviewer had just rejected. task_message_of reads only the original task message.
The groundedness gate discarded successful record lookups. Asking about an appointment fetched the record fine, but with no matching policy text the gate refused and threw the record away — telling a patient nothing was known about their own appointment. The gate now governs only the policy half and returns a partial answer.
A judge that demands literal phrases penalises paraphrase. An answer saying "free of charge" against a context saying "is free" was scored as an invented guarantee. Commitment detection now scores the sentence carrying the phrase.
The composer was deleting facts it had already retrieved. Answers were assembled from the best chunk per document, on the assumption that a second chunk of the same document would restate the first. For four-sentence policy documents that is wrong — chunks are disjoint, so each carries different facts. Asked whether a prescription could be repeated, the system returned the exclusions chunk and omitted the 90-day validity rule sitting one chunk away in the same document. Retrieval had found it; composition threw it away. No retrieval metric can see this, which is a fair argument for why an answer-level evaluation exists alongside Task 5.
Grounding was judged against half the evidence. On a merged answer — appointment record plus policy text — the record half was scored only against retrieved knowledge-base text, which never contained it. Overlap fell below the floor and the system took the partial-answer path, so it was penalised for correctly using the appointments dataset. The record is now part of the evidence set.
A memory rule fired on "the same doctor". The anaphor test for the same matched a self-contained question about follow-up pricing, which was then answered with "I'm not sure what that refers to". The rule now fires only when nothing follows the phrase.
A guardrail that blocks legitimate questions is worse than none. An early injection rule matched "Does the clinic act as a referral centre for cardiology?" as an attempted role reassignment. Every guardrail in this project now has explicit false-positive tests alongside its detection tests.
A diagnosis I got wrong, recorded because it is instructive. The completeness misses were initially attributed to fixed-size chunks cutting mid-sentence, and that reasoning was used to argue for sentence chunking. It was wrong: the cause was the composer bug above, present under either strategy, and switching chunkers made completeness slightly worse. The chunking decision still stands, but on the Task 5 numbers and the fragment problem — not on completeness. A plausible explanation that happens to support the conclusion you already favour is worth checking twice.
Three dependency conflicts resolve differently by install order. See the comments in requirements.txt.
One flaky test, isolated rather than ignored. ChromaDB's persistent backend occasionally fails a query issued shortly after a small upsert (Error creating hnsw segment reader: Nothing found on disk). It is intermittent and specific to collections holding a handful of vectors. The wiring smoke test now runs against an ephemeral store, since it tests the chunker and the metadata round trip rather than persistence; the graded corpus run still uses the persistent backend.
├── README.md this file
├── requirements.txt pinned, pip check clean
├── judge_prompt.txt Task 13 rubric, a deliverable in its own right
├── run_all.py regenerates every transcript
├── .env.example
├── src/practo/
│ ├── config.py offline switches, paths, tunables
│ ├── dataset.py Task 1
│ ├── kb/ Task 2 — knowledge-base documents
│ ├── kb_loader.py frontmatter parsing, abbreviation-aware splitting
│ ├── chunking.py Task 3 — both strategies
│ ├── indexing.py Task 3 — embeddings and two Chroma collections
│ ├── rag.py Task 4 — calibration and grounded generation
│ ├── chunk_eval.py Task 5 — precision/recall comparison
│ ├── tools.py Task 6 — lookup, escalation, tool ACL
│ ├── mock_llm.py CrewAI BaseLLM subclass with both guards
│ ├── crew.py Task 7 — three-agent crew
│ ├── memory.py Task 8 — session memory
│ ├── schemas.py Task 9 — structured output contract
│ ├── guardrails.py Task 10 — PII, injection, groundedness
│ ├── api.py Task 11 — HTTP and WebSocket
│ ├── obs.py Task 12 — JSON-Lines logging
│ ├── judge.py Task 13 — rubric and scorer
│ ├── review.py Task 14 — Autogen review stage
│ ├── governance.py Task 15 — four layers, risk, cost cap
│ ├── cache.py Task 16 — LRU response cache
│ ├── engine.py the pipeline joining all of the above
│ ├── corpus_report.py Tasks 4 and 5 against the real corpus
│ └── live_eval.py Task 13 against the live pipeline
├── tests/ eleven suites
└── transcripts/ graded evidence
Every figure quoted above is reproduced by two commands, both of which need the
knowledge base in src/practo/kb/:
Both exit non-zero and say why if something is wrong — overlapping calibration clusters, an in-scope query the knowledge base cannot answer, or an out-of-scope query that was answered instead of refused. Fix the cause rather than the number.
Every patient name, medical condition, contact number and insurance identifier in this repository is fabricated. No real medical data is present anywhere in the dataset, the knowledge base, the tests or the transcripts.
Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.
Contract coverage
Status
missing
Auth
None
Streaming
No
Data region
Unspecified
Protocol support
Requires: none
Forbidden: none
Guardrails
Operational confidence: low
curl -s "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust"
Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.
Trust signals
Handshake
UNKNOWN
Confidence
unknown
Attempts 30d
unknown
Fallback rate
unknown
Runtime metrics
Observed P50
unknown
Observed P95
unknown
Rate limit
unknown
Estimated cost
unknown
Do not use if
Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.
Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
The Frontend for Agents & Generative UI. React + Angular
Contract JSON
{
"contractStatus": "missing",
"authModes": [],
"requires": [],
"forbidden": [],
"supportsMcp": false,
"supportsA2a": false,
"supportsStreaming": false,
"inputSchemaRef": null,
"outputSchemaRef": null,
"dataRegion": null,
"contractUpdatedAt": null,
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Invocation Guide
{
"preferredApi": {
"snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/snapshot",
"contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract",
"trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust"
},
"curlExamples": [
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/snapshot\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust\""
],
"jsonRequestTemplate": {
"query": "summarize this repo",
"constraints": {
"maxLatencyMs": 2000,
"protocolPreference": [
"OPENCLEW"
]
}
},
"jsonResponseTemplate": {
"ok": true,
"result": {
"summary": "...",
"confidence": 0.9
},
"meta": {
"source": "GITHUB_REPOS",
"generatedAt": "2026-10-09T14:55:05.693Z"
}
},
"retryPolicy": {
"maxAttempts": 3,
"backoffMs": [
500,
1500,
3500
],
"retryableConditions": [
"HTTP_429",
"HTTP_503",
"NETWORK_TIMEOUT"
]
}
}Trust JSON
{
"status": "unavailable",
"handshakeStatus": "UNKNOWN",
"verificationFreshnessHours": null,
"reputationScore": null,
"p95LatencyMs": null,
"successRate30d": null,
"fallbackRate": null,
"attempts30d": null,
"trustUpdatedAt": null,
"trustConfidence": "unknown",
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Capability Matrix
{
"rows": [
{
"key": "OPENCLEW",
"type": "protocol",
"support": "unknown",
"confidenceSource": "profile",
"notes": "Listed on profile"
},
{
"key": "crewai",
"type": "capability",
"support": "supported",
"confidenceSource": "profile",
"notes": "Declared in agent profile metadata"
},
{
"key": "multi-agent",
"type": "capability",
"support": "supported",
"confidenceSource": "profile",
"notes": "Declared in agent profile metadata"
}
],
"flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}Facts JSON
[
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Vk409 Avenger",
"href": "https://github.com/VK409-avenger/practo-support-agent",
"sourceUrl": "https://github.com/VK409-avenger/practo-support-agent",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T12:50:37.933Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T12:50:37.933Z",
"isPublic": true
},
{
"factKey": "docs_crawl",
"category": "integration",
"label": "Crawlable docs",
"value": "6 indexed pages on the official domain",
"href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceType": "search_document",
"confidence": "medium",
"observedAt": "2026-04-15T05:03:46.393Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-vk409-avenger-practo-support-agent/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
]Change Events JSON
[
{
"eventType": "docs_update",
"title": "Docs refreshed: Sign in to GitHub · GitHub",
"description": "Fresh crawlable documentation was indexed for the official domain.",
"href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceType": "search_document",
"confidence": "medium",
"observedAt": "2026-04-15T05:03:46.393Z",
"isPublic": true
}
]Sponsored
Ads related to practo-support-agent and adjacent AI workflows.