{"id":"bd1cc4a8-03b9-481c-a019-86f2ee57cb65","entityType":"agent","slug":"clawhub-docsor1212-cite-holmes","name":"Cite Holmes — Deep Research × Hallucination-Free Citations","canonicalUrl":"https://www.xpersona.co/agent/clawhub-docsor1212-cite-holmes","canonicalPath":"/agent/clawhub-docsor1212-cite-holmes","generatedAt":"2026-10-10T10:43:06.182Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T08:15:53.892Z","emptyReason":null},"description":"Deep research that interrogates its own sources (Verified Deep Research): calibrates scope first (3-5 sharp questions), then searches iteratively and machine-verifies every citation. Also a standalone citation checker: paste any reference list and it runs full citation verification (it will verify citations before you cite) against official registries — hallucinated references, fabricated DOIs, fake PMIDs, arXiv IDs, stitched fakes, retracted papers; a fact check for your bibliography. Five verdicts; unverified references never masquerade as real (AI hallucination detection). Medical mode (Cochrane/BMJ/ChiCTR/NMPA/CDC/NICE presets, PMID check) and bibliography export (BibTeX/GB·T 7714-2025/RIS/CSV + JSON workpaper) built in. Trigger matching is semantic, not exact — mis-triggers are harmless; state your real intent to avoid them. Skill: Cite Holmes — Deep Research × Hallucination-Free Citations Owner: docsor1212 Summary: Deep research that interrogates its own sources (Verified Deep Research): calibrates scope first (3-5 sharp questions), then searches iteratively and machine-verifies every citation. Also a standalone citation checker: paste any reference list and it runs full citation verification (it will verify citations before you cite) a","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.6K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s17dagtwyk21qs6vpz98bzcrh1853t29:cite-holmes","sourceUrl":"https://clawhub.ai/docsor1212/cite-holmes","homepage":"https://clawhub.ai/docsor1212/skills/cite-holmes","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/docsor1212/cite-holmes","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/docsor1212/skills/cite-holmes","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":64,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Deep research that interrogates its own sources (Verified Deep Research): calibrates scope first (3-5 sharp questions), then searches iteratively and machine-ve"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T08:15:53.892Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T08:15:53.892Z","emptyReason":null},"stars":null,"forks":null,"downloads":1564,"packageName":null,"latestVersion":"3.11.0","tractionLabel":"1.6K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T08:15:53.892Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T08:15:53.892Z","lastCrawledAt":"2026-10-10T08:15:53.892Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T08:15:53.892Z","lastVerifiedAt":null,"highlights":[{"version":"3.11.0","createdAt":"2026-10-09T17:22:29.713Z","changelog":"v3.11.0: the bibliography-profiling-and-version-hints release. List-level fabrication fingerprint scan (pure-local, zero network): serial identifier clusters, year over-concentration or future years, single-source concentration, bare-entry share - the statistical fingerprints of batch-fabricated reference lists; flags are human-review leads, never verdicts. Preprint-to-published hints: an arXiv-only reference that has since been registered as a journal article gets annotated with the formal-version DOI (via Semantic Scholar, silent on rate limits). Per-reference completeness score: how many export fields the registry CSL can back-fill, with the missing-field list. 379 tests green, zero new dependencies.","fileCount":29,"zipByteSize":207208},{"version":"3.10.0","createdAt":"2026-10-08T18:43:14.852Z","changelog":"v3.10.0: the cross-lingual-and-clone-detection release. Cross-lingual title bridging: a Chinese original title claimed against an English registry record (the normal case for Chinese-journal DOIs) is resolved via a three-stage bridge - the bilingual original-title registry field, then the DOI landing page, then OpenAlex; a hit upgrades to verified, a miss keeps partial, and a language difference is never treated as fabrication evidence. Clone-pair detection flags same-title/different-DOI and same-DOI/different-title pairs inside one bibliography (the most common fabrication shape in generated text) as paired leads - annotation only, never a verdict change. --doctor (add --net) gives a zero-outbound environment self-check with per-registry connectivity probes, PASS/WARN/FAIL verdicts and CI-friendly exit codes. The SKILL VERIFY section is restructured command-first. 355 tests green, zero new dependencies.","fileCount":29,"zipByteSize":197555},{"version":"3.8.0","createdAt":"2026-10-07T05:30:53.666Z","changelog":"v3.8.0: the documentation & examples completeness release. End-to-end worked example shipped (input refs + document excerpt -> real generated report and all four export formats, under examples/end-to-end/); references/ index and a consolidated anti-patterns list (incl. trigger-word mis-fires and the cross-language mixed-citation boundary); context-check results now render in the md/html reports (previously console/JSON only); author-surname anchor credit (a citing sentence naming a registry author counts as strong binding evidence). 299 tests green, zero new dependencies.","fileCount":30,"zipByteSize":187349},{"version":"3.7.0","createdAt":"2026-10-06T05:31:01.234Z","changelog":"v3.7.0: the context-check release. New --check-document paper.md: parses in-text citation markers ([12], [1-4], (Author, Year), doi.org links), binds each marker to the verified reference list, and scores anchor-word overlap between the citing sentence and the cited title - low anchors are flagged as possible mis-citations (cited paper A, claimed paper B) for human review. Verification moves from the reference list to the context. Also: check_document as a third MCP tool; RIS export (Zotero/EndNote round-trip); --preflight registry status panel. 299 tests green, zero new dependencies. If you find this useful, a Star on GitHub or a bookmark on the SkillHub skill page helps others discover it.","fileCount":21,"zipByteSize":176350},{"version":"3.4.1","createdAt":"2026-10-02T06:34:21.724Z","changelog":"v3.4.1: the dedup-fingerprint release. Batch deduplication now keys on identifier + normalized title fingerprint: two references sharing a URL, DOI, or PMID but carrying different titles (the classic stitched-fake variant) are independently verified instead of silently inheriting each other's verdict. Genuine duplicates (same identifier + same title) still merge with the partial downgrade. Found by external independent testing. 271 tests green, zero new dependencies.","fileCount":15,"zipByteSize":100749},{"version":"3.4.0","createdAt":"2026-10-01T06:25:14.344Z","changelog":"v3.4.0: the metadata-completeness release. PMID-only references now get full title/journal/year cross-checks against PubMed registry data (abbreviation-aware journal matching; genuine-PMID-fake-paper splice attacks are caught as invalid, field mismatches degrade to partial with claimed-vs-registered details), identifier-free references go through OpenAlex + Semantic Scholar title search (confirmation to partial, not-found to honest unverified for human review), and BLUF key_numbers separates retracted from reviewed counts. 271 tests green, zero new dependencies.","fileCount":15,"zipByteSize":100339},{"version":"3.3.1","createdAt":"2026-09-30T06:24:40.189Z","changelog":"v3.3.1: the title-cleaning release. Reference lines are cleaned before metadata matching (numbering prefixes, DOI/PMID/arXiv suffixes, author-year tails, punctuation noise) and the cleaned title feeds all four matching call sites (DOI direct, DOI easy-mode, arXiv, Semantic Scholar confirmation). Demo effect: AlphaFold-style titles match at 0.66 -> 0.9+ similarity. 261 tests green, zero new dependencies.","fileCount":15,"zipByteSize":97182},{"version":"3.3.0","createdAt":"2026-09-29T06:40:46.489Z","changelog":"v3.3.0: the heterogeneous-panel release. An optional NLI third vote","fileCount":15,"zipByteSize":96729}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17dagtwyk21qs6vpz98bzcrh1853t29:cite-holmes","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T10:43:06.178Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-cite-holmes/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T08:15:53.892Z","emptyReason":null},"readme":"Skill: Cite Holmes — Deep Research × Hallucination-Free Citations\n\nOwner: docsor1212\n\nSummary: Deep research that interrogates its own sources (Verified Deep Research): calibrates scope first (3-5 sharp questions), then searches iteratively and machine-verifies every citation. Also a standalone citation checker: paste any reference list and it runs full citation verification (it will verify citations before you cite) against official registries — hallucinated references, fabricated DOIs, fake PMIDs, arXiv IDs, stitched fakes, retracted papers; a fact check for your bibliography. Five verdicts; unverified references never masquerade as real (AI hallucination detection). Medical mode (Cochrane/BMJ/ChiCTR/NMPA/CDC/NICE presets, PMID check) and bibliography export (BibTeX/GB·T 7714-2025/RIS/CSV + JSON workpaper) built in. Trigger matching is semantic, not exact — mis-triggers are harmless; state your real intent to avoid them.\n\nTags: latest:3.11.0\n\nVersion history:\n\nv3.11.0 | 2026-10-09T17:22:29.713Z | user\n\nv3.11.0: the bibliography-profiling-and-version-hints release. List-level fabrication fingerprint scan (pure-local, zero network): serial identifier clusters, year over-concentration or future years, single-source concentration, bare-entry share - the statistical fingerprints of batch-fabricated reference lists; flags are human-review leads, never verdicts. Preprint-to-published hints: an arXiv-only reference that has since been registered as a journal article gets annotated with the formal-version DOI (via Semantic Scholar, silent on rate limits). Per-reference completeness score: how many export fields the registry CSL can back-fill, with the missing-field list. 379 tests green, zero new dependencies.\n\nv3.10.0 | 2026-10-08T18:43:14.852Z | user\n\nv3.10.0: the cross-lingual-and-clone-detection release. Cross-lingual title bridging: a Chinese original title claimed against an English registry record (the normal case for Chinese-journal DOIs) is resolved via a three-stage bridge - the bilingual original-title registry field, then the DOI landing page, then OpenAlex; a hit upgrades to verified, a miss keeps partial, and a language difference is never treated as fabrication evidence. Clone-pair detection flags same-title/different-DOI and same-DOI/different-title pairs inside one bibliography (the most common fabrication shape in generated text) as paired leads - annotation only, never a verdict change. --doctor (add --net) gives a zero-outbound environment self-check with per-registry connectivity probes, PASS/WARN/FAIL verdicts and CI-friendly exit codes. The SKILL VERIFY section is restructured command-first. 355 tests green, zero new dependencies.\n\nv3.8.0 | 2026-10-07T05:30:53.666Z | user\n\nv3.8.0: the documentation & examples completeness release. End-to-end worked example shipped (input refs + document excerpt -> real generated report and all four export formats, under examples/end-to-end/); references/ index and a consolidated anti-patterns list (incl. trigger-word mis-fires and the cross-language mixed-citation boundary); context-check results now render in the md/html reports (previously console/JSON only); author-surname anchor credit (a citing sentence naming a registry author counts as strong binding evidence). 299 tests green, zero new dependencies.\n\nv3.7.0 | 2026-10-06T05:31:01.234Z | user\n\nv3.7.0: the context-check release. New --check-document paper.md: parses in-text citation markers ([12], [1-4], (Author, Year), doi.org links), binds each marker to the verified reference list, and scores anchor-word overlap between the citing sentence and the cited title - low anchors are flagged as possible mis-citations (cited paper A, claimed paper B) for human review. Verification moves from the reference list to the context. Also: check_document as a third MCP tool; RIS export (Zotero/EndNote round-trip); --preflight registry status panel. 299 tests green, zero new dependencies. If you find this useful, a Star on GitHub or a bookmark on the SkillHub skill page helps others discover it.\n\nv3.4.1 | 2026-10-02T06:34:21.724Z | user\n\nv3.4.1: the dedup-fingerprint release. Batch deduplication now keys on identifier + normalized title fingerprint: two references sharing a URL, DOI, or PMID but carrying different titles (the classic stitched-fake variant) are independently verified instead of silently inheriting each other's verdict. Genuine duplicates (same identifier + same title) still merge with the partial downgrade. Found by external independent testing. 271 tests green, zero new dependencies.\n\nv3.4.0 | 2026-10-01T06:25:14.344Z | user\n\nv3.4.0: the metadata-completeness release. PMID-only references now get full title/journal/year cross-checks against PubMed registry data (abbreviation-aware journal matching; genuine-PMID-fake-paper splice attacks are caught as invalid, field mismatches degrade to partial with claimed-vs-registered details), identifier-free references go through OpenAlex + Semantic Scholar title search (confirmation to partial, not-found to honest unverified for human review), and BLUF key_numbers separates retracted from reviewed counts. 271 tests green, zero new dependencies.\n\nv3.3.1 | 2026-09-30T06:24:40.189Z | user\n\nv3.3.1: the title-cleaning release. Reference lines are cleaned before metadata matching (numbering prefixes, DOI/PMID/arXiv suffixes, author-year tails, punctuation noise) and the cleaned title feeds all four matching call sites (DOI direct, DOI easy-mode, arXiv, Semantic Scholar confirmation). Demo effect: AlphaFold-style titles match at 0.66 -> 0.9+ similarity. 261 tests green, zero new dependencies.\n\nv3.3.0 | 2026-09-29T06:40:46.489Z | user\n\nv3.3.0: the heterogeneous-panel release. An optional NLI third vote\n\nv3.2.0 | 2026-09-28T04:45:01.108Z | user\n\nv3.2.0: the conformance release. JSON reports now carry the BLUF block as a first-class object, conformant with the published BLUF Report Specification v1.0 (five core keys + worst-actionable rule; bluf_yaml kept for compatibility) - our own open validator passes our own output, closing the loop from spec to self-check to fix. Optional offline retraction checking: --retraction-cache points at an index built from the official Crossref/Retraction Watch GitLab dump (63k+ DOIs) - hits are flagged with zero network calls, misses fall through to live Crossref. 241 tests green, live acceptance clean.\n\nv3.1.0 | 2026-09-27T06:47:05.672Z | user\n\nv3.1.0: the closed-loop release. The L4 evidence cascade is now wired into the main flow: references the semantic layer could not resolve (not_in_source/unclear) escalate automatically to official PubMed abstracts or arXiv full text, get passage retrieval (bge-m3 embeddings, TF-IDF fallback - zero hard dependencies kept) and a configurable external judge. Judge findings annotate the report and never flip verdicts; without --judge-url the run is byte-identical (zero network calls). Judge client hardened for thinking models (qwen3 family): native ollama uses think:false + JSON-schema enum output, /v1 endpoints fall back from empty content to the reasoning field - all verified live on SciFact/SCitance/CFEVER. 228 tests green, real-network acceptance 8/8, package security self-scan clean.\n\nv3.0.0 | 2026-09-26T04:24:58.564Z | user\n\nv3.0.0: the BLUF release. Every report now opens with a machine-parseable YAML block (verdict / key_numbers / blocker / next_action) plus a 5-line human TL;DR - cite-holmes defines the 3-second dual-reader report standard, a gap no tool fills today. Scholarly reception: verified DOIs fetch Semantic Scholar citation contexts (times-cited with cap annotation, context excerpts, coverage honestly labeled - the free Scite.ai counterpart). L4 evidence cascade: abstract-level judging via configurable external judge endpoints (OpenAI-compatible), NEI escalates to full-text passage retrieval with local bge-m3 embeddings (TF-IDF fallback - zero hard dependencies kept). L1 claim-triplet instruction layer (RefChecker-style) across docs and auditjson. Cross-language title guard: Chinese claims vs English registrations no longer misjudged invalid. 210 tests green, real-network acceptance 6/6, package security self-scan clean.\n\nv2.0.0 | 2026-09-25T04:04:38.667Z | user\n\nv2.0.0: the MCP three-primitives release. mcp/server.py now exposes tools verify_references AND explain_verdict (plain-language verdict explanations with evidence steps), resources (capability matrix, changelog) and a fact-check-workflow prompt - agents that prefer tools over skills get the full verification stack without installing anything. arXiv multi-version awareness: unversioned citations to multi-revision papers and stale version pointers are flagged in notes (zero extra requests - parsed from the same API response). Optional NCBI_API_KEY raises E-utilities throughput 3 to 10 req/s for parallel batches. HTML reports gain a per-reference evidence chain (doi.org, retraction check, S2, OpenAlex, reachability at a glance). First Major version: the transparent multi-source citation verification stack, as shipped. 195 tests green, real-network acceptance 5/5, package security self-scan clean. Zero new dependencies for the skill itself.\n\nv1.13.0 | 2026-09-24T12:12:26.469Z | user\n\nv1.13.0: the discoverability release. EN description rewritten around task-language roots (citation verification, verify citations, citation checker, hallucinated references, fact check) so task-word searches on skills.sh and agent skill markets finally surface it. New MCP server (mcp/server.py, FastMCP): expose the same mechanical verification engine as a verify_references tool for agents that prefer tools over skills - official MCP Registry ready (publish-once to Smithery, mcp.so, PulseMCP). Security & behavior declaration retained. 188 tests green, real-network acceptance 5/5, package security self-scan clean. Zero new dependencies for the skill itself.\n\nv1.12.0 | 2026-09-23T04:05:00.625Z | user\n\nv1.12.0: the trust & transparency release. Persistent disk cache (default on, 7-day TTL, local SQLite): re-running or extending a bibliography reuses stable verdicts with zero network round-trips - the repeat-run pain behind every Trust score. --strict always bypasses the cache; --refresh-cache forces re-verification; unreachable verdicts are never cached. New --proxy flag for restricted-egress environments (HTTPS_PROXY also honored). A Security & behavior declaration now ships in both languages: network access limited to official academic registries over HTTPS, single-run CLI with no background jobs, nothing installed, no telemetry, env vars limited to your own API keys. Capability matrix adds an explicit unsupported-inputs line (PDF/Word docs, Wanfang/CQVIP-only IDs, fact-checking claims). Private maintainer tooling removed from the shipped package. 181 tests green, real-network acceptance 6/6, package security self-scan clean. Zero new dependencies.\n\nv1.11.0 | 2026-09-21T21:05:25.079Z | user\n\nv1.11.0: the determinism release. arXiv API transient 406/403 (its anti-bot window trips even under our 3s global pacing) no longer silently skips content verification - exponential backoff, then fallback to the official arxiv.org/abs landing page for title comparison (same 0.50/0.82 thresholds; landing-page 404 still judges invalid) - same references, same verdicts on re-runs; non-Atom responses (anti-bot HTML) are never misread as 'no record'. Structured semantic audit: the model's per-reference semantic judgment travels as machine-readable workpapers (semantic: {claim, support, quote, note} -> semantic_audit in --export auditjson plus a report section); not_in_source/contradicted cap the verdict at partial with a human-review flag. Pre-submission conclusion corrected to arXiv's actual 2026-05 enforcement start (was 2026-06) and now names ICML 2026 desk-rejections. tools/agentskills_check.py --distro prints skills.sh/agentskills.io readiness. 166 tests green, real-network acceptance 8/8. Zero new dependencies.\n\nv1.10.0 | 2026-09-20T18:33:10.967Z | user\n\nv1.10.0: the speed release, targeting the top evaluation complaint (overseas-database latency on flaky networks). Parallel verification (--workers, default 4; references are independent, results stay index-ordered; circuit breaker and rate limiters are now thread-safe). Degraded-network mode: 3+ different hosts failing at the transport layer flips remaining lookups to fast honest skips with a proxy/retry suggestion. Semantic Scholar title-search confirmation for references with no DOI/PMID/arXiv (confirmation only, never a downgrade; hyphens normalized per S2 API docs). Intra-run cache: duplicate DOI/URL reuses the first verdict. NCBI E-utilities pacing under parallelism (keyless 3 req/s safe). OpenAlex 429/403 hints mention the free key (required for production since 2026-02; 100 credits/day keyless). 147 tests green, real-network acceptance 8/8. Zero new dependencies.\n\nv1.9.0 | 2026-09-19T18:14:17.514Z | user\n\nv1.9.0: multi-source confirmation + lower friction. Semantic Scholar third source (second positive signal when DOI.org metadata is unreachable; confirmation only, never a downgrade; 403/429 landing pages can be rescued on S2 confirmation). BibTeX import: verify Zotero/EndNote .bib exports directly (nested braces, url macros, eprint to arXiv). OpenAlex API key support with actionable 403 hint (OpenAlex requires keys for production since 2026-02). --mailto joins the Crossref polite pool; 429 Retry-After honored. Every report opens with a one-line pre-submission conclusion (arXiv restricts authors over hallucinated references since 2026-06). --export auditjson: machine-readable per-reference check trail for institutional audit. SKILL.md slimmed, details moved to references/. 122 tests green, real-network acceptance 7/7. Zero new dependencies.\n\nv1.8.1 | 2026-09-18T05:15:34.872Z | user\n\nv1.8.1: hotfix - OpenAlex bibliographic confirmation could silently no-op (response read outside its context manager); caught by the new strict real-network acceptance gate and corrected. All v1.8.0 features unchanged (retraction detection via Crossref/Retraction Watch, multi-source checks, anti-patterns docs). 92 tests green.\n\nv1.8.0 | 2026-09-18T04:36:28.560Z | user\n\nv1.8.0: retraction detection (verified DOIs cross-checked against Crossref/Retraction Watch - a retracted paper is real, and uncitable; retracted refs capped at partial + human review), bibliographic existence check via OpenAlex for refs without DOI/PMID/arXiv (positive confirmation only), UA-rotation retry for metadata endpoints (403/406 resilience), common-mistakes anti-patterns section + 30-second quickstart, capability matrix now 10 types. 92 tests green, two rounds of real-network acceptance passed. Zero new dependencies.\n\nv1.7.0 | 2026-09-17T04:19:54.752Z | user\n\nv1.7.0: author-name consistency check vs DOI-registered authors (any claimed surname hit counts, full miss -> partial), capability-boundary matrix in every report (9 machine-caught fabrication types vs semantic-layer items), --format html self-contained single-file report (zero external assets, mobile readable), agentskills.io spec self-check. 77 tests green, two rounds of real-network acceptance passed. Zero new dependencies.\n\nv1.6.0 | 2026-09-16T05:03:34.074Z | user\n\nv1.6.0: host-level circuit breaker (2 consecutive transport failures to the same host -> skip for the rest of the batch with an honest note; broken networks no longer stall the whole run), actionable error notes (timeout suggests --timeout 20, DNS vs refused distinguished, human-review section carries suggested actions), DOI-PMID cross-check (catches stitched fakes where both IDs are real but belong to different papers), journal-name consistency check vs DOI-registered metadata (mismatch -> partial), CN-network FAQ. 62 tests green, two rounds of real-network acceptance passed. Zero new dependencies.\n\nv1.5.0 | 2026-09-15T01:07:22.554Z | user\n\nv1.5.0: arXiv ID metadata validation (official export.arxiv.org API, DOI-grade similarity thresholds; nonexistent IDs -> invalid, never unreachable; resolver-404 fallback), Wayback Machine fallback for unreachable entries (archive link attached), arxiv reference field, eprint in BibTeX + arXiv column in audit CSV. 53 tests green, two rounds of real-network acceptance passed. Zero new dependencies.\n\nv1.4.0 | 2026-09-14T01:02:06.259Z | user\n\nv1.4.0: DOI metadata cross-validation (real-DOI wrong-paper caught as invalid), --easy one-flag mode (auto medical profile + auto BibTeX/CSV export), transient-network retry, human-readable Chinese error messages, FAQ expansion, SSRF literal-IP guard, offline verdicts capped at partial. 30 tests green. Zero new dependencies.\n\nv1.3.0 | 2026-09-12T19:59:37.503Z | user\n\nv1.3.0: CiteScore confidence scorecard in every report (0-100 + A-D grade; verified +10 / partial +4 / unreachable 0 / unverified -2 / invalid -8); tests/ regression suite (offline golden, medical profile, exports, E-utilities); official icon. Zero new dependencies.\n\nv1.2.0 | 2026-09-12T18:46:22.932Z | user\n\nv1.2.0: medical evidence mode (--profile medical: medical source tiers with Chinese journals, community sources flagged as non-supporting for medical conclusions); PMID verification via NCBI E-utilities (catches fabricated PMIDs that pubmed web UI would accept); bibliography export (--export bibtex,csv: verified-only BibTeX + full audit CSV, utf-8-sig); 3-key dedup (URL/DOI/PMID); missing-field fixes for DOI/PMID entries. Zero new dependencies.\n\nv1.1.1 | 2026-08-15T12:11:46.223Z | user\n\nEnglish SKILL.md metadata for international listing (content unchanged); reproducible planted-fakes example\n\nv1.1.0 | 2026-08-15T01:09:39.250Z | user\n\nInitial public release: five-state citation verification, QUICK/FULL research modes, confidence-graded reports, zero-hallucination references\n\nArchive index:\n\nArchive v3.11.0: 29 files, 207208 bytes\n\nFiles: _meta.json (131b), docs/best-practices.md (2552b), examples/demo_refs.json (1780b), examples/end-to-end/exports/report_GB-T7714.txt (156b), examples/end-to-end/exports/report.csv (566b), examples/end-to-end/paper_excerpt.md (56b), examples/end-to-end/README.md (1244b), examples/end-to-end/report.json (4979b), examples/end-to-end/report.md (3102b), examples/end-to-end/research_refs.json (453b), examples/medical_refs.json (2647b), mcp/MCP_SUBMISSION.md (2470b), mcp/pyproject.toml (771b), mcp/README.md (3741b), mcp/server.py (14242b), mcp/sync_engine.sh (919b), mcp/verify_refs.py (213867b), README.md (29073b), references/anti-patterns.md (1857b), references/faq.md (8527b), references/INDEX.md (866b), references/medical-mode.md (3450b), references/report-template.md (3645b), references/search-strategies.md (3120b), references/verification-details.md (11150b), scripts/verify_refs.py (213867b), skill-card.md (2172b), SKILL.md (14865b), tools/agentskills_check.py (4549b)\n\nFile v3.11.0:SKILL.md\n\n---\nname: cite-holmes\nversion: 3.11.0\nauthor: DoctorQ Lab\nlicense: MIT\ndescription: >-\n  Deep research that interrogates its own sources (Verified Deep Research):\n  calibrates scope first (3-5 sharp questions), then searches iteratively and\n  machine-verifies every citation. Also a standalone citation checker: paste\n  any reference list and it runs full citation verification (it will verify\n  citations before you cite) against official registries —\n  hallucinated references, fabricated DOIs, fake PMIDs, arXiv IDs, stitched\n  fakes, retracted papers; a fact check for your bibliography. Five verdicts;\n  unverified references never masquerade as real (AI hallucination detection).\n  Medical mode (Cochrane/BMJ/ChiCTR/NMPA/CDC/NICE presets, PMID check) and\n  bibliography export (BibTeX/GB·T 7714-2025/RIS/CSV + JSON workpaper) built\n  in. Trigger matching is semantic, not exact — mis-triggers are harmless;\n  state your real intent to avoid them.\nwhen_to_use: >-\n  Use when the user says \"deep research\", \"look into\", \"investigate\",\n  \"compare A vs B\", \"fact check\", \"verify this claim\", \"is it true that...\",\n  \"check these references\", \"are these citations real\", wants a research\n  report with sources, a literature review, a medical evidence lookup, or\n  wants references verified before submission — even if they never say the\n  word \"research\".\n---\n\n# cite-holmes (Cite Holmes): deep research with citation verification\n\nOne line: **a question goes in — a verified report comes out.**\n\nThree differences from a plain \"search and summarize\":\n\n1. **Calibrate before working** — ask sharp questions first; the most expensive\n   waste is researching the wrong question.\n2. **Conclusions carry evidence grades** — 🟢 two independent sources agree /\n   🟡 single authority / 🔴 contested.\n3. **Every citation is checked** — mechanical layer (reachability, domain\n   authority, field completeness, dedup) plus semantic layer (does the source\n   actually support the claim?). Unverified references never masquerade as real.\n\n## ⛔ Iron rules (zero exceptions)\n\n1. **Never fabricate**: citations must come from pages actually fetched this\n   session. Re-search rather than write URLs from memory.\n2. **Never pretend**: unchecked references are marked `unverified`; fetch\n   failures are `unreachable` (≠ nonexistent — flagged for human review).\n3. **Don't hide conflicts**: when sources disagree, present the disagreement,\n   mark 🔴, show each side's evidence.\n4. **Budgeted search**: QUICK ≤6 searches, FULL ≤15. Out of budget → state the\n   gaps honestly instead of forcing conclusions.\n5. **Calibrate before searching** (FULL mode): scope / timeframe / audience /\n   output format must be locked first.\n\n## Step 0: mode selection\n\n| Mode | Fits | Calibration | Budget | Output |\n|---|---|---|---|---|\n| **QUICK** | Single fact-check: \"is this claim true\", \"when was X released\" | skipped | ≤6 | short report |\n| **FULL** | Open research: \"state of X\", \"A vs B\", \"do a survey\" | mandatory | ≤15 | full report |\n\nA question answerable by one verifiable fact → QUICK. Needs synthesis or\ntrade-offs → FULL. \"Quick check\" forces QUICK; \"thorough/comprehensive\" forces\nFULL. When unsure, default FULL.\n\n## Five-phase workflow\n\n### 1. CALIBRATE (FULL only)\n\nAsk 3–5 high-leverage questions at once (no drip-feeding): scope, timeframe,\naudience/depth, output format, decision context. Never re-ask what the user\nalready provided. If the user declines (\"your call\"), proceed with stated\ndefaults.\n\n### 2. PLAN\n\nShow a short plan: 3–7 sub-questions, source priority (primary/official >\nmajor media > community/blog as leads only), budget.\n\n### 3. SEARCH (iterative, not one pass)\n\n**Read `references/search-strategies.md` first** (diamond expansion, source\npyramid, query matrix, gap-driven iteration). Essentials: each round targets\none sub-question; evolve queries with discovered terms; search both English\nand Chinese for topics that span both internets; fetch full text of the 2–5\nmost valuable sources (never conclude from search snippets); verify key\nnumbers/dates in the original page before quoting.\n\n### 4. VERIFY (the heart of this skill)\n\nRegister every reference in `research_refs.json` (schema in\n`references/report-template.md`), then run:\n\n```bash\npython scripts/verify_refs.py --refs research_refs.json --out verify_report.md\n# Got a .bib from Zotero/EndNote? Feed it directly (v1.9):\npython scripts/verify_refs.py --refs bibliography.bib --out report.md\n# Environment self-check before a big batch (v3.10) — zero outbound calls\n# by default, add --net to probe every academic registry:\npython scripts/verify_refs.py --doctor --net\n```\n\nFive verdicts: `verified` / `partial` / `unreachable` (needs_human_check) /\n`invalid` / `unverified`. Every identifier (`url` / `doi` / `pmid` / `arxiv`)\nis cross-checked against its official registry (DOI.org metadata, NCBI\nE-utilities for PMID, export.arxiv.org for arXiv), so fabricated IDs are\njudged `invalid` — never silently `unreachable`. Every report opens with a\n**BLUF dual-reader header** (machine-parseable YAML + 5-line human TL;DR), a\n**CiteScore** (0-100 + A-D grade) and a one-line **pre-submission conclusion**\n(arXiv has banned authors over hallucinated references since 2026-05; ICML\n2026 desk-rejects them too).\n\n**Mechanical layer (the script)**: claim↔registry title/year/journal/author\nconsistency; nine machine-readable error codes on every judgment (v3.9);\n**clone-pair detection** (v3.10 — same title under different DOIs, or one DOI\ncarrying different titles, flagged in pairs: the most common fabrication shape\nin generated text); **cross-lingual title bridging** (v3.10 — a Chinese\noriginal title claimed against an English registry record is resolved via the\nbilingual `original-title` field, the DOI landing page, or OpenAlex; a hit\nupgrades to `verified`, a miss keeps `partial`, and a language difference is\nnever treated as fabrication evidence); retraction checks (Crossref online, or\ninstantly offline via an optional local Retraction Watch index\n`--retraction-cache`); Wayback archive links attached to dead links; arXiv\nlanding-page fallback keeps verdicts deterministic when its API flakes;\n**Bibliography profiling** (v3.11 — batch-fabrication\nfingerprints across the whole list: serial identifier clusters, year\nover-concentration or future years, single-source concentration, bare-entry\nshare; flags are human-review leads, never verdicts); **preprint→published\nhints** (v3.11 — an arXiv-only reference that has since been registered as a\njournal article gets annotated with the formal-version DOI, via Semantic\nScholar, silent on rate limits); per-reference **completeness scoring**\n(v3.11 — how many export fields the registry CSL can back-fill, with the\nmissing-field list for authors deciding what to complete before export).\nparallel verification (`--workers`, default 4) plus a local disk cache\n(default on, 7-day TTL) make repeat runs cheap and verdicts stable. Behind a\nfirewall: `--proxy http://host:port`, `--cn` (resilient preset: 25s floor +\nCrossref re-source), or `--preflight` to see registry status before the batch;\nwhen 3+ registries fail at transport level the run flips to degraded-network\nmode — fast, honest skips instead of minutes of waiting.\n\n**Semantic layer (the model)**: decompose each cited claim into atomic\nclaim-triplets (subject–relation–object, RefChecker style) BEFORE judging,\nthen verify each triplet against the source — granularity moves from\nparagraph to triple, so \"which half-sentence is wrong\" becomes answerable.\nRegister the verdict structurally so it becomes auditable workpaper, not a\nfeeling:\n`\"semantic\": {\"claim\": \"...\", \"support\": \"supported|partial|not_in_source|\ncontradicted|unclear\", \"quote\": \"...\", \"note\": \"...\"}`. The verifier carries\nit into `--export auditjson` and a report section; `not_in_source` /\n`contradicted` cap the mechanical verdict at `partial` with a human-review\nflag — a source that exists is not a source that agrees.\n\n**Multi-source confirmation**: verified titles get an OpenAlex bibliographic\ncross-check; DOI references whose registry metadata could not be fetched get\na Semantic Scholar second confirmation; references with no DOI/PMID/arXiv at\nall get a Semantic Scholar title-search confirmation — confirmation only,\nnever a downgrade: databases have coverage gaps, \"not found\" ≠ fabricated.\n\n**Task → flag** (details in `references/verification-details.md`):\n\n| Task | Flags |\n|---|---|\n| Medical topics | `--easy` (auto medical + exports) · `--profile medical` |\n| Share / archive | `--format html` (self-contained single file) |\n| Export | `--export bibtex,gbt7714,ris,csv,auditjson` |\n| CI | `--strict` · `--offline` (structure only, caps at `partial`) |\n| Polite pacing | `--mailto you@lab.edu` |\n| API keys | `--openalex-key` / `--s2-key` / `--ncbi-key` (env of same name) |\n| Weak network | `--proxy` · `--cn` · `--preflight` · `--timeout` |\n| Cache | `--no-cache` / `--refresh-cache` / `--cache-ttl` |\n| Self-check | `--doctor` (add `--net` for registry probes) |\n| External judging | `--judge-url` / `--nli-url` / `--fast-judge-url` (notes only, never flip verdicts) |\n| MCP server | `mcp/server.py` — tools `verify_references` / `explain_verdict` / `check_document`, resources (capability matrix, changelog), `fact_check_workflow` prompt |\n\n**Offline mode (`--offline`)**: structure checks only — honesty first, never\nawards `verified`.\n\n**Medical research mode (v1.2)**: for clinical questions read\n`references/medical-mode.md` first and run with `--profile medical`.\n\n### 5. SYNTHESIZE\n\nFollow the skeleton in `references/report-template.md`: executive summary\nfirst; every key conclusion carries a confidence grade + citation ids; the\nreference table carries verdicts; `unverified/unreachable` items live only in\nthe \"human review\" section; finish with gaps, disagreements, and follow-up\nquestions.\n\n## Files\n\n| File | When |\n|---|---|\n| `scripts/verify_refs.py` | VERIFY phase mechanical check (pure stdlib, cross-platform, rate-limited) |\n| `mcp/server.py` | optional MCP server: expose `verify_references` as an MCP tool (FastMCP; same engine) |\n| `references/search-strategies.md` | read before SEARCH |\n| `references/report-template.md` | skeleton for SYNTHESIZE; refs schema |\n| `references/verification-details.md` | full option semantics, cross-check rules, thresholds |\n| `references/faq.md` | common questions (network, PMID wording, exports) |\n| `references/medical-mode.md` | read before medical/clinical research (v1.2) |\n| `examples/demo_refs.json` | general demo: 8 refs, 3 planted fabrications |\n| `examples/medical_refs.json` | medical demo: PMID/DOI refs + planted fake PMID + planted duplicate |\n\n## ⚠️ Common mistakes (anti-patterns)\n\n- Treating `verified` as \"the content is correct\" — verification confirms the\n  source exists and matches the claim's source, not that the reasoning holds.\n  The semantic check and your own reading still matter.\n- Citing a `partial` source as if it were reliable — `partial` means downgrade\n  or explain in the text; never silently promote it.\n- Retrying `unreachable` links forever — open the Wayback link once, then\n  switch sources if it stays dead.\n- Relying on `--easy` for clinical work — auto-detection is a convenience;\n  for real medical conclusions pass `--profile medical` explicitly.\n- Running without `--strict` in CI/automated pipelines and expecting exit\n  codes to gate anything.\n- Assuming `verified` means \"safe to cite forever\" — retracted papers are\n  flagged (Crossref/Retraction Watch), but retraction status can change after\n  your run.\n\n## Security & behavior declaration\n\n- Single-run CLI: start, verify, write the report, exit. No daemons, no\n  background jobs, nothing downloaded or installed at runtime (pure standard\n  library).\n- Network access is limited to these official academic registries, always over\n  HTTPS: `doi.org`, `api.crossref.org`, `eutils.ncbi.nlm.nih.gov`,\n  `export.arxiv.org` / `arxiv.org`, `archive.org`, `api.openalex.org`,\n  `api.semanticscholar.org`, `pubmed.ncbi.nlm.nih.gov`. No other hosts are\n  contacted; no telemetry, no analytics, no data collection.\n- Your reference lists and reports stay on your machine — the only outbound\n  payloads are the identifiers and titles you asked to verify.\n- Optional environment variables `OPENALEX_API_KEY` / `S2_API_KEY` /\n  `NCBI_API_KEY` authenticate\n  your own requests to those two APIs and are never sent anywhere else. The\n  local verdict cache lives under `~/.cache/cite-holmes/` (`--no-cache` to\n  disable).\n- No OS integration: no subprocesses, no system services, no privilege\n  changes; file access is limited to your inputs/outputs plus the cache and\n  report paths you pass.\n\n## Honest limits\n\n- Not a fit for: PDF/Word documents (extract citations manually first),\n  Chinese-database-only IDs (Wanfang/CQVIP numbers — verify via URL or title\n  instead), and fact-checking what a page *claims* (the mechanical layer\n  verifies existence and consistency; meaning is the semantic layer's job).\n- A fabricated citation pointing to a real, live, plausible page passes the\n  mechanical layer; the semantic layer may catch it — model judgment, not a\n  guarantee.\n- `unreachable` ≠ fake; database \"not found\" ≠ fabricated (coverage lag).\n- No public-registry check replaces reading the source — the July 2026 arXiv\n  survey of citation checkers found none reliable enough to run unsupervised;\n  treat this skill as a transparent multi-source assistant with a full audit\n  trail (`--export auditjson`), not an oracle.\n- Reproducible demo: `python scripts/verify_refs.py --refs examples/demo_refs.json`\n  (8 refs, 3 planted fabrications, all caught).\n- Medical demo: `python scripts/verify_refs.py --refs examples/medical_refs.json\n  --profile medical --export bibtex,csv`.\n\n## FAQ\n\nSee `references/faq.md` (slow networks / blocked endpoints, what PMID wording\nmeans, exports into Zotero/EndNote, how arXiv IDs are judged). Quick answers:\n`verified` is existence+consistency, not truth; `unreachable` ≠ fake;\n`--easy` covers the medical profile auto-detection; `--export bibtex` gives\nZotero/EndNote-ready verified-only bibliography.\n\n## 30-second quickstart\n\n```bash\n# 1. references as JSON (title/url/source/year per item) — or a Zotero .bib\n# 2. verify:\npython scripts/verify_refs.py --refs refs.json --out report.md\n# 3. read report.md — CiteScore + pre-submission conclusion at the top.\n# Medical? add --profile medical.  Paper?  add --export bibtex.\n# Institution/CI?  add --mailto you@lab.edu (Crossref polite pool).\n```\n\n## Environment fallback\n\nWithout web tools: state honestly that only the \"user-supplied material +\nmechanical verification\" mode is possible; never pretend to search.\n\nFile v3.11.0:examples/end-to-end/README.md\n\n# 端到端对照示例（输入 → 命令 → 实际输出）\n\n输入 → 命令 → 实际输出，三件套可直接复跑：\n\n```bash\npython scripts/verify_refs.py --refs end-to-end/research_refs.json \\\n    --check-document end-to-end/paper_excerpt.md \\\n    --export gbt7714,ris,bibtex,csv --out end-to-end/report.md\n```\n\n| 文件 | 是什么 |\n|---|---|\n| research_refs.json | 输入：1 条真实 DOI 引用（含正确标题） |\n| paper_excerpt.md | 输入：含 in-text 引用标记的正文片段（叙述式 + [n] 各一处，其中一处故意跑题） |\n| report.md | 实际输出：五态判定 + 上下文核验区（Kucsko 叙述式=锚定 ✓；[1] 企鹅句=低锚 ⚠️ 演示错配检测） |\n| report.json | 机读版（含 context_check 与逐项 checks） |\n| exports/ | 四种导出的真实产物样例（BibTeX / GB·T 7714-2025 / RIS / CSV） |\n\n> 复跑需网络（DOI.org/PubMed）。判定含时效字段（撤稿/被引），数字可能随时间小幅变化\n> ——以你复跑时的输出为准，结构不变。\n\n> SkillHub 包注：为符合平台文件类型白名单，`exports/` 导出样例文件不随\n> SkillHub 包分发（GitHub/ClawHub 包完整附带）——格式以本页表格描述为准。\n\nFile v3.11.0:mcp/README.md\n\n# cite-holmes-mcp\n\n**Mechanical citation verification as an MCP server** — feed it the references an\nagent is about to cite, get back per-reference verdicts (`verified / partial /\nunreachable / invalid`) against official registries (DOI.org, PubMed\nE-utilities, arXiv API, Crossref/Retraction Watch, Wayback). Zero API keys\nrequired, zero telemetry, local-first.\n\nThree MCP primitives:\n\n| Primitive | Name | What it does |\n|---|---|---|\n| tool | `verify_references` | Verify a batch of references (≤50); returns CiteScore 0–100 + per-item verdicts with evidence chain |\n| tool | `check_document` | Context-level check: parse in-text markers ([12]/[1-4]/(Author, Year)/doi.org links), bind to the reference list, flag low anchor-word overlap as possible mis-citations |\n| tool | `explain_verdict` | Plain-language explanation of one verdict (evidence steps + suggested action); zero network |\n| resource | `cite-holmes://capability-matrix` | What the mechanical layer catches vs. what stays with the semantic layer |\n| resource | `cite-holmes://changelog` | Version history |\n| prompt | `fact_check_workflow` | Three-step research-then-verify workflow for agents |\n\n## Install (three channels)\n\n```bash\n# 1) uvx from GitHub (recommended — always current)\nuvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp\n\n# 2) uvx from a local checkout\ngit clone https://github.com/docsor1212/cite-holmes\nuvx --from ./cite-holmes/mcp cite-holmes-mcp --help\n\n# 3) pip install (module + console script)\npip install \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\"\ncite-holmes-mcp --help\n```\n\n> This is a Python package (uv/pip ecosystem). There is no npm package — the\n> uvx/git channel above is the canonical install for all MCP clients.\n\n## Client configuration\n\n**Claude Code**\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\"]\n    }\n  }\n}\n```\n\n**Codex / any stdio MCP client**\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\", \"--transport\", \"stdio\"]\n    }\n  }\n}\n```\n\n**Cursor** (`~/.cursor/mcp.json`, same shape)\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\"]\n    }\n  }\n}\n```\n\nHTTP mode (local): `cite-holmes-mcp --transport http --port 8760`.\n\n## Bounds & behavior\n\n- Batch ≤ 50 references per `verify_references` call (larger batches get a\n  clear error — split them).\n- Long output fields are clipped (800 chars + truncation marker) so MCP\n  messages stay bounded; structure is preserved.\n- Per-reference network timeout defaults to 15 s (configurable per call).\n- Optional API keys via env: `OPENALEX_API_KEY` / `S2_API_KEY` /\n  `NCBI_API_KEY` (only attach to the caller's own requests).\n\n## Layout\n\n```\nmcp/\n├── server.py          # FastMCP three-primitive server (engine-agnostic impl)\n├── verify_refs.py     # embedded engine copy — synced from ../scripts/ by\n│                      #   sync_engine.sh (sha-gated; never hand-edit here)\n├── sync_engine.sh     # trunk↔mcp engine consistency gate (run before publish)\n├── pyproject.toml\n├── README.md\n└── tests/\n    ├── test_bounds.py        # batch/clip/bad-input unit tests\n    └── e2e_transcript.md     # full MCP client session evidence\n```\n\nLicense: MIT. Author: SorSor. Repo: https://github.com/docsor1212/cite-holmes\n\nFile v3.11.0:README.md\n\n# Cite Holmes 🔍\n\n[![GitHub Stars](https://img.shields.io/github/stars/docsor1212/cite-holmes?style=social&label=Star)](https://github.com/docsor1212/cite-holmes/stargazers)\n\n![icon](assets/icon-512.png)\n\n**Deep research that interrogates its own sources.**\n\nEvery AI research report you've ever read had a dirty secret: some of those polished references were probably fabricated. [A Nature news analysis suggests tens of thousands of 2025 publications might include invalid AI-generated references](https://www.nature.com/articles/d41586-026-00969-z). [GPTZero scanned 4,841 NeurIPS 2025 submissions; as independently reported, at least 100 hallucinated citations were found across 51 accepted papers](https://medium.com/@ljingshan6/100-fake-citations-just-slipped-through-neurips-2025-peer-review-5f34f4436560).\n\nCite Holmes is a deep-research skill with a badge and a magnifying glass: it researches like any deep-research agent — then **arrests its own citations before you can cite them**.\n\n![demo](assets/demo.gif)\n\n*(Demo is real output: 8 references, 3 deliberately planted fabrications — a fake DOI, a dead URL, and a no-URL citation. All 3 were caught and excluded; the 5 real ones passed. Measured: **7.7 s for all 8** — 4.2 s of pure network checks, the rest is deliberate throttling.)*\n\nReproduce it yourself — the planted-fakes file ships with the repo:\n\n```bash\npython scripts/verify_refs.py --refs examples/demo_refs.json\n```\n\n## 30-second quickstart\n\n```bash\n# MCP server (for agents — Claude Code / Codex / Cursor):\nuvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp\n```\n\n\n```bash\npython scripts/verify_refs.py --refs refs.json --out report.md\n# or verify straight from your reference manager (v1.9):\npython scripts/verify_refs.py --refs bibliography.bib --out report.md\n# v3.7.0 context check — verify in-text citations inside a finished draft:\npython scripts/verify_refs.py --refs bibliography.bib --check-document paper.md --out report.md\n```\n\n**China mirror (ModelScope 魔搭)**: <https://modelscope.cn/skills/Docsor/cite-holmes> — if you find this skill useful, a like there helps others find it.\n\nOpen `report.md`: CiteScore + pre-submission conclusion at the top,\nper-reference verdicts below.\nMedical work: add `--profile medical`.  Writing a paper: add `--export bibtex`.\nInstitution/CI: add `--mailto you@lab.edu` (Crossref polite pool).\n\n## What NOT to do (anti-patterns)\n\n- **Don't feed `semantic` fields from the same model that wrote the draft** —\n  the semantic cap trusts structured model judgments; self-review defeats it.\n- **Don't treat `verified` as \"the paper supports my claim\"** — verified means\n  the source exists at an authoritative tier and registry metadata matches the\n  claimed title/authors/journal/year. Whether the *specific sentence* is\n  supported is the semantic layer (agent judgment, or L4 cascade with a judge).\n- **Don't batch >50 references per MCP call** — the server rejects oversized\n  batches by design; split them.\n- **Don't hand-edit `mcp/verify_refs.py`** — it is a sha-gated copy of\n  `scripts/verify_refs.py`; edit the trunk and run `mcp/sync_engine.sh`.\n- **Don't cite from the report without reading `needs_human_check` flags** —\n  partial/unreachable items carry explicit re-check instructions on purpose.\n- **Don't skip `--cn` on lossy CN egress and then report unreachable as\n  \"fake\"** — unreachable is not invalid; re-run with `--cn` or a proxy first.\n\n## How it works\n\nFive phases, two modes:\n\n| Phase | What happens |\n|---|---|\n| **CALIBRATE** | Asks you 3–5 sharp questions first (scope, timeframe, audience) — prevents researching the wrong question |\n| **PLAN** | Breaks your question into 3–7 sub-questions with a search budget |\n| **SEARCH** | Diamond-shaped iterative search: broad → narrow → gap-filling, Chinese + English, source-tier pyramid |\n| **VERIFY** | Two layers: semantic check (does the source actually support the claim?) + mechanical check (reachability, domain authority, field completeness, dedup) |\n| **SYNTHESIZE** | Report where every conclusion carries a confidence grade — 🟢 two independent sources / 🟡 single authority / 🔴 contested |\n\nModes: **QUICK** (single fact-check, ≤6 searches, no interrogation) vs **FULL** (open-ended research, ≤15 searches, calibration mandatory).\n\n## Benchmarks (live runs + quality audit, 2026-09-26; zero-shot pipeline, qwen3:4b judge)\n\n| Benchmark | Setting | n | micro-acc | macro-F1 | Honest reading |\n|---|---|---|---|---|---|\n| SciFact-Open | open retrieval (mxbai top-3) + judge | 279 | 0.484 | 0.381 | the sanctioned zero-shot vehicle; 0.444/0.313 as first run before label-case normalization (both disclosed, sweep-reproducible); top systems (fine-tuned) reach 0.55-0.64 — our gap concentrates in REFUTES recall: abstract-level evidence rarely shows contradiction; remedies on the roadmap are the NLI third vote, claim-triplet judging, and full-text escalation |\n| CFEVER (zh) oracle + panel | gold evidence, zh judge + Erlangshen flip | 1000 | **0.817** | — | panel +8.1 over the 0.736 solo run; REFUTES correct 71→248 | oracle upper bound for the Chinese judge; same REFUTES weakness (R 0.18), excellent NEI discipline (R 1.0) |\n| SciFact-Open + L2 panel | judge + NLI contradiction flip (zero-param) | 279 | **0.617** | 0.563 (R-F1 0.667) | the heterogeneous two-vote panel lifts the same zero-shot pipeline +13 micro points into the fine-tuned top-system range (0.55-0.64); REFUTES F1 from 0 to 0.667 (paired flip 45/8, p=2.4e-7) |\n| SCitance v2.1 (self-built) | citation verification, hard negatives | 2576 | **0.713** (judge solo) / 0.646 (with NLI flip) | — | honest successor of the deprecated v1 0.947; the NLI contradiction flip that helps fact-check framing (+13 on SciFact-Open) HURTS here (-7) — topically-close hard negatives make the NLI over-report contradiction; aggregation is task-dependent, so the product exposes the panel as opt-in | released with quality audits: embedding-only separability 0.726 (v1 was 0.95 = degenerate), PIR 0/40; 751 false negatives filtered (36% of mined negatives actually supported the claim) |\n\n**Ablation (SciFact-Open, zero-shot)**: no-retrieval 0.262 (all-NEI floor) → top-1 **0.488** → top-3 0.444 → top-5 0.441; short-evidence (400ch) 0.405.\nReading: retrieval is decisive, but deeper evidence pools amplify the judge's SUPPORTS\nbias — precision-oriented retrieval (top-1) is optimal for this judge class.\n\nRetrieval: bge-m3 (corpus cache) / mxbai-embed-large (SciFact-Open runs). Judge:\nqwen3:4b via ollama (think:false + JSON-schema output; 8k-ctx resident). Full\nconfusion matrices and the audit gate scripts ship with the benchmark repo.\n\n## The five citation verdicts\n\n| Verdict | Meaning |\n|---|---|\n| `verified` | Reachable + authoritative tier (official/journal/preprint/major media) + complete fields — may support conclusions |\n| `partial` | Reachable but community/blog tier or missing fields — downgraded use |\n| `unreachable` | 404/timeout (≠ nonexistent — flagged for human review) |\n| `invalid` | Unresolvable or nonexistent identifier (bad/fake DOI, arXiv ID, PMID, URL) — never enters the report |\n| `unverified` | Not checked — never masquerades as verified |\n\n## v1.2: medical evidence mode + bibliography export\n\n```bash\n# medical profile: journal-tier sources extend to Cochrane/BMJ/ClinicalTrials/\n# NMPA/CDC/NICE/万方/ChiCTR; community-tier sources get an explicit\n# \"unfit for medical conclusions\" warning\npython scripts/verify_refs.py --refs refs.json --profile medical\n\n# PMID-only references work out of the box — and every PubMed URL gets its\n# PMID existence-checked via NCBI E-utilities. A fabricated PMID is caught\n# even though PubMed's own page happily returns 2xx for it.\n\n# export: verified-only BibTeX (straight into your paper) + full audit CSV\npython scripts/verify_refs.py --refs refs.json --export bibtex,csv\n```\n\nTry the medical demo — real tocilizumab/sJIA references from PubMed, plus one\nplanted fake PMID and one planted duplicate (both caught):\n\n```bash\npython scripts/verify_refs.py --refs examples/medical_refs.json --profile medical --export bibtex,csv\n```\n\n## v1.5: arXiv verification + Wayback fallback\n\n```bash\n# arXiv IDs are now first-class: the official export.arxiv.org API cross-checks\n# the registered title/year with the same thresholds as DOIs. A well-formed but\n# nonexistent ID is judged invalid (fabricated) — never silently \"unreachable\".\n# refs may carry \"arxiv\": \"2310.10631\" (or \"hep-th/9901001\"); arxiv.org/abs|pdf\n# URLs are auto-detected.\npython scripts/verify_refs.py --refs refs.json\n\n# Unreachable entries are automatically checked against the Wayback Machine;\n# when an archived copy exists, the report note carries the archive link so\n# human review has something to compare against.\n```\n\n## Install\n\n**AI agents (skills.sh / one command):**\n\n```bash\nnpx skills add docsor1212/cite-holmes\n```\n\n**OpenClaw users (ClawHub — versioned, auto-updatable):**\n\n```bash\nclawhub search cite-holmes        # find it on the registry\nclawhub install @docsor1212/cite-holmes\n```\n\n**Any Agent Skills-compatible agent** (Claude Code, Codex, Cursor, Gemini CLI, ZCode):\n\n```bash\nclawhub install @docsor1212/cite-holmes        # ClawHub registry\n# or grab the folder directly from SkillHub: https://skillhub.cn/skills/cite-holmes\n# then drop it into ~/.claude/skills/cite-holmes (or your agent's skills directory)\n```\n\n**China mirror (SkillHub 腾讯)**: <https://skillhub.cn/skills/cite-holmes> — fast downloads inside China, 中文说明.\n\n## The DoctorQ Lab academic family\n\nCite Holmes works alongside three sibling skills:\n\n- [academic-figures](https://skillhub.cn/skills/academic-figures) — publication-ready scientific figures (22+ chart types incl. composite panels, PRISMA, forest, KM) from one command\n- [paper-polisher-pro](https://skillhub.cn/skills/paper-polisher-pro) — AI-rate self-check, polish guidance and AIGC-compliance labeling for academic writing, bilingual, 100% local\n- [pubmed-verifier](https://skillhub.cn/skills/pubmed-verifier) — focused PubMed citation verification for medical literature\n\n## Usage\n\nJust talk to your agent — it triggers automatically:\n\n```\n> deep research: what changed in the agent-skills ecosystem this year?\n> is it true that NeurIPS 2025 papers contained 100+ hallucinated citations?\n```\n\nOr use the verifier standalone on any reference list:\n\n```bash\npython scripts/verify_refs.py --refs research_refs.json --out report.md\n# offline structural check / strict CI mode\npython scripts/verify_refs.py --refs refs.json --offline\npython scripts/verify_refs.py --refs refs.json --strict\n```\n\n## What it won't catch (honest limits)\n\n- A fabricated citation that points to a **real, live, plausible page** passes the mechanical check. The semantic layer (the model judging whether the source actually supports the claim) may catch it — it is model judgment, not a guarantee.\n- `unreachable` ≠ fake: pages behind login walls or bot-blocking are flagged for human review, not condemned.\n- The planted fakes in our demo are exactly the catchable types (dead URL / fake DOI / missing URL). We are not claiming it catches everything.\n\n## Why not just use deep research / a citation checker?\n\n- Deep-research skills **research more** but trust their own citations.\n- Standalone citation checkers verify but don't research.\n- Cite Holmes does both in one flow: **every reference in every report is machine-checked before it reaches you.**\n\nZero dependencies (pure Python stdlib), cross-platform (Windows/Linux/macOS), MIT license.\n\n## CiteScore — every report carries a grade\n\nEach verification report opens with a **CiteScore**: a 0-100 confidence score\n(verified +10 / partial +4 / unreachable 0 / unverified −2 / invalid −8,\nnormalized by reference count) with an A–D grade. Screenshot it, quote it,\nor gate your CI on it (`--strict`).\n\n## What's new\n\n- **v3.11.0** — bibliography profiling & version hints: a list-level\n  fabrication fingerprint scan (pure-local, zero network) catches what\n  per-reference checks cannot — serial identifier clusters, year\n  over-concentration or future years, single-source concentration, and the\n  bare-entry share; flags are human-review leads, never verdicts. An\n  arXiv-only reference that has since been registered as a journal article\n  gets a preprint→published hint (formal-version DOI via Semantic Scholar,\n  silent on rate limits). Verified DOI entries now carry a per-reference\n  **completeness score** — how many export fields the registry CSL can\n  back-fill, with the missing-field list, so authors know exactly what to\n  complete before exporting. 379 tests green.\n- **v3.10.0** — cross-lingual titles & clone detection: a Chinese original\n  title claimed against an English registry record (the normal case for\n  Chinese-journal DOIs) is now resolved by a three-stage bridge — the\n  bilingual `original-title` registry field, then the DOI landing page, then\n  OpenAlex — a hit upgrades the verdict to `verified`, a miss keeps\n  `partial`, and a language difference is never treated as fabrication\n  evidence. **Clone-pair detection** flags same-title/different-DOI and\n  same-DOI/different-title pairs inside one bibliography (the most common\n  fabrication shape in generated text) as pairs — a lead for review, never a\n  verdict change. `--doctor` (add `--net`) gives a zero-outbound environment\n  self-check with per-registry connectivity probes, PASS/WARN/FAIL verdicts\n  and CI-friendly exit codes. The SKILL VERIFY section is restructured\n  command-first (run → verdicts → mechanical/semantic layers → task→flag\n  table). 355 tests green.\n- **v3.9.0** — structured error codes + escalation routing: every result now\n  carries machine-readable `error_codes` (E_DOU_NOT_FOUND / E_TITLE_MISMATCH /\n  E_RETRACTED / E_STITCHED / E_UNREACHABLE / ... — five verdicts plus the\n  reason class, derivable offline); the L4 evidence cascade gains a **fast-judge\n  pre-screen route** (with `--fast-judge-url`, suspected mis-citations flagged\n  by the 322M student are escalated first — routing, not skipping: every\n  candidate is still fully cascaded); MCP `verify_references` clipping is now\n  explicit (`max_field_chars` parameter + `output_clipped` marker when fields\n  are shortened — no silent truncation). 306 tests green.\n- **v3.8.0** — documentation & examples completeness release: end-to-end worked example (input refs → real generated report →\n  all four export formats, shipped under `examples/end-to-end/`);\n  `references/INDEX.md` + consolidated anti-patterns list (incl. \"trigger-word\n  mis-fires\" and cross-language mixed-citation boundary); **context-check\n  results now appear in the md/html reports** (previously console/JSON only);\n  author-surname anchor credit (a citing sentence naming a registry author\n  counts as strong binding evidence). 299 tests green.\n- **v3.7.0** — document-level citation checking (`--check-document paper.md`):\n  parse in-text citation markers ([12], [1-4], (Author, Year), doi.org links),\n  bind each marker to the verified reference list, and score anchor-word\n  overlap between the citing sentence and the cited title — low anchors are\n  flagged as possible mis-citations for human review. This moves verification\n  from the reference *list* to the *context*: the list can be clean while the\n  citations still point at the wrong papers. Also: `check_document` as a third\n  MCP tool; RIS export (Zotero/EndNote round-trip); `--preflight` registry\n  status panel. 292 tests green.\n- **v3.6.0** — GB/T 7714-2025 reference-list export (`--export gbt7714`):\n  verified entries formatted per the Chinese national bibliography standard\n  effective 2026-07-01, **built from registry-authoritative CSL metadata**\n  (authors in GB/T abbreviation form, [J]/[EB/OL] type codes,\n  volume(issue):pages, DOI) — not from user-claimed fields. Title-disputed\n  partials are excluded by policy (better absent than wrong). Plus `--cn`\n  resilience preset: 25s timeout floor and a Crossref second-host fallback\n  for DOI metadata (independent host; doi.org outages no longer blind the\n  checker). Anti-pattern guide + scenario best-practices docs. MCP server:\n  batch cap 50, output clipping, packaged entry point (tag mcp-v2.0.0).\n  292 tests green.\n- **v3.5.0** — fast-judge pre-screen (opt-in): point `--fast-judge-url` at a\n  [laya_service.py](https://github.com/docsor1212/cite-holmes) endpoint (322M\n  distilled classifier, ~24ms/item on GPU) and references stuck in\n  semantic-limbo (no official text to escalate, no metadata to verify) get a\n  screening note before a human ever looks at them. Guardrails by design: it\n  NEVER produces or changes a verdict, never flags for review, and only\n  annotates high-confidence SUPPORTS leanings (calibrated threshold 0.95;\n  calibration curve ships in the repo). Off by default — zero behavior change\n  without the flag. Also: label-case normalization disclosed in benchmark\n  table, evaluation FAQ expanded.\n- **v3.4.1** — dedup fingerprint fix: batch deduplication now keys on\n  identifier + normalized title fingerprint, so two references sharing a\n  URL/DOI/PMID but carrying different titles (the classic stitched-fake\n  variant) are independently verified instead of silently inheriting each\n  other's verdict. Genuine duplicates (same identifier + same title) are\n  still merged with the partial downgrade.\n- **v3.4.0** — PMID metadata verification + identifier-free search: PMID-only\n  references now get full title/journal/year cross-checks against PubMed\n  registry data (NLM abbreviation-aware journal matching; genuine-PMID-fake-paper\n  splice attacks are caught as invalid), references without any identifier go\n  through OpenAlex + Semantic Scholar title search (confirmation -> partial,\n  not-found -> honest unverified for human review, never \"fabricated\"), and\n  BLUF key_numbers now reports \"reviewed N (retracted X)\" separately.\n- **v3.3.1** — title cleaning for metadata matching: reference lines are cleaned\n  (strip numbering, DOI/PMID/arXiv suffixes, author+year blocks) before title\n  similarity is computed against official registry metadata — demo-validated to\n  raise AlphaFold-style matches from 0.66 to 0.9+.\n- **v3.3.0** — the heterogeneous-panel release. Optional NLI third vote\n  (`--nli-url`): an independent NLI service watches the generative judge;\n  agreement records high confidence, disagreement marks human review without\n  flipping mechanical verdicts (conservative philosophy, mirroring L4).\n  Experimentally validated: the same panel lifts SciFact-Open zero-shot micro\n  from 0.484 to 0.617 with REFUTES F1 0-to-0.667. 252 tests green.\n- **v3.2.0** — the conformance release. JSON reports carry the BLUF block as a\n  first-class object, conformant with the published BLUF Report Specification v1.0\n  (five core keys, worst-actionable rule; `bluf_yaml` kept for compatibility) —\n  our own validator now passes our own output. Optional offline retraction checking:\n  `--retraction-cache <index.json>` built from the official Crossref/Retraction Watch\n  GitLab dump (63k+ DOIs) turns retraction flags into a local, zero-network lookup.\n  241 tests green.\n- **v3.1.0** — the closed-loop release. The L4 evidence cascade is now wired\n  into the main flow: references the semantic layer could not resolve\n  (not_in_source/unclear) escalate automatically to official PubMed abstracts or\n  arXiv full text, get passage retrieval (bge-m3, TF-IDF fallback) and a\n  configurable external judge. Judge findings annotate the report and never flip\n  verdicts; without --judge-url the run is byte-identical (zero network calls).\n  The judge client handles thinking models (qwen3 family): native ollama calls\n  use think:false + JSON-schema enum output, /v1 endpoints fall back from empty\n  content to the reasoning field (all verified live on SciFact/SCitance/CFEVER).\n- **v3.0.0** — the BLUF release. Every report opens with a machine-parseable\n  YAML block (verdict / key_numbers / blocker / next_action) plus a 5-line\n  human TL;DR — cite-holmes defines the \"3-second dual-reader\" report standard\n  (a gap no tool fills today). Scholarly reception: verified DOIs fetch\n  Semantic Scholar citation contexts (times-cited, context excerpts, coverage\n  honestly labeled — the free Scite.ai counterpart). L4 evidence cascade:\n  abstract-level judging via configurable external judge endpoints (OpenAI-\n  compatible: ollama / llama-server), NEI escalates to full-text passage\n  retrieval with local bge-m3 embeddings (TF-IDF fallback, zero hard deps).\n  Cross-language title guard (Chinese claims vs English registrations no\n  longer misjudged invalid). 210 tests green.\n\n- **v2.0.0** — MCP three-primitives server (`verify_references` + `explain_verdict`\n  tools, capability-matrix/changelog resources, fact-check workflow prompt),\n  arXiv multi-version notes (unversioned citations to multi-revision papers,\n  stale version pointers), optional `NCBI_API_KEY` (E-utilities 3→10 req/s for\n  parallel batches), and a per-reference **evidence chain** in HTML reports\n  (doi.org → retraction → S2 → OpenAlex → reachability, at a glance). First\n  Major: the transparent multi-source citation verification stack, as shipped.\n  195 tests green, real-network acceptance, package security self-scan clean.\n\n- **v1.13.0** — the discoverability release. EN description rewritten around\n  task-language roots (citation verification, verify citations, citation\n  checker, hallucinated references, fact check) so task-word searches on\n  skills.sh and agent skill markets finally surface it. New MCP server\n  (`mcp/server.py`, FastMCP): expose the same mechanical verification engine\n  as a `verify_references` tool for agents that prefer tools over skills —\n  official MCP Registry ready (publish-once to Smithery, mcp.so, PulseMCP).\n  Security & behavior declaration retained. 188 tests green, real-network\n  acceptance 5/5, package security self-scan clean.\n\n- **v1.12.0** — trust & transparency release. Persistent disk cache (default\n  on, 7-day TTL, local SQLite): re-running or extending a bibliography reuses\n  stable verdicts with zero network round-trips — the repeat-run pain behind\n  every Trust score. `--strict` always bypasses the cache; `--refresh-cache`\n  forces re-verification; `unreachable` is never cached. `--proxy` flag for\n  restricted-egress environments (env `HTTPS_PROXY` also honored). A\n  **Security & behavior declaration** section now ships in both languages:\n  network access limited to the official academic registries over HTTPS, no\n  telemetry, single-run CLI, nothing installed, env vars limited to your own\n  API keys. The capability matrix adds an explicit \"unsupported inputs\" line\n  (PDF/Word documents, Wanfang/CQVIP-only IDs, fact-checking claims).\n  Private maintainer tooling removed from the shipped package. 181 tests.\n\n- **v1.11.0** — the determinism release: arXiv API transient 406/403 (its anti-bot\n  window trips even under our 3s global pacing) no longer silently skips content\n  verification — exponential backoff, then a fallback to the official arXiv.org/abs\n  landing page for title comparison (0.50/0.82 thresholds), so the same references\n  produce the same verdicts on re-runs. Structured semantic audit: the model's\n  per-reference semantic judgment now travels as machine-readable workpapers\n  (`semantic: {claim, support, quote, note}` → `semantic_audit` in `--export\n  auditjson` and a dedicated report section); `not_in_source`/`contradicted`\n  judgments cap the verdict at `partial` with a human-review flag. Pre-submission\n  conclusion corrected to arXiv's actual May 2026 enforcement start (was June) and\n  now names ICML 2026 desk-rejections too. `tools/agentskills_check.py --distro`\n  prints a skills.sh/agentskills.io distribution readiness checklist. 166 tests\n  green. Zero new dependencies.\n\n- **v1.10.0** — the speed release, aimed squarely at the top evaluation\n  complaint (\"verification leans on overseas databases; flaky from CN\n  networks\"): **parallel verification** (`--workers`, default 4 — references\n  are independent, so a 20-item bibliography takes roughly a quarter of the\n  serial time; results stay index-ordered and the host circuit breaker /\n  rate-limiters are now thread-safe), **degraded-network mode** (3+ different\n  hosts failing at the transport layer in one run flips remaining lookups to\n  fast honest skips with a proxy/retry suggestion — minutes of waiting become\n  a fast verdict), **Semantic Scholar title-search confirmation** for\n  references with no DOI/PMID/arXiv (confirmation only, never a downgrade;\n  hyphens normalized to spaces per S2 API docs), **intra-run cache** (a\n  duplicate DOI/URL reuses the first verdict instead of re-hitting the\n  network), fast-fail (no retry) once the network is known-degraded, and an\n  OpenAlex 429 hint (key now required for production since 2026-02; free tier\n  is 100 credits/day). 147 tests.\n- **v1.9.0** — Semantic Scholar third source: when DOI.org metadata cannot be\n  fetched, the DOI is re-confirmed at the S2 Graph API (confirmation only,\n  never a downgrade — S2 record quality is uneven; 403/429 landing pages can\n  now be rescued on S2 confirmation too). BibTeX import: `--refs\n  bibliography.bib` verifies Zotero/EndNote exports directly (nested braces,\n  `\\url{}` macros, eprint→arXiv). OpenAlex API key support\n  (`--openalex-key`/`OPENALEX_API_KEY`) with an actionable 403 hint — OpenAlex\n  requires keys for production use since 2026-02. `--mailto` joins the\n  Crossref polite pool; 429 `Retry-After` is honored. Every report now opens\n  with a one-line **pre-submission conclusion** (arXiv has banned authors over\n  hallucinated references since 2026-06). `--export auditjson`: a\n  machine-readable per-reference check trail for institutional audit.\n  SKILL.md slimmed (details moved to `references/verification-details.md` +\n  `references/faq.md`). 120 tests.\n- **v1.8.1** — fix: OpenAlex bibliographic confirmation could silently\n  no-op (response read outside its context manager); caught by the new strict\n  real-network acceptance gate and corrected.\n- **v1.8.0** — retraction detection: verified DOIs are cross-checked against\n  the Crossref/Retraction Watch database (free, updated every working day);\n  a retracted paper is real — and uncitable — so retracted references are\n  capped at `partial` and flagged for human review. Bibliographic existence\n  check via OpenAlex for references without DOI/PMID/arXiv (positive\n  confirmation only; \"not found\" never means \"fabricated\"). UA-rotation retry\n  for metadata endpoints (403/406 resilience). New \"common mistakes\"\n  anti-patterns section and a 30-second quickstart. Capability matrix now\n  covers 10 fabrication/problem types. 89 tests.\n- **v1.7.0** — author-name consistency check against DOI-registered authors\n  (any claimed surname hit counts; a full miss downgrades to `partial`),\n  capability-boundary matrix in every report (9 machine-caught fabrication\n  types vs what stays with the semantic layer), `--format html`: a\n  self-contained single-file HTML report (zero external assets, mobile\n  readable — shareable with advisors/editors), agentskills.io spec self-check,\n  internal competitive landscape. 73 tests.\n- **v1.6.0** — host-level circuit breaker (2 consecutive transport failures to\n  the same host → skip for the rest of the batch with an honest note; broken\n  networks no longer stall the whole run), actionable error notes (timeout\n  suggests `--timeout 20`, DNS vs refused distinguished, human-review section\n  carries suggested actions), DOI↔PMID cross-check (catches \"stitched fakes\"\n  where both IDs are real but belong to different papers), journal-name\n  consistency check against DOI-registered metadata (mismatch → `partial`),\n  CN-network FAQ section, regression suite grown to 62 tests.\n- **v1.5.0** — arXiv ID metadata validation (export.arxiv.org official API,\n  DOI-grade similarity thresholds; nonexistent IDs → `invalid`, never\n  `unreachable`), Wayback Machine fallback for unreachable entries (archive\n  link attached for human review), `arxiv` reference field, `eprint` in BibTeX\n  + arXiv column in the audit CSV, regression suite grown to 53 tests.\n- **v1.4.0** — DOI metadata cross-check (DOI.org CSL JSON vs cited title/year —\n  catches \"real DOI, wrong paper\"), `--easy` one-flag mode (auto medical profile +\n  auto BibTeX/CSV export), transient-network retry, human-readable Chinese error\n  messages, expanded FAQ.\n- **v1.3.0** — CiteScore scorecard in every report (Markdown + JSON), official\n  icon, `tests/` regression suite (offline golden + medical profile + exports).\n- **v1.2.0** — medical evidence mode (`--profile medical`), PMID verification\n  via NCBI E-utilities, BibTeX/CSV bibliography export, 3-key dedup\n  (URL/DOI/PMID).\n\n## License\n\nMIT © DoctorQ Lab\n\nFile v3.11.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b6sjwend7cwwmhf7mdg1wzx82dfbf\",\n  \"slug\": \"cite-holmes\",\n  \"version\": \"3.11.0\",\n  \"publishedAt\": 1791566549713\n}\n\nFile v3.11.0:references/anti-patterns.md\n\n# 反模式统一清单（常见误用方式集中收录）\n\n> 收拢 README 与 SKILL.md 分散的反模式说明 + 新增易踩坑场景。上手前通读一遍。\n\n1. **不要用写草稿的同一个模型喂 `semantic` 字段**——语义封顶信任结构化模型判定，\n   自己审自己等于没审。\n2. **不要把 `verified` 理解为\"这篇论文支撑我的论断\"**——verified 只说明来源存在、\n   层级权威、登记元数据与声称一致；\"句子是否真被支撑\"归语义层。\n3. **MCP 单次批量不要超过 50 条**——超限被拒（分批调用）。\n4. **不要手改 `mcp/verify_refs.py`**——它是 scripts/ 的 sha 门禁副本，改 trunk 后跑\n   sync_engine.sh。\n5. **不要忽略报告中的 `needs_human_check` 旗标**——partial/unreachable 带复核指令\n   是有意设计。\n6. **不要在弱网把 `unreachable` 当\"假引用\"**——先用 `--cn`/代理重跑再下结论。\n7. **不要为求星而自动化点击**（平台红线）——报告页脚的 Star/收藏链接面向人类读者。\n8. **不要在 `--check-document` 里省略 `--refs`**——上下文核验需要清单核验结果做绑定。\n\n## 触发词误触发怎么办\n\n- 触发词（深度研究/引用核查/查证等）按语义匹配，短句如\"查证一下\"若被误触发，\n  直接说明真实意图即可（\"我只要翻译\"）；不产生副作用。\n- 若希望某类请求**不**触发引用核验，避免在请求中同时包含\"研究/查证/引用\"词与\n  文献清单。\n9. **不要对中英混合引用的锚词率过度报警**——跨语言（CJK vs 拉丁）混合引用的\n   锚词率天然偏低（内容词不相交），上下文核验会标\"疑似错配\"；这是预期行为\n   （清单核验侧已有跨语言守卫降 partial 转人工），按人工复核处理而非自动判死。\n\nFile v3.11.0:references/faq.md\n\n# FAQ (moved out of SKILL.md in v1.9 to keep the main guide short)\n\n**Q: Is a `verified` citation guaranteed real?**\nNo. The mechanical layer catches machine-checkable fakes: nonexistent DOIs,\nfabricated PMIDs (via E-utilities), fabricated arXiv IDs (via the official\narXiv API), dead links, duplicates, stitched DOI↔PMID pairs, wrong-paper\nDOIs (title similarity), and retracted papers. A fabricated citation pointing\nto a real, plausible page still relies on the semantic layer — never a\nguarantee.\n\n**Q: How are arXiv references checked?**\nSame treatment as DOIs since v1.5: the official export.arxiv.org API returns\nthe registered title/year; similarity below 0.50 → `invalid` (fabricated or\nwrong ID), 0.50–0.82 → `partial` (human review), above → consistent. An ID the\nAPI has never heard of is `invalid`, not `unreachable`.\n\n**Q: What do I do with `unreachable` references?**\nOpen them manually once — the report automatically attaches a Wayback Machine\narchive link when one exists. If it matters and keeps failing, switch sources.\n\n**Q: Did the PMID check fail because of my network?**\nNo. \"不存在/编造\" verdicts come from the official E-utilities API — that means the\nPMID genuinely isn't in PubMed. Only \"校验失败\" wording means a network/API issue\n(the reference is then treated as reachable and flagged).\n\n**Q: Do I always need `--profile medical`?**\nUse `--easy`: one flag auto-detects medical references, enables the medical\nprofile, and exports BibTeX + CSV automatically. For real clinical\nconclusions, still prefer explicit `--profile medical`.\n\n**Q: Can verified references go straight into my paper?**\nYes: `--export bibtex` exports only `verified` entries (with PMID / arXiv\neprint notes) and imports into Zotero/EndNote; `--export csv` is the full\naudit ledger for advisors and editors; `--export auditjson` (v1.9) is the\nmachine-readable per-check working paper for institutional audit.\n\n**Q: Which services does verification touch, and what if my network is slow or blocked?**\nDOI.org (DOI metadata), NCBI E-utilities (PMID), export.arxiv.org (arXiv\nmetadata), archive.org (archived copies), api.crossref.org (retractions),\napi.openalex.org (bibliographic confirmation), api.semanticscholar.org (third\nsource + title search, v1.9/v1.10). CN networks occasionally throttle several\nof them. The verifier never stalls: transient failures retry once (429 honors\n`Retry-After`), error notes tell you the cause with a fix (`--timeout 20` for\nslow links), and a host-level circuit breaker skips a host after 2 consecutive\ntransport failures — with an honest note — instead of hanging the batch.\nSince v1.10, references are verified in parallel (`--workers 4` by default —\nroughly a quarter of the serial wall time), and when 3+ different hosts fail\nat the transport layer in one run (restricted-egress signature) remaining\nlookups are fast-skipped with a \"全局网络降级\" note and a proxy/retry\nsuggestion. Degraded checks are always labeled, never silently dropped.\n\n**Q: Do I need API keys?**\nNo — everything runs keyless within public rate limits. OpenAlex has required\nkeys for *production* use since 2026-02 (free daily allowance; this skill\nsurvives keyless but benefits from `--openalex-key` / `OPENALEX_API_KEY`).\nSemantic Scholar keyless shares a rate pool; `--s2-key` / `S2_API_KEY` gets a\ndedicated lane. Crossref asks heavy users to identify themselves — pass\n`--mailto you@lab.edu` to join the polite pool.\n\n**Q: MCP 模式是什么？agent 不装 skill 怎么用？**\nv2.0.0 起仓库内置 `mcp/server.py`（FastMCP）：tools `verify_references`（机械验证）/\n`explain_verdict`（判定解释）、resources（能力矩阵/版本史）、prompts（fact-check 工作流）。\n推荐直跑：`uvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp`；\n或自行准备 FastMCP 环境后运行 `python mcp/server.py`。零密钥可用；可选 API key 经环境变量注入。\n\n**Q: arXiv 论文有多个修订版怎么办？**\nv2.0.0 自动备注：引用未指定版本而论文存在多个修订版 → 提示\"引用未指定版本\"；\n引用指向旧版而存在更新版 → 提示最新版号（旧版可能含未修正内容）。同响应内取数，零额外请求。\n\n**Q: 验证要访问哪些外部服务？国内网络慢怎么办？**\n见上方服务清单——全部是官方学术注册库、全部 HTTPS。v1.12 起两件事让它在国内更可用：\n① **持久磁盘缓存**默认开（7 天 TTL）：同一批引用复跑直接复用稳定判定、零外呼，\n`--refresh-cache` 强制重验；② **`--proxy http://host:port`** 显式走代理（或设\n`HTTPS_PROXY` 环境变量），配合网络降级模式，受限出口下也能一次跑完拿到诚实结论。\n\n**Q: 判定结果会被缓存吗？撤稿状态更新了怎么办？**\n缓存只存 verified/partial/invalid 三种稳定判定（TTL 默认 168 小时），\nunreachable 永不入缓存（瞬态）；`--strict` 模式强制绕过缓存读（CI 诚实）。\n撤稿状态这类可变信息以 TTL 为界——对时效敏感的批次用 `--refresh-cache` 强制重验。\n\n**Q: Can I verify references straight from my reference manager?**\nYes, since v1.9: `--refs bibliography.bib` reads Zotero/EndNote/JabRef BibTeX\nexports directly (entry types, nested braces, `\\url{}` macros; DOI/PMID/arXiv\nfields are picked up automatically).\n\n## arXiv 条目 note 里出现「元数据 API 不稳，经官方着陆页标题比对确认」正常吗？\n\n正常。arXiv 的 API 对批内多请求有反机器人窗口（偶发 406/403），与你的网络无关；\n验证器已自动回退官方着陆页完成同等强度的标题比对，判定可信度不受影响。\n\n## semantic 字段写什么？\n\n模型逐条判定\"来源是否真的支撑所引论断\"后，写成结构化对象\n`{\"claim\", \"support\", \"quote\", \"note\"}`（五值：supported / partial /\nnot_in_source / contradicted / unclear）。`not_in_source` 与 `contradicted`\n会被机械验证器封顶为 `partial` 并转入人工复核区——来源存在不等于来源认同。\n旧的字符串写法仍接受（仅注记）。\n\n## Q: fast-judge 预筛会改我的判定吗？（v3.5.0）\n不会。`--fast-judge-url` 外挂的 322M 蒸馏分类器只对\"语义待定且无官方文本可升级\"\n的引用追加「倾向支持（供人工参考）」注记——永不产生/修改五态判定，永不标记人工\n复核，默认关闭。校准曲线随仓库发布（阈值 0.95）。\n\n## Q: GB/T 7714-2025 导出的数据来源是什么？（v3.6.0）\n`--export gbt7714` 仅导出 verified 条目，且每条以**注册库登记的 CSL 元数据**\n为准（作者按国标缩写惯例预格式化、期刊/年/卷期页来自 DOI.org/Crossref 登记），\n不是你手填的字段。无注册元数据的电子资源按 [EB/OL] + 引用日期输出。\n\n## Q: 国内网络直连核验总超时怎么办？（v3.6.0）\n加 `--cn`：超时下限抬到 25 秒，且 DOI.org 三试全败时自动回源\napi.crossref.org（独立主机，同构 CSL 元数据，报告注记会写明「经 Crossref\n回源」）。仍不通时 `--proxy` / `HTTPS_PROXY` 继续兜底；`unreachable` 不等于\n假引用，换网重跑后再下结论。\n\n## Q: 能核验\"正文里的引用\"而不只是参考文献清单吗？（v3.7.0）\n能。`--check-document paper.md --refs refs.json`：解析正文 in-text 标记\n（[12]/[1-4]/(Author, Year)/doi.org 链接），把每处引用绑定到清单条目并计算\n\"引文句 ↔ 被引标题\"锚词率——低于 0.34 标记为疑似错配引文（引 A 文却引了 B 句）。\n锚词率低≠错引，是语义层人工复核候选；语义判断仍由 agent 负责。\n\n## Q: 触发词误触发了怎么办？\n触发词按语义匹配，可能被\"查证一下\"这类短句误触发。直接说明真实意图即可\n（例如\"我只要翻译这段\"），误触发不产生任何副作用。想让某类请求**不**触发\n引用核验：避免同时出现\"研究/查证/引用\"词与文献清单。\n\n## Q: 中英混合引用（如中文声称+英文登记）会误判吗？\n不会判 invalid。v3.0.0 起内置跨语言守卫：声称标题与登记标题跨语言（CJK vs\n拉丁）时相似度不作编造依据，降 `partial` 转人工复核（与期刊名核查同款守卫）。\n已知边界：跨语言混合引用的锚词率（上下文核验，v3.7.0）同样会偏低——属预期，\n按\"疑似错配候选\"人工复核即可。\n\nFile v3.11.0:references/INDEX.md\n\n# references/ 索引（v3.8.0 新增）\n\n| 文件 | 用途 | 何时读 |\n|---|---|---|\n| faq.md | 高频问题解答（网络/PMID/导出/API key/fast-judge/GB·T 导出/--cn/上下文核验） | 遇到具体疑问时 |\n| anti-patterns.md | 反模式统一清单（误用方式集中收录，v3.8.0 从 README/SKILL 收拢） | 上手前通读一遍 |\n| medical-mode.md | 医学证据模式详解（信源金字塔/CEBM 分级/监管信源） | 使用 --profile medical 前 |\n| verification-details.md | 五态判定完整阈值规则与检查明细 | 需要理解判定依据时 |\n| search-strategies.md | 中英双语检索策略矩阵 | 定制检索流程时 |\n| skillhub-mcp.md（如存在） | MCP 集成参考 | 配置 MCP 前 |\n\n> 主文档：SKILL.md（技能定义）/ README.md（用户文档）。本目录文件按需加载，\n> 不进入主执行流。\n\nFile v3.11.0:references/medical-mode.md\n\n# 医学研究模式（MEDICAL 模式必读）\n\n适用触发：临床问题、药物疗效/安全性、疾病患病率、诊疗指南、诊断标准、\nmeta 分析、循证检索、医学文献综述。拿不准是否算\"医学\"时：凡涉及患者\n治疗决策的一律按医学模式处理。\n\n## 启用方式\n\n检索阶段照常（菱形扩展 + 双语矩阵），验证阶段加参数：\n\n```bash\npython scripts/verify_refs.py --refs research_refs.json --profile medical --export bibtex,csv\n```\n\n## 与通用模式的四点差异\n\n1. **期刊层域名扩展**：Cochrane Library、BMJ Best Practice、Embase、\n   ClinicalTrials.gov、ChiCTR、NMPA、CDC/中国疾控、NICE、万方、医脉通指南\n   在医学模式下计入权威期刊层。\n2. **社区层降级警示**：知乎/微博/公众号/维基/博客层来源自动标注\n   \"不得支撑医学结论（仅作线索）\"——医学主张的引用底线比通用问题更高。\n3. **PMID 存在性核实**：PubMed 网页对不存在的 PMID 也返回 2xx，HTTP 检查\n   抓不住编造的 PMID。验证器经 NCBI E-utilities 逐条核实 PMID 真实存在。\n   引用登记时 pmid-only 条目可直接写 `{\"pmid\": \"36443570\", ...}`，无 url/doi\n   也能验证（自动解析为 PubMed 页面）。\n4. **台账导出**：`--export bibtex,csv` 产出 verified-only 参考文献（BibTeX，\n   直接导入论文管理器）+ 全量审计台账（CSV，Excel 打开）。\n\n## 医学信源金字塔\n\n| 层级 | 首选来源 | 用法 |\n|---|---|---|\n| 系统评价/Meta | Cochrane Library；PubMed 上的 SR/MA（`systematic[sb]` 过滤） | 可直接支撑疗效结论 🟢 |\n| 临床指南 | WHO、NICE、中华医学会各分会、国家卫健委、UpToDate / BMJ BP | 支撑诊疗规范类结论 🟢 |\n| 原始研究 | PubMed / Embase（RCT > 队列 > 病例系列）；知网/万方/维普（中文） | 按研究设计定 🟢/🟡 |\n| 预印本 | medRxiv / bioRxiv | 只标 🟡 并注明\"未经同行评审\" |\n| 权威媒体 | Reuters、丁香园（资讯级） | 背景信息，不作疗效证据 |\n| 社区/社交 | 知乎、微博、病友群、公众号 | 仅作线索，禁止支撑任何医学结论 🔴 |\n\n## 检索式要点（PubMed）\n\n- MeSH 主题词 + 自由词组合：`(\"Tocilizumab\"[Mesh]) AND (\"Arthritis, Juvenile\"[Mesh] OR sJIA)`\n- 疗效问题加过滤：`AND (randomized controlled trial[pt])`；病因/患病率换相应过滤\n- 中英双语：英文 PubMed，中文知网/万方，两边合并去重（验证器按\n  URL/DOI/PMID 自动去重）\n- 双源规则在医学场景升级：**疗效结论需 ≥1 个系统评价或 ≥2 个独立 RCT**；\n  仅有病例系列支撑的疗效主张标 🟡 并明示证据等级\n\n## 监管与注册信源\n\n- 药械批件/说明书：NMPA（中国）、FDA、EMA\n- 临床试验注册：ClinicalTrials.gov、ChiCTR（中国临床试验注册中心）\n- \"某药是否获批某适应证\"\"某试验是否注册\"类事实，优先用监管/注册源核实——\n  这类事实比文献更容易被 AI 编造，且编造得最像真的\n\n## 证据分级对照\n\n本 skill 的置信度标记与 CEBM 2009 证据等级的大致映射：\n\n| 标记 | CEBM 等级 | 典型来源 |\n|---|---|---|\n| 🟢 | 1a–2b | SR/MA、高质量 RCT、官方指南推荐 |\n| 🟡 | 3–4 | 队列/病例对照/病例系列、单个专家意见 |\n| 🔴 | 5 或存疑 | 传闻、社区层、利益相关方声明、来源冲突 |\n\nFile v3.11.0:references/report-template.md\n\n# 报告模板与引用 schema（SYNTHESIZE 阶段照此输出）\n\n## research_refs.json schema（VERIFY 阶段的输入）\n\n研究过程中实时登记引用，每条：\n\n```json\n[\n  {\n    \"id\": 1,\n    \"title\": \"页面/文档标题\",\n    \"url\": \"https://…\",\n    \"source\": \"来源站点或机构（如 Anthropic / Nature / Reuters）\",\n    \"year\": 2026,\n    \"claim\": \"本报告用这条引用支撑的那句话\",\n    \"triples\": [{\"subject\": \"...\", \"relation\": \"...\", \"object\": \"...\", \"span\": \"原文出处片段\"}],  # L1 原子声明分解(v3.0.0)\n    \"semantic\": {\n      \"claim\": \"与上层 claim 相同（或更细）\",\n      \"support\": \"supported\",   // supported | partial | not_in_source | contradicted | unclear\n      \"quote\": \"来源页原文关键句（人工复核抓手）\",\n      \"note\": \"可选备注\"\n    },\n    \"found_via\": \"search#3\",          // 哪轮搜索发现的，便于回溯\n    \"tier\": \"official\"                // official | journal | preprint | media | community | blog | social\n  }\n]\n```\n\n`semantic` 字段由模型在语义验证时填写（v1.11 起推荐上面的结构化对象形式；旧的\n字符串形式仍然接受，仅作注记不参与判定）。验证器把结构化判定带进\n`--export auditjson` 的 `semantic_audit` 工作底稿与报告语义区；`not_in_source` /\n`contradicted` 会把机械判定封顶为 `partial` 并转人工复核（来源存在 ≠ 来源认同）。\n机械验证由 `scripts/verify_refs.py` 完成，两层的结果都要体现在最终引用表里。\n\n## 置信度标记规范\n\n| 标记 | 含义 | 判定标准 |\n|---|---|---|\n| 🟢 | 高置信 | ≥2 个独立来源一致，且至少一个是一手/权威源 |\n| 🟡 | 中置信 | 单一权威源，或双源但均为二手 |\n| 🔴 | 存疑 | 来源冲突、仅社区层证据、或数据过旧 |\n\n标记挂在结论句末尾 + 引用编号，如：`缓存命中率提升约 40% 🟢[1][3]`。\n\n## 报告骨架\n\n```markdown\n# {研究问题}\n\n> 研究模式：FULL/QUICK · 检索 N 轮 · 引用 M 条（verified X / partial Y / …）\n> 完成日期：YYYY-MM-DD · 覆盖时间窗：…\n\n## 执行摘要（≤200 字，先给答案）\n{直接回答研究问题的主要发现，2–4 条，每条带置信度标记}\n\n## 主要发现\n### 子问题 ①：{…}\n{结论句 🟢[引用号]。证据展开：数字/事实 + 出处定位。}\n{与结论相悖的证据如有，必须写。}\n\n### 子问题 ②：{…}\n…\n\n## 分歧与存疑\n{来源打架的地方：各自说法 + 各自来源 + 可能的成因（数据新旧/口径不同/立场差异）}\n\n## 研究缺口\n{预算内没能回答的问题、抓取失败的来源、需要付费/权限才能核实的数据}\n\n## 后续值得追问\n{2–3 个基于本次发现自然延伸的问题}\n\n## 引用清单\n| # | 标题 | 来源 | 年份 | 语义 | 机械验证 | 支撑论断 |\n|---|---|---|---|---|---|---|\n| 1 | … | … | … | supports | verified | … |\n\n### 待人工复核（unverified / unreachable）\n| # | 标题 | URL | 状态 | 原因 |\n```\n\n## 输出要求\n\n- 正文结论**只能**引用 `semantic.support=supported`（旧字符串 `supports`）且机械验证非 `invalid` 的条目；`not_in_source`/`contradicted` 条目已由验证器封顶 `partial`，只可作对照线索。\n- `partial_support` 只能支撑结论中它确实支持的那半句，并在句中注明。\n- `unverified/unreachable` 一律只出现在\"待人工复核\"分区。\n- QUICK 模式可省略\"子问题\"分节，直接：执行摘要 → 核查结论（含反证）→ 引用清单。\n- 报告语言跟随用户提问语言；术语首次出现给中英对照。\n\nFile v3.11.0:references/search-strategies.md\n\n# 检索策略（SEARCH 阶段必读）\n\n## 菱形扩展模式\n\n检索节奏是\"菱形\"：**宽 → 窄 → 补缺**。\n\n1. **宽**（第 1–2 轮）：用问题的自然表述直接搜，目的是摸清领域地形——主流说法、关键术语、\n   主要玩家。此时发现的**术语**比**答案**更重要（下一轮的搜索词从这里来）。\n2. **窄**（第 3–N 轮）：用上轮挖出的精确术语/产品名/人名/事件名构造查询，逐个子问题收网。\n   搜索词要具体到\"能命中单一事实\"的程度。\n3. **补缺**（最后 1–2 轮）：对照 PLAN 里的子问题清单，只搜还没有双源确认的缺口。\n   缺口补不上就停止——剩余缺口写进报告，不硬凑。\n\n## 搜索词矩阵\n\n对核心子问题，按三个维度变化搜索词，避免同义反复地浪费预算：\n\n- **语言维**：中文术语 ↔ 英文术语（同一概念的中英文名经常不同，如\"深度研究\" vs \"deep research\"）\n- **角度维**：正面（\"XX 优势\"）↔ 反面（\"XX 问题/翻车/争议/批评\"）——只搜正面会得到软文视角\n- **粒度维**：综述型（\"XX 进展综述 2026\"）↔ 事实型（\"XX 具体数字/日期/版本\"）\n\n## 信源金字塔（可信度从高到低）\n\n| 层级 | 例子 | 用法 |\n|---|---|---|\n| 一手/官方 | 官方文档、发布公告、论文原文、财报、法规原文 | 可直接支撑结论 |\n| 权威二手 | 权威媒体（Reuters/AP/新华等）、行业头部分析 | 可支撑，尽量回溯到一手 |\n| 社区/百科 | GitHub、Stack Overflow、Wikipedia、知乎高赞 | **只作线索**：从这里找到一手源再验证 |\n| 社交/自媒体 | X/Twitter、微博、个人博客、公众号 | 仅作趋势感知，不得单独支撑任何结论 |\n\n- 社区层发现的关键论断 → 找到它的一手出处 → 用一手源作为引用。\n- 找不到一手出处的社区论断 → 如标注明\"仅社区层证据 🔴\"。\n\n## 关键事实的双源规则\n\n以下进入报告前必须有两个独立来源（不同站点，不是互相转载）确认：\n\n- 具体数字（价格、百分比、性能数据、装机量）\n- 日期与版本号\n- 引语（谁说了什么）\n- \"最/第一/唯一\"类断言（这类断言即使双源也标 🟡，除非来源是一手）\n\n两源打架 → 报告里如实写分歧，标 🔴，注明各自来源与时间（旧数据 vs 新数据是常见原因）。\n\n## 页面抓取（WebFetch）纪律\n\n- 每次抓取前明确要找什么；抓完立即记录\"找到了什么/没找到什么\"。\n- 搜索结果摘要 ≠ 页面内容。凡是要写进报告的具体细节，必须抓到页面里确认。\n- 单页抓取失败（超时/反爬）→ 换一个来源覆盖同一子问题，不无限重试同一 URL。\n- 优先抓：官方文档 > 公告原文 > 权威媒体正文。\n\n## 搜索预算台账（边搜边记）\n\n```\n| 轮 | 查询 | 目标子问题 | 收获 | 新缺口 |\n|---|---|---|---|---|\n| 1 | … | ① | 关键术语 X、Y | Y 的数据缺 |\n```\n\n到达预算上限即停，剩余缺口如实写进报告\"研究缺口\"一节。\n\nFile v3.11.0:references/verification-details.md\n\n# Verification details (v1.10)\n\nFull semantics of every check `verify_refs.py` performs. SKILL.md keeps the\nshort version; this file is the complete reference.\n\n## Cross-check rules by identifier\n\n- **DOI** (v1.4+): DOI.org content negotiation returns the registered CSL\n  metadata. Title similarity < 0.50 → `invalid` (\"real DOI, wrong paper\" —\n  fabricated or mis-cited); 0.50–0.82 → `partial` (human review); ≥ 0.82 →\n  consistent. Year difference ≥ 2 adds a warning note. A DOI the registry has\n  never heard of (404) → `invalid`, never `unreachable`.\n- **PMID**: NCBI E-utilities esummary. PubMed's own web pages return 2xx even\n  for nonexistent PMIDs, so only the API actually catches fabrications.\n- **DOI ↔ PMID cross-check** (v1.6): when both keys are present, the PMID's\n  registered DOI must match the claimed DOI — a mismatch is `invalid`\n  (\"stitched fake\": each key real, pointing at different papers).\n- **Journal name** (v1.6): claimed source vs DOI-registered container-title.\n  NLM-style abbreviations are treated as equivalent (word-initial sequential\n  match, stop-words skipped); cross-language names are not comparable and are\n  skipped. A real mismatch downgrades to `partial` (journals do rename).\n- **Author names** (v1.7): claimed surnames matched against the CSL author\n  list; any hit counts (spelling variants tolerated); single-letter initials\n  never substring-match (APA \"Torvalds, L.\" would otherwise match any \"G.\").\n  CJK-vs-Latin names are skipped (not comparable). A full miss downgrades to\n  `partial`.\n- **arXiv** (v1.5): official export.arxiv.org API (≥3 s/request politeness\n  pause). Unknown ID or API error element → `invalid`. Same 0.50/0.82\n  similarity thresholds as DOI.\n- **Retraction** (v1.8): Crossref REST API `updated-by[]` (Retraction Watch\n  data, updated daily). Any `type == \"retraction\"` → capped at `partial` +\n  human-review flag. Check failures are silently skipped — never penalize a\n  reference for the checker's network.\n- **OpenAlex bibliographic check** (v1.8, v1.9 keys): for references without\n  DOI/PMID/arXiv, title search against OpenAlex. Best of full/prefix title\n  similarity ≥ 0.82 → \"confirmed exists\"; 0.60–0.82 → \"near match, review\";\n  below → honest \"no obvious match, not evidence of fabrication\" (coverage\n  lag). HTTP 403 → actionable hint to configure `--openalex-key` /\n  `OPENALEX_API_KEY` (OpenAlex requires API keys for production use since\n  2026-02; keyless calls still work within a free allowance).\n- **Semantic Scholar third source** (v1.9): when a DOI reference's DOI.org\n  metadata could not be fetched (network restricted / transient 5xx), the DOI\n  is re-queried at the S2 Graph API. Confirmation (title similarity ≥ 0.82)\n  adds a second-source note and enables the anti-bot rescue below. S2 record\n  quality is uneven (real papers registered as journal ToC pages have been\n  observed), so **S2 never speaks negatively**: mismatches and misses are\n  silent. Optional `--s2-key` / `S2_API_KEY` (keyless shares a rate pool;\n  client paces ≥1.5 s between S2 calls).\n- **Semantic Scholar title search** (v1.10): references carrying neither\n  DOI/PMID nor arXiv get one extra positive signal — the title is queried at\n  the S2 `/paper/search` endpoint (best-of full/prefix similarity ≥ 0.82 →\n  \"confirmed exists\"). Same discipline as above: confirmation only, misses\n  and mismatches silent (titles < 18 chars are skipped — search noise).\n  Per S2 API docs, hyphenated query terms return no results, so hyphens are\n  normalized to spaces before encoding.\n\n## Parallel verification & degraded-network mode (v1.10)\n\n- `--workers N` (default 4): references are independent, so up to N of them\n  are verified concurrently (offline mode and `--workers 1` stay serial).\n  Progress lines print in completion order; the report/JSON keep strict\n  index order. Shared state (host circuit breaker, arXiv/S2 rate pacing) is\n  lock-protected, so politeness guarantees hold under parallelism. The\n  per-reference `--interval` pacing applies to the serial path; the parallel\n  path throttles through concurrency itself.\n- **Degraded-network mode**: when ≥ 3 *different* hosts fail at the transport\n  layer within one run (a restricted-egress signature), remaining lookups to\n  not-yet-confirmed hosts are skipped immediately with an honest\n  \"global network degraded\" note plus a one-line suggestion (configure a\n  proxy / switch networks and rerun), and the run prints a summary advice\n  line. Site-level HTTP responses (403/429 rate limits, anti-bot) never\n  count toward degradation — the egress path itself is proven working when\n  any server answers. Once degraded, transport retries stop (fast-fail).\n  The per-host v1.6 circuit breaker still applies independently.\n- **Intra-run cache**: a reference whose DOI/URL/PMID/arXiv matches an\n  already-verified one reuses that verdict (note says so) instead of\n  re-hitting the network. `mark_duplicates` still applies afterwards — the\n  cache never promotes a duplicate.\n\n## Rescue rules (anti-bot, anti-flaky-network)\n\n- Landing page 403/429 but registry metadata confirmed (DOI.org, arXiv, or\n  S2): existence is proven, so the reference is NOT `unreachable` — it is\n  judged on tier/fields like any 200 page (trusted tier + complete fields →\n  `verified`; any partial adjustment or missing field → `partial`). This\n  prevents \"metadata confirmed + 403 blog page\" from slipping through.\n- arXiv resolver 404 with unconfirmed metadata → `invalid` (official registry\n  says no such ID).\n- Wayback Machine (v1.5): every `unreachable` link is checked against\n  archive.org; an archived copy adds a dated comparison link for human review.\n\n## Rate limiting & resilience\n\n- Host circuit breaker (v1.6): 2 consecutive *transport* failures\n  (timeout/DNS/refused) → remaining calls to that host are skipped for the\n  batch with an honest note. HTTP responses (403/404/429…) never trip it.\n  Retries within one logical call count once.\n- Transient failures retry once; error notes carry the cause (timeout / DNS /\n  refused) and the fix (`--timeout 20` for slow links).\n- 429 with `Retry-After` (v1.9): the DOI metadata check waits the requested\n  duration (capped at 5 s) before its retry.\n- `--mailto you@lab.edu` (v1.9): appended to Crossref requests — enters the\n  Crossref polite pool with more generous rate limits. Recommended for CI and\n  institutional batches. arXiv politeness (3 s/request) and S2 pacing (1.5 s\n  keyless) are always on — both are global under parallel verification.\n- OpenAlex hints (v1.10): 403 and 429 both explain the key situation —\n  OpenAlex has required API keys for production use since 2026-02 (keyless:\n  100 credits/day); configure `--openalex-key` / `OPENALEX_API_KEY`.\n\n## Exports\n\n- `--export bibtex`: verified-only bibliography, Zotero/EndNote-ready (PMID /\n  arXiv eprint noted). Field values sanitized (brace/newline stripped).\n- `--export csv`: full audit ledger (all refs, all fields, formula-injection\n  neutralized, utf-8-sig for Excel).\n- `--export auditjson` (v1.9): machine-readable working paper — per reference,\n  which checks ran and what they found (`checks_catalog` documents every check\n  type), plus a generation timestamp. Purpose: transparency and institutional\n  audit. The July 2026 arXiv survey of citation checkers (arXiv 2607.22693)\n  found none of five evaluated tools reliable enough to run unsupervised and\n  calls for \"transparent multi-source detection systems\" — the audit trail is\n  this skill's answer: a human can re-trace every mechanical verdict.\n\n## Persistent disk cache (v1.12)\n\n`--cache` is on by default. Stable verdicts (`verified` / `partial` /\n`invalid`) are stored in a local SQLite database\n(`~/.cache/cite-holmes/cache.sqlite3`, standard library) keyed by every\njudgment-relevant input: identifier, claimed title/year/source/authors, tier,\nverifier version, and profile — the same DOI with a different claimed title is\na different cache key, because title-similarity verdicts depend on the claim.\nTTL is 168 h (`--cache-ttl`); `unreachable` is never cached (transient);\n`--offline` results are never cached; `--strict` bypasses cache reads so CI\nalways runs live; `--refresh-cache` bypasses reads but still writes;\n`--no-cache` disables the store entirely. Cache hits append an explicit\n\"磁盘缓存命中\" note with the age, so an audited report always shows when a\nverdict came from the network and when from the local store. What this buys:\nre-running or extending a bibliography no longer re-travels the ocean —\nthe repeat-run pain flagged in every evaluation's Trust score.\n\n## Proxy support (v1.12)\n\n`--proxy http://host:port` (scheme optional) routes all verification traffic\nthrough the given proxy; without the flag, urllib's standard\n`HTTP_PROXY`/`HTTPS_PROXY` environment handling applies. Combined with the\nv1.10 degraded-network mode, restricted-egress environments get one fast,\nhonest pass instead of minutes of stalls.\n\n## BibTeX input (v1.9)\n\n`--refs bibliography.bib` is auto-detected (extension `.bib`, or content\nstarting with `@` when JSON parsing fails). Minimal pure-stdlib parser:\n`@article`/`@inproceedings`/`@book`/`@misc` and other entry types, nested\nbraces, quoted and bare values, `@comment`/`@preamble`/`@string` skipped.\nMapped fields: title, author (` and ` → `, `), year, journal/booktitle/\njournaltitle/publisher → source, doi, url (including `\\url{...}` macros in\nhowpublished/note), pmid, eprint+archivePrefix → arxiv. Unparseable entries\ndegrade to missing-field references and are judged honestly by the normal\npipeline; zero entries → actionable error.\n\n## Verdict meanings (unchanged since v1.3)\n\n| Verdict | Meaning | Score |\n|---|---|---|\n| `verified` | reachable + trusted tier + fields complete | +10 |\n| `partial` | reachable but community tier / field or metadata concerns | +4 |\n| `unreachable` | fetch failed — needs human check, ≠ nonexistent | 0 |\n| `unverified` | skipped (`--offline`) | −2 |\n| `invalid` | no identifier, registry says no, or stitched fake | −8 |\n\n## arXiv 第二源 fallback（v1.11）\n\nexport.arxiv.org API 对批内多请求偶发 406/403（反机器人窗口；3 秒全局控频可降频\n率、不能归零，实测单发全 200、批内可稳定 406）。v1.10 及之前 API 失败即静默跳过\n内容核验，同一引用两次运行会漂移（一次 verified 一次 partial）。v1.11 起：\n\n1. API 首败后指数退避（3s 时隙 + 5s 退避）重试一次；\n2. 仍失败 → fallback 拉官方着陆页 `arxiv.org/abs/<id>`，剥掉 `[ID] ` 前缀后做\n   标题相似度比对（0.50/0.82 阈值与 API 路径一致）；\n3. 着陆页 404 → `invalid`（官方库查无）；着陆页也不可达 → 维持可达性路径判定，\n   note 明示「元数据核验未完成」，不静默。\n\nAPI 明确回答（Atom feed 正常返回）时行为不变：查无记录/格式错误 → `invalid`。\n非 Atom 响应（反爬 HTML 页，常是合法 XML）**不作为查无依据**，走 fallback——\n防止瞬态反爬把真论文误杀成编造。\n\nArchive v3.10.0: 29 files, 197555 bytes\n\nFiles: _meta.json (131b), docs/best-practices.md (2552b), examples/demo_refs.json (1780b), examples/end-to-end/exports/report_GB-T7714.txt (156b), examples/end-to-end/exports/report.csv (566b), examples/end-to-end/paper_excerpt.md (56b), examples/end-to-end/README.md (1244b), examples/end-to-end/report.json (4979b), examples/end-to-end/report.md (3102b), examples/end-to-end/research_refs.json (453b), examples/medical_refs.json (2647b), mcp/MCP_SUBMISSION.md (2470b), mcp/pyproject.toml (771b), mcp/README.md (3741b), mcp/server.py (14242b), mcp/sync_engine.sh (919b), mcp/verify_refs.py (202016b), README.md (28302b), references/anti-patterns.md (1857b), references/faq.md (8527b), references/INDEX.md (866b), references/medical-mode.md (3450b), references/report-template.md (3645b), references/search-strategies.md (3120b), references/verification-details.md (11150b), scripts/verify_refs.py (202016b), skill-card.md (2267b), SKILL.md (14215b), tools/agentskills_check.py (4549b)\n\nFile v3.10.0:SKILL.md\n\n---\nname: cite-holmes\nversion: 3.10.0\nauthor: DoctorQ Lab\nlicense: MIT\ndescription: >-\n  Deep research that interrogates its own sources (Verified Deep Research):\n  calibrates scope first (3-5 sharp questions), then searches iteratively and\n  machine-verifies every citation. Also a standalone citation checker: paste\n  any reference list and it runs full citation verification (it will verify\n  citations before you cite) against official registries —\n  hallucinated references, fabricated DOIs, fake PMIDs, arXiv IDs, stitched\n  fakes, retracted papers; a fact check for your bibliography. Five verdicts;\n  unverified references never masquerade as real (AI hallucination detection).\n  Medical mode (Cochrane/BMJ/ChiCTR/NMPA/CDC/NICE presets, PMID check) and\n  bibliography export (BibTeX/GB·T 7714-2025/RIS/CSV + JSON workpaper) built\n  in. Trigger matching is semantic, not exact — mis-triggers are harmless;\n  state your real intent to avoid them.\nwhen_to_use: >-\n  Use when the user says \"deep research\", \"look into\", \"investigate\",\n  \"compare A vs B\", \"fact check\", \"verify this claim\", \"is it true that...\",\n  \"check these references\", \"are these citations real\", wants a research\n  report with sources, a literature review, a medical evidence lookup, or\n  wants references verified before submission — even if they never say the\n  word \"research\".\n---\n\n# cite-holmes (Cite Holmes): deep research with citation verification\n\nOne line: **a question goes in — a verified report comes out.**\n\nThree differences from a plain \"search and summarize\":\n\n1. **Calibrate before working** — ask sharp questions first; the most expensive\n   waste is researching the wrong question.\n2. **Conclusions carry evidence grades** — 🟢 two independent sources agree /\n   🟡 single authority / 🔴 contested.\n3. **Every citation is checked** — mechanical layer (reachability, domain\n   authority, field completeness, dedup) plus semantic layer (does the source\n   actually support the claim?). Unverified references never masquerade as real.\n\n## ⛔ Iron rules (zero exceptions)\n\n1. **Never fabricate**: citations must come from pages actually fetched this\n   session. Re-search rather than write URLs from memory.\n2. **Never pretend**: unchecked references are marked `unverified`; fetch\n   failures are `unreachable` (≠ nonexistent — flagged for human review).\n3. **Don't hide conflicts**: when sources disagree, present the disagreement,\n   mark 🔴, show each side's evidence.\n4. **Budgeted search**: QUICK ≤6 searches, FULL ≤15. Out of budget → state the\n   gaps honestly instead of forcing conclusions.\n5. **Calibrate before searching** (FULL mode): scope / timeframe / audience /\n   output format must be locked first.\n\n## Step 0: mode selection\n\n| Mode | Fits | Calibration | Budget | Output |\n|---|---|---|---|---|\n| **QUICK** | Single fact-check: \"is this claim true\", \"when was X released\" | skipped | ≤6 | short report |\n| **FULL** | Open research: \"state of X\", \"A vs B\", \"do a survey\" | mandatory | ≤15 | full report |\n\nA question answerable by one verifiable fact → QUICK. Needs synthesis or\ntrade-offs → FULL. \"Quick check\" forces QUICK; \"thorough/comprehensive\" forces\nFULL. When unsure, default FULL.\n\n## Five-phase workflow\n\n### 1. CALIBRATE (FULL only)\n\nAsk 3–5 high-leverage questions at once (no drip-feeding): scope, timeframe,\naudience/depth, output format, decision context. Never re-ask what the user\nalready provided. If the user declines (\"your call\"), proceed with stated\ndefaults.\n\n### 2. PLAN\n\nShow a short plan: 3–7 sub-questions, source priority (primary/official >\nmajor media > community/blog as leads only), budget.\n\n### 3. SEARCH (iterative, not one pass)\n\n**Read `references/search-strategies.md` first** (diamond expansion, source\npyramid, query matrix, gap-driven iteration). Essentials: each round targets\none sub-question; evolve queries with discovered terms; search both English\nand Chinese for topics that span both internets; fetch full text of the 2–5\nmost valuable sources (never conclude from search snippets); verify key\nnumbers/dates in the original page before quoting.\n\n### 4. VERIFY (the heart of this skill)\n\nRegister every reference in `research_refs.json` (schema in\n`references/report-template.md`), then run:\n\n```bash\npython scripts/verify_refs.py --refs research_refs.json --out verify_report.md\n# Got a .bib from Zotero/EndNote? Feed it directly (v1.9):\npython scripts/verify_refs.py --refs bibliography.bib --out report.md\n# Environment self-check before a big batch (v3.10) — zero outbound calls\n# by default, add --net to probe every academic registry:\npython scripts/verify_refs.py --doctor --net\n```\n\nFive verdicts: `verified` / `partial` / `unreachable` (needs_human_check) /\n`invalid` / `unverified`. Every identifier (`url` / `doi` / `pmid` / `arxiv`)\nis cross-checked against its official registry (DOI.org metadata, NCBI\nE-utilities for PMID, export.arxiv.org for arXiv), so fabricated IDs are\njudged `invalid` — never silently `unreachable`. Every report opens with a\n**BLUF dual-reader header** (machine-parseable YAML + 5-line human TL;DR), a\n**CiteScore** (0-100 + A-D grade) and a one-line **pre-submission conclusion**\n(arXiv has banned authors over hallucinated references since 2026-05; ICML\n2026 desk-rejects them too).\n\n**Mechanical layer (the script)**: claim↔registry title/year/journal/author\nconsistency; nine machine-readable error codes on every judgment (v3.9);\n**clone-pair detection** (v3.10 — same title under different DOIs, or one DOI\ncarrying different titles, flagged in pairs: the most common fabrication shape\nin generated text); **cross-lingual title bridging** (v3.10 — a Chinese\noriginal title claimed against an English registry record is resolved via the\nbilingual `original-title` field, the DOI landing page, or OpenAlex; a hit\nupgrades to `verified`, a miss keeps `partial`, and a language difference is\nnever treated as fabrication evidence); retraction checks (Crossref online, or\ninstantly offline via an optional local Retraction Watch index\n`--retraction-cache`); Wayback archive links attached to dead links; arXiv\nlanding-page fallback keeps verdicts deterministic when its API flakes;\nparallel verification (`--workers`, default 4) plus a local disk cache\n(default on, 7-day TTL) make repeat runs cheap and verdicts stable. Behind a\nfirewall: `--proxy http://host:port`, `--cn` (resilient preset: 25s floor +\nCrossref re-source), or `--preflight` to see registry status before the batch;\nwhen 3+ registries fail at transport level the run flips to degraded-network\nmode — fast, honest skips instead of minutes of waiting.\n\n**Semantic layer (the model)**: decompose each cited claim into atomic\nclaim-triplets (subject–relation–object, RefChecker style) BEFORE judging,\nthen verify each triplet against the source — granularity moves from\nparagraph to triple, so \"which half-sentence is wrong\" becomes answerable.\nRegister the verdict structurally so it becomes auditable workpaper, not a\nfeeling:\n`\"semantic\": {\"claim\": \"...\", \"support\": \"supported|partial|not_in_source|\ncontradicted|unclear\", \"quote\": \"...\", \"note\": \"...\"}`. The verifier carries\nit into `--export auditjson` and a report section; `not_in_source` /\n`contradicted` cap the mechanical verdict at `partial` with a human-review\nflag — a source that exists is not a source that agrees.\n\n**Multi-source confirmation**: verified titles get an OpenAlex bibliographic\ncross-check; DOI references whose registry metadata could not be fetched get\na Semantic Scholar second confirmation; references with no DOI/PMID/arXiv at\nall get a Semantic Scholar title-search confirmation — confirmation only,\nnever a downgrade: databases have coverage gaps, \"not found\" ≠ fabricated.\n\n**Task → flag** (details in `references/verification-details.md`):\n\n| Task | Flags |\n|---|---|\n| Medical topics | `--easy` (auto medical + exports) · `--profile medical` |\n| Share / archive | `--format html` (self-contained single file) |\n| Export | `--export bibtex,gbt7714,ris,csv,auditjson` |\n| CI | `--strict` · `--offline` (structure only, caps at `partial`) |\n| Polite pacing | `--mailto you@lab.edu` |\n| API keys | `--openalex-key` / `--s2-key` / `--ncbi-key` (env of same name) |\n| Weak network | `--proxy` · `--cn` · `--preflight` · `--timeout` |\n| Cache | `--no-cache` / `--refresh-cache` / `--cache-ttl` |\n| Self-check | `--doctor` (add `--net` for registry probes) |\n| External judging | `--judge-url` / `--nli-url` / `--fast-judge-url` (notes only, never flip verdicts) |\n| MCP server | `mcp/server.py` — tools `verify_references` / `explain_verdict` / `check_document`, resources (capability matrix, changelog), `fact_check_workflow` prompt |\n\n**Offline mode (`--offline`)**: structure checks only — honesty first, never\nawards `verified`.\n\n**Medical research mode (v1.2)**: for clinical questions read\n`references/medical-mode.md` first and run with `--profile medical`.\n\n### 5. SYNTHESIZE\n\nFollow the skeleton in `references/report-template.md`: executive summary\nfirst; every key conclusion carries a confidence grade + citation ids; the\nreference table carries verdicts; `unverified/unreachable` items live only in\nthe \"human review\" section; finish with gaps, disagreements, and follow-up\nquestions.\n\n## Files\n\n| File | When |\n|---|---|\n| `scripts/verify_refs.py` | VERIFY phase mechanical check (pure stdlib, cross-platform, rate-limited) |\n| `mcp/server.py` | optional MCP server: expose `verify_references` as an MCP tool (FastMCP; same engine) |\n| `references/search-strategies.md` | read before SEARCH |\n| `references/report-template.md` | skeleton for SYNTHESIZE; refs schema |\n| `references/verification-details.md` | full option semantics, cross-check rules, thresholds |\n| `references/faq.md` | common questions (network, PMID wording, exports) |\n| `references/medical-mode.md` | read before medical/clinical research (v1.2) |\n| `examples/demo_refs.json` | general demo: 8 refs, 3 planted fabrications |\n| `examples/medical_refs.json` | medical demo: PMID/DOI refs + planted fake PMID + planted duplicate |\n\n## ⚠️ Common mistakes (anti-patterns)\n\n- Treating `verified` as \"the content is correct\" — verification confirms the\n  source exists and matches the claim's source, not that the reasoning holds.\n  The semantic check and your own reading still matter.\n- Citing a `partial` source as if it were reliable — `partial` means downgrade\n  or explain in the text; never silently promote it.\n- Retrying `unreachable` links forever — open the Wayback link once, then\n  switch sources if it stays dead.\n- Relying on `--easy` for clinical work — auto-detection is a convenience;\n  for real medical conclusions pass `--profile medical` explicitly.\n- Running without `--strict` in CI/automated pipelines and expecting exit\n  codes to gate anything.\n- Assuming `verified` means \"safe to cite forever\" — retracted papers are\n  flagged (Crossref/Retraction Watch), but retraction status can change after\n  your run.\n\n## Security & behavior declaration\n\n- Single-run CLI: start, verify, write the report, exit. No daemons, no\n  background jobs, nothing downloaded or installed at runtime (pure standard\n  library).\n- Network access is limited to these official academic registries, always over\n  HTTPS: `doi.org`, `api.crossref.org`, `eutils.ncbi.nlm.nih.gov`,\n  `export.arxiv.org` / `arxiv.org`, `archive.org`, `api.openalex.org`,\n  `api.semanticscholar.org`, `pubmed.ncbi.nlm.nih.gov`. No other hosts are\n  contacted; no telemetry, no analytics, no data collection.\n- Your reference lists and reports stay on your machine — the only outbound\n  payloads are the identifiers and titles you asked to verify.\n- Optional environment variables `OPENALEX_API_KEY` / `S2_API_KEY` /\n  `NCBI_API_KEY` authenticate\n  your own requests to those two APIs and are never sent anywhere else. The\n  local verdict cache lives under `~/.cache/cite-holmes/` (`--no-cache` to\n  disable).\n- No OS integration: no subprocesses, no system services, no privilege\n  changes; file access is limited to your inputs/outputs plus the cache and\n  report paths you pass.\n\n## Honest limits\n\n- Not a fit for: PDF/Word documents (extract citations manually first),\n  Chinese-database-only IDs (Wanfang/CQVIP numbers — verify via URL or title\n  instead), and fact-checking what a page *claims* (the mechanical layer\n  verifies existence and consistency; meaning is the semantic layer's job).\n- A fabricated citation pointing to a real, live, plausible page passes the\n  mechanical layer; the semantic layer may catch it — model judgment, not a\n  guarantee.\n- `unreachable` ≠ fake; database \"not found\" ≠ fabricated (coverage lag).\n- No public-registry check replaces reading the source — the July 2026 arXiv\n  survey of citation checkers found none reliable enough to run unsupervised;\n  treat this skill as a transparent multi-source assistant with a full audit\n  trail (`--export auditjson`), not an oracle.\n- Reproducible demo: `python scripts/verify_refs.py --refs examples/demo_refs.json`\n  (8 refs, 3 planted fabrications, all caught).\n- Medical demo: `python scripts/verify_refs.py --refs examples/medical_refs.json\n  --profile medical --export bibtex,csv`.\n\n## FAQ\n\nSee `references/faq.md` (slow networks / blocked endpoints, what PMID wording\nmeans, exports into Zotero/EndNote, how arXiv IDs are judged). Quick answers:\n`verified` is existence+consistency, not truth; `unreachable` ≠ fake;\n`--easy` covers the medical profile auto-detection; `--export bibtex` gives\nZotero/EndNote-ready verified-only bibliography.\n\n## 30-second quickstart\n\n```bash\n# 1. references as JSON (title/url/source/year per item) — or a Zotero .bib\n# 2. verify:\npython scripts/verify_refs.py --refs refs.json --out report.md\n# 3. read report.md — CiteScore + pre-submission conclusion at the top.\n# Medical? add --profile medical.  Paper?  add --export bibtex.\n# Institution/CI?  add --mailto you@lab.edu (Crossref polite pool).\n```\n\n## Environment fallback\n\nWithout web tools: state honestly that only the \"user-supplied material +\nmechanical verification\" mode is possible; never pretend to search.\n\nFile v3.10.0:examples/end-to-end/README.md\n\n# 端到端对照示例（输入 → 命令 → 实际输出）\n\n输入 → 命令 → 实际输出，三件套可直接复跑：\n\n```bash\npython scripts/verify_refs.py --refs end-to-end/research_refs.json \\\n    --check-document end-to-end/paper_excerpt.md \\\n    --export gbt7714,ris,bibtex,csv --out end-to-end/report.md\n```\n\n| 文件 | 是什么 |\n|---|---|\n| research_refs.json | 输入：1 条真实 DOI 引用（含正确标题） |\n| paper_excerpt.md | 输入：含 in-text 引用标记的正文片段（叙述式 + [n] 各一处，其中一处故意跑题） |\n| report.md | 实际输出：五态判定 + 上下文核验区（Kucsko 叙述式=锚定 ✓；[1] 企鹅句=低锚 ⚠️ 演示错配检测） |\n| report.json | 机读版（含 context_check 与逐项 checks） |\n| exports/ | 四种导出的真实产物样例（BibTeX / GB·T 7714-2025 / RIS / CSV） |\n\n> 复跑需网络（DOI.org/PubMed）。判定含时效字段（撤稿/被引），数字可能随时间小幅变化\n> ——以你复跑时的输出为准，结构不变。\n\n> SkillHub 包注：为符合平台文件类型白名单，`exports/` 导出样例文件不随\n> SkillHub 包分发（GitHub/ClawHub 包完整附带）——格式以本页表格描述为准。\n\nFile v3.10.0:mcp/README.md\n\n# cite-holmes-mcp\n\n**Mechanical citation verification as an MCP server** — feed it the references an\nagent is about to cite, get back per-reference verdicts (`verified / partial /\nunreachable / invalid`) against official registries (DOI.org, PubMed\nE-utilities, arXiv API, Crossref/Retraction Watch, Wayback). Zero API keys\nrequired, zero telemetry, local-first.\n\nThree MCP primitives:\n\n| Primitive | Name | What it does |\n|---|---|---|\n| tool | `verify_references` | Verify a batch of references (≤50); returns CiteScore 0–100 + per-item verdicts with evidence chain |\n| tool | `check_document` | Context-level check: parse in-text markers ([12]/[1-4]/(Author, Year)/doi.org links), bind to the reference list, flag low anchor-word overlap as possible mis-citations |\n| tool | `explain_verdict` | Plain-language explanation of one verdict (evidence steps + suggested action); zero network |\n| resource | `cite-holmes://capability-matrix` | What the mechanical layer catches vs. what stays with the semantic layer |\n| resource | `cite-holmes://changelog` | Version history |\n| prompt | `fact_check_workflow` | Three-step research-then-verify workflow for agents |\n\n## Install (three channels)\n\n```bash\n# 1) uvx from GitHub (recommended — always current)\nuvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp\n\n# 2) uvx from a local checkout\ngit clone https://github.com/docsor1212/cite-holmes\nuvx --from ./cite-holmes/mcp cite-holmes-mcp --help\n\n# 3) pip install (module + console script)\npip install \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\"\ncite-holmes-mcp --help\n```\n\n> This is a Python package (uv/pip ecosystem). There is no npm package — the\n> uvx/git channel above is the canonical install for all MCP clients.\n\n## Client configuration\n\n**Claude Code**\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\"]\n    }\n  }\n}\n```\n\n**Codex / any stdio MCP client**\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\", \"--transport\", \"stdio\"]\n    }\n  }\n}\n```\n\n**Cursor** (`~/.cursor/mcp.json`, same shape)\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\"]\n    }\n  }\n}\n```\n\nHTTP mode (local): `cite-holmes-mcp --transport http --port 8760`.\n\n## Bounds & behavior\n\n- Batch ≤ 50 references per `verify_references` call (larger batches get a\n  clear error — split them).\n- Long output fields are clipped (800 chars + truncation marker) so MCP\n  messages stay bounded; structure is preserved.\n- Per-reference network timeout defaults to 15 s (configurable per call).\n- Optional API keys via env: `OPENALEX_API_KEY` / `S2_API_KEY` /\n  `NCBI_API_KEY` (only attach to the caller's own requests).\n\n## Layout\n\n```\nmcp/\n├── server.py          # FastMCP three-primitive server (engine-agnostic impl)\n├── verify_refs.py     # embedded engine copy — synced from ../scripts/ by\n│                      #   sync_engine.sh (sha-gated; never hand-edit here)\n├── sync_engine.sh     # trunk↔mcp engine consistency gate (run before publish)\n├── pyproject.toml\n├── README.md\n└── tests/\n    ├── test_bounds.py        # batch/clip/bad-input unit tests\n    └── e2e_transcript.md     # full MCP client session evidence\n```\n\nLicense: MIT. Author: SorSor. Repo: https://github.com/docsor1212/cite-holmes\n\nFile v3.10.0:README.md\n\n# Cite Holmes 🔍\n\n[![GitHub Stars](https://img.shields.io/github/stars/docsor1212/cite-holmes?style=social&label=Star)](https://github.com/docsor1212/cite-holmes/stargazers)\n\n![icon](assets/icon-512.png)\n\n**Deep research that interrogates its own sources.**\n\nEvery AI research report you've ever read had a dirty secret: some of those polished references were probably fabricated. [A Nature news analysis suggests tens of thousands of 2025 publications might include invalid AI-generated references](https://www.nature.com/articles/d41586-026-00969-z). [GPTZero scanned 4,841 NeurIPS 2025 submissions; as independently reported, at least 100 hallucinated citations were found across 51 accepted papers](https://medium.com/@ljingshan6/100-fake-citations-just-slipped-through-neurips-2025-peer-review-5f34f4436560).\n\nCite Holmes is a deep-research skill with a badge and a magnifying glass: it researches like any deep-research agent — then **arrests its own citations before you can cite them**.\n\n![demo](assets/demo.gif)\n\n*(Demo is real output: 8 references, 3 deliberately planted fabrications — a fake DOI, a dead URL, and a no-URL citation. All 3 were caught and excluded; the 5 real ones passed. Measured: **7.7 s for all 8** — 4.2 s of pure network checks, the rest is deliberate throttling.)*\n\nReproduce it yourself — the planted-fakes file ships with the repo:\n\n```bash\npython scripts/verify_refs.py --refs examples/demo_refs.json\n```\n\n## 30-second quickstart\n\n```bash\n# MCP server (for agents — Claude Code / Codex / Cursor):\nuvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp\n```\n\n\n```bash\npython scripts/verify_refs.py --refs refs.json --out report.md\n# or verify straight from your reference manager (v1.9):\npython scripts/verify_refs.py --refs bibliography.bib --out report.md\n# v3.7.0 context check — verify in-text citations inside a finished draft:\npython scripts/verify_refs.py --refs bibliography.bib --check-document paper.md --out report.md\n```\n\n**China mirror (ModelScope 魔搭)**: <https://modelscope.cn/skills/Docsor/cite-holmes> — if you find this skill useful, a like there helps others find it.\n\nOpen `report.md`: CiteScore + pre-submission conclusion at the top,\nper-reference verdicts below.\nMedical work: add `--profile medical`.  Writing a paper: add `--export bibtex`.\nInstitution/CI: add `--mailto you@lab.edu` (Crossref polite pool).\n\n## What NOT to do (anti-patterns)\n\n- **Don't feed `semantic` fields from the same model that wrote the draft** —\n  the semantic cap trusts structured model judgments; self-review defeats it.\n- **Don't treat `verified` as \"the paper supports my claim\"** — verified means\n  the source exists at an authoritative tier and registry metadata matches the\n  claimed title/authors/journal/year. Whether the *specific sentence* is\n  supported is the semantic layer (agent judgment, or L4 cascade with a judge).\n- **Don't batch >50 references per MCP call** — the server rejects oversized\n  batches by design; split them.\n- **Don't hand-edit `mcp/verify_refs.py`** — it is a sha-gated copy of\n  `scripts/verify_refs.py`; edit the trunk and run `mcp/sync_engine.sh`.\n- **Don't cite from the report without reading `needs_human_check` flags** —\n  partial/unreachable items carry explicit re-check instructions on purpose.\n- **Don't skip `--cn` on lossy CN egress and then report unreachable as\n  \"fake\"** — unreachable is not invalid; re-run with `--cn` or a proxy first.\n\n## How it works\n\nFive phases, two modes:\n\n| Phase | What happens |\n|---|---|\n| **CALIBRATE** | Asks you 3–5 sharp questions first (scope, timeframe, audience) — prevents researching the wrong question |\n| **PLAN** | Breaks your question into 3–7 sub-questions with a search budget |\n| **SEARCH** | Diamond-shaped iterative search: broad → narrow → gap-filling, Chinese + English, source-tier pyramid |\n| **VERIFY** | Two layers: semantic check (does the source actually support the claim?) + mechanical check (reachability, domain authority, field completeness, dedup) |\n| **SYNTHESIZE** | Report where every conclusion carries a confidence grade — 🟢 two independent sources / 🟡 single authority / 🔴 contested |\n\nModes: **QUICK** (single fact-check, ≤6 searches, no interrogation) vs **FULL** (open-ended research, ≤15 searches, calibration mandatory).\n\n## Benchmarks (live runs + quality audit, 2026-09-26; zero-shot pipeline, qwen3:4b judge)\n\n| Benchmark | Setting | n | micro-acc | macro-F1 | Honest reading |\n|---|---|---|---|---|---|\n| SciFact-Open | open retrieval (mxbai top-3) + judge | 279 | 0.484 | 0.381 | the sanctioned zero-shot vehicle; 0.444/0.313 as first run before label-case normalization (both disclosed, sweep-reproducible); top systems (fine-tuned) reach 0.55-0.64 — our gap concentrates in REFUTES recall: abstract-level evidence rarely shows contradiction; remedies on the roadmap are the NLI third vote, claim-triplet judging, and full-text escalation |\n| CFEVER (zh) oracle + panel | gold evidence, zh judge + Erlangshen flip | 1000 | **0.817** | — | panel +8.1 over the 0.736 solo run; REFUTES correct 71→248 | oracle upper bound for the Chinese judge; same REFUTES weakness (R 0.18), excellent NEI discipline (R 1.0) |\n| SciFact-Open + L2 panel | judge + NLI contradiction flip (zero-param) | 279 | **0.617** | 0.563 (R-F1 0.667) | the heterogeneous two-vote panel lifts the same zero-shot pipeline +13 micro points into the fine-tuned top-system range (0.55-0.64); REFUTES F1 from 0 to 0.667 (paired flip 45/8, p=2.4e-7) |\n| SCitance v2.1 (self-built) | citation verification, hard negatives | 2576 | **0.713** (judge solo) / 0.646 (with NLI flip) | — | honest successor of the deprecated v1 0.947; the NLI contradiction flip that helps fact-check framing (+13 on SciFact-Open) HURTS here (-7) — topically-close hard negatives make the NLI over-report contradiction; aggregation is task-dependent, so the product exposes the panel as opt-in | released with quality audits: embedding-only separability 0.726 (v1 was 0.95 = degenerate), PIR 0/40; 751 false negatives filtered (36% of mined negatives actually supported the claim) |\n\n**Ablation (SciFact-Open, zero-shot)**: no-retrieval 0.262 (all-NEI floor) → top-1 **0.488** → top-3 0.444 → top-5 0.441; short-evidence (400ch) 0.405.\nReading: retrieval is decisive, but deeper evidence pools amplify the judge's SUPPORTS\nbias — precision-oriented retrieval (top-1) is optimal for this judge class.\n\nRetrieval: bge-m3 (corpus cache) / mxbai-embed-large (SciFact-Open runs). Judge:\nqwen3:4b via ollama (think:false + JSON-schema output; 8k-ctx resident). Full\nconfusion matrices and the audit gate scripts ship with the benchmark repo.\n\n## The five citation verdicts\n\n| Verdict | Meaning |\n|---|---|\n| `verified` | Reachable + authoritative tier (official/journal/preprint/major media) + complete fields — may support conclusions |\n| `partial` | Reachable but community/blog tier or missing fields — downgraded use |\n| `unreachable` | 404/timeout (≠ nonexistent — flagged for human review) |\n| `invalid` | Unresolvable or nonexistent identifier (bad/fake DOI, arXiv ID, PMID, URL) — never enters the report |\n| `unverified` | Not checked — never masquerades as verified |\n\n## v1.2: medical evidence mode + bibliography export\n\n```bash\n# medical profile: journal-tier sources extend to Cochrane/BMJ/ClinicalTrials/\n# NMPA/CDC/NICE/万方/ChiCTR; community-tier sources get an explicit\n# \"unfit for medical conclusions\" warning\npython scripts/verify_refs.py --refs refs.json --profile medical\n\n# PMID-only references work out of the box — and every PubMed URL gets its\n# PMID existence-checked via NCBI E-utilities. A fabricated PMID is caught\n# even though PubMed's own page happily returns 2xx for it.\n\n# export: verified-only BibTeX (straight into your paper) + full audit CSV\npython scripts/verify_refs.py --refs refs.json --export bibtex,csv\n```\n\nTry the medical demo — real tocilizumab/sJIA references from PubMed, plus one\nplanted fake PMID and one planted duplicate (both caught):\n\n```bash\npython scripts/verify_refs.py --refs examples/medical_refs.json --profile medical --export bibtex,csv\n```\n\n## v1.5: arXiv verification + Wayback fallback\n\n```bash\n# arXiv IDs are now first-class: the official export.arxiv.org API cross-checks\n# the registered title/year with the same thresholds as DOIs. A well-formed but\n# nonexistent ID is judged invalid (fabricated) — never silently \"unreachable\".\n# refs may carry \"arxiv\": \"2310.10631\" (or \"hep-th/9901001\"); arxiv.org/abs|pdf\n# URLs are auto-detected.\npython scripts/verify_refs.py --refs refs.json\n\n# Unreachable entries are automatically checked against the Wayback Machine;\n# when an archived copy exists, the report note carries the archive link so\n# human review has something to compare against.\n```\n\n## Install\n\n**AI agents (skills.sh / one command):**\n\n```bash\nnpx skills add docsor1212/cite-holmes\n```\n\n**OpenClaw users (ClawHub — versioned, auto-updatable):**\n\n```bash\nclawhub search cite-holmes        # find it on the registry\nclawhub install @docsor1212/cite-holmes\n```\n\n**Any Agent Skills-compatible agent** (Claude Code, Codex, Cursor, Gemini CLI, ZCode):\n\n```bash\nclawhub install @docsor1212/cite-holmes        # ClawHub registry\n# or grab the folder directly from SkillHub: https://skillhub.cn/skills/cite-holmes\n# then drop it into ~/.claude/skills/cite-holmes (or your agent's skills directory)\n```\n\n**China mirror (SkillHub 腾讯)**: <https://skillhub.cn/skills/cite-holmes> — fast downloads inside China, 中文说明.\n\n## The DoctorQ Lab academic family\n\nCite Holmes works alongside three sibling skills:\n\n- [academic-figures](https://skillhub.cn/skills/academic-figures) — publication-ready scientific figures (22+ chart types incl. composite panels, PRISMA, forest, KM) from one command\n- [paper-polisher-pro](https://skillhub.cn/skills/paper-polisher-pro) — AI-rate self-check, polish guidance and AIGC-compliance labeling for academic writing, bilingual, 100% local\n- [pubmed-verifier](https://skillhub.cn/skills/pubmed-verifier) — focused PubMed citation verification for medical literature\n\n## Usage\n\nJust talk to your agent — it triggers automatically:\n\n```\n> deep research: what changed in the agent-skills ecosystem this year?\n> is it true that NeurIPS 2025 papers contained 100+ hallucinated citations?\n```\n\nOr use the verifier standalone on any reference list:\n\n```bash\npython scripts/verify_refs.py --refs research_refs.json --out report.md\n# offline structural check / strict CI mode\npython scripts/verify_refs.py --refs refs.json --offline\npython scripts/verify_refs.py --refs refs.json --strict\n```\n\n## What it won't catch (honest limits)\n\n- A fabricated citation that points to a **real, live, plausible page** passes the mechanical check. The semantic layer (the model judging whether the source actually supports the claim) may catch it — it is model judgment, not a guarantee.\n- `unreachable` ≠ fake: pages behind login walls or bot-blocking are flagged for human review, not condemned.\n- The planted fakes in our demo are exactly the catchable types (dead URL / fake DOI / missing URL). We are not claiming it catches everything.\n\n## Why not just use deep research / a citation checker?\n\n- Deep-research skills **research more** but trust their own citations.\n- Standalone citation checkers verify but don't research.\n- Cite Holmes does both in one flow: **every reference in every report is machine-checked before it reaches you.**\n\nZero dependencies (pure Python stdlib), cross-platform (Windows/Linux/macOS), MIT license.\n\n## CiteScore — every report carries a grade\n\nEach verification report opens with a **CiteScore**: a 0-100 confidence score\n(verified +10 / partial +4 / unreachable 0 / unverified −2 / invalid −8,\nnormalized by reference count) with an A–D grade. Screenshot it, quote it,\nor gate your CI on it (`--strict`).\n\n## What's new\n\n- **v3.10.0** — cross-lingual titles & clone detection: a Chinese original\n  title claimed against an English registry record (the normal case for\n  Chinese-journal DOIs) is now resolved by a three-stage bridge — the\n  bilingual `original-title` registry field, then the DOI landing page, then\n  OpenAlex — a hit upgrades the verdict to `verified`, a miss keeps\n  `partial`, and a language difference is never treated as fabrication\n  evidence. **Clone-pair detection** flags same-title/different-DOI and\n  same-DOI/different-title pairs inside one bibliography (the most common\n  fabrication shape in generated text) as pairs — a lead for review, never a\n  verdict change. `--doctor` (add `--net`) gives a zero-outbound environment\n  self-check with per-registry connectivity probes, PASS/WARN/FAIL verdicts\n  and CI-friendly exit codes. The SKILL VERIFY section is restructured\n  command-first (run → verdicts → mechanical/semantic layers → task→flag\n  table). 355 tests green.\n- **v3.9.0** — structured error codes + escalation routing: every result now\n  carries machine-readable `error_codes` (E_DOU_NOT_FOUND / E_TITLE_MISMATCH /\n  E_RETRACTED / E_STITCHED / E_UNREACHABLE / ... — five verdicts plus the\n  reason class, derivable offline); the L4 evidence cascade gains a **fast-judge\n  pre-screen route** (with `--fast-judge-url`, suspected mis-citations flagged\n  by the 322M student are escalated first — routing, not skipping: every\n  candidate is still fully cascaded); MCP `verify_references` clipping is now\n  explicit (`max_field_chars` parameter + `output_clipped` marker when fields\n  are shortened — no silent truncation). 306 tests green.\n- **v3.8.0** — documentation & examples completeness release: end-to-end worked example (input refs → real generated report →\n  all four export formats, shipped under `examples/end-to-end/`);\n  `references/INDEX.md` + consolidated anti-patterns list (incl. \"trigger-word\n  mis-fires\" and cross-language mixed-citation boundary); **context-check\n  results now appear in the md/html reports** (previously console/JSON only);\n  author-surname anchor credit (a citing sentence naming a registry author\n  counts as strong binding evidence). 299 tests green.\n- **v3.7.0** — document-level citation checking (`--check-document paper.md`):\n  parse in-text citation markers ([12], [1-4], (Author, Year), doi.org links),\n  bind each marker to the verified reference list, and score anchor-word\n  overlap between the citing sentence and the cited title — low anchors are\n  flagged as possible mis-citations for human review. This moves verification\n  from the reference *list* to the *context*: the list can be clean while the\n  citations still point at the wrong papers. Also: `check_document` as a third\n  MCP tool; RIS export (Zotero/EndNote round-trip); `--preflight` registry\n  status panel. 292 tests green.\n- **v3.6.0** — GB/T 7714-2025 reference-list export (`--export gbt7714`):\n  verified entries formatted per the Chinese national bibliography standard\n  effective 2026-07-01, **built from registry-authoritative CSL metadata**\n  (authors in GB/T abbreviation form, [J]/[EB/OL] type codes,\n  volume(issue):pages, DOI) — not from user-claimed fields. Title-disputed\n  partials are excluded by policy (better absent than wrong). Plus `--cn`\n  resilience preset: 25s timeout floor and a Crossref second-host fallback\n  for DOI metadata (independent host; doi.org outages no longer blind the\n  checker). Anti-pattern guide + scenario best-practices docs. MCP server:\n  batch cap 50, output clipping, packaged entry point (tag mcp-v2.0.0).\n  292 tests green.\n- **v3.5.0** — fast-judge pre-screen (opt-in): point `--fast-judge-url` at a\n  [laya_service.py](https://github.com/docsor1212/cite-holmes) endpoint (322M\n  distilled classifier, ~24ms/item on GPU) and references stuck in\n  semantic-limbo (no official text to escalate, no metadata to verify) get a\n  screening note before a human ever looks at them. Guardrails by design: it\n  NEVER produces or changes a verdict, never flags for review, and only\n  annotates high-confidence SUPPORTS leanings (calibrated threshold 0.95;\n  calibration curve ships in the repo). Off by default — zero behavior change\n  without the flag. Also: label-case normalization disclosed in benchmark\n  table, evaluation FAQ expanded.\n- **v3.4.1** — dedup fingerprint fix: batch deduplication now keys on\n  identifier + normalized title fingerprint, so two references sharing a\n  URL/DOI/PMID but carrying different titles (the classic stitched-fake\n  variant) are independently verified instead of silently inheriting each\n  other's verdict. Genuine duplicates (same identifier + same title) are\n  still merged with the partial downgrade.\n- **v3.4.0** — PMID metadata verification + identifier-free search: PMID-only\n  references now get full title/journal/year cross-checks against PubMed\n  registry data (NLM abbreviation-aware journal matching; genuine-PMID-fake-paper\n  splice attacks are caught as invalid), references without any identifier go\n  through OpenAlex + Semantic Scholar title search (confirmation -> partial,\n  not-found -> honest unverified for human review, never \"fabricated\"), and\n  BLUF key_numbers now reports \"reviewed N (retracted X)\" separately.\n- **v3.3.1** — title cleaning for metadata matching: reference lines are cleaned\n  (strip numbering, DOI/PMID/arXiv suffixes, author+year blocks) before title\n  similarity is computed against official registry metadata — demo-validated to\n  raise AlphaFold-style matches from 0.66 to 0.9+.\n- **v3.3.0** — the heterogeneous-panel release. Optional NLI third vote\n  (`--nli-url`): an independent NLI service watches the generative judge;\n  agreement records high confidence, disagreement marks human review without\n  flipping mechanical verdicts (conservative philosophy, mirroring L4).\n  Experimentally validated: the same panel lifts SciFact-Open zero-shot micro\n  from 0.484 to 0.617 with REFUTES F1 0-to-0.667. 252 tests green.\n- **v3.2.0** — the conformance release. JSON reports carry the BLUF block as a\n  first-class object, conformant with the published BLUF Report Specification v1.0\n  (five core keys, worst-actionable rule; `bluf_yaml` kept for compatibility) —\n  our own validator now passes our own output. Optional offline retraction checking:\n  `--retraction-cache <index.json>` built from the official Crossref/Retraction Watch\n  GitLab dump (63k+ DOIs) turns retraction flags into a local, zero-network lookup.\n  241 tests green.\n- **v3.1.0** — the closed-loop release. The L4 evidence cascade is now wired\n  into the main flow: references the semantic layer could not resolve\n  (not_in_source/unclear) escalate automatically to official PubMed abstracts or\n  arXiv full text, get passage retrieval (bge-m3, TF-IDF fallback) and a\n  configurable external judge. Judge findings annotate the report and never flip\n  verdicts; without --judge-url the run is byte-identical (zero network calls).\n  The judge client handles thinking models (qwen3 family): native ollama calls\n  use think:false + JSON-schema enum output, /v1 endpoints fall back from empty\n  content to the reasoning field (all verified live on SciFact/SCitance/CFEVER).\n- **v3.0.0** — the BLUF release. Every report opens with a machine-parseable\n  YAML block (verdict / key_numbers / blocker / next_action) plus a 5-line\n  human TL;DR — cite-holmes defines the \"3-second dual-reader\" report standard\n  (a gap no tool fills today). Scholarly reception: verified DOIs fetch\n  Semantic Scholar citation contexts (times-cited, context excerpts, coverage\n  honestly labeled — the free Scite.ai counterpart). L4 evidence cascade:\n  abstract-level judging via configurable external judge endpoints (OpenAI-\n  compatible: ollama / llama-server), NEI escalates to full-text passage\n  retrieval with local bge-m3 embeddings (TF-IDF fallback, zero hard deps).\n  Cross-language title guard (Chinese claims vs English registrations no\n  longer misjudged invalid). 210 tests green.\n\n- **v2.0.0** — MCP three-primitives server (`verify_references` + `explain_verdict`\n  tools, capability-matrix/changelog resources, fact-check workflow prompt),\n  arXiv multi-version notes (unversioned citations to multi-revision papers,\n  stale version pointers), optional `NCBI_API_KEY` (E-utilities 3→10 req/s for\n  parallel batches), and a per-reference **evidence chain** in HTML reports\n  (doi.org → retraction → S2 → OpenAlex → reachability, at a glance). First\n  Major: the transparent multi-source citation verification stack, as shipped.\n  195 tests green, real-network acceptance, package security self-scan clean.\n\n- **v1.13.0** — the discoverability release. EN description rewritten around\n  task-language roots (citation verification, verify citations, citation\n  checker, hallucinated references, fact check) so task-word searches on\n  skills.sh and agent skill markets finally surface it. New MCP server\n  (`mcp/server.py`, FastMCP): expose the same mechanical verification engine\n  as a `verify_references` tool for agents that prefer tools over skills —\n  official MCP Registry ready (publish-once to Smithery, mcp.so, PulseMCP).\n  Security & behavior declaration retained. 188 tests green, real-network\n  acceptance 5/5, package security self-scan clean.\n\n- **v1.12.0** — trust & transparency release. Persistent disk cache (default\n  on, 7-day TTL, local SQLite): re-running or extending a bibliography reuses\n  stable verdicts with zero network round-trips — the repeat-run pain behind\n  every Trust score. `--strict` always bypasses the cache; `--refresh-cache`\n  forces re-verification; `unreachable` is never cached. `--proxy` flag for\n  restricted-egress environments (env `HTTPS_PROXY` also honored). A\n  **Security & behavior declaration** section now ships in both languages:\n  network access limited to the official academic registries over HTTPS, no\n  telemetry, single-run CLI, nothing installed, env vars limited to your own\n  API keys. The capability matrix adds an explicit \"unsupported inputs\" line\n  (PDF/Word documents, Wanfang/CQVIP-only IDs, fact-checking claims).\n  Private maintainer tooling removed from the shipped package. 181 tests.\n\n- **v1.11.0** — the determinism release: arXiv API transient 406/403 (its anti-bot\n  window trips even under our 3s global pacing) no longer silently skips content\n  verification — exponential backoff, then a fallback to the official arXiv.org/abs\n  landing page for title comparison (0.50/0.82 thresholds), so the same references\n  produce the same verdicts on re-runs. Structured semantic audit: the model's\n  per-reference semantic judgment now travels as machine-readable workpapers\n  (`semantic: {claim, support, quote, note}` → `semantic_audit` in `--export\n  auditjson` and a dedicated report section); `not_in_source`/`contradicted`\n  judgments cap the verdict at `partial` with a human-review flag. Pre-submission\n  conclusion corrected to arXiv's actual May 2026 enforcement start (was June) and\n  now names ICML 2026 desk-rejections too. `tools/agentskills_check.py --distro`\n  prints a skills.sh/agentskills.io distribution readiness checklist. 166 tests\n  green. Zero new dependencies.\n\n- **v1.10.0** — the speed release, aimed squarely at the top evaluation\n  complaint (\"verification leans on overseas databases; flaky from CN\n  networks\"): **parallel verification** (`--workers`, default 4 — references\n  are independent, so a 20-item bibliography takes roughly a quarter of the\n  serial time; results stay index-ordered and the host circuit breaker /\n  rate-limiters are now thread-safe), **degraded-network mode** (3+ different\n  hosts failing at the transport layer in one run flips remaining lookups to\n  fast honest skips with a proxy/retry suggestion — minutes of waiting become\n  a fast verdict), **Semantic Scholar title-search confirmation** for\n  references with no DOI/PMID/arXiv (confirmation only, never a downgrade;\n  hyphens normalized to spaces per S2 API docs), **intra-run cache** (a\n  duplicate DOI/URL reuses the first verdict instead of re-hitting the\n  network), fast-fail (no retry) once the network is known-degraded, and an\n  OpenAlex 429 hint (key now required for production since 2026-02; free tier\n  is 100 credits/day). 147 tests.\n- **v1.9.0** — Semantic Scholar third source: when DOI.org metadata cannot be\n  fetched, the DOI is re-confirmed at the S2 Graph API (confirmation only,\n  never a downgrade — S2 record quality is uneven; 403/429 landing pages can\n  now be rescued on S2 confirmation too). BibTeX import: `--refs\n  bibliography.bib` verifies Zotero/EndNote exports directly (nested braces,\n  `\\url{}` macros, eprint→arXiv). OpenAlex API key support\n  (`--openalex-key`/`OPENALEX_API_KEY`) with an actionable 403 hint — OpenAlex\n  requires keys for production use since 2026-02. `--mailto` joins the\n  Crossref polite pool; 429 `Retry-After` is honored. Every report now opens\n  with a one-line **pre-submission conclusion** (arXiv has banned authors over\n  hallucinated references since 2026-06). `--export auditjson`: a\n  machine-readable per-reference check trail for institutional audit.\n  SKILL.md slimmed (details moved to `references/verification-details.md` +\n  `references/faq.md`). 120 tests.\n- **v1.8.1** — fix: OpenAlex bibliographic confirmation could silently\n  no-op (response read outside its context manager); caught by the new strict\n  real-network acceptance gate and corrected.\n- **v1.8.0** — retraction detection: verified DOIs are cross-checked against\n  the Crossref/Retraction Watch database (free, updated every working day);\n  a retracted paper is real — and uncitable — so retracted references are\n  capped at `partial` and flagged for human review. Bibliographic existence\n  check via OpenAlex for references without DOI/PMID/arXiv (positive\n  confirmation only; \"not found\" never means \"fabricated\"). UA-rotation retry\n  for metadata endpoints (403/406 resilience). New \"common mistakes\"\n  anti-patterns section and a 30-second quickstart. Capability matrix now\n  covers 10 fabrication/problem types. 89 tests.\n- **v1.7.0** — author-name consistency check against DOI-registered authors\n  (any claimed surname hit counts; a full miss downgrades to `partial`),\n  capability-boundary matrix in every report (9 machine-caught fabrication\n  types vs what stays with the semantic layer), `--format html`: a\n  self-contained single-file HTML report (zero external assets, mobile\n  readable — shareable with advisors/editors), agentskills.io spec self-check,\n  internal competitive landscape. 73 tests.\n- **v1.6.0** — host-level circuit breaker (2 consecutive transport failures to\n  the same host → skip for the rest of the batch with an honest note; broken\n  networks no longer stall the whole run), actionable error notes (timeout\n  suggests `--timeout 20`, DNS vs refused distinguished, human-review section\n  carries suggested actions), DOI↔PMID cross-check (catches \"stitched fakes\"\n  where both IDs are real but belong to different papers), journal-name\n  consistency check against DOI-registered metadata (mismatch → `partial`),\n  CN-network FAQ section, regression suite grown to 62 tests.\n- **v1.5.0** — arXiv ID metadata validation (export.arxiv.org official API,\n  DOI-grade similarity thresholds; nonexistent IDs → `invalid`, never\n  `unreachable`), Wayback Machine fallback for unreachable entries (archive\n  link attached for human review), `arxiv` reference field, `eprint` in BibTeX\n  + arXiv column in the audit CSV, regression suite grown to 53 tests.\n- **v1.4.0** — DOI metadata cross-check (DOI.org CSL JSON vs cited title/year —\n  catches \"real DOI, wrong paper\"), `--easy` one-flag mode (auto medical profile +\n  auto BibTeX/CSV export), transient-network retry, human-readable Chinese error\n  messages, expanded FAQ.\n- **v1.3.0** — CiteScore scorecard in every report (Markdown + JSON), official\n  icon, `tests/` regression suite (offline golden + medical profile + exports).\n- **v1.2.0** — medical evidence mode (`--profile medical`), PMID verification\n  via NCBI E-utilities, BibTeX/CSV bibliography export, 3-key dedup\n  (URL/DOI/PMID).\n\n## License\n\nMIT © DoctorQ Lab\n\nFile v3.10.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b6sjwend7cwwmhf7mdg1wzx82dfbf\",\n  \"slug\": \"cite-holmes\",\n  \"version\": \"3.10.0\",\n  \"publishedAt\": 1791484994852\n}\n\nFile v3.10.0:references/anti-patterns.md\n\n# 反模式统一清单（常见误用方式集中收录）\n\n> 收拢 README 与 SKILL.md 分散的反模式说明 + 新增易踩坑场景。上手前通读一遍。\n\n1. **不要用写草稿的同一个模型喂 `semantic` 字段**——语义封顶信任结构化模型判定，\n   自己审自己等于没审。\n2. **不要把 `verified` 理解为\"这篇论文支撑我的论断\"**——verified 只说明来源存在、\n   层级权威、登记元数据与声称一致；\"句子是否真被支撑\"归语义层。\n3. **MCP 单次批量不要超过 50 条**——超限被拒（分批调用）。\n4. **不要手改 `mcp/verify_refs.py`**——它是 scripts/ 的 sha 门禁副本，改 trunk 后跑\n   sync_engine.sh。\n5. **不要忽略报告中的 `needs_human_check` 旗标**——partial/unreachable 带复核指令\n   是有意设计。\n6. **不要在弱网把 `unreachable` 当\"假引用\"**——先用 `--cn`/代理重跑再下结论。\n7. **不要为求星而自动化点击**（平台红线）——报告页脚的 Star/收藏链接面向人类读者。\n8. **不要在 `--check-document` 里省略 `--refs`**——上下文核验需要清单核验结果做绑定。\n\n## 触发词误触发怎么办\n\n- 触发词（深度研究/引用核查/查证等）按语义匹配，短句如\"查证一下\"若被误触发，\n  直接说明真实意图即可（\"我只要翻译\"）；不产生副作用。\n- 若希望某类请求**不**触发引用核验，避免在请求中同时包含\"研究/查证/引用\"词与\n  文献清单。\n9. **不要对中英混合引用的锚词率过度报警**——跨语言（CJK vs 拉丁）混合引用的\n   锚词率天然偏低（内容词不相交），上下文核验会标\"疑似错配\"；这是预期行为\n   （清单核验侧已有跨语言守卫降 partial 转人工），按人工复核处理而非自动判死。\n\nFile v3.10.0:references/faq.md\n\n# FAQ (moved out of SKILL.md in v1.9 to keep the main guide short)\n\n**Q: Is a `verified` citation guaranteed real?**\nNo. The mechanical layer catches machine-checkable fakes: nonexistent DOIs,\nfabricated PMIDs (via E-utilities), fabricated arXiv IDs (via the official\narXiv API), dead links, duplicates, stitched DOI↔PMID pairs, wrong-paper\nDOIs (title similarity), and retracted papers. A fabricated citation pointing\nto a real, plausible page still relies on the semantic layer — never a\nguarantee.\n\n**Q: How are arXiv references checked?**\nSame treatment as DOIs since v1.5: the official export.arxiv.org API returns\nthe registered title/year; similarity below 0.50 → `invalid` (fabricated or\nwrong ID), 0.50–0.82 → `partial` (human review), above → consistent. An ID the\nAPI has never heard of is `invalid`, not `unreachable`.\n\n**Q: What do I do with `unreachable` references?**\nOpen them manually once — the report automatically attaches a Wayback Machine\narchive link when one exists. If it matters and keeps failing, switch sources.\n\n**Q: Did the PMID check fail because of my network?**\nNo. \"不存在/编造\" verdicts come from the official E-utilities API — that means the\nPMID genuinely isn't in PubMed. Only \"校验失败\" wording means a network/API issue\n(the reference is then treated as reachable and flagged).\n\n**Q: Do I always need `--profile medical`?**\nUse `--easy`: one flag auto-detects medical references, enables the medical\nprofile, and exports BibTeX + CSV automatically. For real clinical\nconclusions, still prefer explicit `--profile medical`.\n\n**Q: Can verified references go straight into my paper?**\nYes: `--export bibtex` exports only `verified` entries (with PMID / arXiv\neprint notes) and imports into Zotero/EndNote; `--export csv` is the full\naudit ledger for advisors and editors; `--export auditjson` (v1.9) is the\nmachine-readable per-check working paper for institutional audit.\n\n**Q: Which services does verification touch, and what if my network is slow or blocked?**\nDOI.org (DOI metadata), NCBI E-utilities (PMID), export.arxiv.org (arXiv\nmetadata), archive.org (archived copies), api.crossref.org (retractions),\napi.openalex.org (bibliographic confirmation), api.semanticscholar.org (third\nsource + title search, v1.9/v1.10). CN networks occasionally throttle several\nof them. The verifier never stalls: transient failures retry once (429 honors\n`Retry-After`), error notes tell you the cause with a fix (`--timeout 20` for\nslow links), and a host-level circuit breaker skips a host after 2 consecutive\ntransport failures — with an honest note — instead of hanging the batch.\nSince v1.10, references are verified in parallel (`--workers 4` by default —\nroughly a quarter of the serial wall time), and when 3+ different hosts fail\nat the transport layer in one run (restricted-egress signature) remaining\nlookups are fast-skipped with a \"全局网络降级\" note and a proxy/retry\nsuggestion. Degraded checks are always labeled, never silently dropped.\n\n**Q: Do I need API keys?**\nNo — everything runs keyless within public rate limits. OpenAlex has required\nkeys for *production* use since 2026-02 (free daily allowance; this skill\nsurvives keyless but benefits from `--openalex-key` / `OPENALEX_API_KEY`).\nSemantic Scholar keyless shares a rate pool; `--s2-key` / `S2_API_KEY` gets a\ndedicated lane. Crossref asks heavy users to identify themselves — pass\n`--mailto you@lab.edu` to join the polite pool.\n\n**Q: MCP 模式是什么？agent 不装 skill 怎么用？**\nv2.0.0 起仓库内置 `mcp/server.py`（FastMCP）：tools `verify_references`（机械验证）/\n`explain_verdict`（判定解释）、resources（能力矩阵/版本史）、prompts（fact-check 工作流）。\n推荐直跑：`uvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp`；\n或自行准备 FastMCP 环境后运行 `python mcp/server.py`。零密钥可用；可选 API key 经环境变量注入。\n\n**Q: arXiv 论文有多个修订版怎么办？**\nv2.0.0 自动备注：引用未指定版本而论文存在多个修订版 → 提示\"引用未指定版本\"；\n引用指向旧版而存在更新版 → 提示最新版号（旧版可能含未修正内容）。同响应内取数，零额外请求。\n\n**Q: 验证要访问哪些外部服务？国内网络慢怎么办？**\n见上方服务清单——全部是官方学术注册库、全部 HTTPS。v1.12 起两件事让它在国内更可用：\n① **持久磁盘缓存**默认开（7 天 TTL）：同一批引用复跑直接复用稳定判定、零外呼，\n`--refresh-cache` 强制重验；② **`--proxy http://host:port`** 显式走代理（或设\n`HTTPS_PROXY` 环境变量），配合网络降级模式，受限出口下也能一次跑完拿到诚实结论。\n\n**Q: 判定结果会被缓存吗？撤稿状态更新了怎么办？**\n缓存只存 verified/partial/invalid 三种稳定判定（TTL 默认 168 小时），\nunreachable 永不入缓存（瞬态）；`--strict` 模式强制绕过缓存读（CI 诚实）。\n撤稿状态这类可变信息以 TTL 为界——对时效敏感的批次用 `--refresh-cache` 强制重验。\n\n**Q: Can I verify references straight from my reference manager?**\nYes, since v1.9: `--refs bibliography.bib` reads Zotero/EndNote/JabRef BibTeX\nexports directly (entry types, nested braces, `\\url{}` macros; DOI/PMID/arXiv\nfields are picked up automatically).\n\n## arXiv 条目 note 里出现「元数据 API 不稳，经官方着陆页标题比对确认」正常吗？\n\n正常。arXiv 的 API 对批内多请求有反机器人窗口（偶发 406/403），与你的网络无关；\n验证器已自动回退官方着陆页完成同等强度的标题比对，判定可信度不受影响。\n\n## semantic 字段写什么？\n\n模型逐条判定\"来源是否真的支撑所引论断\"后，写成结构化对象\n`{\"claim\", \"support\", \"quote\", \"note\"}`（五值：supported / partial /\nnot_in_source / contradicted / unclear）。`not_in_source` 与 `contradicted`\n会被机械验证器封顶为 `partial` 并转入人工复核区——来源存在不等于来源认同。\n旧的字符串写法仍接受（仅注记）。\n\n## Q: fast-judge 预筛会改我的判定吗？（v3.5.0）\n不会。`--fast-judge-url` 外挂的 322M 蒸馏分类器只对\"语义待定且无官方文本可升级\"\n的引用追加「倾向支持（供人工参考）」注记——永不产生/修改五态判定，永不标记人工\n复核，默认关闭。校准曲线随仓库发布（阈值 0.95）。\n\n## Q: GB/T 7714-2025 导出的数据来源是什么？（v3.6.0）\n`--export gbt7714` 仅导出 verified 条目，且每条以**注册库登记的 CSL 元数据**\n为准（作者按国标缩写惯例预格式化、期刊/年/卷期页来自 DOI.org/Crossref 登记），\n不是你手填的字段。无注册元数据的电子资源按 [EB/OL] + 引用日期输出。\n\n## Q: 国内网络直连核验总超时怎么办？（v3.6.0）\n加 `--cn`：超时下限抬到 25 秒，且 DOI.org 三试全败时自动回源\napi.crossref.org（独立主机，同构 CSL 元数据，报告注记会写明「经 Crossref\n回源」）。仍不通时 `--proxy` / `HTTPS_PROXY` 继续兜底；`unreachable` 不等于\n假引用，换网重跑后再下结论。\n\n## Q: 能核验\"正文里的引用\"而不只是参考文献清单吗？（v3.7.0）\n能。`--check-document paper.md --refs refs.json`：解析正文 in-text 标记\n（[12]/[1-4]/(Author, Year)/doi.org 链接），把每处引用绑定到清单条目并计算\n\"引文句 ↔ 被引标题\"锚词率——低于 0.34 标记为疑似错配引文（引 A 文却引了 B 句）。\n锚词率低≠错引，是语义层人工复核候选；语义判断仍由 agent 负责。\n\n## Q: 触发词误触发了怎么办？\n触发词按语义匹配，可能被\"查证一下\"这类短句误触发。直接说明真实意图即可\n（例如\"我只要翻译这段\"），误触发不产生任何副作用。想让某类请求**不**触发\n引用核验：避免同时出现\"研究/查证/引用\"词与文献清单。\n\n## Q: 中英混合引用（如中文声称+英文登记）会误判吗？\n不会判 invalid。v3.0.0 起内置跨语言守卫：声称标题与登记标题跨语言（CJK vs\n拉丁）时相似度不作编造依据，降 `partial` 转人工复核（与期刊名核查同款守卫）。\n已知边界：跨语言混合引用的锚词率（上下文核验，v3.7.0）同样会偏低——属预期，\n按\"疑似错配候选\"人工复核即可。\n\nFile v3.10.0:references/INDEX.md\n\n# references/ 索引（v3.8.0 新增）\n\n| 文件 | 用途 | 何时读 |\n|---|---|---|\n| faq.md | 高频问题解答（网络/PMID/导出/API key/fast-judge/GB·T 导出/--cn/上下文核验） | 遇到具体疑问时 |\n| anti-patterns.md | 反模式统一清单（误用方式集中收录，v3.8.0 从 README/SKILL 收拢） | 上手前通读一遍 |\n| medical-mode.md | 医学证据模式详解（信源金字塔/CEBM 分级/监管信源） | 使用 --profile medical 前 |\n| verification-details.md | 五态判定完整阈值规则与检查明细 | 需要理解判定依据时 |\n| search-strategies.md | 中英双语检索策略矩阵 | 定制检索流程时 |\n| skillhub-mcp.md（如存在） | MCP 集成参考 | 配置 MCP 前 |\n\n> 主文档：SKILL.md（技能定义）/ README.md（用户文档）。本目录文件按需加载，\n> 不进入主执行流。\n\nFile v3.10.0:references/medical-mode.md\n\n# 医学研究模式（MEDICAL 模式必读）\n\n适用触发：临床问题、药物疗效/安全性、疾病患病率、诊疗指南、诊断标准、\nmeta 分析、循证检索、医学文献综述。拿不准是否算\"医学\"时：凡涉及患者\n治疗决策的一律按医学模式处理。\n\n## 启用方式\n\n检索阶段照常（菱形扩展 + 双语矩阵），验证阶段加参数：\n\n```bash\npython scripts/verify_refs.py --refs research_refs.json --profile medical --export bibtex,csv\n```\n\n## 与通用模式的四点差异\n\n1. **期刊层域名扩展**：Cochrane Library、BMJ Best Practice、Embase、\n   ClinicalTrials.gov、ChiCTR、NMPA、CDC/中国疾控、NICE、万方、医脉通指南\n   在医学模式下计入权威期刊层。\n2. **社区层降级警示**：知乎/微博/公众号/维基/博客层来源自动标注\n   \"不得支撑医学结论（仅作线索）\"——医学主张的引用底线比通用问题更高。\n3. **PMID 存在性核实**：PubMed 网页对不存在的 PMID 也返回 2xx，HTTP 检查\n   抓不住编造的 PMID。验证器经 NCBI E-utilities 逐条核实 PMID 真实存在。\n   引用登记时 pmid-only 条目可直接写 `{\"pmid\": \"36443570\", ...}`，无 url/doi\n   也能验证（自动解析为 PubMed 页面）。\n4. **台账导出**：`--export bibtex,csv` 产出 verified-only 参考文献（BibTeX，\n   直接导入论文管理器）+ 全量审计台账（CSV，Excel 打开）。\n\n## 医学信源金字塔\n\n| 层级 | 首选来源 | 用法 |\n|---|---|---|\n| 系统评价/Meta | Cochrane Library；PubMed 上的 SR/MA（`systematic[sb]` 过滤） | 可直接支撑疗效结论 🟢 |\n| 临床指南 | WHO、NICE、中华医学会各分会、国家卫健委、UpToDate / BMJ BP | 支撑诊疗规范类结论 🟢 |\n| 原始研究 | PubMed / Embase（RCT > 队列 > 病例系列）；知网/万方/维普（中文） | 按研究设计定 🟢/🟡 |\n| 预印本 | medRxiv / bioRxiv | 只标 🟡 并注明\"未经同行评审\" |\n| 权威媒体 | Reuters、丁香园（资讯级） | 背景信息，不作疗效证据 |\n| 社区/社交 | 知乎、微博、病友群、公众号 | 仅作线索，禁止支撑任何医学结论 🔴 |\n\n## 检索式要点（PubMed）\n\n- MeSH 主题词 + 自由词组合：`(\"Tocilizumab\"[Mesh]) AND (\"Arthritis, Juvenile\"[Mesh] OR sJIA)`\n- 疗效问题加过滤：`AND (randomized controlled trial[pt])`；病因/患病率换相应过滤\n- 中英双语：英文 PubMed，中文知网/万方，两边合并去重（验证器按\n  URL/DOI/PMID 自动去重）\n- 双源规则在医学场景升级：**疗效结论需 ≥1 个系统评价或 ≥2 个独立 RCT**；\n  仅有病例系列支撑的疗效主张标 🟡 并明示证据等级\n\n## 监管与注册信源\n\n- 药械批件/说明书：NMPA（中国）、FDA、EMA\n- 临床试验注册：ClinicalTrials.gov、ChiCTR（中国临床试验注册中心）\n- \"某药是否获批某适应证\"\"某试验是否注册\"类事实，优先用监管/注册源核实——\n  这类事实比文献更容易被 AI 编造，且编造得最像真的\n\n## 证据分级对照\n\n本 skill 的置信度标记与 CEBM 2009 证据等级的大致映射：\n\n| 标记 | CEBM 等级 | 典型来源 |\n|---|---|---|\n| 🟢 | 1a–2b | SR/MA、高质量 RCT、官方指南推荐 |\n| 🟡 | 3–4 | 队列/病例对照/病例系列、单个专家意见 |\n| 🔴 | 5 或存疑 | 传闻、社区层、利益相关方声明、来源冲突 |\n\nFile v3.10.0:references/report-template.md\n\n# 报告模板与引用 schema（SYNTHESIZE 阶段照此输出）\n\n## research_refs.json schema（VERIFY 阶段的输入）\n\n研究过程中实时登记引用，每条：\n\n```json\n[\n  {\n    \"id\": 1,\n    \"title\": \"页面/文档标题\",\n    \"url\": \"https://…\",\n    \"source\": \"来源站点或机构（如 Anthropic / Nature / Reuters）\",\n    \"year\": 2026,\n    \"claim\": \"本报告用这条引用支撑的那句话\",\n    \"triples\": [{\"subject\": \"...\", \"relation\": \"...\", \"object\": \"...\", \"span\": \"原文出处片段\"}],  # L1 原子声明分解(v3.0.0)\n    \"semantic\": {\n      \"claim\": \"与上层 claim 相同（或更细）\",\n      \"support\": \"supported\",   // supported | partial | not_in_source | contradicted | unclear\n      \"quote\": \"来源页原文关键句（人工复核抓手）\",\n      \"note\": \"可选备注\"\n    },\n    \"found_via\": \"search#3\",          // 哪轮搜索发现的，便于回溯\n    \"tier\": \"official\"                // official | journal | preprint | media | community | blog | social\n  }\n]\n```\n\n`semantic` 字段由模型在语义验证时填写（v1.11 起推荐上面的结构化对象形式；旧的\n字符串形式仍然接受，仅作注记不参与判定）。验证器把结构化判定带进\n`--export auditjson` 的 `semantic_audit` 工作底稿与报告语义区；`not_in_source` /\n`contradicted` 会把机械判定封顶为 `partial` 并转人工复核（来源存在 ≠ 来源认同）。\n机械验证由 `scripts/verify_refs.py` 完成，两层的结果都要体现在最终引用表里。\n\n## 置信度标记规范\n\n| 标记 | 含义 | 判定标准 |\n|---|---|---|\n| 🟢 | 高置信 | ≥2 个独立来源一致，且至少一个是一手/权威源 |\n| 🟡 | 中置信 | 单一权威源，或双源但均为二手 |\n| 🔴 | 存疑 | 来源冲突、仅社区层证据、或数据过旧 |\n\n标记挂在结论句末尾 + 引用编号，如：`缓存命中率提升约 40% 🟢[1][3]`。\n\n## 报告骨架\n\n```markdown\n# {研究问题}\n\n> 研究模式：FULL/QUICK · 检索 N 轮 · 引用 M 条（verified X / partial Y / …）\n> 完成日期：YYYY-MM-DD · 覆盖时间窗：…\n\n## 执行摘要（≤200 字，先给答案）\n{直接回答研究问题的主要发现，2–4 条，每条带置信度标记}\n\n## 主要发现\n### 子问题 ①：{…}\n{结论句 🟢[引用号]。证据展开：数字/事实 + 出处定位。}\n{与结论相悖的证据如有，必须写。}\n\n### 子问题 ②：{…}\n…\n\n## 分歧与存疑\n{来源打架的地方：各自说法 + 各自来源 + 可能的成因（数据新旧/口径不同/立场差异）}\n\n## 研究缺口\n{预算内没能回答的问题、抓取失败的来源、需要付费/权限才能核实的数据}\n\n## 后续值得追问\n{2–3 个基于本次发现自然延伸的问题}\n\n## 引用清单\n| # | 标题 | 来源 | 年份 | 语义 | 机械验证 | 支撑论断 |\n|---|---|---|---|---|---|---|\n| 1 | … | … | … | supports | verified | … |\n\n### 待人工复核（unverified / unreachable）\n| # | 标题 | URL | 状态 | 原因 |\n```\n\n## 输出要求\n\n- 正文结论**只能**引用 `semantic.support=supported`（旧字符串 `supports`）且机械验证非 `invalid` 的条目；`not_in_source`/`contradicted` 条目已由验证器封顶 `partial`，只可作对照线索。\n- `partial_support` 只能支撑结论中它确实支持的那半句，并在句中注明。\n- `unverified/unreachable` 一律只出现在\"待人工复核\"分区。\n- QUICK 模式可省略\"子问题\"分节，直接：执行摘要 → 核查结论（含反证）→ 引用清单。\n- 报告语言跟随用户提问语言；术语首次出现给中英对照。\n\nFile v3.10.0:references/search-strategies.md\n\n# 检索策略（SEARCH 阶段必读）\n\n## 菱形扩展模式\n\n检索节奏是\"菱形\"：**宽 → 窄 → 补缺**。\n\n1. **宽**（第 1–2 轮）：用问题的自然表述直接搜，目的是摸清领域地形——主流说法、关键术语、\n   主要玩家。此时发现的**术语**比**答案**更重要（下一轮的搜索词从这里来）。\n2. **窄**（第 3–N 轮）：用上轮挖出的精确术语/产品名/人名/事件名构造查询，逐个子问题收网。\n   搜索词要具体到\"能命中单一事实\"的程度。\n3. **补缺**（最后 1–2 轮）：对照 PLAN 里的子问题清单，只搜还没有双源确认的缺口。\n   缺口补不上就停止——剩余缺口写进报告，不硬凑。\n\n## 搜索词矩阵\n\n对核心子问题，按三个维度变化搜索词，避免同义反复地浪费预算：\n\n- **语言维**：中文术语 ↔ 英文术语（同一概念的中英文名经常不同，如\"深度研究\" vs \"deep research\"）\n- **角度维**：正面（\"XX 优势\"）↔ 反面（\"XX 问题/翻车/争议/批评\"）——只搜正面会得到软文视角\n- **粒度维**：综述型（\"XX 进展综述 2026\"）↔ 事实型（\"XX 具体数字/日期/版本\"）\n\n## 信源金字塔（可信度从高到低）\n\n| 层级 | 例子 | 用法 |\n|---|---|---|\n| 一手/官方 | 官方文档、发布公告、论文原文、财报、法规原文 | 可直接支撑结论 |\n| 权威二手 | 权威媒体（Reuters/AP/新华等）、行业头部分析 | 可支撑，尽量回溯到一手 |\n| 社区/百科 | GitHub、Stack Overflow、Wikipedia、知乎高赞 | **只作线索**：从这里找到一手源再验证 |\n| 社交/自媒体 | X/Twitter、微博、个人博客、公众号 | 仅作趋势感知，不得单独支撑任何结论 |\n\n- 社区层发现的关键论断 → 找到它的一手出处 → 用一手源作为引用。\n- 找不到一手出处的社区论断 → 如标注明\"仅社区层证据 🔴\"。\n\n## 关键事实的双源规则\n\n以下进入报告前必须有两个独立来源（不同站点，不是互相转载）确认：\n\n- 具体数字（价格、百分比、性能数据、装机量）\n- 日期与版本号\n- 引语（谁说了什么）\n- \"最/第一/唯一\"类断言（这类断言即使双源也标 🟡，除非来源是一手）\n\n两源打架 → 报告里如实写分歧，标 🔴，注明各自来源与时间（旧数据 vs 新数据是常见原因）。\n\n## 页面抓取（WebFetch）纪律\n\n- 每次抓取前明确要找什么；抓完立即记录\"找到了什么/没找到什么\"。\n- 搜索结果摘要 ≠ 页面内容。凡是要写进报告的具体细节，必须抓到页面里确认。\n- 单页抓取失败（超时/反爬）→ 换一个来源覆盖同一子问题，不无限重试同一 URL。\n- 优先抓：官方文档 > 公告原文 > 权威媒体正文。\n\n## 搜索预算台账（边搜边记）\n\n```\n| 轮 | 查询 | 目标子问题 | 收获 | 新缺口 |\n|---|---|---|---|---|\n| 1 | … | ① | 关键术语 X、Y | Y 的数据缺 |\n```\n\n到达预算上限即停，剩余缺口如实写进报告\"研究缺口\"一节。\n\nFile v3.10.0:references/verification-details.md\n\n# Verification details (v1.10)\n\nFull semantics of every check `verify_refs.py` performs. SKILL.md keeps the\nshort version; this file is the complete reference.\n\n## Cross-check rules by identifier\n\n- **DOI** (v1.4+): DOI.org content negotiation returns the registered CSL\n  metadata. Title similarity < 0.50 → `invalid` (\"real DOI, wrong paper\" —\n  fabricated or mis-cited); 0.50–0.82 → `partial` (human review); ≥ 0.82 →\n  consistent. Year difference ≥ 2 adds a warning note. A DOI the registry has\n  never heard of (404) → `invalid`, never `unreachable`.\n- **PMID**: NCBI E-utilities esummary. PubMed's own web pages return 2xx even\n  for nonexistent PMIDs, so only the API actually catches fabrications.\n- **DOI ↔ PMID cross-check** (v1.6): when both keys are present, the PMID's\n  registered DOI must match the claimed DOI — a mismatch is `invalid`\n  (\"stitched fake\": each key real, pointing at different papers).\n- **Journal name** (v1.6): claimed source vs DOI-registered container-title.\n  NLM-style abbreviations are treated as equivalent (word-initial sequential\n  match, stop-words skipped); cross-language names are not comparable and are\n  skipped. A real mismatch downgrades to `partial` (journals do rename).\n- **Author names** (v1.7): claimed surnames matched against the CSL author\n  list; any hit counts (spelling variants tolerated); single-letter initials\n  never substring-match (APA \"Torvalds, L.\" would otherwise match any \"G.\").\n  CJK-vs-Latin names are skipped (not comparable). A full miss downgrades to\n  `partial`.\n- **arXiv** (v1.5): official export.arxiv.org API (≥3 s/request politeness\n  pause). Unknown ID or API error element → `invalid`. Same 0.50/0.82\n  similarity thresholds as DOI.\n- **Retraction** (v1.8): Crossref REST API `updated-by[]` (Retraction Watch\n  data, updated daily). Any `type == \"retraction\"` → capped at `partial` +\n  human-review flag. Check failures are silently skipped — never penalize a\n  reference for the checker's network.\n- **OpenAlex bibliographic check** (v1.8, v1.9 keys): for references without\n  DOI/PMID/arXiv, title search against OpenAlex. Best of full/prefix title\n  similarity ≥ 0.82 → \"confirmed exists\"; 0.60–0.82 → \"near match, review\";\n  below → honest \"no obvious match, not evidence of fabrication\" (coverage\n  lag). HTTP 403 → actionable hint to configure `--openalex-key` /\n  `OPENALEX_API_KEY` (OpenAlex requires API keys for production use since\n  2026-02; keyless calls still work within a free allowance).\n- **Semantic Scholar third source** (v1.9): when a DOI reference's DOI.org\n  metadata could not be fetched (network restricted / transient 5xx), the DOI\n  is re-queried at the S2 Graph API. Confirmation (title similarity ≥ 0.82)\n  adds a second-source note and enables the anti-bot rescue below. S2 record\n  quality is uneven (real papers registered as journal ToC pages have been\n  observed), so **S2 never speaks negatively**: mismatches and misses are\n  silent. Optional `--s2-key` / `S2_API_KEY` (keyless shares a rate pool;\n  client paces ≥1.5 s between S2 calls).\n- **Semantic Scholar title search** (v1.10): references carrying neither\n  DOI/PMID nor arXiv get one extra positive signal — the title is queried at\n  the S2 `/paper/search` endpoint (best-of full/prefix similarity ≥ 0.82 →\n  \"confirmed exists\"). Same discipline as above: confirmation only, misses\n  and mismatches silent (titles < 18 chars are skipped — search noise).\n  Per S2 API docs, hyphenated query terms return no results, so hyphens are\n  normalized to spaces before encoding.\n\n## Parallel verification & degraded-network mode (v1.10)\n\n- `--workers N` (default 4): references are independent, so up to N of them\n  are verified concurrently (offline mode and `--workers 1` stay serial).\n  Progress lines print in completion order; the report/JSON keep strict\n  index order. Shared state (host circuit breaker, arXiv/S2 rate pacing) is\n  lock-protected, so politeness guarantees hold under parallelism. The\n  per-reference `--interval` pacing applies to the serial path; the parallel\n  path throttles through concurrency itself.\n- **Degraded-network mode**: when ≥ 3 *different* hosts fail at the transport\n  layer within one run (a restricted-egress signature), remaining lookups to\n  not-yet-confirmed hosts are skipped immediately with an honest\n  \"global network degraded\" note plus a one-line suggestion (configure a\n  proxy / switch networks and rerun), and the run prints a summary advice\n  line. Site-level HTTP responses (403/429 rate limits, anti-bot) never\n  count toward degradation — the egress path itself is proven working when\n  any server answers. Once degraded, transport retries stop (fast-fail).\n  The per-host v1.6 circuit breaker still applies independently.\n- **Intra-run cache**: a reference whose DOI/URL/PMID/arXiv matches an\n  already-verified one reuses that verdict (note says so) instead of\n  re-hitting the network. `mark_duplicates` still applies afterwards — the\n  cache never promotes a duplicate.\n\n## Rescue rules (anti-bot, anti-flaky-network)\n\n- Landing page 403/429 but registry metadata confirmed (DOI.org, arXiv, or\n  S2): existence is proven, so the reference is NOT `unreachable` — it is\n  judged on tier/fields like any 200 page (trusted tier + complete fields →\n  `verified`; any partial adjustment or missing field → `partial`). This\n  prevents \"metadata confirmed + 403 blog page\" from slipping through.\n- arXiv resolver 404 with unconfirmed metadata → `invalid` (official registry\n  says no such ID).\n- Wayback Machine (v1.5): every `unreachable` link is checked against\n  archive.org; an archived copy adds a dated comparison link for human review.\n\n## Rate limiting & resilience\n\n- Host circuit breaker (v1.6): 2 consecutive *transport* failures\n  (timeout/DNS/refused) → remaining calls to that host are skipped for the\n  batch with an honest note. HTTP responses (403/404/429…) never trip it.\n  Retries within one logical call count once.\n- Transient failures retry once; error notes carry the cause (timeout / DNS /\n  refused) and the fix (`--timeout 20` for slow links).\n- 429 with `Retry-After` (v1.9): the DOI metadata check waits the requested\n  duration (capped at 5 s) before its retry.\n- `--mailto you@lab.edu` (v1.9): appended to Crossref requests — enters the\n  Crossref polite pool with more generous rate limits. Recommended for CI and\n  institutional batches. arXiv politeness (3 s/request) and S2 pacing (1.5 s\n  keyless) are always on — both are global under parallel verification.\n- OpenAlex hints (v1.10): 403 and 429 both explain the key situation —\n  OpenAlex has required API keys for production use since 2026-02 (keyless:\n  100 credits/day); configure `--openalex-key` / `OPENALEX_API_KEY`.\n\n## Exports\n\n- `--export bibtex`: verified-only bibliography, Zotero/EndNote-ready (PMID /\n  arXiv eprint noted). Field values sanitized (brace/newline stripped).\n- `--export csv`: full audit ledger (all refs, all fields, formula-injection\n  neutralized, utf-8-sig for Excel).\n- `--export auditjson` (v1.9): machine-readable working paper — per reference,\n  which checks ran and what they found (`checks_catalog` documents every check\n  type), plus a generation timestamp. Purpose: transparency and institutional\n  audit. The July 2026 arXiv survey of citation checkers (arXiv 2607.22693)\n  found none of five evaluated tools reliable enough to run unsupervised and\n  calls for \"transparent multi-source detection systems\" — the audit trail is\n  this skill's answer: a human can re-trace every mechanical verdict.\n\n## Persistent disk cache (v1.12)\n\n`--cache` is on by default. Stable verdicts (`verified` / `partial` /\n`invalid`) are stored in a local SQLite database\n(`~/.cache/cite-holmes/cache.sqlite3`, standard library) keyed by every\njudgment-relevant input: identifier, claimed title/year/source/authors, tier,\nverifier version, and profile — the same DOI with a different claimed title is\na different cache key, because title-similarity verdicts depend on the claim.\nTTL is 168 h (`--cache-ttl`); `unreachable` is never cached (transient);\n`--offline` results are never cached; `--strict` bypasses cache reads so CI\nalways runs live; `--refresh-cache` bypasses reads but still writes;\n`--no-cache` disables the store entirely. Cache hits append an explicit\n\"磁盘缓存命中\" note with...","readmeExcerpt":"Skill: Cite Holmes — Deep Research × Hallucination-Free Citations Owner: docsor1212 Summary: Deep research that interrogates its own sources (Verified Deep Research): calibrates scope first (3-5 sharp questions), then searches iteratively and machine-verifies every citation. Also a standalone citation checker: paste any reference list and it runs full citation verification (it will verify citations before you cite) a","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"python scripts/verify_refs.py --refs research_refs.json --out verify_report.md\n# Got a .bib from Zotero/EndNote? Feed it directly (v1.9):\npython scripts/verify_refs.py --refs bibliography.bib --out report.md\n# Environment self-check before a big batch (v3.10) — zero outbound calls\n# by default, add --net to probe every academic registry:\npython scripts/verify_refs.py --doctor --net"},{"language":"bash","snippet":"# 1. references as JSON (title/url/source/year per item) — or a Zotero .bib\n# 2. verify:\npython scripts/verify_refs.py --refs refs.json --out report.md\n# 3. read report.md — CiteScore + pre-submission conclusion at the top.\n# Medical? add --profile medical.  Paper?  add --export bibtex.\n# Institution/CI?  add --mailto you@lab.edu (Crossref polite pool)."},{"language":"bash","snippet":"python scripts/verify_refs.py --refs end-to-end/research_refs.json \\\n    --check-document end-to-end/paper_excerpt.md \\\n    --export gbt7714,ris,bibtex,csv --out end-to-end/report.md"},{"language":"bash","snippet":"# 1) uvx from GitHub (recommended — always current)\nuvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp\n\n# 2) uvx from a local checkout\ngit clone https://github.com/docsor1212/cite-holmes\nuvx --from ./cite-holmes/mcp cite-holmes-mcp --help\n\n# 3) pip install (module + console script)\npip install \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\"\ncite-holmes-mcp --help"},{"language":"json","snippet":"{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\"]\n    }\n  }\n}"},{"language":"json","snippet":"{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\", \"--transport\", \"stdio\"]\n    }\n  }\n}"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: cite-holmes\nversion: 3.11.0\nauthor: DoctorQ Lab\nlicense: MIT\ndescription: >-\n  Deep research that interrogates its own sources (Verified Deep Research):\n  calibrates scope first (3-5 sharp questions), then searches iteratively and\n  machine-verifies every citation. Also a standalone citation checker: paste\n  any reference list and it runs full citation verification (it will verify\n  citations before you cite) against official registries —\n  hallucinated references, fabricated DOIs, fake PMIDs, arXiv IDs, stitched\n  fakes, retracted papers; a fact check for your bibliography. Five verdicts;\n  unverified references never masquerade as real (AI hallucination detection).\n  Medical mode (Cochrane/BMJ/ChiCTR/NMPA/CDC/NICE presets, PMID check) and\n  bibliography export (BibTeX/GB·T 7714-2025/RIS/CSV + JSON workpaper) built\n  in. Trigger matching is semantic, not exact — mis-triggers are harmless;\n  state your real intent to avoid them.\nwhen_to_use: >-\n  Use when the user says \"deep research\", \"look into\", \"investigate\",\n  \"compare A vs B\", \"fact check\", \"verify this claim\", \"is it true that...\",\n  \"check these references\", \"are these citations real\", wants a research\n  report with sources, a literature review, a medical evidence lookup, or\n  wants references verified before submission — even if they never say the\n  word \"research\".\n---\n\n# cite-holmes (Cite Holmes): deep research with citation verification\n\nOne line: **a question goes in — a verified report comes out.**\n\nThree differences from a plain \"search and summarize\":\n\n1. **Calibrate before working** — ask sharp questions first; the most expensive\n   waste is researching the wrong question.\n2. **Conclusions carry evidence grades** — 🟢 two independent sources agree /\n   🟡 single authority / 🔴 contested.\n3. **Every citation is checked** — mechanical layer (reachability, domain\n   authority, field completeness, dedup) plus semantic layer (does the source\n   actually support the claim?). Unverified references never masquerade as real.\n\n## ⛔ Iron rules (zero exceptions)\n\n1. **Never fabricate**: citations must come from pages actually fetched this\n   session. Re-search rather than write URLs from memory.\n2. **Never pretend**: unchecked references are marked `unverified`; fetch\n   failures are `unreachable` (≠ nonexistent — flagged for human review).\n3. **Don't hide conflicts**: when sources disagree, present the disagreement,\n   mark 🔴, show each side's evidence.\n4. **Budgeted search**: QUICK ≤6 searches, FULL ≤15. Out of budget → state the\n   gaps honestly instead of forcing conclusions.\n5. **Calibrate before searching** (FULL mode): scope / timeframe / audience /\n   output format must be locked first.\n\n## Step 0: mode selection\n\n| Mode | Fits | Calibration | Budget | Output |\n|---|---|---|---|---|\n| **QUICK** | Single fact-check: \"is this claim true\", \"when was X released\" | skipped | ≤6 | short report |\n| **FULL** | Open research: \"state of X\", \"A vs B\", \"do a survey\" | mandatory | ≤15 "},{"path":"examples/end-to-end/README.md","content":"# 端到端对照示例（输入 → 命令 → 实际输出）\n\n输入 → 命令 → 实际输出，三件套可直接复跑：\n\n```bash\npython scripts/verify_refs.py --refs end-to-end/research_refs.json \\\n    --check-document end-to-end/paper_excerpt.md \\\n    --export gbt7714,ris,bibtex,csv --out end-to-end/report.md\n```\n\n| 文件 | 是什么 |\n|---|---|\n| research_refs.json | 输入：1 条真实 DOI 引用（含正确标题） |\n| paper_excerpt.md | 输入：含 in-text 引用标记的正文片段（叙述式 + [n] 各一处，其中一处故意跑题） |\n| report.md | 实际输出：五态判定 + 上下文核验区（Kucsko 叙述式=锚定 ✓；[1] 企鹅句=低锚 ⚠️ 演示错配检测） |\n| report.json | 机读版（含 context_check 与逐项 checks） |\n| exports/ | 四种导出的真实产物样例（BibTeX / GB·T 7714-2025 / RIS / CSV） |\n\n> 复跑需网络（DOI.org/PubMed）。判定含时效字段（撤稿/被引），数字可能随时间小幅变化\n> ——以你复跑时的输出为准，结构不变。\n\n> SkillHub 包注：为符合平台文件类型白名单，`exports/` 导出样例文件不随\n> SkillHub 包分发（GitHub/ClawHub 包完整附带）——格式以本页表格描述为准。"},{"path":"mcp/README.md","content":"# cite-holmes-mcp\n\n**Mechanical citation verification as an MCP server** — feed it the references an\nagent is about to cite, get back per-reference verdicts (`verified / partial /\nunreachable / invalid`) against official registries (DOI.org, PubMed\nE-utilities, arXiv API, Crossref/Retraction Watch, Wayback). Zero API keys\nrequired, zero telemetry, local-first.\n\nThree MCP primitives:\n\n| Primitive | Name | What it does |\n|---|---|---|\n| tool | `verify_references` | Verify a batch of references (≤50); returns CiteScore 0–100 + per-item verdicts with evidence chain |\n| tool | `check_document` | Context-level check: parse in-text markers ([12]/[1-4]/(Author, Year)/doi.org links), bind to the reference list, flag low anchor-word overlap as possible mis-citations |\n| tool | `explain_verdict` | Plain-language explanation of one verdict (evidence steps + suggested action); zero network |\n| resource | `cite-holmes://capability-matrix` | What the mechanical layer catches vs. what stays with the semantic layer |\n| resource | `cite-holmes://changelog` | Version history |\n| prompt | `fact_check_workflow` | Three-step research-then-verify workflow for agents |\n\n## Install (three channels)\n\n```bash\n# 1) uvx from GitHub (recommended — always current)\nuvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp\n\n# 2) uvx from a local checkout\ngit clone https://github.com/docsor1212/cite-holmes\nuvx --from ./cite-holmes/mcp cite-holmes-mcp --help\n\n# 3) pip install (module + console script)\npip install \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\"\ncite-holmes-mcp --help\n```\n\n> This is a Python package (uv/pip ecosystem). There is no npm package — the\n> uvx/git channel above is the canonical install for all MCP clients.\n\n## Client configuration\n\n**Claude Code**\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\"]\n    }\n  }\n}\n```\n\n**Codex / any stdio MCP client**\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\", \"--transport\", \"stdio\"]\n    }\n  }\n}\n```\n\n**Cursor** (`~/.cursor/mcp.json`, same shape)\n\n```json\n{\n  \"mcpServers\": {\n    \"cite-holmes\": {\n      \"command\": \"uvx\",\n      \"args\": [\"--from\", \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\",\n               \"cite-holmes-mcp\"]\n    }\n  }\n}\n```\n\nHTTP mode (local): `cite-holmes-mcp --transport http --port 8760`.\n\n## Bounds & behavior\n\n- Batch ≤ 50 references per `verify_references` call (larger batches get a\n  clear error — split them).\n- Long output fields are clipped (800 chars + truncation marker) so MCP\n  messages stay bounded; structure is preserved.\n- Per-reference network timeout defaults to 15 s (configurable per call).\n- Optional API keys via env: `OPENALEX_A"},{"path":"README.md","content":"# Cite Holmes 🔍\n\n[![GitHub Stars](https://img.shields.io/github/stars/docsor1212/cite-holmes?style=social&label=Star)](https://github.com/docsor1212/cite-holmes/stargazers)\n\n![icon](assets/icon-512.png)\n\n**Deep research that interrogates its own sources.**\n\nEvery AI research report you've ever read had a dirty secret: some of those polished references were probably fabricated. [A Nature news analysis suggests tens of thousands of 2025 publications might include invalid AI-generated references](https://www.nature.com/articles/d41586-026-00969-z). [GPTZero scanned 4,841 NeurIPS 2025 submissions; as independently reported, at least 100 hallucinated citations were found across 51 accepted papers](https://medium.com/@ljingshan6/100-fake-citations-just-slipped-through-neurips-2025-peer-review-5f34f4436560).\n\nCite Holmes is a deep-research skill with a badge and a magnifying glass: it researches like any deep-research agent — then **arrests its own citations before you can cite them**.\n\n![demo](assets/demo.gif)\n\n*(Demo is real output: 8 references, 3 deliberately planted fabrications — a fake DOI, a dead URL, and a no-URL citation. All 3 were caught and excluded; the 5 real ones passed. Measured: **7.7 s for all 8** — 4.2 s of pure network checks, the rest is deliberate throttling.)*\n\nReproduce it yourself — the planted-fakes file ships with the repo:\n\n```bash\npython scripts/verify_refs.py --refs examples/demo_refs.json\n```\n\n## 30-second quickstart\n\n```bash\n# MCP server (for agents — Claude Code / Codex / Cursor):\nuvx --from \"git+https://github.com/docsor1212/cite-holmes#subdirectory=mcp\" cite-holmes-mcp\n```\n\n\n```bash\npython scripts/verify_refs.py --refs refs.json --out report.md\n# or verify straight from your reference manager (v1.9):\npython scripts/verify_refs.py --refs bibliography.bib --out report.md\n# v3.7.0 context check — verify in-text citations inside a finished draft:\npython scripts/verify_refs.py --refs bibliography.bib --check-document paper.md --out report.md\n```\n\n**China mirror (ModelScope 魔搭)**: <https://modelscope.cn/skills/Docsor/cite-holmes> — if you find this skill useful, a like there helps others find it.\n\nOpen `report.md`: CiteScore + pre-submission conclusion at the top,\nper-reference verdicts below.\nMedical work: add `--profile medical`.  Writing a paper: add `--export bibtex`.\nInstitution/CI: add `--mailto you@lab.edu` (Crossref polite pool).\n\n## What NOT to do (anti-patterns)\n\n- **Don't feed `semantic` fields from the same model that wrote the draft** —\n  the semantic cap trusts structured model judgments; self-review defeats it.\n- **Don't treat `verified` as \"the paper supports my claim\"** — verified means\n  the source exists at an authoritative tier and registry metadata matches the\n  claimed title/authors/journal/year. Whether the *specific sentence* is\n  supported is the semantic layer (agent judgment, or L4 cascade with a judge).\n- **Don't batch >50 references per MCP call** — the server rejects oversized\n  batches by des"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b6sjwend7cwwmhf7mdg1wzx82dfbf\",\n  \"slug\": \"cite-holmes\",\n  \"version\": \"3.11.0\",\n  \"publishedAt\": 1791566549713\n}"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"Deep research that interrogates its own sources (Verified Deep Research): calibrates scope first (3-5 sharp questions), then searches iteratively and machine-verifies every citation. Also a standalone citation checker: paste any reference list and it runs full citation verification (it will verify citations before you cite) against official registries — hallucinated references, fabricated DOIs, fake PMIDs, arXiv IDs, stitched fakes, retracted papers; a fact check for your bibliography. Five verdicts; unverified references never masquerade as real (AI hallucination detection). Medical mode (Cochrane/BMJ/ChiCTR/NMPA/CDC/NICE presets, PMID check) and bibliography export (BibTeX/GB·T 7714-2025/RIS/CSV + JSON workpaper) built in. Trigger matching is semantic, not exact — mis-triggers are harmless; state your real intent to avoid them. Skill: Cite Holmes — Deep Research × Hallucination-Free Citations Owner: docsor1212 Summary: Deep research that interrogates its own sources (Verified Deep Research): calibrates scope first (3-5 sharp questions), then searches iteratively and machine-verifies every citation. Also a standalone citation checker: paste any reference list and it runs full citation verification (it will verify citations before you cite) a","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":2140,"uniquenessScore":47,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T08:15:53.892Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T08:15:53.892Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T10:43:06.182Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}