{"id":"653c61fe-cef9-494e-b24a-bc7e3cd9da5d","entityType":"agent","slug":"clawhub-docsor1212-pubmed-verifier","name":"pubmed-verifier","canonicalUrl":"https://www.xpersona.co/agent/clawhub-docsor1212-pubmed-verifier","canonicalPath":"/agent/clawhub-docsor1212-pubmed-verifier","generatedAt":"2026-10-09T23:50:28.317Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T20:21:50.745Z","emptyReason":null},"description":"Reference checker for AI-fabricated citations: batch-verify PMIDs against PubMed and catch the hallucination existence checks miss — a REAL PMID pointing to a DIFFERENT paper. Five-state citation verification (correct / mismatch / partial / invalid / unknown), citation-context parsing, Crossref DOI cross-check, retraction detection (capped at partial), correct-PMID suggestion, arXiv ID verification, plain-text reference list parsing, formatted reference lists (GB/T 7714 / Vancouver / APA / AMA), OpenAlex DOI fallback for non-Crossref registries, Chinese-reference title search, CSV/JSON claims, HTML/JSON/text reports. Four data sources: NCBI, Europe PMC fallback, Crossref, OpenAlex. Network failures are honestly reported as unverified, never as \"not found\". Zero dependencies, runs fully local. Triggers: verify PMIDs, check citations, citation audit, reference check, PMID check, batch verify references, AI hallucination detection, verify DOI, DOI check, citation formatting, reference formatter, GB/T 7714.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 2K downloads reported by the source. Last updated 10/9/2026.","installCommand":"clawhub skill install s17dagtwyk21qs6vpz98bzcrh1853t29:pubmed-verifier","sourceUrl":"https://clawhub.ai/docsor1212/pubmed-verifier","homepage":"https://clawhub.ai/docsor1212/skills/pubmed-verifier","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/docsor1212/pubmed-verifier","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/docsor1212/skills/pubmed-verifier","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":66,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"pubmed-verifier technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-09T20:21:50.745Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T20:21:50.745Z","emptyReason":null},"stars":null,"forks":null,"downloads":2030,"packageName":null,"latestVersion":"4.1.0","tractionLabel":"2K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T20:21:50.745Z","emptyReason":null},"lastUpdatedAt":"2026-10-09T20:21:50.745Z","lastCrawledAt":"2026-10-09T20:21:50.745Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-10T20:21:50.745Z","lastVerifiedAt":null,"highlights":[{"version":"4.1.0","createdAt":"2026-10-09T16:55:39.147Z","changelog":"OpenAlex is now the fourth metadata source: a DOI that Crossref reports 404 gets a second look at OpenAlex before any verdict — DataCite/Zenodo/Chinese-registry DOIs become fully decidable (only unknown-in-BOTH-registries keeps the 404 semantics). Chinese references without any ID are now verifiable: parse-text CJK entries search OpenAlex by their title-shaped segment (verbatim-substring + year gates, resolved_by=openalex_cjk_search, never a mismatch source). --check-net probes five sources; docs add output-shape samples and three worked FAQ scenarios; README gains a Chinese intro. 212 offline + 31 real-network tests green. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。","fileCount":11,"zipByteSize":92393},{"version":"4.0.0","createdAt":"2026-10-08T19:10:59.265Z","changelog":"Plain-text reference lists and formatted reference list export. --parse-text refs.txt verifies a reference list copied straight out of a manuscript draft: numbered entries split cleanly, inline PMID/DOI/arXiv IDs route exactly (full five-state cross-check), title-only entries resolve via PubMed title search honestly labeled resolved_by=title_search (never a mismatch source, network failures stay unknown). --format-references out.txt --citation-style gbt|vancouver|apa|ama renders verified entries as a ready-to-paste list (GB/T 7714-2015, Vancouver, APA 7th, AMA 11th) — partial entries go to a manual-review section, retracted ones are excluded with a warning. 200 offline + 28 real-network tests green. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。","fileCount":11,"zipByteSize":86860},{"version":"3.9.0","createdAt":"2026-10-08T04:01:32.256Z","changelog":"Chinese references (GB/T 7714) now parse: title/authors/journal/year/DOI extraction from contexts like ……标题[J]. 刊名, 年… PMID: xxx, with CJK-vs-Latin title/journal comparisons skipped cross-language (never false mismatches); --check-net pre-flight probes the four data sources and prints practical next steps; references/python_api.md documents stable Python integration surfaces. 181 offline + 26 real-network tests green.","fileCount":10,"zipByteSize":78638},{"version":"3.8.0","createdAt":"2026-10-07T04:01:08.235Z","changelog":"Verdict-ladder corrections from an independent real-data audit: scan-mode author lines are no longer taken as claimed titles (B1 false positives); an entirely different claimed author set now caps at partial (B2 mis-attribution); title-less claims cap at partial exactly as documented (B3); merged duplicates announce on stderr; Chinese-context guidance added. 173 offline + 26 real-network tests green.","fileCount":9,"zipByteSize":73959},{"version":"3.7.0","createdAt":"2026-10-06T05:05:49.738Z","changelog":"RIS bibliography support: --bibliography refs.ris (Zotero/EndNote/Mendeley exports) routes records by PMID (AN tag or note) > DOI (DO) > arXiv (UR/eprint) with full claimed-metadata cross-checks; --export-ris completes the loop (correct = TY JOUR, partial = TY DATA with PARTIAL note); --lint-claims extends to .ris. 164 offline + 25 real-network tests green. desc: add Trigger on routing words.","fileCount":9,"zipByteSize":71573},{"version":"3.6.0","createdAt":"2026-10-05T04:00:35.324Z","changelog":"BibTeX bibliography audit: --bibliography refs.bib verifies a .bib file directly — entries route by PMID (field or note, including notes --export-bibtex itself writes) > DOI > arXiv eprint/ID pattern, each with claimed metadata for the full cross-check; --lint-claims extends to .bib (unroutable entries, missing titles, duplicates, unclosed blocks); round-trips with --export-bibtex. 158 offline + 25 real-network tests green.","fileCount":9,"zipByteSize":68941},{"version":"3.5.0","createdAt":"2026-10-03T19:24:55.873Z","changelog":"Claims lint: --lint-claims validates a claims file offline before any run (ID shapes, missing titles, unknown/typo columns, duplicates, DOI prefix — zero network, exit-coded). Documentation restructured around usage: FAQ, anti-patterns and declared boundaries right after Quick start with a Top-10 NOT-to-do table, new Best practices & tuning section (API-key batching, workers, two-phase deep-verification, cache policy, scale expectations), trigger priority & tool-choice guidance, claims reference format. 148 offline + 24 real-network tests green.","fileCount":9,"zipByteSize":65500},{"version":"3.4.0","createdAt":"2026-10-03T05:00:56.008Z","changelog":"DOI claims become first-class: claims rows keyed by DOI (no PMID/arXiv) now get the full verdict ladder (correct/mismatch/partial) against the linked PubMed record or Crossref metadata, with retraction capping; claims-sourced DOI 404 counts as a fabrication signal (user-endorsed, like --dois); audit working-papers record the claimed input for third-party replay. 140 offline + 24 real-network tests green.","fileCount":9,"zipByteSize":61608}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17dagtwyk21qs6vpz98bzcrh1853t29:pubmed-verifier","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17dagtwyk21qs6vpz98bzcrh1853t29:pubmed-verifier` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/docsor1212/pubmed-verifier before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-09T23:50:28.313Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-pubmed-verifier/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T20:21:50.745Z","emptyReason":null},"readme":"Skill: pubmed-verifier\n\nOwner: docsor1212\n\nSummary: Reference checker for AI-fabricated citations: batch-verify PMIDs against PubMed and catch the hallucination existence checks miss — a REAL PMID pointing to a DIFFERENT paper. Five-state citation verification (correct / mismatch / partial / invalid / unknown), citation-context parsing, Crossref DOI cross-check, retraction detection (capped at partial), correct-PMID suggestion, arXiv ID verification, plain-text reference list parsing, formatted reference lists (GB/T 7714 / Vancouver / APA / AMA), OpenAlex DOI fallback for non-Crossref registries, Chinese-reference title search, CSV/JSON claims, HTML/JSON/text reports. Four data sources: NCBI, Europe PMC fallback, Crossref, OpenAlex. Network failures are honestly reported as unverified, never as \"not found\". Zero dependencies, runs fully local. Triggers: verify PMIDs, check citations, citation audit, reference check, PMID check, batch verify references, AI hallucination detection, verify DOI, DOI check, citation formatting, reference formatter, GB/T 7714.\n\nTags: latest:4.1.0\n\nVersion history:\n\nv4.1.0 | 2026-10-09T16:55:39.147Z | user\n\nOpenAlex is now the fourth metadata source: a DOI that Crossref reports 404 gets a second look at OpenAlex before any verdict — DataCite/Zenodo/Chinese-registry DOIs become fully decidable (only unknown-in-BOTH-registries keeps the 404 semantics). Chinese references without any ID are now verifiable: parse-text CJK entries search OpenAlex by their title-shaped segment (verbatim-substring + year gates, resolved_by=openalex_cjk_search, never a mismatch source). --check-net probes five sources; docs add output-shape samples and three worked FAQ scenarios; README gains a Chinese intro. 212 offline + 31 real-network tests green. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。\n\nv4.0.0 | 2026-10-08T19:10:59.265Z | user\n\nPlain-text reference lists and formatted reference list export. --parse-text refs.txt verifies a reference list copied straight out of a manuscript draft: numbered entries split cleanly, inline PMID/DOI/arXiv IDs route exactly (full five-state cross-check), title-only entries resolve via PubMed title search honestly labeled resolved_by=title_search (never a mismatch source, network failures stay unknown). --format-references out.txt --citation-style gbt|vancouver|apa|ama renders verified entries as a ready-to-paste list (GB/T 7714-2015, Vancouver, APA 7th, AMA 11th) — partial entries go to a manual-review section, retracted ones are excluded with a warning. 200 offline + 28 real-network tests green. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。\n\nv3.9.0 | 2026-10-08T04:01:32.256Z | user\n\nChinese references (GB/T 7714) now parse: title/authors/journal/year/DOI extraction from contexts like ……标题[J]. 刊名, 年… PMID: xxx, with CJK-vs-Latin title/journal comparisons skipped cross-language (never false mismatches); --check-net pre-flight probes the four data sources and prints practical next steps; references/python_api.md documents stable Python integration surfaces. 181 offline + 26 real-network tests green.\n\nv3.8.0 | 2026-10-07T04:01:08.235Z | user\n\nVerdict-ladder corrections from an independent real-data audit: scan-mode author lines are no longer taken as claimed titles (B1 false positives); an entirely different claimed author set now caps at partial (B2 mis-attribution); title-less claims cap at partial exactly as documented (B3); merged duplicates announce on stderr; Chinese-context guidance added. 173 offline + 26 real-network tests green.\n\nv3.7.0 | 2026-10-06T05:05:49.738Z | user\n\nRIS bibliography support: --bibliography refs.ris (Zotero/EndNote/Mendeley exports) routes records by PMID (AN tag or note) > DOI (DO) > arXiv (UR/eprint) with full claimed-metadata cross-checks; --export-ris completes the loop (correct = TY JOUR, partial = TY DATA with PARTIAL note); --lint-claims extends to .ris. 164 offline + 25 real-network tests green. desc: add Trigger on routing words.\n\nv3.6.0 | 2026-10-05T04:00:35.324Z | user\n\nBibTeX bibliography audit: --bibliography refs.bib verifies a .bib file directly — entries route by PMID (field or note, including notes --export-bibtex itself writes) > DOI > arXiv eprint/ID pattern, each with claimed metadata for the full cross-check; --lint-claims extends to .bib (unroutable entries, missing titles, duplicates, unclosed blocks); round-trips with --export-bibtex. 158 offline + 25 real-network tests green.\n\nv3.5.0 | 2026-10-03T19:24:55.873Z | user\n\nClaims lint: --lint-claims validates a claims file offline before any run (ID shapes, missing titles, unknown/typo columns, duplicates, DOI prefix — zero network, exit-coded). Documentation restructured around usage: FAQ, anti-patterns and declared boundaries right after Quick start with a Top-10 NOT-to-do table, new Best practices & tuning section (API-key batching, workers, two-phase deep-verification, cache policy, scale expectations), trigger priority & tool-choice guidance, claims reference format. 148 offline + 24 real-network tests green.\n\nv3.4.0 | 2026-10-03T05:00:56.008Z | user\n\nDOI claims become first-class: claims rows keyed by DOI (no PMID/arXiv) now get the full verdict ladder (correct/mismatch/partial) against the linked PubMed record or Crossref metadata, with retraction capping; claims-sourced DOI 404 counts as a fabrication signal (user-endorsed, like --dois); audit working-papers record the claimed input for third-party replay. 140 offline + 24 real-network tests green.\n\nv3.3.0 | 2026-10-02T05:01:07.650Z | user\n\narXiv preprint-to-published-version cross-check: claimed DOI vs the arXiv-registered version-of-record DOI (agreement = evidence, mismatch caps at partial); bare preprint citations surface the registered DOI and linked PMID; abbreviation-safe context parsing; ETA progress on stderr; readiness/exit-code count arXiv DOI-pairing mismatches; fixed markdown --diff crash (P0) and latent arXiv stats double-count (P1, since v3.0.0). 129 offline + 20 real-network tests green.\n\nv3.2.0 | 2026-10-01T04:16:35.797Z | user\n\narXiv ID verification + three citation types in one audit. 109 offline + 16 realnet green.\n\nv3.1.0 | 2026-09-30T05:47:40.392Z | user\n\narXiv claims verification closes the loop: claims now accept arxiv_id as the key (--claims/--claims-file/CSV); a resolved arXiv ID is compared against the claimed title (correct/mismatch). Anti-patterns section (6 wrong-usages with fixes) and claims.csv reference format added per reviewer feedback. PMID batch progress output. 100+ offline / 16 real-network tests green.\n\nv3.0.0 | 2026-09-29T03:30:28.907Z | user\n\nMilestone: arXiv ID verification — arXiv:2401.12345 and arxiv.org/abs patterns are extracted from scans (or --arxivs) and checked against the official arXiv API; a nonexistent ID is a fabrication signal (invalid). PMIDs, DOIs and arXiv IDs now audited in one pass. Timely: arXiv penalizes submissions containing hallucinated references (2026-05 policy). 109 offline + 16 real-network tests green.\n\nv2.9.0 | 2026-09-28T04:30:22.571Z | user\n\nDOI entries become first-class: a resolved DOI is linked back to its PMID via the Europe PMC DOI query, attaching the full PubMed record (five-state metadata, retraction pubtype, cache). Parallel DOI resolution --workers (default 4) — 1,000 DOIs drop to minutes with live progress. New --export-csv audit table (formula-injection hardened). 100 offline + 16 real-network tests green.\n\nv2.8.0 | 2026-09-27T04:16:09.211Z | user\n\nDOI-native verification: --dois verifies DOIs without PMIDs; --source scans auto-extract DOIs from your files; Crossref not-found on an explicitly provided DOI = fabrication signal (auto-extracted ones stay honest suspects — DataCite DOIs do not live in Crossref). Delta audits: --diff against a previous working-paper reports newly retracted, degraded, improved, new and dropped. 89 offline + 16 real-network tests green.\n\nv2.7.0 | 2026-09-26T04:30:36.356Z | user\n\nRetraction detection for EVERY PMID: the registry publication type (Retracted Publication; NCBI + Europe PMC) flags retracted papers with no DOI and no flags needed — and survives the cache (schema v3). Crossref updated-by stays as the detail source. Adds --export-bibtex verified bibliography (correct included, partial commented, retracted excluded) and a submission-readiness verdict heading every report. 75 offline + 14 real-network tests green.\n\nv2.6.0 | 2026-09-25T03:50:56.782Z | user\n\nAudit working-paper: --export-audit writes a self-contained JSON trail (tool identity, redacted invocation, per-citation evidence chain, verdict trace) a third party can replay. HTML report v2: verdict filter tabs, severity sort, field-level evidence column, retraction/splice highlighting, reproducibility footer. Reliability: negative cache expires in 3 days; circuit breaker self-heals after a 30s cooldown. 67 offline + 12 real-network tests green.\n\nv2.5.0 | 2026-09-24T04:00:12.060Z | user\n\nAuthor-name verification: single-letter initials never take part in surname matching (no more false author hits); CJK/Latin cross-language author names are honestly skipped, not counted as mismatches; nothing comparable falls back to unknown. Adds a Security & behavior declaration (100% local, three academic registries only, zero telemetry). 61 offline + 11 real-network tests green.\n\nv2.4.0 | 2026-09-23T03:30:55.840Z | user\n\nDOI-splice cross-check: claims now accept a doi field; a claimed DOI differing from the registered one is flagged (capped at partial, exit 1) — no extra API calls. Bidirectional NLM journal abbreviation matching (N Engl J Med ~ New England Journal of Medicine). FAQ + when-to-invoke docs. 56 offline + 11 real-network tests green.\n\nv2.3.0 | 2026-09-20T21:48:57.963Z | user\n\nRetraction detection: with --verify-doi, papers Crossref lists as RETRACTED are flagged and their verdict capped at partial with a human-review note (retracted citations exit 1). RETRACTED-title prefixes are stripped before DOI title comparison; defensive parsing of Crossref update records. 45 offline + 9 real-network tests green, including a real retracted paper (STAP 2014). Zero new dependencies.\n\nv2.2.1 | 2026-09-19T21:43:09.512Z | user\n\nCompliance hygiene: removed an evasion-flavored phrase (de-AI rewriting) from the family cross-link section; product functionality unchanged from 2.2.0 (dual-source verification, network hardening, honest unknown). 36 offline + 8 real-network tests green.\n\nv2.2.0 | 2026-09-19T19:50:07.173Z | user\n\nNetwork hardening: NCBI API key support (faster batches), automatic Europe PMC fallback when NCBI is unreachable (--meta-source), Crossref polite pool via --mailto, 429 Retry-After backoff, 403/406 UA rotation, per-host circuit breaker. Honest verdicts: network failures are reported as unknown (exit code 2), never as not-found, and never cached. DOI verification now compares the Crossref title (doi_title_match). Bilingual docs (SKILL_ZH) + full test suites (36 offline + 8 real-network acceptance).\n\nv2.1.4 | 2026-04-20T16:16:48.188Z | user\n\nShort bilingual description\n\nv2.1.3 | 2026-04-20T16:11:21.713Z | user\n\nClean English description\n\nv2.1.2 | 2026-04-20T16:09:04.311Z | user\n\nCompact bilingual description\n\nv2.1.1 | 2026-04-20T16:02:42.045Z | user\n\nBilingual description (English + 中文)\n\nv2.1.0 | 2026-04-20T15:25:58.200Z | user\n\nv2.1: Five-state verdict system (correct/mismatch/partial/invalid/unknown), citation context parser with fuzzy matching, suggest correct PMID, SQLite cache, CSV claims-file, Crossref DOI verification, retry with exponential backoff\n\nv1.3.0 | 2026-04-18T04:13:40.932Z | user\n\n中文化简介描述，面向国内用户优化\n\nv1.2.0 | 2026-04-18T02:31:17.581Z | user\n\n医学缩写展开增强内容匹配，同领域匹配率35%→95%\n\nv1.1.0 | 2026-04-18T00:44:47.759Z | user\n\nSEO: enriched description with bilingual keywords, added README.md with use cases and real-world results\n\nv1.0.0 | 2026-04-18T00:39:52.707Z | user\n\nInitial release: batch PMID verification via PubMed E-utilities API, content keyword matching, HTML/JSON/text reports\n\nArchive index:\n\nArchive v4.1.0: 11 files, 92393 bytes\n\nFiles: examples/claims.sample.csv (469b), examples/README.txt (1480b), examples/refs_list.sample.txt (512b), README.md (8122b), references/api_examples.md (4849b), references/python_api.md (3617b), scripts/verify_pmids.py (219298b), skill-card.md (2009b), SKILL.md (42076b), skillhub-meta.json (1397b), _meta.json (134b)\n\nFile v4.1.0:SKILL.md\n\n---\nname: pubmed-verifier\nlicense: MIT-0\nversion: 4.1.0\ndescription: >-\n  Reference checker for AI-fabricated citations: batch-verify PMIDs against\n  PubMed and catch the hallucination existence checks miss — a REAL PMID\n  pointing to a DIFFERENT paper. Five-state citation verification (correct /\n  mismatch / partial / invalid / unknown), citation-context parsing, Crossref\n  DOI cross-check, retraction detection (capped at partial), correct-PMID\n  suggestion, arXiv ID verification, plain-text reference list parsing,\n  formatted reference lists (GB/T 7714 / Vancouver / APA / AMA), OpenAlex DOI\n  fallback for non-Crossref registries, Chinese-reference title search,\n  CSV/JSON claims, HTML/JSON/text reports. Four data sources: NCBI, Europe PMC\n  fallback, Crossref, OpenAlex. Network failures are honestly reported as\n  unverified, never as \"not found\". Zero dependencies, runs fully local.\n  Triggers: verify PMIDs, check citations, citation audit, reference check,\n  PMID check, batch verify references, AI hallucination detection, verify DOI,\n  DOI check, citation formatting, reference formatter, GB/T 7714.\n---\n\n# PubMed Citation Verifier v4.1.0\n\nBatch verification of PMID citations via the PubMed E-utilities API. Not just\n\"does this PMID exist\" — **does this PMID point to the paper you claim?**\nZero dependencies, pure standard library, fully local.\n\n## When to use this skill\n\nInvoke it whenever citation truth matters:\n\n- \"Verify / check these PMIDs / this reference list\" (before submission or release)\n- Auditing citations in AI-generated text (fabricated or mismatched references)\n- Spot-checking a systematic review's bibliography\n- \"Does PMID 12345678 really say X?\" — point-of-doubt verification\n- Batch QA of a knowledge base's PMID/DOI citations\n- Automatic invocation: trigger this skill whenever the user asks to\n  \"verify/check/audit\" citations, PMIDs, DOIs, arXiv preprints or\n  references — no flags needed for a first pass; add `--claims-file`\n  when they supply the expected titles.\n\n**Trigger priority & tool choice** — explicit \"verify / check / audit\ncitations, references, PMIDs, DOIs\" requests invoke this skill first.\ncite-holmes is for deep research with machine-verified citations; when a\nrequest mixes research and verification, run the research first, then\nthis tool for the final reference audit.\n\n## The five-state verdict\n\n| Verdict | Meaning |\n|---------|---------|\n| ✅ Correct | PMID exists AND matches the claimed paper |\n| ⚠️ Mismatch | PMID exists but points to a **different** paper (the most common AI hallucination!) |\n| 🔶 Partial | Some metadata matches (e.g. author+journal but title differs) |\n| ❌ Invalid | PMID does not exist in PubMed |\n| ❓ Unknown | Not enough claimed metadata to cross-check — or both data sources unreachable (never misreported as invalid) |\n\n**Why existence checks are not enough:** a large share of fabricated\ncitations use REAL PMIDs that point to a different paper from the same\nyear/journal/field — in one of our own audits, 4 of 5 \"valid\" PMIDs were\nwrong this way. A binary exists/not-exists check misses them all.\n\n## Quick start\n\n```bash\n# Scan a project directory for PMIDs (parses citation context automatically)\npython3 scripts/verify_pmids.py --source /path/to/project --output report.html\n\n# Verify specific PMIDs\npython3 scripts/verify_pmids.py --pmids 31018962,22213727\n\n# Mismatch demo: PMID 34078778 is actually a dental-materials paper, so the\n# JIA claims below will NOT match it — expect ⚠️ mismatch verdicts\npython3 scripts/verify_pmids.py --claims '[{\"pmid\":\"34078778\",\"title\":\"JIA pathogenesis\",\"authors\":[\"Zaripova\"],\"journal\":\"Pediatr Rheumatol Online J\",\"year\":\"2021\"}]' --output report.html\n\n# Claims from a CSV file + suggest correct PMIDs for mismatches\npython3 scripts/verify_pmids.py --claims-file claims.csv --suggest --output report.html\n\n# Crossref DOI cross-verification + audit working-paper + BibTeX + full pipeline\npython3 scripts/verify_pmids.py --source /path/to/files --verify-doi --suggest --output report.html --export-audit audit.json --export-bibtex refs.bib\n\n# Verify DOIs directly (no PMIDs) + delta audit vs a previous run\npython3 scripts/verify_pmids.py --dois \"10.1038/nature12968,10.4012/dmj.2020-408\" --workers 4 --export-audit audit.json --export-csv table.csv\npython3 scripts/verify_pmids.py --source /path/to/project --diff audit.json --output report.html\n\n# Verify arXiv IDs (preprints) — mixed audits supported\npython3 scripts/verify_pmids.py --arxivs \"2401.12345,cs/0211004\" --no-cache\n\n# Pre-flight: are the five data sources reachable right now?\npython3 scripts/verify_pmids.py --check-net\n\n# Audit a BibTeX or RIS bibliography file directly (PMID > DOI > arXiv routing)\npython3 scripts/verify_pmids.py --bibliography refs.bib --no-cache\npython3 scripts/verify_pmids.py --bibliography refs.ris --no-cache --export-ris verified.ris\n\n# Verify a reference list copied from a paper draft (inline IDs route\n# exactly; title-only entries resolve via PubMed title search)\npython3 scripts/verify_pmids.py --parse-text references.txt --no-cache\n\n# Format the verified entries as a ready-to-paste reference list\n# (--citation-style: gbt | vancouver | apa | ama; default gbt)\npython3 scripts/verify_pmids.py --parse-text references.txt --format-references refs_gbt.txt --citation-style gbt\n\n# Institutional niceties (recommended): NCBI API key + contact email\npython3 scripts/verify_pmids.py --source . --verify-doi --ncbi-api-key $NCBI_API_KEY --mailto you@lab.org\n```\n\n## What the output looks like\n\nTerminal text report (verdict line, then per-citation evidence):\n\n```\nReadiness: NOT SUBMISSION-READY — 1 invalid, 1 mismatched\nResults: 1/3 correct, 1 mismatch, 1 invalid, 0 partial, 0 unknown\n============================================================\n\n✅ PMID 31018962 (claims) [correct]\n   Actual: Classification criteria for autoinflammatory recurrent fevers\n   Claimed: Classification criteria for autoinflammatory recurrent fevers\n   Evidence: title ✓ · author ✓ · journal ✓ · year ✓\n\n⚠️ PMID 34078778 (claims) [mismatch]\n   Actual: Effect of CAD/CAM materials on the marginal fit of crowds\n   Claimed: JIA pathogenesis and treatment\n   → Suggest: PMID 34425842 - Juvenile idiopathic arthritis: from aetio...\n   Evidence: title ✗ · author ✗ · journal ✗ · year —\n\n❌ PMID 99999999 (cli) [invalid]\n   Error: PMID not found in API response\n```\n\n`--output report.json` carries the same verdicts plus the full evidence\nchain per citation (`evidence.title_match / author_match / journal_match /\nyear_match`, `fields` tri-state, `meta_source` naming which registry\nanswered, `resolved_by` when the title-search/OpenAlex legs matched) and a\n`stats` block with the submission-readiness verdict. `--format-references`\nwrites the ready-to-paste citation list shown in the v4.0.0 section.\n\n## FAQ & common mistakes\n\n**Top 10 things NOT to do** (each is detailed below or in Anti-patterns):\n\n| # | Don't | Do instead |\n|---|-------|------------|\n| 1 | Treat `--pmids` existence output as \"verified\" | Feed `--claims-file` with titles for real verification |\n| 2 | Submit claims without `title` | Always include titles — the verdict caps at partial without one |\n| 3 | Trust cached verdicts on publication day | Final check with `--no-cache` |\n| 4 | Read \"not found\" as \"fabricated\" for auto-extracted DOIs | Check doi.org / arxiv.org by hand first |\n| 5 | Treat the leading `'` in CSV cells as corruption | It is the formula-injection guard — strip after import |\n| 6 | Read the READY line as a quality score | It means \"no problems among the checks that ran\" |\n| 7 | Pass `--source` together with `--pmids` | `--source` is ignored entirely when `--pmids` is given |\n| 8 | Deep-verify (`--verify-doi` / `--suggest`) a thousand-entry sweep | Sweep first, deep-verify the flagged subset |\n| 9 | Expect author matching across CJK↔Latin names | They are skipped honestly (`author_check: skipped`) |\n| 10 | Ship a reference list without the audit trail | `--export-audit` writes a replayable working paper |\n\n**Large batch (hundreds of PMIDs) is slow — how to speed it up?**\nMetadata-only verification queries in batches of 50 with 0.4 s spacing\n(0.12 s with `--ncbi-api-key`); cached re-runs are ~5 s. `--verify-doi` adds\none Crossref call *per citation* and `--suggest` adds one search *per\nmismatch* — skip them for bulk sweeps, run them on the flagged subset.\n\n**When must I use `--claims-file` instead of scanning?**\nContext parsing is heuristic (v3.3.0 guards common abbreviations, exotic formatting can still mis-split).\nFor precise verification — or DOIs in claims (splice detection needs `doi`)\n— feed structured JSON/CSV claims.\n\n**My citation text is in Chinese — the scan parses little?**\nGB/T 7714-style references (……标题[J]. 刊名, 年… PMID: xxx) parse since\nv3.9.0 — title/journal comparisons against Latin registries are skipped\ncross-language, so verdicts rest on year/DOI evidence. Free Chinese prose\nwithout a PMID marker still yields little — feed `--claims-file` for full\nverdicts regardless of language.\n\n**Slow or unstable network (China)?**\nStandard `HTTPS_PROXY`/`HTTP_PROXY` env vars are honored natively; raise\n`--timeout`; `--meta-source europepmc` routes via Europe PMC when NCBI is\nunreachable (per-entry `meta_source` shows which was used); cached results\nare reused for 30 days.\n\n**❓ unknown vs ❌ invalid?**\n`unknown` (exit 2) = \"could not verify, sources unreachable\" — retry later;\n`invalid` (exit 1) = \"verified not-found\". Network failures are never\nreported as not-found and never cached.\n\n**What does RETRACTED mean in a report?**\nThe registry itself lists the paper's publication type as \"Retracted\nPublication\" (checked for every citation since v2.7.0 — no DOI or flags\nneeded), and/or Crossref records a retraction. The verdict is capped at\npartial and a human review note is attached — citing it would propagate\nwithdrawn science. A retraction *notice* is never flagged; papers under\n*Expression of Concern* (an editorial note, not a retraction) are not\nflagged either. Retraction status reflects the registry at cache time — for\na final pre-submission check, run with `--no-cache`.\n\n**Mismatch reported but the title looks similar?**\nCheck `details` for which field diverged; thresholds are strict on purpose.\nFeed the full citation via `--claims-file` for a precise verdict.\n\n**Can I audit my .bib or .ris file directly?**\nYes — `--bibliography refs.bib` (BibTeX) and `--bibliography refs.ris`\n(RIS/Zotero/EndNote/Mendeley) route each entry by PMID > DOI > arXiv\nand cross-check the claimed metadata. `--lint-claims` validates either\nformat offline first. The round trip works both ways: `--export-bibtex`\nand `--export-ris` output can be fed back after edits.\n\n**My claims file seems to lose rows / verdicts look weaker than expected?**\nLint it offline first: `python3 scripts/verify_pmids.py --lint-claims\nclaims.csv` reports unusable rows, ID shape errors, unknown columns\n(typo'd headers like \"titel\"), missing titles and DOI prefix problems —\nno network, exit 1 on errors.\n\n**Can I verify a DOI with claimed metadata (full verdict)?**\nYes — since v3.4.0 a claims row keyed by `doi` (with `title`, optionally\n`authors`/`journal`/`year`) gets the same cross-check as PMID claims:\ncorrect / mismatch / partial against the registered metadata. A DOI row\nwithout claims stays `unknown` (existence only).\n\n**Why did my arXiv citation drop from correct to partial?**\nYour claims row paired an `arxiv_id` with a `doi`, and the DOI does not\nmatch the version-of-record DOI registered on that arXiv entry — a typo,\nor a DOI from a different paper. The registered DOI is in `details`;\nfix the claim or drop the `doi` cell.\n\n### Worked scenarios\n\n- **\"The reviewer asked how I checked my references.\"** — run the audit\n  with `--export-audit audit.json --verify-doi`: the JSON working-paper\n  contains tool identity, version, the exact (API-key-redacted) invocation,\n  and the per-citation evidence chain a reviewer or editor can replay.\n  Submit it alongside the manuscript.\n- **\"Checking the reference list of my thesis.\"** — copy the list into a\n  plain-text file (one reference per line, or keep the `[1]` numbering) and\n  run `--parse-text refs.txt`: inline PMIDs/DOIs route exactly; English\n  entries resolve via PubMed title search; Chinese entries resolve via the\n  OpenAlex CJK leg. Finish with\n  `--format-references refs_clean.txt --citation-style gbt` for a corrected,\n  uniformly formatted list.\n- **\"Screening a systematic review's bibliography (PRISMA).\"** — verify the\n  full list once, export the audit as the baseline, then re-run\n  `--diff audit.json` before each resubmission: newly retracted, degraded\n  and new citations are reported as deltas, so you only re-read what\n  changed.\n\n## Anti-patterns — things done WRONG\n\nEach entry: the mistake → why it fails → the right way.\n\n1. **Treating `--pmids` output as \"fully verified\"** — existence-only.\n   → Wrong: \"all 5 PMIDs exist, so the citations are correct.\"\n   → Right: existence-checked only; feed `--claims-file` with titles for\n   real verification (the READY line says so explicitly).\n2. **Claims without `title`** — author/journal/year alone can never reach\n   `correct`; the report caps at `partial`. → Always include titles.\n3. **Trusting a cached verdict right after publication day** — a brand-new\n   PMID may have been cached as not-found by an earlier run, and retraction\n   status is as of cache time. → Final pre-submission check: `--no-cache`.\n4. **Assuming \"not found\" always means fabricated** — a Crossref 404 is\n   re-checked at OpenAlex (DataCite/Zenodo/Chinese-registry DOIs live\n   there), so \"unknown in both registries\" is now the bar; auto-extracted\n   DOIs that 404 in both stay *suspects*. arXiv IDs removed by moderators\n   also return empty. → Check doi.org / arxiv.org by hand before accusing.\n5. **Copying the leading `'` from CSV cells** — that apostrophe is the\n   formula-injection guard, not data corruption. → Strip it after import.\n6. **Reading the READY line as a quality score** — it only means \"no\n   problems found among the checks that ran\", not \"this paper is good\".\n\n## Boundaries — declared limits\n\nWhat this tool can NOT do, consolidated in one place:\n\n- **Splice/mismatch signals report disagreement, never pick a side** —\n  when claim and registry disagree, a human reads the evidence line.\n- **Cross-language authors are skipped, not failed** — CJK↔Latin author\n  names are never compared (`author_check: skipped`); the verdict rests\n  on title/journal/year alone.\n- **Context parsing is heuristic** — v3.3.0 keeps common abbreviations\n  (U.S., e.g., vs., St., Vol., No.) from splitting a title, but exotic\n  formatting can still mis-split; for exact metadata use `--claims-file`.\n- **DataCite/repository DOIs are not in Crossref — OpenAlex covers them**\n  (v4.1.0): a Crossref 404 now gets a second look at OpenAlex; only when\n  BOTH registries report the DOI unknown do the 404 semantics apply\n  (`--dois`/claims = invalid, auto-extracted = suspect). Resolved-via-\n  OpenAlex entries carry `meta_source: openalex`.\n- **arXiv moderator removals also return \"not found\"** — the invalid\n  verdict carries that caveat in its details.\n- **Retraction status is as-of-cache-time** — final pre-submission\n  checks should run with `--no-cache`.\n- **Single-letter initials never match** — \"Smith J\" vs \"Smith John\" is\n  not counted as a miss.\n- **unknown ≠ invalid** — unreachable sources yield exit 2 and\n  `unknown`; network failures are never reported as \"not found\" and\n  never cached.\n- **arXiv pacing is deliberate** — the official API asks for ≥3 s\n  between calls; large arXiv batches are slow by design (progress + ETA\n  on stderr). Entries that register a version-of-record DOI add one\n  Europe PMC lookup each for the PMID link.\n- **DOI claims compare against the registry that actually answered** —\n  the linked PubMed record when the DOI resolves to one, Crossref\n  otherwise (Crossref author fields are sparser, so the author mark is\n  more often \"—\"); a DOI row without a claimed title stays `unknown`,\n  not partial; an explicitly user-provided DOI (`--dois` or claims) that\n  is missing from Crossref counts as invalid.\n\n## Claims reference format\n\n`--claims-file` accepts JSON (an array of objects) or CSV. Recognized\ncolumns: `pmid`, `title`, `authors` (semicolon/pipe-separated),\n`journal`, `year`, `doi`, `arxiv_id`. A row needs one of `pmid`,\n`arxiv_id` or `doi`; a missing `title` caps the verdict at partial.\n\n```csv\npmid,title,authors,journal,year,doi,arxiv_id\n31018962,Candidate criteria for diagnosis of familial...,Gattorno,Ann Rheum Dis,2019,10.1136/annrheumdis-2019-215048,\n,Attention Is All You Need,Vaswani,NeurIPS,2017,,1706.03762\n,City size and the spreading of COVID-19 in Brazil,Silva Junior;Other,PLOS ONE,2020,10.1371/journal.pone.0239699,\n```\n\nValidate any file offline first: `--lint-claims file.csv` reports\nunusable rows, ID shape errors, unknown columns, duplicates and missing\ntitles (no network). Lint wins when combined with verification flags —\nonly the lint runs.\n\n## Best practices & tuning\n\n- **Speed up large batches** — request an NCBI API key (see\n  https://ncbiinsights.ncbi.nlm.nih.gov/api-keys/): batches of 50 IDs run\n  at 0.12 s spacing instead of 0.4 s; cached re-runs take seconds.\n- **Parallel DOI verification** — `--workers` (default 4, cap 8) applies to\n  Crossref resolution and Europe PMC linking; arXiv stays serial by\n  official etiquette (≥3 s between calls).\n- **Two-phase workflow** — sweep with metadata-only verification first\n  (no `--verify-doi`, no `--suggest`), then deep-verify only the flagged\n  subset; each deep flag adds one API call per citation.\n- **Round-trip bibliographies** — `--export-bibtex` / `--export-ris`\n  write verified entries; after edits, `--bibliography refs.bib` /\n  `refs.ris` re-audits the file (the PMID re-links the full record).\n- **Claims over context parsing** — whenever you know the expected titles,\n  feed `--claims-file`: it enables the full verdict ladder and the DOI /\n  arXiv pairing checks. Validate the file offline first:\n  `python3 scripts/verify_pmids.py --lint-claims claims.csv` reports ID\n  shape errors, missing titles, unknown columns and duplicates without any\n  network access.\n- **Flaky networks** — raise `--timeout`; HTTPS_PROXY/HTTP_PROXY are\n  honored natively; unreachable NCBI falls back to Europe PMC\n  automatically (`meta_source` shows which answered).\n- **Cache policy** — results cache 30 days, negative entries 3 days;\n  `--cache-days` to tune; `--no-cache` for the final pre-submission pass.\n- **Scale expectations** — metadata-only throughput is API-bound\n  (~1–2 min per 1000 PMIDs with an API key); DOI resolution adds one\n  Crossref call per DOI. One deliberate trade-off: the verifier is a\n  single stdlib-only file — copy `scripts/verify_pmids.py` anywhere with\n  Python 3.9+ and it runs, no pip, no venv (that portability is why the\n  code is not split into modules).\n\n## v2.2.0 — network hardening\n\n| Feature | Flag | Effect |\n|---------|------|--------|\n| NCBI API key | `--ncbi-api-key` / env `NCBI_API_KEY` | Rate ceiling 3→10 req/s, batch interval 0.4s→0.12s (~3x faster) |\n| Europe PMC fallback | `--meta-source auto\\|ncbi\\|europepmc` | NCBI batch failure automatically retries via Europe PMC (free, no key); per-entry origin in JSON (`meta_source`) |\n| Crossref polite pool | `--mailto` / env `PUBMED_VERIFIER_MAILTO` | `?mailto=` on Crossref + tool/email params on NCBI — more generous limits |\n| Retry-After backoff | automatic | 429 responses honored (clamped 1–5 s) instead of failing |\n| UA rotation | automatic | 403/406 retried with a browser User-Agent |\n| Host circuit breaker | automatic | After 2 call-level transport failures a host is skipped with an actionable message; success resets; HTTP errors never trip it |\n| Honest unknown | automatic | Network failures report as ❓ unknown + exit code 2, never as \"PMID not found\", and are never cached |\n\n**Exit codes:** `0` clean · `1` problems found (invalid / mismatch / retracted /\nDOI-splice, incl. arXiv claimed-DOI pairing mismatches) · `2` could not verify\n(data sources unreachable) — automation can tell \"all good\" from \"no answer\".\n\n## v2.3.0 — retraction detection\n\nWith `--verify-doi`, each cited DOI is also checked against Crossref's\nwithdrawal records (`updated-by`). A paper Crossref lists as RETRACTED is:\n\n- flagged in JSON (`retracted: true` + `retraction_note`) and in reports,\n- **capped at 🔶 partial** even when every metadata field matches — citing a\n  retracted paper is never \"correct\"; the report says *human review required*.\n\nCorrections and other update types do not trigger the cap. Crossref outages\nnever flag anything (a missing check is not a retraction).\n\n## v2.4.0 — DOI↔PMID cross-check & journal abbreviations\n\n- **DOI splice detection**: add a `doi` field to your claims (JSON or CSV).\n  The claimed DOI is compared with the DOI registered for that PMID — a\n  mismatch is a splice/fabrication signal (a real DOI attached to the wrong\n  paper): flagged in JSON (`doi_splice_suspect`), verdict capped at 🔶\n  partial, counted in exit 1. Uses the PubMed record only — no extra API call.\n- **Journal abbreviation equivalence**: journal matching now understands\n  NLM-style abbreviations in both directions — \"N Engl J Med\" matches \"New\n  England Journal of Medicine\", \"Pediatr Rheumatol\" matches \"Pediatric\n  Rheumatology\" (in-order word prefixes, function words skipped). No more\n  false \"journal differs\" for abbreviated citations.\n\nKnown limits: highly ambiguous abbreviations can over-match at the\njournal-only level (\"J Immunol\" ~ \"Journal of Immunology Research\") — the\ntitle remains the decisive field. A DOI-splice flag can also appear on an\notherwise-unverifiable citation (the DOI mismatch is an independent fact).\n\n## v2.5.0 — author-name verification\n\n- **Initials never match**: single-letter tokens (\"A.\", \"L.\") on either side\n  are excluded from surname matching — an initial is not evidence, and\n  substring-matching one produced false author hits.\n- **Cross-language honesty**: CJK author names claimed against Latin\n  registry records (or the reverse) are skipped, not counted as a mismatch —\n  the report marks them `author_check: skipped` and the verdict falls back\n  to what was actually comparable (title/journal/year), or to ❓ unknown\n  when nothing else is checkable.\n\n## Security & behavior declaration\n\n- Single-run CLI: scan, verify, write the report, exit. No daemons, no\n  background jobs, nothing downloaded or installed at runtime (pure standard\n  library, zero dependencies).\n- Network access is limited to these official academic registries, always\n  over HTTPS: `eutils.ncbi.nlm.nih.gov`, `www.ebi.ac.uk` (Europe PMC),\n  `api.crossref.org`, `api.openalex.org`, `export.arxiv.org` — no other\n  hosts are contacted; no telemetry, no\n  analytics, no data collection — the only outbound payloads are the PMIDs,\n  DOIs and titles you asked to verify.\n- Your files and reports stay on your machine. Writes are limited to the\n  report paths you pass and the SQLite cache under\n  `~/.cache/pubmed-verifier/` (`--no-cache` to disable).\n- Optional environment variables `NCBI_API_KEY` / `PUBMED_VERIFIER_MAILTO`\n  authenticate or attribute your own API requests and are never sent\n  anywhere else.\n- No OS integration: no subprocesses, no system services, no privilege\n  changes, no scheduled tasks.\n\n## v2.6.0 — audit working-paper & report v2\n\n- **`--export-audit audit.json`** — a self-contained JSON working-paper for\n  transparent review: tool identity and version, the exact (API-key-redacted)\n  invocation, per-citation evidence chains (claimed vs registered fields,\n  title match scores from both algorithms, author match with cross-language\n  skip records, DOI cross-check, retraction signals) and the\n  verdict-ladder trace for every citation. A reviewer can replay the entire\n  verification from this file alone.\n- **HTML report v2** — verdict filter tabs, severity-sorted rows (retracted\n  and DOI-splice first, highlighted), a field-level evidence column\n  (title/author/journal/year ✓✗—) and a reproducibility footer (redacted\n  command line + version + data sources).\n- **Reliability** — negative cache entries now expire after 3 days (a\n  legitimately new, ahead-of-print PMID is no longer reported \"not found\"\n  for a month), and the circuit breaker self-heals: after a 30 s cooldown it\n  admits one probe call and resets on success.\n\n## v2.7.0 — retraction for every PMID, BibTeX export, readiness verdict\n\n- **Retraction detection, source-independent** — the registry's own\n  publication type (\"Retracted Publication\"; present in both NCBI esummary\n  and Europe PMC) now flags retracted papers for EVERY citation: no DOI\n  required, no `--verify-doi` required, and the flag survives the cache\n  (schema v3). Crossref `updated-by` remains the detail source (the\n  retraction-notice DOI) when `--verify-doi` is on. A retraction *notice*\n  itself is never flagged.\n- **`--export-bibtex refs.bib`** — export the verified bibliography: correct\n  entries as `@article`, partial entries commented out with their divergence\n  note, mismatched/invalid/unknown/retracted entries excluded and counted.\n- **Submission-readiness verdict** — every report now leads with one line:\n  `SUBMISSION READY` or `NOT SUBMISSION-READY — <per-problem counts>`.\n- Cache schema v3 (adds a `retracted` column, auto-migrated).\n\n## v2.8.0 — DOI-native verification & delta audits\n\n- **`--dois \"10.x/a, 10.y/b\"`** — verify DOIs natively, no PMID required.\n  `--source` scans now also extract DOIs from your files automatically.\n  Each DOI is resolved via the Crossref works API: not-found on an explicitly\n  provided DOI = fabrication signal (invalid, exit 1); on one auto-extracted\n  from scanned text it stays a suspect (unknown) — scanned strings are never\n  user-endorsed, and DataCite/repository DOIs do not live in Crossref, so\n  always double-check at doi.org. Resolved = existence confirmed with the\n  registered metadata attached for manual comparison — *existence is never\n  dressed up as a match*.\n- **`--diff previous-audit.json`** — delta audit against a previous working\n  paper: **newly retracted** (the safety signal — a paper retracted after\n  your last audit; act on it: swap or drop the citation, cite the retraction\n  notice instead, and re-check any conclusion that relied on it), degraded,\n  improved, new and dropped citations, with counts in every report format.\n  Built for periodic knowledge-base audits: \"what changed since last time?\"\n\n## v2.9.0 — DOI entries become first-class\n\n- **DOI→PMID linking** — a resolved DOI is linked back to its PMID via the\n  Europe PMC DOI field query, pulling the full PubMed record: complete\n  metadata, retraction pubtype signal, and cache coverage. A DOI citation\n  now gets the same five-state record as a PMID citation (existence\n  confirmation only — the verdict remains unknown until claims are\n  provided).\n- **Parallel DOI resolution** — `--workers N` (default 4, max 8) resolves\n  DOI batches on a thread pool (roughly 3x faster on large lists), with\n  live progress output. For large `--dois` batches, set `--mailto` to stay\n  in Crossref's polite pool.\n- **`--export-csv table.csv`** — spreadsheet-friendly audit table\n  (key/verdict/flags/fields/details; formula-injection hardened).\n\n## v3.0.0 — arXiv ID verification (three citation types, one audit)\n\nReference lists carry preprints. v3.0.0 verifies **arXiv IDs** alongside\nPMIDs and DOIs: `arXiv:2401.12345` and `arxiv.org/abs/...` patterns are\nextracted from scans (or passed via `--arxivs`), checked against the\nofficial arXiv API, and judged — nonexistent ID = fabrication signal\n(invalid, exit 1); resolving ID = registered title/year attached, verdict\nstays unknown. Malformed IDs (bad YYMM month) are flagged by shape.\nTimely: arXiv penalizes submissions containing hallucinated or unverified\nreferences (2026-05 policy) — audit before you submit.\n\n## v3.3.0 — preprint ↔ published-version cross-check\n\narXiv entries carry the version-of-record DOI their authors registered at\npublication (`arxiv:doi`). v3.3.0 puts it to work:\n\n- **Claimed DOI vs registered DOI** — a claims row with both `arxiv_id`\n  and `doi` is cross-checked: agreement is reported as evidence\n  (`fields.doi ✓`); disagreement caps the verdict at `partial` — the DOI\n  belongs to a different paper (same failure class as PMID DOI-splice).\n- **Version of record surfaced** — verifying a bare preprint ID now shows\n  the registered DOI and, when the published version is PubMed-indexed,\n  its linked PMID — cite and verify the final version, not just the\n  preprint.\n- **Honest accounting** — the readiness line counts arXiv DOI-pairing\n  mismatches as problems; DOI/arXiv phase progress (stderr) now includes\n  elapsed time and an ETA for large batches.\n- Context parsing no longer truncates titles at sentence-internal\n  abbreviations (\"U.S. population\", \"e.g.\", \"vs.\", \"Vol.\").\n\nPairing example (match → correct with DOI evidence; wrong DOI → partial):\n\n```bash\npython3 scripts/verify_pmids.py --claims '[{\"arxiv_id\":\"2005.13892\",\n  \"title\":\"City size and the spreading of COVID-19 in Brazil\",\n  \"doi\":\"10.1371/journal.pone.0239699\"}]'\n```\n\n## v3.4.0 — DOI claims become first-class\n\nClaims rows could carry a PMID or an arXiv ID — a row keyed by DOI alone\nwas silently ignored, and DOI entries always stayed `unknown` (\"no claimed\nmetadata to cross-verify\"). v3.4.0 closes the matrix: all three citation\ntypes now accept claimed metadata.\n\n- A claims row with a `doi` (no PMID, no arXiv ID) is cross-checked against\n  the registered metadata — the linked PubMed record when the DOI resolves\n  to one, Crossref otherwise — and gets the full verdict ladder:\n  correct / mismatch / partial.\n- Retraction capping applies as everywhere: a claimed-correct match on a\n  retracted paper is capped at partial with the retraction note.\n- Without claims, DOI entries stay unknown — existence is never dressed up\n  as a match.\n\n```bash\npython3 scripts/verify_pmids.py --claims '[{\"doi\":\"10.1371/journal.pone.0239699\",\n  \"title\":\"City size and the spreading of COVID-19 in Brazil\",\n  \"journal\":\"PLoS ONE\",\"year\":\"2020\"}]'\n```\n\n## v3.5.0 — claims lint & usage-first restructuring\n\n- **`--lint-claims FILE`** — offline pre-flight for claims files\n  (JSON/CSV, zero network): ID shape errors, missing titles (the verdict\n  would cap at partial), unknown/typo'd columns, DOI prefix checks,\n  unusable rows — exit 1 on errors. Fix the format before the run\n  instead of guessing from weak verdicts.\n- **Documentation restructured around usage**: FAQ, anti-patterns and\n  declared boundaries now sit right after Quick start, led by a Top-10\n  \"don't do this\" table; new **Best practices & tuning** section (API-key\n  batching, worker tuning, two-phase deep-verification, cache policy,\n  scale expectations — and why the verifier is deliberately one file).\n\n## v3.6.0 — BibTeX bibliography audit\n\n**`--bibliography refs.bib`** audits a .bib file directly: entries route by\nPMID (the `pmid` field, or a \"PMID: NNNN\" in the note — including the\nnotes `--export-bibtex` itself writes) > DOI > arXiv (`eprint`, or an\n\"arXiv:XXXX.XXXXX\" in the journal/note), each carrying its claimed\ntitle/authors/journal/year for the full cross-check. Entries without any\nroutable ID are counted and skipped. This closes the loop with\n`--export-bibtex`: a verified bibliography can be re-audited after edits.\n`--lint-claims refs.bib` validates the file offline (unroutable entries,\nmissing titles, unclosed blocks).\n\n## v3.7.0 — RIS bibliography support\n\n`--bibliography` now accepts **RIS files** (`refs.ris` — Zotero/EndNote/\nMendeley exports) alongside BibTeX: records route by PMID (AN tag, or a\n\"PMID: NNNN\" note) > DOI (DO) > arXiv (UR/eprint), each with claimed\nmetadata for the full cross-check. **`--export-ris`** completes the loop —\nverified entries as `TY JOUR`, partials as `TY DATA` with a PARTIAL note.\n`--lint-claims refs.ris` validates offline (unroutable records, missing\ntitles, duplicates, unterminated records).\n\n## v3.8.0 — verdict-ladder corrections (independent external test audit)\n\nAn independent real-data audit (24 scenarios, registry-ground-truthed) found\nthree judgment-ladder defects; all three are fixed and locked:\n\n- **B1 scan false positives** — when a citation is not the first sentence of\n  its block, the author line was taken as the claimed title (correct\n  references reported as mismatch). The parser now shifts past an\n  author-shaped segment.\n- **B2 author mis-attribution** — a claimed author set entirely different\n  from the registry (zero surname overlap) was still reported as correct\n  when title and journal matched; it now caps at partial. Partial surname\n  overlap (abbreviated or reordered names) keeps correct.\n- **B3 title-less claims** — author+journal+year matches without a claimed\n  title now cap at partial, exactly as the FAQ/anti-patterns/boundaries\n  always promised (the same author/journal/year can cover several papers).\n  Cross-language author skips follow the same cap.\n- Also: duplicate arXiv IDs announce their merge on stderr (run\n  `--lint-claims` to catch duplicate claim rows offline); the\n  cross-language skip no longer coexists with contradictory details text;\n  empty `--claims` gets a one-line hint; docs state that Chinese-language\n  citation contexts parse weakly — use `--claims-file`.\n\n## v3.9.0 — Chinese references (GB/T 7714), network pre-flight, Python API\n\n- **GB/T 7714 Chinese references parse** — `……标题[J]. 刊名, 年… PMID: xxx`\n  contexts now yield the title (via the `[J]/[M]/[R]` type marker),\n  authors, journal, year and any DOI that follows the PMID. The title\n  comparison against a Latin registry title is skipped cross-language\n  (never a mismatch source); a DOI in the reference is cross-checked\n  against the registry like any claims DOI. Requires a PMID marker per\n  reference.\n- **`--check-net`** — probes the data sources (5 s each), shows your\n  proxy state and prints practical next steps when something is\n  unreachable. Run it when results come back unknown on a constrained\n  network.\n- **Python API reference** — `references/python_api.md` documents the\n  stable importable surfaces (`fetch_summaries`, `cross_check_citation`,\n  `verify_doi_entry`, `parse_citation_context`, `suggest_correct_pmid`)\n  with copy-paste examples.\n\n## v4.0.0 — plain-text reference lists & formatted reference list export\n\n- **`--parse-text refs.txt`** — verify a reference list copied straight out\n  of a manuscript draft. Numbered entries (`[1]`, `1.`) split cleanly;\n  inline PMID / DOI / arXiv IDs route exactly (full five-state\n  cross-check); a markerless entry with a parseable title resolves via\n  PubMed title search and is marked `resolved_by: title_search` — weaker\n  evidence than a supplied ID, honestly labeled, and never a mismatch\n  source. A network failure stays unknown with a retry hint, never\n  \"not found\". Entries with no routable ID and no usable title are\n  skipped with a note instead of being guessed.\n- **`--format-references out.txt --citation-style gbt|vancouver|apa|ama`**\n  — the formatting leg of verify-then-format: verified entries rendered as\n  a ready-to-paste numbered reference list (GB/T 7714-2015 numeric for\n  Chinese submissions, Vancouver, APA 7th, AMA 11th). Formatting stands on\n  registry-verified records only: correct entries form the list, partial\n  entries go to a manual-review section (never silently dropped), and\n  retracted entries are excluded with a warning. Author lists are\n  formatted per style without inventing names — a registry-truncated list\n  closes with the style's et-al form, and APA discloses the truncation.\n- Markerless Latin references now parse through the full sentence\n  segmentation (titles feed the title-search leg); CJK entries without a\n  PMID/DOI marker keep the documented weak-parsing boundary — supply a\n  PMID or DOI for exact routing.\n- `examples/refs_list.sample.txt` ships a ready-made parse-text fixture.\n\n## v4.1.0 — OpenAlex fallback & Chinese-reference title search\n\n- **OpenAlex is now the fourth data source** — a DOI that Crossref reports\n  404 gets a second look at OpenAlex before any verdict: DataCite, Zenodo,\n  SSRN and Chinese-registry DOIs (which do not live in Crossref) become\n  fully decidable — resolved, metadata attached, claims cross-checked,\n  retractions flagged (`is_retracted`). Only when OpenAlex also comes up\n  empty do the original 404 semantics stand (explicit DOI = invalid,\n  scanned DOI = suspect). `--check-net` probes five sources now.\n- **Chinese references without any ID can now be verified** — a markerless\n  CJK entry in `--parse-text` searches OpenAlex by its title (the query is\n  the title-shaped segment of the entry, not the whole line); a candidate\n  counts only when its registered title appears **verbatim** in your entry\n  text with an agreeing year, and the verdict carries\n  `resolved_by: openalex_cjk_search`. No hit = honestly unknown; this leg\n  never manufactures a mismatch, and a network failure stays unknown.\n- Docs: the reports' actual shape is now shown in \"What the output looks\n  like\"; three worked scenarios added to the FAQ (reviewer evidence,\n  thesis citation check, PRISMA screening).\n\n## How it works\n\n1. **Extract + parse context** — finds `PMID: 12345678` / PubMed URLs in\n   `.html .md .txt .htm .json`, and parses the surrounding reference into\n   claimed authors / title / journal / year.\n2. **Fetch metadata (cached)** — PubMed esummary in batches of 50, SQLite\n   cache (30 days, `--cache-days`), 3 retries with backoff.\n   Europe PMC steps in per failed batch when NCBI is unreachable.\n3. **Cross-check claimed vs actual** — dual fuzzy matching: word-level\n   Jaccard overlap ≥ 50% OR SequenceMatcher ≥ 90% on titles; author surname\n   hits; journal containment or NLM abbreviation equivalence; exact year.\n4. **DOI↔PMID cross-check** (automatic when claims include `doi`) — a claimed\n   DOI differing from the PMID's registered DOI is a splice/fabrication\n   signal (capped at partial).\n5. **Crossref DOI verification** (optional `--verify-doi`) — resolves each\n   cited DOI via Crossref, compares the registered title with the PubMed\n   record (`doi_title_match` in JSON), and detects RETRACTED papers (verdict\n   capped at partial). A `doi_verified: false` with note\n   \"crossref unreachable\" is a network fact, not a verdict.\n6. **Suggest the right PMID** (optional `--suggest`) — for mismatches,\n   searches PubMed with the claimed metadata and proposes top-3 candidates.\n   (Suggestion search always uses NCBI, even with `--meta-source europepmc`.)\n\nContext parsing is heuristic — since v3.3.0, common abbreviations\n(\"U.S.\", \"e.g.\", \"vs.\") no longer split a title, but exotic formatting\nstill can. For precise verification, feed structured claims via\n`--claims-file`.\n\n| Report | Flag | Use |\n|--------|------|-----|\n| HTML | `--output report.html` | Human review: claimed vs actual side by side |\n| JSON | `--output report.json` | Programmatic processing (includes `meta_source` per entry) |\n| Text | default | Quick terminal look |\n\n## Performance\n\nMeasured on a 225-PMID audit (5 esummary batches): metadata-only verification\nruns in seconds; cached re-runs take ~5 s. Each batch waits 0.4 s between\ncalls (0.12 s with `--ncbi-api-key`). Optional extras are per-citation:\n`--verify-doi` adds one Crossref call (~0.5–1 s) per cited DOI, and\n`--suggest` adds one PubMed search per mismatch.\n\n## When to use which tool\n\n- **pubmed-verifier (this skill)** — fast, batch, targeted: I have a list of\n  PMIDs/DOIs and need to know if they are real and correctly cited.\n- **cite-holmes** — deep research: interrogate every citation of a whole\n  document across multiple databases, with graded confidence reports.\n\nThey share the same five-state philosophy and are safe to use together.\n\n## Use cases\n\n- Systematic review / meta-analysis reference audits\n- Verifying citations in AI-generated content\n- Pre-submission self-check of a manuscript's reference list\n- Medical knowledge base / teaching material QA\n- Pharmacovigilance literature verification\n\n## Related tools\n\nEach tool solves one step of reference work; use whichever fits the task.\n\n- **cn-med-oa** — free Chinese medical literature full-text download & metadata\n- **cite-holmes** — deep research with machine-verified citations\n- **paper-polisher-pro** — academic polishing, terminology & journal precheck\n- **academic-figures** — publication-ready scientific figures in one command\n- **doc-holmes** — layout-preserving PDF translation (in testing)\n\nTypical order: get papers (cn-med-oa), verify citations (this tool or\ncite-holmes), polish (paper-polisher-pro), make figures (academic-figures),\ntranslate PDFs (doc-holmes) — pick whichever step you need.\n\n## Files\n\n| File | Purpose |\n|------|---------|\n| `scripts/verify_pmids.py` | Main verifier (v4.1.0, stdlib-only) |\n| `references/api_examples.md` | PubMed / Europe PMC / Crossref / arXiv API notes |\n| `references/python_api.md` | Calling the verifier from Python (stable surfaces + examples) |\n| `examples/claims.sample.csv` | Reference format for `--claims-file` (incl. a DOI-only row) |\n| `examples/refs.sample.bib` | Sample bibliography for `--bibliography` (DOI/PMID/arXiv routing) |\n| `examples/refs.sample.ris` | Sample RIS bibliography (Zotero/EndNote/Mendeley) |\n| `tests/` | Offline matrix + real-network acceptance (repo only, not in the package) |\n\n## License\n\nMIT-0 — free to use, modify and redistribute, no attribution required.\n\nFile v4.1.0:examples/README.txt\n\nclaims.sample.csv — reference format for --claims-file\n\n- pmid column: PubMed pipeline (full five-state verification)\n- rows with only arxiv_id (+title): arXiv pipeline (official API check)\n- doi column: DOI-splice cross-check against the PubMed-registered DOI\n  (PMID rows) / version-of-record pairing check (arXiv rows, v3.3.0) /\n  full five-state verdicts (doi-only rows, v3.4.0)\n- title: required for verdicts beyond existence-check\n- expected verdicts (as of v3.4.0): the 31018962 row carries a\n  deliberately inexact title → partial (title mismatch, everything else\n  matches — demonstrates the partial rung, not a data error); the\n  1706.03762 row → correct; the doi-only 10.1371 row → correct (its\n  second author \"Other\" is a synthetic placeholder; v3.8.0's\n  mis-attribution cap means a wrong author set would drop it to partial); the 24476887 row\n  (STAP cells) is RETRACTED on purpose → RETRACTED flag with a partial\n  cap — it demonstrates retraction detection, not a data error\n\nRun:  python3 scripts/verify_pmids.py --claims-file examples/claims.sample.csv\n\nLint first (offline, v3.5.0): python3 scripts/verify_pmids.py --lint-claims examples/claims.sample.csv\n\nBibliography (v3.6.0): refs.sample.bib audits a .bib file directly —\npython3 scripts/verify_pmids.py --bibliography examples/refs.sample.bib\n\nRIS (v3.7.0): refs.sample.ris audits Zotero/EndNote/Mendeley exports —\npython3 scripts/verify_pmids.py --bibliography examples/refs.sample.ris\n\nFile v4.1.0:README.md\n\n# PubMed Citation Verifier 🔬\n\n[![GitHub Stars](https://img.shields.io/github/stars/docsor1212/pubmed-verifier?style=social&label=Star)](https://github.com/docsor1212/pubmed-verifier/stargazers)\n\nBatch-verify PMID citations against PubMed API. Built for researchers, medical writers, and evidence-based medicine teams.\n\n## 简介（中文）\n\n**PubMed 文献引用批量验证工具**——把引用列表交给它，逐条核验 PMID/DOI/arXiv ID 是否真实存在、是否指向你声称的那篇论文（AI 幻觉引用最常见的形态就是「真 PMID 指向别的论文」）。五态判定（正确/不匹配/部分匹配/无效/待确认），撤稿检测全覆盖，网络故障如实标注「未判定」、绝不误报「不存在」。零依赖、纯本地运行。\n\n- **中文文献**：GB/T 7714 格式参考文献可解析（`……标题[J]. 刊名, 年… PMID: xxx`）；`--parse-text` 粘贴即核验整张文献表——行内 PMID/DOI 精确路由，中文条目经 OpenAlex 标题检索判定（无编号 ID 也能核）\n- **参考文献格式化**：`--format-references --citation-style gbt`（另支持 Vancouver/APA/AMA），先验证后排版，只排已验真的条目\n- **四数据源**：NCBI PubMed、Europe PMC 自动兜底、Crossref、OpenAlex（DataCite/Zenodo 等非 Crossref 注册的 DOI 也可判定）\n- **投稿就绪判定 + 审计底稿**：每份报告自带证据链，`--export-audit` 输出可复放的 JSON 工作底稿\n\n## Why?\n\nAcademic projects routinely contain hundreds of PMID citations. Manual verification is tedious and error-prone. During our own 225-reference audit, we found 3 invalid PMIDs and 6 cross-domain mismatches — errors that would have undermined the entire project.\n\n## Features\n\n- **Batch verification** — Scan entire project directories, extract all PMIDs, verify against PubMed in one run\n- **Mismatch detection** — Five-state verdicts catch a REAL PMID pointing to a DIFFERENT paper (the most common AI hallucination)\n- **Metadata validation** — Title, authors, journal and date fuzzy-compared against your claims (DOI cross-checked via Crossref with `--verify-doi`)\n- **Dual sources** — NCBI E-utilities primary, Europe PMC automatic fallback when NCBI is unreachable\n- **Network hardening** — Optional NCBI API key (3x faster batches), Crossref polite pool, 429 Retry-After backoff, UA rotation, per-host circuit breaker\n- **Content matching** — Keyword overlap scoring flags potentially irrelevant citations\n- **Three citation types in one audit** — PMIDs, DOIs and arXiv IDs (preprints) verified in a single scan; arXiv IDs checked against the official API (nonexistent = fabrication signal); claims rows keyed by DOI get full verdicts, and preprints surface their published version of record\n- **Claims lint** — `--lint-claims file.csv` validates a claims file offline (ID shapes, missing titles, unknown columns, duplicates) before any verification run\n- **Bibliography audit** — `--bibliography refs.bib` / `refs.ris` verifies BibTeX and RIS files directly (Zotero/EndNote/Mendeley exports; entries route by PMID > DOI > arXiv with full cross-checks); round-trips with `--export-bibtex` / `--export-ris`\n- **Plain-text reference list parsing** — `--parse-text refs.txt` verifies a reference list copied straight out of a manuscript draft: numbered entries (`[1]`, `1.`) split cleanly, inline PMID/DOI/arXiv IDs route exactly, and a title-only entry resolves via PubMed title search (honestly labeled `resolved_by: title_search`, never a mismatch source)\n- **Formatted reference list export** — `--format-references out.txt --citation-style gbt|vancouver|apa|ama` renders verified entries as a ready-to-paste numbered list (GB/T 7714-2015, Vancouver, APA 7th, AMA 11th); partial entries go to a manual-review section, retracted ones are excluded with a warning — formatting that stands on registry-verified records\n- **Retraction detection for every citation** — the registry publication type (\"Retracted Publication\") flags retracted papers with no DOI or extra flags needed; Crossref `updated-by` adds the retraction-notice DOI\n- **Verified-bibliography export** — `--export-bibtex` writes correct entries as BibTeX, comments partial ones, excludes and counts the rest\n- **Submission-readiness verdict** — every report leads with `SUBMISSION READY` / `NOT SUBMISSION-READY` and per-problem counts\n- **Audit working-paper** — `--export-audit` writes a self-contained JSON trail (tool identity, redacted invocation, evidence chain, verdict trace) a third party can replay\n- **DOI-native verification** — `--dois` verifies DOIs directly (Crossref resolve, 404 = fabrication signal on explicit input); `--source` scans auto-extract DOIs; linked back to PMIDs via Europe PMC\n- **Delta audits** — `--diff previous-audit.json` reports newly retracted, degraded, improved, new and dropped citations\n- **Exports** — `--export-audit` (self-contained JSON trail), `--export-bibtex` (verified bibliography), `--export-csv` (spreadsheet audit table)\n- **Replacement search** — Find correct PMIDs for broken citations via PubMed search\n- **Multiple output formats** — HTML report, JSON (with per-entry metadata origin), or terminal summary\n\n## Quick Start\n\n```bash\n# Install\nopenclaw skills install docsor1212/pubmed-verifier\n\n# 中文场景速览\npython3 scripts/verify_pmids.py --parse-text references.txt --no-cache          # 核验整张参考文献表\npython3 scripts/verify_pmids.py --parse-text references.txt --format-references out.txt --citation-style gbt\npython3 scripts/verify_pmids.py --pmids 31018962,22213727                       # 批量验证 PMID\npython3 scripts/verify_pmids.py --check-net                                     # 网络预检（五源）\n\n# Verify all PMIDs in a project\npython3 scripts/verify_pmids.py --source /path/to/project --output report.html\n\n# Verify specific PMIDs\npython3 scripts/verify_pmids.py --pmids 31018962,22213727,999999999\n\n# Verify a reference list copied from a paper draft\npython3 scripts/verify_pmids.py --parse-text references.txt --no-cache\n\n# Format verified entries as a citation list (gbt | vancouver | apa | ama)\npython3 scripts/verify_pmids.py --parse-text references.txt --format-references refs_gbt.txt --citation-style gbt\n\n# Institutional mode: API key + polite pool\npython3 scripts/verify_pmids.py --source ./papers --verify-doi --ncbi-api-key $NCBI_API_KEY --mailto you@lab.org\n```\n\n**China mirror (ModelScope 魔搭)**: <https://modelscope.cn/skills/Docsor/pubmed-verifier> — if you find this skill useful, a like there helps others find it.\n\n## Use Cases\n\n| Scenario | Example |\n|----------|---------|\n| **Systematic review QA** | Verify all 200+ references before submission |\n| **Medical website audit** | Check evidence citations across clinical case library |\n| **Paper manuscript check** | Validate every PMID in your draft |\n| **Teaching material review** | Ensure lecture citations are accurate |\n| **Evidence library maintenance** | Periodic batch verification of reference databases |\n| **Draft reference list check & formatting** | Paste a manuscript's reference list, verify it, get a GB/T 7714 / APA list back |\n\n## Real-World Results\n\nAudited a 35-file pediatric rheumatology evidence library (225 PMID citations):\n- **222** citations: valid and content-matched ✅\n- **3** citations: invalid PMIDs found and corrected\n- **6** cross-domain citations: correctly flagged, reviewed, confirmed appropriate\n- Total time: ~5 minutes for full audit\n\n## Technical Details\n\n- **APIs**: PubMed E-utilities (esummary, esearch); Europe PMC fallback; Crossref DOI check\n- **Rate limit**: 3 req/s free (0.4s batch delay), 10 req/s with `--ncbi-api-key` (0.12s)\n- **Resilience**: 429 Retry-After backoff, 403/406 UA rotation, per-host circuit breaker\n- **Batch size**: 50 PMIDs per request\n- **File types**: `.html`, `.md`, `.txt`, `.json`, `.htm`\n- **PMID patterns**: `PMID: 12345678`, `PubMed: 12345678`, `pubmed.ncbi.nlm.nih.gov/12345678/`\n- **Exit codes**: 0 clean / 1 problems found / 2 network-incomplete\n\n## License\n\nMIT-0\n\nFile v4.1.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b6sjwend7cwwmhf7mdg1wzx82dfbf\",\n  \"slug\": \"pubmed-verifier\",\n  \"version\": \"4.1.0\",\n  \"publishedAt\": 1791564939147\n}\n\nFile v4.1.0:references/api_examples.md\n\n# PubMed E-utilities API Quick Reference\n\n## esummary — Article Metadata\n\n```bash\n# Single PMID\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=31018962&retmode=json\"\n\n# Batch (comma-separated)\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=31018962,22213727&retmode=json\"\n```\n\nResponse fields: `title`, `authors[].name`, `source` (journal), `pubdate`, `volume`, `pages`, `elocationid` (DOI).\n\nInvalid PMID → `{\"result\": {\"pmid\": {\"error\": \"cannot get document summary\"}}}` (the exact string has varied over time — the verifier only checks for the presence of an `error` key).\n\n## esearch — Find Articles by Query\n\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=Gattorno+classification+autoinflammatory&retmode=json&retmax=5\"\n```\n\nReturns `esearchresult.idlist` → array of PMIDs.\n\n## efetch — Full Abstracts\n\n```bash\n# efetch supports TEXT and XML only — retmode=json is NOT valid for efetch\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=37635643&rettype=abstract&retmode=text\"\n```\n\n## Deep Research Workflow (Systematic PubMed Search)\n\nWhen the user needs comprehensive medical literature research beyond simple PMID verification:\n\n### Step 1: esearch → keyword search → get PMIDs\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=<keywords>&retmax=10&sort=relevance\" | grep -oP '<Id>\\K\\d+'\n```\n\n### Step 2: esummary → metadata summaries (fast, compact)\n```bash\nPMIDS=\"pmid1,pmid2,pmid3\"\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=$PMIDS&retmode=json\" > /tmp/pubmed-results.json\npython3 -c \"\nimport json\ndata = json.load(open('/tmp/pubmed-results.json'))\nfor uid, art in data.get('result', {}).items():\n    if uid == 'uids': continue\n    authors = ', '.join([a.get('name','') for a in art.get('authors',[])[:6]])\n    print(f'PMID {uid}: {art.get(\\\"title\\\",\\\"\\\")}')\n    print(f'  {art.get(\\\"fulljournalname\\\",\\\"\\\")} {art.get(\\\"pubdate\\\",\\\"\\\")};{art.get(\\\"volume\\\",\\\"\\\")}:{art.get(\\\"pages\\\",\\\"\\\")}')\n    print(f'  DOI: {art.get(\\\"elocationid\\\",\\\"\\\")}')\n    print()\n\"\n```\n\n### Step 3: efetch → full abstracts for key papers\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=$PMIDS&rettype=abstract&retmode=text\"\n```\n\n### Pitfalls\n- **efetch has no JSON mode** — use `retmode=text` (or XML) only\n- **esearch may return unrelated results** — always cross-check with esummary titles\n- **Rate limit**: 3 req/s without API key. Add `sleep 0.5` between batches\n- **Large author lists**: esummary `authors` array can be 50+ — truncate to first 6 for display\n\n## Rate Limits\n\n- Without API key: 3 requests/second\n- With API key (`&api_key=YOUR_KEY`): 10 requests/second\n- API key obtained from NCBI Settings page\n- Etiquette params: `&tool=pubmed-verifier&email=you@lab.org` (added automatically with `--mailto`)\n\n## Europe PMC (fallback source, v2.2.0)\n\nFree, no key, mirrors PubMed; used automatically when NCBI batches fail\n(`--meta-source auto`) or forced with `--meta-source europepmc`.\n\n```bash\n# Batch lookup by PMID (EXT_ID), SRC:MED restricts to PubMed records\ncurl -s \"https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=(EXT_ID:31018962%20OR%20EXT_ID:22213727)%20AND%20SRC:MED&format=json&resultType=lite&pageSize=25\"\n```\n\nUseful fields: `id` (PMID), `title`, `authorString` (comma-separated),\n`journalTitle`, `pubYear`, `doi`, `journalVolume`, `pageInfo`.\n\n## Crossref DOI check (polite pool, v2.2.0)\n\n```bash\n# ?mailto= joins the polite pool — more generous rate limits\ncurl -s \"https://api.crossref.org/works/10.1038/nature12968?mailto=you@lab.org\"\n```\n\n- 429 responses carry `Retry-After` — back off accordingly (v2.2.0 clamps 1–5 s)\n- 403/406 usually mean rate limiting — back off and retry later (the tool rotates client identifiers, including a standard browser UA on retry, per the SKILL.md network-hardening table)\n- A correct DOI with WRONG paper metadata = spliced/fake citation signature\n\n## arXiv API (export.arxiv.org)\n\n```bash\n# one ID per lookup, Atom feed back; arxiv:doi = the version-of-record\n# DOI the authors registered — the preprint↔published cross-check uses it\ncurl -s \"https://export.arxiv.org/api/query?id_list=2005.13892&max_results=1\"\n```\n\n- No `<entry>` in the feed = the ID does not exist (fabrication signal);\n  a 200 response that is not an Atom feed (portal/maintenance page) is\n  treated as \"could not verify\", never as \"not found\"\n- Official etiquette: ≥3 s between calls — large batches are paced\n  deliberately (progress with ETA goes to stderr)\n- `arxiv:journal_ref` / `arxiv:doi` are author-registered fields, shown\n  in the audit trail; see SKILL.md v3.3.0 section for the pairing check\n\nFile v4.1.0:references/python_api.md\n\n# Python API — calling pubmed-verifier from your code\n\nThe whole tool is one stdlib-only module. Import it directly — no pip, no\nvenv (Python 3.9+):\n\n```python\nimport sys\nsys.path.insert(0, \"/path/to/pubmed-verifier/scripts\")\nimport verify_pmids as vp\n```\n\nAll functions below are stable surfaces used by the CLI itself (behavior\nis pinned by the offline test matrix in the source repository).\n\n## 1. Batch metadata for PMIDs\n\n```python\ninfo = vp.fetch_summaries([\"31018962\", \"22213727\"])   # NCBI, EPMC fallback\nr = info[\"31018962\"]\nprint(r[\"valid\"], r[\"title\"], r[\"journal\"], r.get(\"retracted\"))\n```\n\n## 2. Cross-check claimed vs registered metadata (the five-state ladder)\n\n```python\ncross = vp.cross_check_citation(\n    {\"claimed_title\": \"Classification criteria for autoinflammatory recurrent fevers\",\n     \"claimed_authors\": [\"Gattorno\"], \"claimed_journal\": \"Ann Rheum Dis\",\n     \"claimed_year\": \"2019\"},\n    {\"title\": r[\"title\"], \"authors\": r[\"authors\"],\n     \"journal\": r[\"journal\"], \"pubdate\": r[\"pubdate\"]})\nprint(cross[\"verdict\"], cross[\"details\"])\n# verdict ∈ correct / mismatch / partial / unknown\n```\n\n`claimed` / `actual` keys are plain strings/lists — feed them from any\nsource (database, form, LLM extraction). Cross-language CJK↔Latin author\nnames (and, since v3.9.0, CJK↔Latin titles) are skipped honestly, never\ncounted as misses.\n\n## 3. DOI-native verification\n\n```python\nres = vp.resolve_doi(\"10.1038/nature12968\")     # Crossref resolve (+retraction)\nentry, audit = vp.verify_doi_entry(\n    \"10.1038/nature12968\", \"cli\", resolution=res,\n    claimed={\"claimed_title\": \"...\", \"claimed_year\": \"2014\"})\nprint(entry[\"verdict\"], entry.get(\"retracted\"))\n```\n\n## 4. Context parsing (extract claims from free text)\n\n```python\nclaimed = vp.parse_citation_context(\n    \"Gattorno A, Van Dijk M. Classification criteria for autoinflammatory \"\n    \"recurrent fevers. Ann Rheum Dis. 2019. PMID: 31018962\")\nprint(claimed)\n# GB/T 7714 Chinese references are recognized since v3.9.0 (title via the\n# [J]/[M] type marker; the title comparison itself is cross-language-skipped)\n```\n\n## 5. Suggestions for broken citations\n\n```python\ncands = vp.suggest_correct_pmid({\"claimed_title\": \"Juvenile idiopathic arthritis\",\n                                 \"claimed_journal\": \"Pediatric Rheumatology\"})\nfor c in cands[:3]:\n    print(c[\"pmid\"], c[\"title\"][:60])   # abbreviated journal names can return [] — use the full NLM name\n```\n\n## 6. Plain-text reference lists & formatted output (v4.0.0)\n\n```python\nentries, issues = vp.parse_plaintext_references(\"refs_list.txt\")\nfor e in entries:\n    print(e[\"key\"], e[\"pmid\"] or e[\"doi\"] or e[\"arxiv_id\"] or\n          \"(title search)\", e[\"claimed_title\"][:50])\n\n# One markerless entry, verified by title search (honest route label):\nres, audit = vp.verify_title_entry(entries[0], \"refs_list.txt\")\nprint(res[\"verdict\"], res.get(\"resolved_by\"))   # correct title_search\n\n# Render verified results as a numbered reference list:\nprint(vp.generate_reference_list(results, \"gbt\"))       # also: vancouver / apa / ama\n```\n\n`parse_citation_context` accepts markerless Latin references since v4.0.0\n(returns `claimed_source_format: \"plaintext\"`); CJK entries without a\nPMID marker stay with the weak-parsing boundary.\n\n## Notes\n\n- Network failures surface as `network_error=True` entries / honest\n  `unknown` verdicts — never as \"not found\", never cached. Run\n  `python3 scripts/verify_pmids.py --check-net` to diagnose connectivity.\n- Retraction status follows the registry at call time; for final\n  pre-submission passes run without the cache (`--no-cache` on the CLI).\n\nFile v4.1.0:skill-card.md\n\n## Description:\n\nChecks whether cited PMIDs, DOIs, and arXiv IDs resolve to the papers a reference claims, and reports citation mismatches and uncertain results.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[docsor1212](https://clawhub.ai/user/docsor1212)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nResearchers, medical writers, and evidence-based medicine teams use this skill to audit reference lists and AI-generated citations before publication, including checks for real identifiers that point to the wrong paper.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Citation identifiers, title queries, and optional contact email or API key may be sent to academic registry services.\n\nMitigation: Review the data and credentials before running networked checks; do not use this as a private offline verifier.\n\nRisk: Broad scans may read unrelated private drafts, and cached citation data may persist locally.\n\nMitigation: Use specific input paths, avoid unrelated private files, and consider --no-cache for sensitive work.\n\n## Reference(s):\n\n- [PubMed Verifier release](https://clawhub.ai/docsor1212/skills/pubmed-verifier)\n- [API examples](references/api_examples.md)\n- [Python API guide](references/python_api.md)\n\n## Skill Output:\n\n**Output Type(s):** [Text, Markdown, Shell commands, Files]\n\n**Output Format:** [Citation-verification guidance and text, HTML, JSON, or CSV reports; optional formatted references and audit exports]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Reports distinguish matched, mismatched, partial, invalid, and unknown citations.]\n\n## Skill Version(s):\n\n4.1.0 (source: skill frontmatter and server-resolved release)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v4.1.0:skillhub-meta.json\n\n{\n  \"name\": \"pubmed-verifier\",\n  \"version\": \"4.1.0\",\n  \"description\": \"Reference checker for AI-fabricated citations: batch-verify PMIDs against PubMed and catch the hallucination existence checks miss — a REAL PMID pointing to a DIFFERENT paper. Five-state citation verification (correct / mismatch / partial / invalid / unknown), citation-context parsing, Crossref DOI cross-check, retraction detection (capped at partial), correct-PMID suggestion, arXiv ID verification, plain-text reference list parsing, formatted reference lists (GB/T 7714 / Vancouver / APA / AMA), OpenAlex DOI fallback for non-Crossref registries, Chinese-reference title search, CSV/JSON claims, HTML/JSON/text reports. Four data sources: NCBI, Europe PMC fallback, Crossref, OpenAlex. Network failures are honestly reported as unverified, never as \\\"not found\\\". Zero dependencies, runs fully local. Triggers: verify PMIDs, check citations, citation audit, reference check, PMID check, batch verify references, AI hallucination detection, verify DOI, DOI check, citation formatting, reference formatter, GB/T 7714.\",\n  \"author\": \"docsor1212\",\n  \"tags\": [\n    \"pubmed\",\n    \"pmid\",\n    \"文献验证\",\n    \"引用核查\",\n    \"AI幻觉\",\n    \"学术写作\",\n    \"医学文献\",\n    \"citation\",\n    \"verification\",\n    \"academic\"\n  ],\n  \"license\": \"MIT-0\",\n  \"homepage\": \"https://github.com/docsor1212/pubmed-verifier\"\n}\n\nFile v4.1.0:examples/refs_list.sample.txt\n\n[1] Gattorno M, Sanni K. Development and validation of the autoinflammatory diseases effect scale. Annals of the Rheumatic Diseases. 2019;78(5):618-622. PMID: 31018962\n[2] Zaripova LN, Midgley A, Christmas SE. Autoinflammatory diseases in practice. Pediatric Rheumatology. 2021;19(1):75. https://doi.org/10.1186/s12969-021-00566-6\n[3]Obokata H. Stimulus-triggered fate conversion of somatic cells into pluripotency. Nature. 2014;511(7510):540-545.\n[4] And a deliberately broken line without any usable metadata.\n\nArchive v4.0.0: 11 files, 86860 bytes\n\nFiles: examples/claims.sample.csv (469b), examples/README.txt (1480b), examples/refs_list.sample.txt (512b), README.md (6523b), references/api_examples.md (4849b), references/python_api.md (3617b), scripts/verify_pmids.py (207425b), skill-card.md (2045b), SKILL.md (38183b), skillhub-meta.json (1378b), _meta.json (134b)\n\nFile v4.0.0:SKILL.md\n\n---\nname: pubmed-verifier\nlicense: MIT-0\nversion: 4.0.0\ndescription: >-\n  Reference checker for AI-fabricated citations: batch-verify PMIDs against\n  PubMed and catch the hallucination existence checks miss — a REAL PMID\n  pointing to a DIFFERENT paper. Five-state citation verification (correct /\n  mismatch / partial / invalid / unknown), citation-context parsing, dual\n  fuzzy matching, Crossref DOI cross-check, retraction detection (capped at\n  partial), correct-PMID suggestion, arXiv ID verification, plain-text\n  reference list parsing, formatted reference lists (GB/T 7714 / Vancouver /\n  APA / AMA), SQLite cache, CSV/JSON claims, HTML/JSON/text reports. Dual data\n  sources with automatic Europe PMC fallback. Network failures are honestly\n  reported as unverified, never as \"not found\". Zero dependencies, runs fully\n  local. Triggers: verify PMIDs, check citations, validate references,\n  citation audit, reference check, PMID check, batch verify references, AI\n  hallucination detection, verify DOI, DOI check, validate citations, citation\n  formatting, reference formatter, GB/T 7714.\n---\n\n# PubMed Citation Verifier v4.0.0\n\nBatch verification of PMID citations via the PubMed E-utilities API. Not just\n\"does this PMID exist\" — **does this PMID point to the paper you claim?**\nZero dependencies, pure standard library, fully local.\n\n## When to use this skill\n\nInvoke it whenever citation truth matters:\n\n- \"Verify / check these PMIDs / this reference list\" (before submission or release)\n- Auditing citations in AI-generated text (fabricated or mismatched references)\n- Spot-checking a systematic review's bibliography\n- \"Does PMID 12345678 really say X?\" — point-of-doubt verification\n- Batch QA of a knowledge base's PMID/DOI citations\n- Automatic invocation: trigger this skill whenever the user asks to\n  \"verify/check/audit\" citations, PMIDs, DOIs, arXiv preprints or\n  references — no flags needed for a first pass; add `--claims-file`\n  when they supply the expected titles.\n\n**Trigger priority & tool choice** — explicit \"verify / check / audit\ncitations, references, PMIDs, DOIs\" requests invoke this skill first.\ncite-holmes is for deep research with machine-verified citations; when a\nrequest mixes research and verification, run the research first, then\nthis tool for the final reference audit.\n\n## The five-state verdict\n\n| Verdict | Meaning |\n|---------|---------|\n| ✅ Correct | PMID exists AND matches the claimed paper |\n| ⚠️ Mismatch | PMID exists but points to a **different** paper (the most common AI hallucination!) |\n| 🔶 Partial | Some metadata matches (e.g. author+journal but title differs) |\n| ❌ Invalid | PMID does not exist in PubMed |\n| ❓ Unknown | Not enough claimed metadata to cross-check — or both data sources unreachable (never misreported as invalid) |\n\n**Why existence checks are not enough:** a large share of fabricated\ncitations use REAL PMIDs that point to a different paper from the same\nyear/journal/field — in one of our own audits, 4 of 5 \"valid\" PMIDs were\nwrong this way. A binary exists/not-exists check misses them all.\n\n## Quick start\n\n```bash\n# Scan a project directory for PMIDs (parses citation context automatically)\npython3 scripts/verify_pmids.py --source /path/to/project --output report.html\n\n# Verify specific PMIDs\npython3 scripts/verify_pmids.py --pmids 31018962,22213727\n\n# Mismatch demo: PMID 34078778 is actually a dental-materials paper, so the\n# JIA claims below will NOT match it — expect ⚠️ mismatch verdicts\npython3 scripts/verify_pmids.py --claims '[{\"pmid\":\"34078778\",\"title\":\"JIA pathogenesis\",\"authors\":[\"Zaripova\"],\"journal\":\"Pediatr Rheumatol Online J\",\"year\":\"2021\"}]' --output report.html\n\n# Claims from a CSV file + suggest correct PMIDs for mismatches\npython3 scripts/verify_pmids.py --claims-file claims.csv --suggest --output report.html\n\n# Crossref DOI cross-verification + audit working-paper + BibTeX + full pipeline\npython3 scripts/verify_pmids.py --source /path/to/files --verify-doi --suggest --output report.html --export-audit audit.json --export-bibtex refs.bib\n\n# Verify DOIs directly (no PMIDs) + delta audit vs a previous run\npython3 scripts/verify_pmids.py --dois \"10.1038/nature12968,10.4012/dmj.2020-408\" --workers 4 --export-audit audit.json --export-csv table.csv\npython3 scripts/verify_pmids.py --source /path/to/project --diff audit.json --output report.html\n\n# Verify arXiv IDs (preprints) — mixed audits supported\npython3 scripts/verify_pmids.py --arxivs \"2401.12345,cs/0211004\" --no-cache\n\n# Pre-flight: are the four data sources reachable right now?\npython3 scripts/verify_pmids.py --check-net\n\n# Audit a BibTeX or RIS bibliography file directly (PMID > DOI > arXiv routing)\npython3 scripts/verify_pmids.py --bibliography refs.bib --no-cache\npython3 scripts/verify_pmids.py --bibliography refs.ris --no-cache --export-ris verified.ris\n\n# Verify a reference list copied from a paper draft (inline IDs route\n# exactly; title-only entries resolve via PubMed title search)\npython3 scripts/verify_pmids.py --parse-text references.txt --no-cache\n\n# Format the verified entries as a ready-to-paste reference list\n# (--citation-style: gbt | vancouver | apa | ama; default gbt)\npython3 scripts/verify_pmids.py --parse-text references.txt --format-references refs_gbt.txt --citation-style gbt\n\n# Institutional niceties (recommended): NCBI API key + contact email\npython3 scripts/verify_pmids.py --source . --verify-doi --ncbi-api-key $NCBI_API_KEY --mailto you@lab.org\n```\n\n## FAQ & common mistakes\n\n**Top 10 things NOT to do** (each is detailed below or in Anti-patterns):\n\n| # | Don't | Do instead |\n|---|-------|------------|\n| 1 | Treat `--pmids` existence output as \"verified\" | Feed `--claims-file` with titles for real verification |\n| 2 | Submit claims without `title` | Always include titles — the verdict caps at partial without one |\n| 3 | Trust cached verdicts on publication day | Final check with `--no-cache` |\n| 4 | Read \"not found\" as \"fabricated\" for auto-extracted DOIs | Check doi.org / arxiv.org by hand first |\n| 5 | Treat the leading `'` in CSV cells as corruption | It is the formula-injection guard — strip after import |\n| 6 | Read the READY line as a quality score | It means \"no problems among the checks that ran\" |\n| 7 | Pass `--source` together with `--pmids` | `--source` is ignored entirely when `--pmids` is given |\n| 8 | Deep-verify (`--verify-doi` / `--suggest`) a thousand-entry sweep | Sweep first, deep-verify the flagged subset |\n| 9 | Expect author matching across CJK↔Latin names | They are skipped honestly (`author_check: skipped`) |\n| 10 | Ship a reference list without the audit trail | `--export-audit` writes a replayable working paper |\n\n**Large batch (hundreds of PMIDs) is slow — how to speed it up?**\nMetadata-only verification queries in batches of 50 with 0.4 s spacing\n(0.12 s with `--ncbi-api-key`); cached re-runs are ~5 s. `--verify-doi` adds\none Crossref call *per citation* and `--suggest` adds one search *per\nmismatch* — skip them for bulk sweeps, run them on the flagged subset.\n\n**When must I use `--claims-file` instead of scanning?**\nContext parsing is heuristic (v3.3.0 guards common abbreviations, exotic formatting can still mis-split).\nFor precise verification — or DOIs in claims (splice detection needs `doi`)\n— feed structured JSON/CSV claims.\n\n**My citation text is in Chinese — the scan parses little?**\nGB/T 7714-style references (……标题[J]. 刊名, 年… PMID: xxx) parse since\nv3.9.0 — title/journal comparisons against Latin registries are skipped\ncross-language, so verdicts rest on year/DOI evidence. Free Chinese prose\nwithout a PMID marker still yields little — feed `--claims-file` for full\nverdicts regardless of language.\n\n**Slow or unstable network (China)?**\nStandard `HTTPS_PROXY`/`HTTP_PROXY` env vars are honored natively; raise\n`--timeout`; `--meta-source europepmc` routes via Europe PMC when NCBI is\nunreachable (per-entry `meta_source` shows which was used); cached results\nare reused for 30 days.\n\n**❓ unknown vs ❌ invalid?**\n`unknown` (exit 2) = \"could not verify, sources unreachable\" — retry later;\n`invalid` (exit 1) = \"verified not-found\". Network failures are never\nreported as not-found and never cached.\n\n**What does RETRACTED mean in a report?**\nThe registry itself lists the paper's publication type as \"Retracted\nPublication\" (checked for every citation since v2.7.0 — no DOI or flags\nneeded), and/or Crossref records a retraction. The verdict is capped at\npartial and a human review note is attached — citing it would propagate\nwithdrawn science. A retraction *notice* is never flagged; papers under\n*Expression of Concern* (an editorial note, not a retraction) are not\nflagged either. Retraction status reflects the registry at cache time — for\na final pre-submission check, run with `--no-cache`.\n\n**Mismatch reported but the title looks similar?**\nCheck `details` for which field diverged; thresholds are strict on purpose.\nFeed the full citation via `--claims-file` for a precise verdict.\n\n**Can I audit my .bib or .ris file directly?**\nYes — `--bibliography refs.bib` (BibTeX) and `--bibliography refs.ris`\n(RIS/Zotero/EndNote/Mendeley) route each entry by PMID > DOI > arXiv\nand cross-check the claimed metadata. `--lint-claims` validates either\nformat offline first. The round trip works both ways: `--export-bibtex`\nand `--export-ris` output can be fed back after edits.\n\n**My claims file seems to lose rows / verdicts look weaker than expected?**\nLint it offline first: `python3 scripts/verify_pmids.py --lint-claims\nclaims.csv` reports unusable rows, ID shape errors, unknown columns\n(typo'd headers like \"titel\"), missing titles and DOI prefix problems —\nno network, exit 1 on errors.\n\n**Can I verify a DOI with claimed metadata (full verdict)?**\nYes — since v3.4.0 a claims row keyed by `doi` (with `title`, optionally\n`authors`/`journal`/`year`) gets the same cross-check as PMID claims:\ncorrect / mismatch / partial against the registered metadata. A DOI row\nwithout claims stays `unknown` (existence only).\n\n**Why did my arXiv citation drop from correct to partial?**\nYour claims row paired an `arxiv_id` with a `doi`, and the DOI does not\nmatch the version-of-record DOI registered on that arXiv entry — a typo,\nor a DOI from a different paper. The registered DOI is in `details`;\nfix the claim or drop the `doi` cell.\n\n## Anti-patterns — things done WRONG\n\nEach entry: the mistake → why it fails → the right way.\n\n1. **Treating `--pmids` output as \"fully verified\"** — existence-only.\n   → Wrong: \"all 5 PMIDs exist, so the citations are correct.\"\n   → Right: existence-checked only; feed `--claims-file` with titles for\n   real verification (the READY line says so explicitly).\n2. **Claims without `title`** — author/journal/year alone can never reach\n   `correct`; the report caps at `partial`. → Always include titles.\n3. **Trusting a cached verdict right after publication day** — a brand-new\n   PMID may have been cached as not-found by an earlier run, and retraction\n   status is as of cache time. → Final pre-submission check: `--no-cache`.\n4. **Assuming \"not found\" always means fabricated** — auto-extracted DOIs\n   that 404 stay *suspects* (DataCite DOIs don't live in Crossref); arXiv\n   IDs removed by moderators also return empty. → Check doi.org / arxiv.org\n   by hand before accusing.\n5. **Copying the leading `'` from CSV cells** — that apostrophe is the\n   formula-injection guard, not data corruption. → Strip it after import.\n6. **Reading the READY line as a quality score** — it only means \"no\n   problems found among the checks that ran\", not \"this paper is good\".\n\n## Boundaries — declared limits\n\nWhat this tool can NOT do, consolidated in one place:\n\n- **Splice/mismatch signals report disagreement, never pick a side** —\n  when claim and registry disagree, a human reads the evidence line.\n- **Cross-language authors are skipped, not failed** — CJK↔Latin author\n  names are never compared (`author_check: skipped`); the verdict rests\n  on title/journal/year alone.\n- **Context parsing is heuristic** — v3.3.0 keeps common abbreviations\n  (U.S., e.g., vs., St., Vol., No.) from splitting a title, but exotic\n  formatting can still mis-split; for exact metadata use `--claims-file`.\n- **DataCite/repository DOIs are not in Crossref** — an auto-extracted\n  or bibliography-sourced DOI missing from Crossref stays a *suspect*\n  (`.bib` files commonly hold DataCite/repository DOIs — check doi.org\n  by hand); only `--dois`/claims DOIs count a Crossref 404 as invalid.\n- **arXiv moderator removals also return \"not found\"** — the invalid\n  verdict carries that caveat in its details.\n- **Retraction status is as-of-cache-time** — final pre-submission\n  checks should run with `--no-cache`.\n- **Single-letter initials never match** — \"Smith J\" vs \"Smith John\" is\n  not counted as a miss.\n- **unknown ≠ invalid** — unreachable sources yield exit 2 and\n  `unknown`; network failures are never reported as \"not found\" and\n  never cached.\n- **arXiv pacing is deliberate** — the official API asks for ≥3 s\n  between calls; large arXiv batches are slow by design (progress + ETA\n  on stderr). Entries that register a version-of-record DOI add one\n  Europe PMC lookup each for the PMID link.\n- **DOI claims compare against the registry that actually answered** —\n  the linked PubMed record when the DOI resolves to one, Crossref\n  otherwise (Crossref author fields are sparser, so the author mark is\n  more often \"—\"); a DOI row without a claimed title stays `unknown`,\n  not partial; an explicitly user-provided DOI (`--dois` or claims) that\n  is missing from Crossref counts as invalid.\n\n## Claims reference format\n\n`--claims-file` accepts JSON (an array of objects) or CSV. Recognized\ncolumns: `pmid`, `title`, `authors` (semicolon/pipe-separated),\n`journal`, `year`, `doi`, `arxiv_id`. A row needs one of `pmid`,\n`arxiv_id` or `doi`; a missing `title` caps the verdict at partial.\n\n```csv\npmid,title,authors,journal,year,doi,arxiv_id\n31018962,Candidate criteria for diagnosis of familial...,Gattorno,Ann Rheum Dis,2019,10.1136/annrheumdis-2019-215048,\n,Attention Is All You Need,Vaswani,NeurIPS,2017,,1706.03762\n,City size and the spreading of COVID-19 in Brazil,Silva Junior;Other,PLOS ONE,2020,10.1371/journal.pone.0239699,\n```\n\nValidate any file offline first: `--lint-claims file.csv` reports\nunusable rows, ID shape errors, unknown columns, duplicates and missing\ntitles (no network). Lint wins when combined with verification flags —\nonly the lint runs.\n\n## Best practices & tuning\n\n- **Speed up large batches** — request an NCBI API key (see\n  https://ncbiinsights.ncbi.nlm.nih.gov/api-keys/): batches of 50 IDs run\n  at 0.12 s spacing instead of 0.4 s; cached re-runs take seconds.\n- **Parallel DOI verification** — `--workers` (default 4, cap 8) applies to\n  Crossref resolution and Europe PMC linking; arXiv stays serial by\n  official etiquette (≥3 s between calls).\n- **Two-phase workflow** — sweep with metadata-only verification first\n  (no `--verify-doi`, no `--suggest`), then deep-verify only the flagged\n  subset; each deep flag adds one API call per citation.\n- **Round-trip bibliographies** — `--export-bibtex` / `--export-ris`\n  write verified entries; after edits, `--bibliography refs.bib` /\n  `refs.ris` re-audits the file (the PMID re-links the full record).\n- **Claims over context parsing** — whenever you know the expected titles,\n  feed `--claims-file`: it enables the full verdict ladder and the DOI /\n  arXiv pairing checks. Validate the file offline first:\n  `python3 scripts/verify_pmids.py --lint-claims claims.csv` reports ID\n  shape errors, missing titles, unknown columns and duplicates without any\n  network access.\n- **Flaky networks** — raise `--timeout`; HTTPS_PROXY/HTTP_PROXY are\n  honored natively; unreachable NCBI falls back to Europe PMC\n  automatically (`meta_source` shows which answered).\n- **Cache policy** — results cache 30 days, negative entries 3 days;\n  `--cache-days` to tune; `--no-cache` for the final pre-submission pass.\n- **Scale expectations** — metadata-only throughput is API-bound\n  (~1–2 min per 1000 PMIDs with an API key); DOI resolution adds one\n  Crossref call per DOI. One deliberate trade-off: the verifier is a\n  single stdlib-only file — copy `scripts/verify_pmids.py` anywhere with\n  Python 3.9+ and it runs, no pip, no venv (that portability is why the\n  code is not split into modules).\n\n## v2.2.0 — network hardening\n\n| Feature | Flag | Effect |\n|---------|------|--------|\n| NCBI API key | `--ncbi-api-key` / env `NCBI_API_KEY` | Rate ceiling 3→10 req/s, batch interval 0.4s→0.12s (~3x faster) |\n| Europe PMC fallback | `--meta-source auto\\|ncbi\\|europepmc` | NCBI batch failure automatically retries via Europe PMC (free, no key); per-entry origin in JSON (`meta_source`) |\n| Crossref polite pool | `--mailto` / env `PUBMED_VERIFIER_MAILTO` | `?mailto=` on Crossref + tool/email params on NCBI — more generous limits |\n| Retry-After backoff | automatic | 429 responses honored (clamped 1–5 s) instead of failing |\n| UA rotation | automatic | 403/406 retried with a browser User-Agent |\n| Host circuit breaker | automatic | After 2 call-level transport failures a host is skipped with an actionable message; success resets; HTTP errors never trip it |\n| Honest unknown | automatic | Network failures report as ❓ unknown + exit code 2, never as \"PMID not found\", and are never cached |\n\n**Exit codes:** `0` clean · `1` problems found (invalid / mismatch / retracted /\nDOI-splice, incl. arXiv claimed-DOI pairing mismatches) · `2` could not verify\n(data sources unreachable) — automation can tell \"all good\" from \"no answer\".\n\n## v2.3.0 — retraction detection\n\nWith `--verify-doi`, each cited DOI is also checked against Crossref's\nwithdrawal records (`updated-by`). A paper Crossref lists as RETRACTED is:\n\n- flagged in JSON (`retracted: true` + `retraction_note`) and in reports,\n- **capped at 🔶 partial** even when every metadata field matches — citing a\n  retracted paper is never \"correct\"; the report says *human review required*.\n\nCorrections and other update types do not trigger the cap. Crossref outages\nnever flag anything (a missing check is not a retraction).\n\n## v2.4.0 — DOI↔PMID cross-check & journal abbreviations\n\n- **DOI splice detection**: add a `doi` field to your claims (JSON or CSV).\n  The claimed DOI is compared with the DOI registered for that PMID — a\n  mismatch is a splice/fabrication signal (a real DOI attached to the wrong\n  paper): flagged in JSON (`doi_splice_suspect`), verdict capped at 🔶\n  partial, counted in exit 1. Uses the PubMed record only — no extra API call.\n- **Journal abbreviation equivalence**: journal matching now understands\n  NLM-style abbreviations in both directions — \"N Engl J Med\" matches \"New\n  England Journal of Medicine\", \"Pediatr Rheumatol\" matches \"Pediatric\n  Rheumatology\" (in-order word prefixes, function words skipped). No more\n  false \"journal differs\" for abbreviated citations.\n\nKnown limits: highly ambiguous abbreviations can over-match at the\njournal-only level (\"J Immunol\" ~ \"Journal of Immunology Research\") — the\ntitle remains the decisive field. A DOI-splice flag can also appear on an\notherwise-unverifiable citation (the DOI mismatch is an independent fact).\n\n## v2.5.0 — author-name verification\n\n- **Initials never match**: single-letter tokens (\"A.\", \"L.\") on either side\n  are excluded from surname matching — an initial is not evidence, and\n  substring-matching one produced false author hits.\n- **Cross-language honesty**: CJK author names claimed against Latin\n  registry records (or the reverse) are skipped, not counted as a mismatch —\n  the report marks them `author_check: skipped` and the verdict falls back\n  to what was actually comparable (title/journal/year), or to ❓ unknown\n  when nothing else is checkable.\n\n## Security & behavior declaration\n\n- Single-run CLI: scan, verify, write the report, exit. No daemons, no\n  background jobs, nothing downloaded or installed at runtime (pure standard\n  library, zero dependencies).\n- Network access is limited to these official academic registries, always\n  over HTTPS: `eutils.ncbi.nlm.nih.gov`, `www.ebi.ac.uk` (Europe PMC),\n  `api.crossref.org`, `export.arxiv.org`. No other hosts are contacted; no telemetry, no\n  analytics, no data collection — the only outbound payloads are the PMIDs,\n  DOIs and titles you asked to verify.\n- Your files and reports stay on your machine. Writes are limited to the\n  report paths you pass and the SQLite cache under\n  `~/.cache/pubmed-verifier/` (`--no-cache` to disable).\n- Optional environment variables `NCBI_API_KEY` / `PUBMED_VERIFIER_MAILTO`\n  authenticate or attribute your own API requests and are never sent\n  anywhere else.\n- No OS integration: no subprocesses, no system services, no privilege\n  changes, no scheduled tasks.\n\n## v2.6.0 — audit working-paper & report v2\n\n- **`--export-audit audit.json`** — a self-contained JSON working-paper for\n  transparent review: tool identity and version, the exact (API-key-redacted)\n  invocation, per-citation evidence chains (claimed vs registered fields,\n  title match scores from both algorithms, author match with cross-language\n  skip records, DOI cross-check, retraction signals) and the\n  verdict-ladder trace for every citation. A reviewer can replay the entire\n  verification from this file alone.\n- **HTML report v2** — verdict filter tabs, severity-sorted rows (retracted\n  and DOI-splice first, highlighted), a field-level evidence column\n  (title/author/journal/year ✓✗—) and a reproducibility footer (redacted\n  command line + version + data sources).\n- **Reliability** — negative cache entries now expire after 3 days (a\n  legitimately new, ahead-of-print PMID is no longer reported \"not found\"\n  for a month), and the circuit breaker self-heals: after a 30 s cooldown it\n  admits one probe call and resets on success.\n\n## v2.7.0 — retraction for every PMID, BibTeX export, readiness verdict\n\n- **Retraction detection, source-independent** — the registry's own\n  publication type (\"Retracted Publication\"; present in both NCBI esummary\n  and Europe PMC) now flags retracted papers for EVERY citation: no DOI\n  required, no `--verify-doi` required, and the flag survives the cache\n  (schema v3). Crossref `updated-by` remains the detail source (the\n  retraction-notice DOI) when `--verify-doi` is on. A retraction *notice*\n  itself is never flagged.\n- **`--export-bibtex refs.bib`** — export the verified bibliography: correct\n  entries as `@article`, partial entries commented out with their divergence\n  note, mismatched/invalid/unknown/retracted entries excluded and counted.\n- **Submission-readiness verdict** — every report now leads with one line:\n  `SUBMISSION READY` or `NOT SUBMISSION-READY — <per-problem counts>`.\n- Cache schema v3 (adds a `retracted` column, auto-migrated).\n\n## v2.8.0 — DOI-native verification & delta audits\n\n- **`--dois \"10.x/a, 10.y/b\"`** — verify DOIs natively, no PMID required.\n  `--source` scans now also extract DOIs from your files automatically.\n  Each DOI is resolved via the Crossref works API: not-found on an explicitly\n  provided DOI = fabrication signal (invalid, exit 1); on one auto-extracted\n  from scanned text it stays a suspect (unknown) — scanned strings are never\n  user-endorsed, and DataCite/repository DOIs do not live in Crossref, so\n  always double-check at doi.org. Resolved = existence confirmed with the\n  registered metadata attached for manual comparison — *existence is never\n  dressed up as a match*.\n- **`--diff previous-audit.json`** — delta audit against a previous working\n  paper: **newly retracted** (the safety signal — a paper retracted after\n  your last audit; act on it: swap or drop the citation, cite the retraction\n  notice instead, and re-check any conclusion that relied on it), degraded,\n  improved, new and dropped citations, with counts in every report format.\n  Built for periodic knowledge-base audits: \"what changed since last time?\"\n\n## v2.9.0 — DOI entries become first-class\n\n- **DOI→PMID linking** — a resolved DOI is linked back to its PMID via the\n  Europe PMC DOI field query, pulling the full PubMed record: complete\n  metadata, retraction pubtype signal, and cache coverage. A DOI citation\n  now gets the same five-state record as a PMID citation (existence\n  confirmation only — the verdict remains unknown until claims are\n  provided).\n- **Parallel DOI resolution** — `--workers N` (default 4, max 8) resolves\n  DOI batches on a thread pool (roughly 3x faster on large lists), with\n  live progress output. For large `--dois` batches, set `--mailto` to stay\n  in Crossref's polite pool.\n- **`--export-csv table.csv`** — spreadsheet-friendly audit table\n  (key/verdict/flags/fields/details; formula-injection hardened).\n\n## v3.0.0 — arXiv ID verification (three citation types, one audit)\n\nReference lists carry preprints. v3.0.0 verifies **arXiv IDs** alongside\nPMIDs and DOIs: `arXiv:2401.12345` and `arxiv.org/abs/...` patterns are\nextracted from scans (or passed via `--arxivs`), checked against the\nofficial arXiv API, and judged — nonexistent ID = fabrication signal\n(invalid, exit 1); resolving ID = registered title/year attached, verdict\nstays unknown. Malformed IDs (bad YYMM month) are flagged by shape.\nTimely: arXiv penalizes submissions containing hallucinated or unverified\nreferences (2026-05 policy) — audit before you submit.\n\n## v3.3.0 — preprint ↔ published-version cross-check\n\narXiv entries carry the version-of-record DOI their authors registered at\npublication (`arxiv:doi`). v3.3.0 puts it to work:\n\n- **Claimed DOI vs registered DOI** — a claims row with both `arxiv_id`\n  and `doi` is cross-checked: agreement is reported as evidence\n  (`fields.doi ✓`); disagreement caps the verdict at `partial` — the DOI\n  belongs to a different paper (same failure class as PMID DOI-splice).\n- **Version of record surfaced** — verifying a bare preprint ID now shows\n  the registered DOI and, when the published version is PubMed-indexed,\n  its linked PMID — cite and verify the final version, not just the\n  preprint.\n- **Honest accounting** — the readiness line counts arXiv DOI-pairing\n  mismatches as problems; DOI/arXiv phase progress (stderr) now includes\n  elapsed time and an ETA for large batches.\n- Context parsing no longer truncates titles at sentence-internal\n  abbreviations (\"U.S. population\", \"e.g.\", \"vs.\", \"Vol.\").\n\nPairing example (match → correct with DOI evidence; wrong DOI → partial):\n\n```bash\npython3 scripts/verify_pmids.py --claims '[{\"arxiv_id\":\"2005.13892\",\n  \"title\":\"City size and the spreading of COVID-19 in Brazil\",\n  \"doi\":\"10.1371/journal.pone.0239699\"}]'\n```\n\n## v3.4.0 — DOI claims become first-class\n\nClaims rows could carry a PMID or an arXiv ID — a row keyed by DOI alone\nwas silently ignored, and DOI entries always stayed `unknown` (\"no claimed\nmetadata to cross-verify\"). v3.4.0 closes the matrix: all three citation\ntypes now accept claimed metadata.\n\n- A claims row with a `doi` (no PMID, no arXiv ID) is cross-checked against\n  the registered metadata — the linked PubMed record when the DOI resolves\n  to one, Crossref otherwise — and gets the full verdict ladder:\n  correct / mismatch / partial.\n- Retraction capping applies as everywhere: a claimed-correct match on a\n  retracted paper is capped at partial with the retraction note.\n- Without claims, DOI entries stay unknown — existence is never dressed up\n  as a match.\n\n```bash\npython3 scripts/verify_pmids.py --claims '[{\"doi\":\"10.1371/journal.pone.0239699\",\n  \"title\":\"City size and the spreading of COVID-19 in Brazil\",\n  \"journal\":\"PLoS ONE\",\"year\":\"2020\"}]'\n```\n\n## v3.5.0 — claims lint & usage-first restructuring\n\n- **`--lint-claims FILE`** — offline pre-flight for claims files\n  (JSON/CSV, zero network): ID shape errors, missing titles (the verdict\n  would cap at partial), unknown/typo'd columns, DOI prefix checks,\n  unusable rows — exit 1 on errors. Fix the format before the run\n  instead of guessing from weak verdicts.\n- **Documentation restructured around usage**: FAQ, anti-patterns and\n  declared boundaries now sit right after Quick start, led by a Top-10\n  \"don't do this\" table; new **Best practices & tuning** section (API-key\n  batching, worker tuning, two-phase deep-verification, cache policy,\n  scale expectations — and why the verifier is deliberately one file).\n\n## v3.6.0 — BibTeX bibliography audit\n\n**`--bibliography refs.bib`** audits a .bib file directly: entries route by\nPMID (the `pmid` field, or a \"PMID: NNNN\" in the note — including the\nnotes `--export-bibtex` itself writes) > DOI > arXiv (`eprint`, or an\n\"arXiv:XXXX.XXXXX\" in the journal/note), each carrying its claimed\ntitle/authors/journal/year for the full cross-check. Entries without any\nroutable ID are counted and skipped. This closes the loop with\n`--export-bibtex`: a verified bibliography can be re-audited after edits.\n`--lint-claims refs.bib` validates the file offline (unroutable entries,\nmissing titles, unclosed blocks).\n\n## v3.7.0 — RIS bibliography support\n\n`--bibliography` now accepts **RIS files** (`refs.ris` — Zotero/EndNote/\nMendeley exports) alongside BibTeX: records route by PMID (AN tag, or a\n\"PMID: NNNN\" note) > DOI (DO) > arXiv (UR/eprint), each with claimed\nmetadata for the full cross-check. **`--export-ris`** completes the loop —\nverified entries as `TY JOUR`, partials as `TY DATA` with a PARTIAL note.\n`--lint-claims refs.ris` validates offline (unroutable records, missing\ntitles, duplicates, unterminated records).\n\n## v3.8.0 — verdict-ladder corrections (independent external test audit)\n\nAn independent real-data audit (24 scenarios, registry-ground-truthed) found\nthree judgment-ladder defects; all three are fixed and locked:\n\n- **B1 scan false positives** — when a citation is not the first sentence of\n  its block, the author line was taken as the claimed title (correct\n  references reported as mismatch). The parser now shifts past an\n  author-shaped segment.\n- **B2 author mis-attribution** — a claimed author set entirely different\n  from the registry (zero surname overlap) was still reported as correct\n  when title and journal matched; it now caps at partial. Partial surname\n  overlap (abbreviated or reordered names) keeps correct.\n- **B3 title-less claims** — author+journal+year matches without a claimed\n  title now cap at partial, exactly as the FAQ/anti-patterns/boundaries\n  always promised (the same author/journal/year can cover several papers).\n  Cross-language author skips follow the same cap.\n- Also: duplicate arXiv IDs announce their merge on stderr (run\n  `--lint-claims` to catch duplicate claim rows offline); the\n  cross-language skip no longer coexists with contradictory details text;\n  empty `--claims` gets a one-line hint; docs state that Chinese-language\n  citation contexts parse weakly — use `--claims-file`.\n\n## v3.9.0 — Chinese references (GB/T 7714), network pre-flight, Python API\n\n- **GB/T 7714 Chinese references parse** — `……标题[J]. 刊名, 年… PMID: xxx`\n  contexts now yield the title (via the `[J]/[M]/[R]` type marker),\n  authors, journal, year and any DOI that follows the PMID. The title\n  comparison against a Latin registry title is skipped cross-language\n  (never a mismatch source); a DOI in the reference is cross-checked\n  against the registry like any claims DOI. Requires a PMID marker per\n  reference.\n- **`--check-net`** — probes the four data sources (5 s each), shows your\n  proxy state and prints practical next steps when something is\n  unreachable. Run it when results come back unknown on a constrained\n  network.\n- **Python API reference** — `references/python_api.md` documents the\n  stable importable surfaces (`fetch_summaries`, `cross_check_citation`,\n  `verify_doi_entry`, `parse_citation_context`, `suggest_correct_pmid`)\n  with copy-paste examples.\n\n## v4.0.0 — plain-text reference lists & formatted reference list export\n\n- **`--parse-text refs.txt`** — verify a reference list copied straight out\n  of a manuscript draft. Numbered entries (`[1]`, `1.`) split cleanly;\n  inline PMID / DOI / arXiv IDs route exactly (full five-state\n  cross-check); a markerless entry with a parseable title resolves via\n  PubMed title search and is marked `resolved_by: title_search` — weaker\n  evidence than a supplied ID, honestly labeled, and never a mismatch\n  source. A network failure stays unknown with a retry hint, never\n  \"not found\". Entries with no routable ID and no usable title are\n  skipped with a note instead of being guessed.\n- **`--format-references out.txt --citation-style gbt|vancouver|apa|ama`**\n  — the formatting leg of verify-then-format: verified entries rendered as\n  a ready-to-paste numbered reference list (GB/T 7714-2015 numeric for\n  Chinese submissions, Vancouver, APA 7th, AMA 11th). Formatting stands on\n  registry-verified records only: correct entries form the list, partial\n  entries go to a manual-review section (never silently dropped), and\n  retracted entries are excluded with a warning. Author lists are\n  formatted per style without inventing names — a registry-truncated list\n  closes with the style's et-al form, and APA discloses the truncation.\n- Markerless Latin references now parse through the full sentence\n  segmentation (titles feed the title-search leg); CJK entries without a\n  PMID/DOI marker keep the documented weak-parsing boundary — supply a\n  PMID or DOI for exact routing.\n- `examples/refs_list.sample.txt` ships a ready-made parse-text fixture.\n\n## How it works\n\n1. **Extract + parse context** — finds `PMID: 12345678` / PubMed URLs in\n   `.html .md .txt .htm .json`, and parses the surrounding reference into\n   claimed authors / title / journal / year.\n2. **Fetch metadata (cached)** — PubMed esummary in batches of 50, SQLite\n   cache (30 days, `--cache-days`), 3 retries with backoff.\n   Europe PMC steps in per failed batch when NCBI is unreachable.\n3. **Cross-check claimed vs actual** — dual fuzzy matching: word-level\n   Jaccard overlap ≥ 50% OR SequenceMatcher ≥ 90% on titles; author surname\n   hits; journal containment or NLM abbreviation equivalence; exact year.\n4. **DOI↔PMID cross-check** (automatic when claims include `doi`) — a claimed\n   DOI differing from the PMID's registered DOI is a splice/fabrication\n   signal (capped at partial).\n5. **Crossref DOI verification** (optional `--verify-doi`) — resolves each\n   cited DOI via Crossref, compares the registered title with the PubMed\n   record (`doi_title_match` in JSON), and detects RETRACTED papers (verdict\n   capped at partial). A `doi_verified: false` with note\n   \"crossref unreachable\" is a network fact, not a verdict.\n6. **Suggest the right PMID** (optional `--suggest`) — for mismatches,\n   searches PubMed with the claimed metadata and proposes top-3 candidates.\n   (Suggestion search always uses NCBI, even with `--meta-source europepmc`.)\n\nContext parsing is heuristic — since v3.3.0, common abbreviations\n(\"U.S.\", \"e.g.\", \"vs.\") no longer split a title, but exotic formatting\nstill can. For precise verification, feed structured claims via\n`--claims-file`.\n\n| Report | Flag | Use |\n|--------|------|-----|\n| HTML | `--output report.html` | Human review: claimed vs actual side by side |\n| JSON | `--output report.json` | Programmatic processing (includes `meta_source` per entry) |\n| Text | default | Quick terminal look |\n\n## Performance\n\nMeasured on a 225-PMID audit (5 esummary batches): metadata-only verification\nruns in seconds; cached re-runs take ~5 s. Each batch waits 0.4 s between\ncalls (0.12 s with `--ncbi-api-key`). Optional extras are per-citation:\n`--verify-doi` adds one Crossref call (~0.5–1 s) per cited DOI, and\n`--suggest` adds one PubMed search per mismatch.\n\n## When to use which tool\n\n- **pubmed-verifier (this skill)** — fast, batch, targeted: I have a list of\n  PMIDs/DOIs and need to know if they are real and correctly cited.\n- **cite-holmes** — deep research: interrogate every citation of a whole\n  document across multiple databases, with graded confidence reports.\n\nThey share the same five-state philosophy and are safe to use together.\n\n## Use cases\n\n- Systematic review / meta-analysis reference audits\n- Verifying citations in AI-generated content\n- Pre-submission self-check of a manuscript's reference list\n- Medical knowledge base / teaching material QA\n- Pharmacovigilance literature verification\n\n## Related tools\n\nEach tool solves one step of reference work; use whichever fits the task.\n\n- **cn-med-oa** — free Chinese medical literature full-text download & metadata\n- **cite-holmes** — deep research with machine-verified citations\n- **paper-polisher-pro** — academic polishing, terminology & journal precheck\n- **academic-figures** — publication-ready scientific figures in one command\n- **doc-holmes** — layout-preserving PDF translation (in testing)\n\nTypical order: get papers (cn-med-oa), verify citations (this tool or\ncite-holmes), polish (paper-polisher-pro), make figures (academic-figures),\ntranslate PDFs (doc-holmes) — pick whichever step you need.\n\n## Files\n\n| File | Purpose |\n|------|---------|\n| `scripts/verify_pmids.py` | Main verifier (v4.0.0, stdlib-only) |\n| `references/api_examples.md` | PubMed / Europe PMC / Crossref / arXiv API notes |\n| `references/python_api.md` | Calling the verifier from Python (stable surfaces + examples) |\n| `examples/claims.sample.csv` | Reference format for `--claims-file` (incl. a DOI-only row) |\n| `examples/refs.sample.bib` | Sample bibliography for `--bibliography` (DOI/PMID/arXiv routing) |\n| `examples/refs.sample.ris` | Sample RIS bibliography (Zotero/EndNote/Mendeley) |\n| `tests/` | Offline matrix + real-network acceptance (repo only, not in the package) |\n\n## License\n\nMIT-0 — free to use, modify and redistribute, no attribution required.\n\nFile v4.0.0:examples/README.txt\n\nclaims.sample.csv — reference format for --claims-file\n\n- pmid column: PubMed pipeline (full five-state verification)\n- rows with only arxiv_id (+title): arXiv pipeline (official API check)\n- doi column: DOI-splice cross-check against the PubMed-registered DOI\n  (PMID rows) / version-of-record pairing check (arXiv rows, v3.3.0) /\n  full five-state verdicts (doi-only rows, v3.4.0)\n- title: required for verdicts beyond existence-check\n- expected verdicts (as of v3.4.0): the 31018962 row carries a\n  deliberately inexact title → partial (title mismatch, everything else\n  matches — demonstrates the partial rung, not a data error); the\n  1706.03762 row → correct; the doi-only 10.1371 row → correct (its\n  second author \"Other\" is a synthetic placeholder; v3.8.0's\n  mis-attribution cap means a wrong author set would drop it to partial); the 24476887 row\n  (STAP cells) is RETRACTED on purpose → RETRACTED flag with a partial\n  cap — it demonstrates retraction detection, not a data error\n\nRun:  python3 scripts/verify_pmids.py --claims-file examples/claims.sample.csv\n\nLint first (offline, v3.5.0): python3 scripts/verify_pmids.py --lint-claims examples/claims.sample.csv\n\nBibliography (v3.6.0): refs.sample.bib audits a .bib file directly —\npython3 scripts/verify_pmids.py --bibliography examples/refs.sample.bib\n\nRIS (v3.7.0): refs.sample.ris audits Zotero/EndNote/Mendeley exports —\npython3 scripts/verify_pmids.py --bibliography examples/refs.sample.ris\n\nFile v4.0.0:README.md\n\n# PubMed Citation Verifier 🔬\n\n[![GitHub Stars](https://img.shields.io/github/stars/docsor1212/pubmed-verifier?style=social&label=Star)](https://github.com/docsor1212/pubmed-verifier/stargazers)\n\nBatch-verify PMID citations against PubMed API. Built for researchers, medical writers, and evidence-based medicine teams.\n\n## Why?\n\nAcademic projects routinely contain hundreds of PMID citations. Manual verification is tedious and error-prone. During our own 225-reference audit, we found 3 invalid PMIDs and 6 cross-domain mismatches — errors that would have undermined the entire project.\n\n## Features\n\n- **Batch verification** — Scan entire project directories, extract all PMIDs, verify against PubMed in one run\n- **Mismatch detection** — Five-state verdicts catch a REAL PMID pointing to a DIFFERENT paper (the most common AI hallucination)\n- **Metadata validation** — Title, authors, journal and date fuzzy-compared against your claims (DOI cross-checked via Crossref with `--verify-doi`)\n- **Dual sources** — NCBI E-utilities primary, Europe PMC automatic fallback when NCBI is unreachable\n- **Network hardening** — Optional NCBI API key (3x faster batches), Crossref polite pool, 429 Retry-After backoff, UA rotation, per-host circuit breaker\n- **Content matching** — Keyword overlap scoring flags potentially irrelevant citations\n- **Three citation types in one audit** — PMIDs, DOIs and arXiv IDs (preprints) verified in a single scan; arXiv IDs checked against the official API (nonexistent = fabrication signal); claims rows keyed by DOI get full verdicts, and preprints surface their published version of record\n- **Claims lint** — `--lint-claims file.csv` validates a claims file offline (ID shapes, missing titles, unknown columns, duplicates) before any verification run\n- **Bibliography audit** — `--bibliography refs.bib` / `refs.ris` verifies BibTeX and RIS files directly (Zotero/EndNote/Mendeley exports; entries route by PMID > DOI > arXiv with full cross-checks); round-trips with `--export-bibtex` / `--export-ris`\n- **Plain-text reference list parsing** — `--parse-text refs.txt` verifies a reference list copied straight out of a manuscript draft: numbered entries (`[1]`, `1.`) split cleanly, inline PMID/DOI/arXiv IDs route exactly, and a title-only entry resolves via PubMed title search (honestly labeled `resolved_by: title_search`, never a mismatch source)\n- **Formatted reference list export** — `--format-references out.txt --citation-style gbt|vancouver|apa|ama` renders verified entries as a ready-to-paste numbered list (GB/T 7714-2015, Vancouver, APA 7th, AMA 11th); partial entries go to a manual-review section, retracted ones are excluded with a warning — formatting that stands on registry-verified records\n- **Retraction detection for every citation** — the registry publication type (\"Retracted Publication\") flags retracted papers with no DOI or extra flags needed; Crossref `updated-by` adds the retraction-notice DOI\n- **Verified-bibliography export** — `--export-bibtex` writes correct entries as BibTeX, comments partial ones, excludes and counts the rest\n- **Submission-readiness verdict** — every report leads with `SUBMISSION READY` / `NOT SUBMISSION-READY` and per-problem counts\n- **Audit working-paper** — `--export-audit` writes a self-contained JSON trail (tool identity, redacted invocation, evidence chain, verdict trace) a third party can replay\n- **DOI-native verification** — `--dois` verifies DOIs directly (Crossref resolve, 404 = fabrication signal on explicit input); `--source` scans auto-extract DOIs; linked back to PMIDs via Europe PMC\n- **Delta audits** — `--diff previous-audit.json` reports newly retracted, degraded, improved, new and dropped citations\n- **Exports** — `--export-audit` (self-contained JSON trail), `--export-bibtex` (verified bibliography), `--export-csv` (spreadsheet audit table)\n- **Replacement search** — Find correct PMIDs for broken citations via PubMed search\n- **Multiple output formats** — HTML report, JSON (with per-entry metadata origin), or terminal summary\n\n## Quick Start\n\n```bash\n# Install\nopenclaw skills install docsor1212/pubmed-verifier\n\n# Verify all PMIDs in a project\npython3 scripts/verify_pmids.py --source /path/to/project --output report.html\n\n# Verify specific PMIDs\npython3 scripts/verify_pmids.py --pmids 31018962,22213727,999999999\n\n# Verify a reference list copied from a paper draft\npython3 scripts/verify_pmids.py --parse-text references.txt --no-cache\n\n# Format verified entries as a citation list (gbt | vancouver | apa | ama)\npython3 scripts/verify_pmids.py --parse-text references.txt --format-references refs_gbt.txt --citation-style gbt\n\n# Institutional mode: API key + polite pool\npython3 scripts/verify_pmids.py --source ./papers --verify-doi --ncbi-api-key $NCBI_API_KEY --mailto you@lab.org\n```\n\n**China mirror (ModelScope 魔搭)**: <https://modelscope.cn/skills/Docsor/pubmed-verifier> — if you find this skill useful, a like there helps others find it.\n\n## Use Cases\n\n| Scenario | Example |\n|----------|---------|\n| **Systematic review QA** | Verify all 200+ references before submission |\n| **Medical website audit** | Check evidence citations across clinical case library |\n| **Paper manuscript check** | Validate every PMID in your draft |\n| **Teaching material review** | Ensure lecture citations are accurate |\n| **Evidence library maintenance** | Periodic batch verification of reference databases |\n| **Draft reference list check & formatting** | Paste a manuscript's reference list, verify it, get a GB/T 7714 / APA list back |\n\n## Real-World Results\n\nAudited a 35-file pediatric rheumatology evidence library (225 PMID citations):\n- **222** citations: valid and content-matched ✅\n- **3** citations: invalid PMIDs found and corrected\n- **6** cross-domain citations: correctly flagged, reviewed, confirmed appropriate\n- Total time: ~5 minutes for full audit\n\n## Technical Details\n\n- **APIs**: PubMed E-utilities (esummary, esearch); Europe PMC fallback; Crossref DOI check\n- **Rate limit**: 3 req/s free (0.4s batch delay), 10 req/s with `--ncbi-api-key` (0.12s)\n- **Resilience**: 429 Retry-After backoff, 403/406 UA rotation, per-host circuit breaker\n- **Batch size**: 50 PMIDs per request\n- **File types**: `.html`, `.md`, `.txt`, `.json`, `.htm`\n- **PMID patterns**: `PMID: 12345678`, `PubMed: 12345678`, `pubmed.ncbi.nlm.nih.gov/12345678/`\n- **Exit codes**: 0 clean / 1 problems found / 2 network-incomplete\n\n## License\n\nMIT-0\n\nFile v4.0.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b6sjwend7cwwmhf7mdg1wzx82dfbf\",\n  \"slug\": \"pubmed-verifier\",\n  \"version\": \"4.0.0\",\n  \"publishedAt\": 1791486659265\n}\n\nFile v4.0.0:references/api_examples.md\n\n# PubMed E-utilities API Quick Reference\n\n## esummary — Article Metadata\n\n```bash\n# Single PMID\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=31018962&retmode=json\"\n\n# Batch (comma-separated)\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=31018962,22213727&retmode=json\"\n```\n\nResponse fields: `title`, `authors[].name`, `source` (journal), `pubdate`, `volume`, `pages`, `elocationid` (DOI).\n\nInvalid PMID → `{\"result\": {\"pmid\": {\"error\": \"cannot get document summary\"}}}` (the exact string has varied over time — the verifier only checks for the presence of an `error` key).\n\n## esearch — Find Articles by Query\n\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=Gattorno+classification+autoinflammatory&retmode=json&retmax=5\"\n```\n\nReturns `esearchresult.idlist` → array of PMIDs.\n\n## efetch — Full Abstracts\n\n```bash\n# efetch supports TEXT and XML only — retmode=json is NOT valid for efetch\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=37635643&rettype=abstract&retmode=text\"\n```\n\n## Deep Research Workflow (Systematic PubMed Search)\n\nWhen the user needs comprehensive medical literature research beyond simple PMID verification:\n\n### Step 1: esearch → keyword search → get PMIDs\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=<keywords>&retmax=10&sort=relevance\" | grep -oP '<Id>\\K\\d+'\n```\n\n### Step 2: esummary → metadata summaries (fast, compact)\n```bash\nPMIDS=\"pmid1,pmid2,pmid3\"\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=$PMIDS&retmode=json\" > /tmp/pubmed-results.json\npython3 -c \"\nimport json\ndata = json.load(open('/tmp/pubmed-results.json'))\nfor uid, art in data.get('result', {}).items():\n    if uid == 'uids': continue\n    authors = ', '.join([a.get('name','') for a in art.get('authors',[])[:6]])\n    print(f'PMID {uid}: {art.get(\\\"title\\\",\\\"\\\")}')\n    print(f'  {art.get(\\\"fulljournalname\\\",\\\"\\\")} {art.get(\\\"pubdate\\\",\\\"\\\")};{art.get(\\\"volume\\\",\\\"\\\")}:{art.get(\\\"pages\\\",\\\"\\\")}')\n    print(f'  DOI: {art.get(\\\"elocationid\\\",\\\"\\\")}')\n    print()\n\"\n```\n\n### Step 3: efetch → full abstracts for key papers\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=$PMIDS&rettype=abstract&retmode=text\"\n```\n\n### Pitfalls\n- **efetch has no JSON mode** — use `retmode=text` (or XML) only\n- **esearch may return unrelated results** — always cross-check with esummary titles\n- **Rate limit**: 3 req/s without API key. Add `sleep 0.5` between batches\n- **Large author lists**: esummary `authors` array can be 50+ — truncate to first 6 for display\n\n## Rate Limits\n\n- Without API key: 3 requests/second\n- With API key (`&api_key=YOUR_KEY`): 10 requests/second\n- API key obtained from NCBI Settings page\n- Etiquette params: `&tool=pubmed-verifier&email=you@lab.org` (added automatically with `--mailto`)\n\n## Europe PMC (fallback source, v2.2.0)\n\nFree, no key, mirrors PubMed; used automatically when NCBI batches fail\n(`--meta-source auto`) or forced with `--meta-source europepmc`.\n\n```bash\n# Batch lookup by PMID (EXT_ID), SRC:MED restricts to PubMed records\ncurl -s \"https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=(EXT_ID:31018962%20OR%20EXT_ID:22213727)%20AND%20SRC:MED&format=json&resultType=lite&pageSize=25\"\n```\n\nUseful fields: `id` (PMID), `title`, `authorString` (comma-separated),\n`journalTitle`, `pubYear`, `doi`, `journalVolume`, `pageInfo`.\n\n## Crossref DOI check (polite pool, v2.2.0)\n\n```bash\n# ?mailto= joins the polite pool — more generous rate limits\ncurl -s \"https://api.crossref.org/works/10.1038/nature12968?mailto=you@lab.org\"\n```\n\n- 429 responses carry `Retry-After` — back off accordingly (v2.2.0 clamps 1–5 s)\n- 403/406 usually mean rate limiting — back off and retry later (the tool rotates client identifiers, including a standard browser UA on retry, per the SKILL.md network-hardening table)\n- A correct DOI with WRONG paper metadata = spliced/fake citation signature\n\n## arXiv API (export.arxiv.org)\n\n```bash\n# one ID per lookup, Atom feed back; arxiv:doi = the version-of-record\n# DOI the authors registered — the preprint↔published cross-check uses it\ncurl -s \"https://export.arxiv.org/api/query?id_list=2005.13892&max_results=1\"\n```\n\n- No `<entry>` in the feed = the ID does not exist (fabrication signal);\n  a 200 response that is not an Atom feed (portal/maintenance page) is\n  treated as \"could not verify\", never as \"not found\"\n- Official etiquette: ≥3 s between calls — large batches are paced\n  deliberately (progress with ETA goes to stderr)\n- `arxiv:journal_ref` / `arxiv:doi` are author-registered fields, shown\n  in the audit trail; see SKILL.md v3.3.0 section for the pairing check\n\nFile v4.0.0:references/python_api.md\n\n# Python API — calling pubmed-verifier from your code\n\nThe whole tool is one stdlib-only module. Import it directly — no pip, no\nvenv (Python 3.9+):\n\n```python\nimport sys\nsys.path.insert(0, \"/path/to/pubmed-verifier/scripts\")\nimport verify_pmids as vp\n```\n\nAll functions below are stable surfaces used by the CLI itself (behavior\nis pinned by the offline test matrix in the source repository).\n\n## 1. Batch metadata for PMIDs\n\n```python\ninfo = vp.fetch_summaries([\"31018962\", \"22213727\"])   # NCBI, EPMC fallback\nr = info[\"31018962\"]\nprint(r[\"valid\"], r[\"title\"], r[\"journal\"], r.get(\"retracted\"))\n```\n\n## 2. Cross-check claimed vs registered metadata (the five-state ladder)\n\n```python\ncross = vp.cross_check_citation(\n    {\"claimed_title\": \"Classification criteria for autoinflammatory recurrent fevers\",\n     \"claimed_authors\": [\"Gattorno\"], \"claimed_journal\": \"Ann Rheum Dis\",\n     \"claimed_year\": \"2019\"},\n    {\"title\": r[\"title\"], \"authors\": r[\"authors\"],\n     \"journal\": r[\"journal\"], \"pubdate\": r[\"pubdate\"]})\nprint(cross[\"verdict\"], cross[\"details\"])\n# verdict ∈ correct / mismatch / partial / unknown\n```\n\n`claimed` / `actual` keys are plain strings/lists — feed them from any\nsource (database, form, LLM extraction). Cross-language CJK↔Latin author\nnames (and, since v3.9.0, CJK↔Latin titles) are skipped honestly, never\ncounted as misses.\n\n## 3. DOI-native verification\n\n```python\nres = vp.resolve_doi(\"10.1038/nature12968\")     # Crossref resolve (+retraction)\nentry, audit = vp.verify_doi_entry(\n    \"10.1038/nature12968\", \"cli\", resolution=res,\n    claimed={\"claimed_title\": \"...\", \"claimed_year\": \"2014\"})\nprint(entry[\"verdict\"], entry.get(\"retracted\"))\n```\n\n## 4. Context parsing (extract claims from free text)\n\n```python\nclaimed = vp.parse_citation_context(\n    \"Gattorno A, Van Dijk M. Classification criteria for autoinflammatory \"\n    \"recurrent fevers. Ann Rheum Dis. 2019. PMID: 31018962\")\nprint(claimed)\n# GB/T 7714 Chinese references are recognized since v3.9.0 (title via the\n# [J]/[M] type marker; the title comparison itself is cross-language-skipped)\n```\n\n## 5. Suggestions for broken citations\n\n```python\ncands = vp.suggest_correct_pmid({\"claimed_title\": \"Juvenile idiopathic arthritis\",\n                                 \"claimed_journal\": \"Pediatric Rheumatology\"})\nfor c in cands[:3]:\n    print(c[\"pmid\"], c[\"title\"][:60])   # abbreviated journal names can return [] — use the full NLM name\n```\n\n## 6. Plain-text reference lists & formatted output (v4.0.0)\n\n```python\nentries, issues = vp.parse_plaintext_references(\"refs_list.txt\")\nfor e in entries:\n    print(e[\"key\"], e[\"pmid\"] or e[\"doi\"] or e[\"arxiv_id\"] or\n          \"(title search)\", e[\"claimed_title\"][:50])\n\n# One markerless entry, verified by title search (honest route label):\nres, audit = vp.verify_title_entry(entries[0], \"refs_list.txt\")\nprint(res[\"verdict\"], res.get(\"resolved_by\"))   # correct title_search\n\n# Render verified results as a numbered reference list:\nprint(vp.generate_reference_list(results, \"gbt\"))       # also: vancouver / apa / ama\n```\n\n`parse_citation_context` accepts markerless Latin references since v4.0.0\n(returns `claimed_source_format: \"plaintext\"`); CJK entries without a\nPMID marker stay with the weak-parsing boundary.\n\n## Notes\n\n- Network failures surface as `network_error=True` entries / honest\n  `unknown` verdicts — never as \"not found\", never cached. Run\n  `python3 scripts/verify_pmids.py --check-net` to diagnose connectivity.\n- Retraction status follows the registry at call time; for final\n  pre-submission passes run without the cache (`--no-cache` on the CLI).\n\nFile v4.0.0:skill-card.md\n\n## Description:\n\nChecks cited PMIDs, DOIs, and arXiv IDs against academic registries, compares publication details, and produces citation audit reports or formatted reference lists.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[docsor1212](https://clawhub.ai/user/docsor1212)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nResearchers, authors, and developers use this skill to audit reference lists and AI-generated citations before publication, identify mismatched or unverified records, and prepare formatted references for review.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Citation identifiers, titles, and optional contact email or API-key parameters are sent to academic registry services during verification.\n\nMitigation: Use only with material approved for external lookup; use a role email for --mailto and choose formatting without verification when lookups are not intended.\n\nRisk: Cached lookup results may persist on the local machine or be stale for a final manuscript check.\n\nMitigation: Use --no-cache for sensitive manuscripts and final pre-submission verification.\n\n## Reference(s):\n\n- [ClawHub skill listing](https://clawhub.ai/docsor1212/skills/pubmed-verifier)\n- [API examples](artifact/references/api_examples.md)\n- [Python API reference](artifact/references/python_api.md)\n\n## Skill Output:\n\n**Output Type(s):** [Text, Files, Citation verification guidance]\n\n**Output Format:** [Text, HTML, JSON, CSV, BibTeX, RIS, or formatted reference lists]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Five-state citation verdicts; unresolved and partial matches require human review.]\n\n## Skill Version(s):\n\n4.0.0 (source: skill frontmatter and server-resolved release)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v4.0.0:skillhub-meta.json\n\n{\n  \"name\": \"pubmed-verifier\",\n  \"version\": \"4.0.0\",\n  \"description\": \"Reference checker for AI-fabricated citations: batch-verify PMIDs against PubMed and catch the hallucination existence checks miss — a REAL PMID pointing to a DIFFERENT paper. Five-state citation verification (correct / mismatch / partial / invalid / unknown), citation-context parsing, dual fuzzy matching, Crossref DOI cross-check, retraction detection (capped at partial), correct-PMID suggestion, arXiv ID verification, plain-text reference list parsing, formatted reference lists (GB/T 7714 / Vancouver / APA / AMA), SQLite cache, CSV/JSON claims, HTML/JSON/text reports. Dual data sources with automatic Europe PMC fallback. Network failures are honestly reported as unverified, never as \\\"not found\\\". Zero dependencies, runs fully local. Triggers: verify PMIDs, check citations, validate references, citation audit, reference check, PMID check, batch verify references, AI hallucination detection, verify DOI, DOI check, validate citations, citation formatting, reference formatter, GB/T 7714.\",\n  \"author\": \"docsor1212\",\n  \"tags\": [\n    \"pubmed\",\n    \"pmid\",\n    \"文献验证\",\n    \"引用核查\",\n    \"AI幻觉\",\n    \"学术写作\",\n    \"医学文献\",\n    \"citation\",\n    \"verification\",\n    \"academic\"\n  ],\n  \"license\": \"MIT-0\",\n  \"homepage\": \"https://github.com/docsor1212/pubmed-verifier\"\n}\n\nFile v4.0.0:examples/refs_list.sample.txt\n\n[1] Gattorno M, Sanni K. Development and validation of the autoinflammatory diseases effect scale. Annals of the Rheumatic Diseases. 2019;78(5):618-622. PMID: 31018962\n[2] Zaripova LN, Midgley A, Christmas SE. Autoinflammatory diseases in practice. Pediatric Rheumatology. 2021;19(1):75. https://doi.org/10.1186/s12969-021-00566-6\n[3]Obokata H. Stimulus-triggered fate conversion of somatic cells into pluripotency. Nature. 2014;511(7510):540-545.\n[4] And a deliberately broken line without any usable metadata.\n\nArchive v3.9.0: 10 files, 78638 bytes\n\nFiles: examples/claims.sample.csv (469b), examples/README.txt (1480b), README.md (5363b), references/api_examples.md (4849b), references/python_api.md (2798b), scripts/verify_pmids.py (182587b), skill-card.md (2148b), SKILL.md (36083b), skillhub-meta.json (1386b), _meta.json (134b)\n\nFile v3.9.0:SKILL.md\n\n---\nname: pubmed-verifier\nlicense: MIT-0\nversion: 3.9.0\ndescription: >-\n  Reference checker for AI-fabricated citations: batch-verify PMIDs against\n  PubMed and catch the hallucination existence checks miss — a REAL PMID\n  pointing to a DIFFERENT paper. Five-state citation verification (correct /\n  mismatch / partial / invalid / unknown), citation-context parsing, dual\n  fuzzy matching, Crossref DOI cross-check, retraction detection (capped at\n  partial), correct-PMID suggestion, arXiv ID verification, SQLite cache,\n  CSV/JSON claims, HTML/JSON/text reports. Dual data sources with automatic\n  Europe PMC fallback, optional NCBI API key, Crossref polite pool,\n  Retry-After backoff, UA rotation, host circuit breaker. Network failures\n  are honestly reported as unverified, never as \"not found\". Zero\n  dependencies, runs fully local. Triggers: verify PMIDs, check citations,\n  validate references, citation audit, reference check, PMID check, audit\n  references, batch verify references, AI hallucination detection, verify\n  DOI, DOI check, validate citations, PubMed citation verifier, BibTeX audit.\n---\n\n# PubMed Citation Verifier v3.9.0\n\nBatch verification of PMID citations via the PubMed E-utilities API. Not just\n\"does this PMID exist\" — **does this PMID point to the paper you claim?**\nZero dependencies, pure standard library, fully local.\n\n## When to use this skill\n\nInvoke it whenever citation truth matters:\n\n- \"Verify / check these PMIDs / this reference list\" (before submission or release)\n- Auditing citations in AI-generated text (fabricated or mismatched references)\n- Spot-checking a systematic review's bibliography\n- \"Does PMID 12345678 really say X?\" — point-of-doubt verification\n- Batch QA of a knowledge base's PMID/DOI citations\n- Automatic invocation: trigger this skill whenever the user asks to\n  \"verify/check/audit\" citations, PMIDs, DOIs, arXiv preprints or\n  references — no flags needed for a first pass; add `--claims-file`\n  when they supply the expected titles.\n\n**Trigger priority & tool choice** — explicit \"verify / check / audit\ncitations, references, PMIDs, DOIs\" requests invoke this skill first.\ncite-holmes is for deep research with machine-verified citations; when a\nrequest mixes research and verification, run the research first, then\nthis tool for the final reference audit.\n\n## The five-state verdict\n\n| Verdict | Meaning |\n|---------|---------|\n| ✅ Correct | PMID exists AND matches the claimed paper |\n| ⚠️ Mismatch | PMID exists but points to a **different** paper (the most common AI hallucination!) |\n| 🔶 Partial | Some metadata matches (e.g. author+journal but title differs) |\n| ❌ Invalid | PMID does not exist in PubMed |\n| ❓ Unknown | Not enough claimed metadata to cross-check — or both data sources unreachable (never misreported as invalid) |\n\n**Why existence checks are not enough:** a large share of fabricated\ncitations use REAL PMIDs that point to a different paper from the same\nyear/journal/field — in one of our own audits, 4 of 5 \"valid\" PMIDs were\nwrong this way. A binary exists/not-exists check misses them all.\n\n## Quick start\n\n```bash\n# Scan a project directory for PMIDs (parses citation context automatically)\npython3 scripts/verify_pmids.py --source /path/to/project --output report.html\n\n# Verify specific PMIDs\npython3 scripts/verify_pmids.py --pmids 31018962,22213727\n\n# Mismatch demo: PMID 34078778 is actually a dental-materials paper, so the\n# JIA claims below will NOT match it — expect ⚠️ mismatch verdicts\npython3 scripts/verify_pmids.py --claims '[{\"pmid\":\"34078778\",\"title\":\"JIA pathogenesis\",\"authors\":[\"Zaripova\"],\"journal\":\"Pediatr Rheumatol Online J\",\"year\":\"2021\"}]' --output report.html\n\n# Claims from a CSV file + suggest correct PMIDs for mismatches\npython3 scripts/verify_pmids.py --claims-file claims.csv --suggest --output report.html\n\n# Crossref DOI cross-verification + audit working-paper + BibTeX + full pipeline\npython3 scripts/verify_pmids.py --source /path/to/files --verify-doi --suggest --output report.html --export-audit audit.json --export-bibtex refs.bib\n\n# Verify DOIs directly (no PMIDs) + delta audit vs a previous run\npython3 scripts/verify_pmids.py --dois \"10.1038/nature12968,10.4012/dmj.2020-408\" --workers 4 --export-audit audit.json --export-csv table.csv\npython3 scripts/verify_pmids.py --source /path/to/project --diff audit.json --output report.html\n\n# Verify arXiv IDs (preprints) — mixed audits supported\npython3 scripts/verify_pmids.py --arxivs \"2401.12345,cs/0211004\" --no-cache\n\n# Pre-flight: are the four data sources reachable right now?\npython3 scripts/verify_pmids.py --check-net\n\n# Audit a BibTeX or RIS bibliography file directly (PMID > DOI > arXiv routing)\npython3 scripts/verify_pmids.py --bibliography refs.bib --no-cache\npython3 scripts/verify_pmids.py --bibliography refs.ris --no-cache --export-ris verified.ris\n\n# Institutional niceties (recommended): NCBI API key + contact email\npython3 scripts/verify_pmids.py --source . --verify-doi --ncbi-api-key $NCBI_API_KEY --mailto you@lab.org\n```\n\n## FAQ & common mistakes\n\n**Top 10 things NOT to do** (each is detailed below or in Anti-patterns):\n\n| # | Don't | Do instead |\n|---|-------|------------|\n| 1 | Treat `--pmids` existence output as \"verified\" | Feed `--claims-file` with titles for real verification |\n| 2 | Submit claims without `title` | Always include titles — the verdict caps at partial without one |\n| 3 | Trust cached verdicts on publication day | Final check with `--no-cache` |\n| 4 | Read \"not found\" as \"fabricated\" for auto-extracted DOIs | Check doi.org / arxiv.org by hand first |\n| 5 | Treat the leading `'` in CSV cells as corruption | It is the formula-injection guard — strip after import |\n| 6 | Read the READY line as a quality score | It means \"no problems among the checks that ran\" |\n| 7 | Pass `--source` together with `--pmids` | `--source` is ignored entirely when `--pmids` is given |\n| 8 | Deep-verify (`--verify-doi` / `--suggest`) a thousand-entry sweep | Sweep first, deep-verify the flagged subset |\n| 9 | Expect author matching across CJK↔Latin names | They are skipped honestly (`author_check: skipped`) |\n| 10 | Ship a reference list without the audit trail | `--export-audit` writes a replayable working paper |\n\n**Large batch (hundreds of PMIDs) is slow — how to speed it up?**\nMetadata-only verification queries in batches of 50 with 0.4 s spacing\n(0.12 s with `--ncbi-api-key`); cached re-runs are ~5 s. `--verify-doi` adds\none Crossref call *per citation* and `--suggest` adds one search *per\nmismatch* — skip them for bulk sweeps, run them on the flagged subset.\n\n**When must I use `--claims-file` instead of scanning?**\nContext parsing is heuristic (v3.3.0 guards common abbreviations, exotic formatting can still mis-split).\nFor precise verification — or DOIs in claims (splice detection needs `doi`)\n— feed structured JSON/CSV claims.\n\n**My citation text is in Chinese — the scan parses little?**\nGB/T 7714-style references (……标题[J]. 刊名, 年… PMID: xxx) parse since\nv3.9.0 — title/journal comparisons against Latin registries are skipped\ncross-language, so verdicts rest on year/DOI evidence. Free Chinese prose\nwithout a PMID marker still yields little — feed `--claims-file` for full\nverdicts regardless of language.\n\n**Slow or unstable network (China)?**\nStandard `HTTPS_PROXY`/`HTTP_PROXY` env vars are honored natively; raise\n`--timeout`; `--meta-source europepmc` routes via Europe PMC when NCBI is\nunreachable (per-entry `meta_source` shows which was used); cached results\nare reused for 30 days.\n\n**❓ unknown vs ❌ invalid?**\n`unknown` (exit 2) = \"could not verify, sources unreachable\" — retry later;\n`invalid` (exit 1) = \"verified not-found\". Network failures are never\nreported as not-found and never cached.\n\n**What does RETRACTED mean in a report?**\nThe registry itself lists the paper's publication type as \"Retracted\nPublication\" (checked for every citation since v2.7.0 — no DOI or flags\nneeded), and/or Crossref records a retraction. The verdict is capped at\npartial and a human review note is attached — citing it would propagate\nwithdrawn science. A retraction *notice* is never flagged; papers under\n*Expression of Concern* (an editorial note, not a retraction) are not\nflagged either. Retraction status reflects the registry at cache time — for\na final pre-submission check, run with `--no-cache`.\n\n**Mismatch reported but the title looks similar?**\nCheck `details` for which field diverged; thresholds are strict on purpose.\nFeed the full citation via `--claims-file` for a precise verdict.\n\n**Can I audit my .bib or .ris file directly?**\nYes — `--bibliography refs.bib` (BibTeX) and `--bibliography refs.ris`\n(RIS/Zotero/EndNote/Mendeley) route each entry by PMID > DOI > arXiv\nand cross-check the claimed metadata. `--lint-claims` validates either\nformat offline first. The round trip works both ways: `--export-bibtex`\nand `--export-ris` output can be fed back after edits.\n\n**My claims file seems to lose rows / verdicts look weaker than expected?**\nLint it offline first: `python3 scripts/verify_pmids.py --lint-claims\nclaims.csv` reports unusable rows, ID shape errors, unknown columns\n(typo'd headers like \"titel\"), missing titles and DOI prefix problems —\nno network, exit 1 on errors.\n\n**Can I verify a DOI with claimed metadata (full verdict)?**\nYes — since v3.4.0 a claims row keyed by `doi` (with `title`, optionally\n`authors`/`journal`/`year`) gets the same cross-check as PMID claims:\ncorrect / mismatch / partial against the registered metadata. A DOI row\nwithout claims stays `unknown` (existence only).\n\n**Why did my arXiv citation drop from correct to partial?**\nYour claims row paired an `arxiv_id` with a `doi`, and the DOI does not\nmatch the version-of-record DOI registered on that arXiv entry — a typo,\nor a DOI from a different paper. The registered DOI is in `details`;\nfix the claim or drop the `doi` cell.\n\n## Anti-patterns — things done WRONG\n\nEach entry: the mistake → why it fails → the right way.\n\n1. **Treating `--pmids` output as \"fully verified\"** — existence-only.\n   → Wrong: \"all 5 PMIDs exist, so the citations are correct.\"\n   → Right: existence-checked only; feed `--claims-file` with titles for\n   real verification (the READY line says so explicitly).\n2. **Claims without `title`** — author/journal/year alone can never reach\n   `correct`; the report caps at `partial`. → Always include titles.\n3. **Trusting a cached verdict right after publication day** — a brand-new\n   PMID may have been cached as not-found by an earlier run, and retraction\n   status is as of cache time. → Final pre-submission check: `--no-cache`.\n4. **Assuming \"not found\" always means fabricated** — auto-extracted DOIs\n   that 404 stay *suspects* (DataCite DOIs don't live in Crossref); arXiv\n   IDs removed by moderators also return empty. → Check doi.org / arxiv.org\n   by hand before accusing.\n5. **Copying the leading `'` from CSV cells** — that apostrophe is the\n   formula-injection guard, not data corruption. → Strip it after import.\n6. **Reading the READY line as a quality score** — it only means \"no\n   problems found among the checks that ran\", not \"this paper is good\".\n\n## Boundaries — declared limits\n\nWhat this tool can NOT do, consolidated in one place:\n\n- **Splice/mismatch signals report disagreement, never pick a side** —\n  when claim and registry disagree, a human reads the evidence line.\n- **Cross-language authors are skipped, not failed** — CJK↔Latin author\n  names are never compared (`author_check: skipped`); the verdict rests\n  on title/journal/year alone.\n- **Context parsing is heuristic** — v3.3.0 keeps common abbreviations\n  (U.S., e.g., vs., St., Vol., No.) from splitting a title, but exotic\n  formatting can still mis-split; for exact metadata use `--claims-file`.\n- **DataCite/repository DOIs are not in Crossref** — an auto-extracted\n  or bibliography-sourced DOI missing from Crossref stays a *suspect*\n  (`.bib` files commonly hold DataCite/repository DOIs — check doi.org\n  by hand); only `--dois`/claims DOIs count a Crossref 404 as invalid.\n- **arXiv moderator removals also return \"not found\"** — the invalid\n  verdict carries that caveat in its details.\n- **Retraction status is as-of-cache-time** — final pre-submission\n  checks should run with `--no-cache`.\n- **Single-letter initials never match** — \"Smith J\" vs \"Smith John\" is\n  not counted as a miss.\n- **unknown ≠ invalid** — unreachable sources yield exit 2 and\n  `unknown`; network failures are never reported as \"not found\" and\n  never cached.\n- **arXiv pacing is deliberate** — the official API asks for ≥3 s\n  between calls; large arXiv batches are slow by design (progress + ETA\n  on stderr). Entries that register a version-of-record DOI add one\n  Europe PMC lookup each for the PMID link.\n- **DOI claims compare against the registry that actually answered** —\n  the linked PubMed record when the DOI resolves to one, Crossref\n  otherwise (Crossref author fields are sparser, so the author mark is\n  more often \"—\"); a DOI row without a claimed title stays `unknown`,\n  not partial; an explicitly user-provided DOI (`--dois` or claims) that\n  is missing from Crossref counts as invalid.\n\n## Claims reference format\n\n`--claims-file` accepts JSON (an array of objects) or CSV. Recognized\ncolumns: `pmid`, `title`, `authors` (semicolon/pipe-separated),\n`journal`, `year`, `doi`, `arxiv_id`. A row needs one of `pmid`,\n`arxiv_id` or `doi`; a missing `title` caps the verdict at partial.\n\n```csv\npmid,title,authors,journal,year,doi,arxiv_id\n31018962,Candidate criteria for diagnosis of familial...,Gattorno,Ann Rheum Dis,2019,10.1136/annrheumdis-2019-215048,\n,Attention Is All You Need,Vaswani,NeurIPS,2017,,1706.03762\n,City size and the spreading of COVID-19 in Brazil,Silva Junior;Other,PLOS ONE,2020,10.1371/journal.pone.0239699,\n```\n\nValidate any file offline first: `--lint-claims file.csv` reports\nunusable rows, ID shape errors, unknown columns, duplicates and missing\ntitles (no network). Lint wins when combined with verification flags —\nonly the lint runs.\n\n## Best practices & tuning\n\n- **Speed up large batches** — request an NCBI API key (see\n  https://ncbiinsights.ncbi.nlm.nih.gov/api-keys/): batches of 50 IDs run\n  at 0.12 s spacing instead of 0.4 s; cached re-runs take seconds.\n- **Parallel DOI verification** — `--workers` (default 4, cap 8) applies to\n  Crossref resolution and Europe PMC linking; arXiv stays serial by\n  official etiquette (≥3 s between calls).\n- **Two-phase workflow** — sweep with metadata-only verification first\n  (no `--verify-doi`, no `--suggest`), then deep-verify only the flagged\n  subset; each deep flag adds one API call per citation.\n- **Round-trip bibliographies** — `--export-bibtex` / `--export-ris`\n  write verified entries; after edits, `--bibliography refs.bib` /\n  `refs.ris` re-audits the file (the PMID re-links the full record).\n- **Claims over context parsing** — whenever you know the expected titles,\n  feed `--claims-file`: it enables the full verdict ladder and the DOI /\n  arXiv pairing checks. Validate the file offline first:\n  `python3 scripts/verify_pmids.py --lint-claims claims.csv` reports ID\n  shape errors, missing titles, unknown columns and duplicates without any\n  network access.\n- **Flaky networks** — raise `--timeout`; HTTPS_PROXY/HTTP_PROXY are\n  honored natively; unreachable NCBI falls back to Europe PMC\n  automatically (`meta_source` shows which answered).\n- **Cache policy** — results cache 30 days, negative entries 3 days;\n  `--cache-days` to tune; `--no-cache` for the final pre-submission pass.\n- **Scale expectations** — metadata-only throughput is API-bound\n  (~1–2 min per 1000 PMIDs with an API key); DOI resolution adds one\n  Crossref call per DOI. One deliberate trade-off: the verifier is a\n  single stdlib-only file — copy `scripts/verify_pmids.py` anywhere with\n  Python 3.8+ and it runs, no pip, no venv (that portability is why the\n  code is not split into modules).\n\n## v2.2.0 — network hardening\n\n| Feature | Flag | Effect |\n|---------|------|--------|\n| NCBI API key | `--ncbi-api-key` / env `NCBI_API_KEY` | Rate ceiling 3→10 req/s, batch interval 0.4s→0.12s (~3x faster) |\n| Europe PMC fallback | `--meta-source auto\\|ncbi\\|europepmc` | NCBI batch failure automatically retries via Europe PMC (free, no key); per-entry origin in JSON (`meta_source`) |\n| Crossref polite pool | `--mailto` / env `PUBMED_VERIFIER_MAILTO` | `?mailto=` on Crossref + tool/email params on NCBI — more generous limits |\n| Retry-After backoff | automatic | 429 responses honored (clamped 1–5 s) instead of failing |\n| UA rotation | automatic | 403/406 retried with a browser User-Agent |\n| Host circuit breaker | automatic | After 2 call-level transport failures a host is skipped with an actionable message; success resets; HTTP errors never trip it |\n| Honest unknown | automatic | Network failures report as ❓ unknown + exit code 2, never as \"PMID not found\", and are never cached |\n\n**Exit codes:** `0` clean · `1` problems found (invalid / mismatch / retracted /\nDOI-splice, incl. arXiv claimed-DOI pairing mismatches) · `2` could not verify\n(data sources unreachable) — automation can tell \"all good\" from \"no answer\".\n\n## v2.3.0 — retraction detection\n\nWith `--verify-doi`, each cited DOI is also checked against Crossref's\nwithdrawal records (`updated-by`). A paper Crossref lists as RETRACTED is:\n\n- flagged in JSON (`retracted: true` + `retraction_note`) and in reports,\n- **capped at 🔶 partial** even when every metadata field matches — citing a\n  retracted paper is never \"correct\"; the report says *human review required*.\n\nCorrections and other update types do not trigger the cap. Crossref outages\nnever flag anything (a missing check is not a retraction).\n\n## v2.4.0 — DOI↔PMID cross-check & journal abbreviations\n\n- **DOI splice detection**: add a `doi` field to your claims (JSON or CSV).\n  The claimed DOI is compared with the DOI registered for that PMID — a\n  mismatch is a splice/fabrication signal (a real DOI attached to the wrong\n  paper): flagged in JSON (`doi_splice_suspect`), verdict capped at 🔶\n  partial, counted in exit 1. Uses the PubMed record only — no extra API call.\n- **Journal abbreviation equivalence**: journal matching now understands\n  NLM-style abbreviations in both directions — \"N Engl J Med\" matches \"New\n  England Journal of Medicine\", \"Pediatr Rheumatol\" matches \"Pediatric\n  Rheumatology\" (in-order word prefixes, function words skipped). No more\n  false \"journal differs\" for abbreviated citations.\n\nKnown limits: highly ambiguous abbreviations can over-match at the\njournal-only level (\"J Immunol\" ~ \"Journal of Immunology Research\") — the\ntitle remains the decisive field. A DOI-splice flag can also appear on an\notherwise-unverifiable citation (the DOI mismatch is an independent fact).\n\n## v2.5.0 — author-name verification\n\n- **Initials never match**: single-letter tokens (\"A.\", \"L.\") on\n\nArchive v3.8.0: 9 files, 73959 bytes\n\nFiles: examples/claims.sample.csv (469b), examples/README.txt (1480b), README.md (5363b), references/api_examples.md (4849b), scripts/verify_pmids.py (174179b), skill-card.md (2235b), SKILL.md (34819b), skillhub-meta.json (1386b), _meta.json (134b)\n\nArchive v3.7.0: 9 files, 71573 bytes\n\nFiles: examples/claims.sample.csv (474b), examples/README.txt (1398b), README.md (5363b), references/api_examples.md (4849b), scripts/verify_pmids.py (170509b), skill-card.md (2124b), SKILL.md (33055b), skillhub-meta.json (1386b), _meta.json (134b)\n\nArchive v3.6.0: 9 files, 68941 bytes\n\nFiles: examples/claims.sample.csv (474b), examples/README.txt (1252b), README.md (4970b), references/api_examples.md (4849b), scripts/verify_pmids.py (161613b), skill-card.md (2271b), SKILL.md (32315b), skillhub-meta.json (1386b), _meta.json (134b)\n\nArchive v3.5.0: 9 files, 65500 bytes\n\nFiles: examples/claims.sample.csv (474b), examples/README.txt (1108b), README.md (4781b), references/api_examples.md (4849b), scripts/verify_pmids.py (153035b), skill-card.md (1970b), SKILL.md (30831b), skillhub-meta.json (1372b), _meta.json (134b)\n\nArchive v3.4.0: 9 files, 61608 bytes\n\nFiles: examples/claims.sample.csv (474b), examples/README.txt (1004b), README.md (4516b), references/api_examples.md (4814b), scripts/verify_pmids.py (145498b), skill-card.md (2050b), SKILL.md (25664b), skillhub-meta.json (1372b), _meta.json (134b)\n\nArchive v3.3.0: 9 files, 59609 bytes\n\nFiles: examples/claims.sample.csv (360b), examples/README.txt (657b), README.md (4516b), references/api_examples.md (4814b), scripts/verify_pmids.py (139202b), skill-card.md (1828b), SKILL.md (23848b), skillhub-meta.json (1372b), _meta.json (134b)\n\nArchive v3.2.0: 9 files, 55852 bytes\n\nFiles: examples/claims.sample.csv (360b), examples/README.txt (397b), README.md (4863b), references/api_examples.md (4006b), scripts/verify_pmids.py (134048b), skill-card.md (2051b), SKILL.md (20783b), skillhub-meta.json (1372b), _meta.json (134b)","readmeExcerpt":"Skill: pubmed-verifier Owner: docsor1212 Summary: Reference checker for AI-fabricated citations: batch-verify PMIDs against PubMed and catch the hallucination existence checks miss — a REAL PMID pointing to a DIFFERENT paper. Five-state citation verification (correct / mismatch / partial / invalid / unknown), citation-context parsing, Crossref DOI cross-check, retraction detection (capped at partial), correct-PMID su","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"# Scan a project directory for PMIDs (parses citation context automatically)\npython3 scripts/verify_pmids.py --source /path/to/project --output report.html\n\n# Verify specific PMIDs\npython3 scripts/verify_pmids.py --pmids 31018962,22213727\n\n# Mismatch demo: PMID 34078778 is actually a dental-materials paper, so the\n# JIA claims below will NOT match it — expect ⚠️ mismatch verdicts\npython3 scripts/verify_pmids.py --claims '[{\"pmid\":\"34078778\",\"title\":\"JIA pathogenesis\",\"authors\":[\"Zaripova\"],\"journal\":\"Pediatr Rheumatol Online J\",\"year\":\"2021\"}]' --output report.html\n\n# Claims from a CSV file + suggest correct PMIDs for mismatches\npython3 scripts/verify_pmids.py --claims-file claims.csv --suggest --output report.html\n\n# Crossref DOI cross-verification + audit working-paper + BibTeX + full pipeline\npython3 scripts/verify_pmids.py --source /path/to/files --verify-doi --suggest --output report.html --export-audit audit.json --export-bibtex refs.bib\n\n# Verify DOIs directly (no PMIDs) + delta audit vs a previous run\npython3 scripts/verify_pmids.py --dois \"10.1038/nature12968,10.4012/dmj.2020-408\" --workers 4 --export-audit audit.json --export-csv table.csv\npython3 scripts/verify_pmids.py --source /path/to/project --diff audit.json --output report.html\n\n# Verify arXiv IDs (preprints) — mixed audits supported\npython3 scripts/verify_pmids.py --arxivs \"2401.12345,cs/0211004\" --no-cache\n\n# Pre-flight: are the five data sources reachable right now?\npython3 scripts/verify_pmids.py --check-net\n\n# Audit a BibTeX or RIS bibliography file directly (PMID > DOI > arXiv routing)\npython3 scripts/verify_pmids.py --bibliography refs.bib --no-cache\npython3 scripts/verify_pmids.py --bibliography refs.ris --no-cache --export-ris verified.ris\n\n# Verify a reference list copied from a paper draft (inline IDs route\n# exactly; title-only entries resolve via PubMed title search)\npython3 scripts/verify_pmids.py --parse-text references.txt --no-cache\n\n# Format the verified entries as a ready-to-paste"},{"language":"text","snippet":"Readiness: NOT SUBMISSION-READY — 1 invalid, 1 mismatched\nResults: 1/3 correct, 1 mismatch, 1 invalid, 0 partial, 0 unknown\n============================================================\n\n✅ PMID 31018962 (claims) [correct]\n   Actual: Classification criteria for autoinflammatory recurrent fevers\n   Claimed: Classification criteria for autoinflammatory recurrent fevers\n   Evidence: title ✓ · author ✓ · journal ✓ · year ✓\n\n⚠️ PMID 34078778 (claims) [mismatch]\n   Actual: Effect of CAD/CAM materials on the marginal fit of crowds\n   Claimed: JIA pathogenesis and treatment\n   → Suggest: PMID 34425842 - Juvenile idiopathic arthritis: from aetio...\n   Evidence: title ✗ · author ✗ · journal ✗ · year —\n\n❌ PMID 99999999 (cli) [invalid]\n   Error: PMID not found in API response"},{"language":"csv","snippet":"pmid,title,authors,journal,year,doi,arxiv_id\n31018962,Candidate criteria for diagnosis of familial...,Gattorno,Ann Rheum Dis,2019,10.1136/annrheumdis-2019-215048,\n,Attention Is All You Need,Vaswani,NeurIPS,2017,,1706.03762\n,City size and the spreading of COVID-19 in Brazil,Silva Junior;Other,PLOS ONE,2020,10.1371/journal.pone.0239699,"},{"language":"bash","snippet":"python3 scripts/verify_pmids.py --claims '[{\"arxiv_id\":\"2005.13892\",\n  \"title\":\"City size and the spreading of COVID-19 in Brazil\",\n  \"doi\":\"10.1371/journal.pone.0239699\"}]'"},{"language":"bash","snippet":"python3 scripts/verify_pmids.py --claims '[{\"doi\":\"10.1371/journal.pone.0239699\",\n  \"title\":\"City size and the spreading of COVID-19 in Brazil\",\n  \"journal\":\"PLoS ONE\",\"year\":\"2020\"}]'"},{"language":"bash","snippet":"# Install\nopenclaw skills install docsor1212/pubmed-verifier\n\n# 中文场景速览\npython3 scripts/verify_pmids.py --parse-text references.txt --no-cache          # 核验整张参考文献表\npython3 scripts/verify_pmids.py --parse-text references.txt --format-references out.txt --citation-style gbt\npython3 scripts/verify_pmids.py --pmids 31018962,22213727                       # 批量验证 PMID\npython3 scripts/verify_pmids.py --check-net                                     # 网络预检（五源）\n\n# Verify all PMIDs in a project\npython3 scripts/verify_pmids.py --source /path/to/project --output report.html\n\n# Verify specific PMIDs\npython3 scripts/verify_pmids.py --pmids 31018962,22213727,999999999\n\n# Verify a reference list copied from a paper draft\npython3 scripts/verify_pmids.py --parse-text references.txt --no-cache\n\n# Format verified entries as a citation list (gbt | vancouver | apa | ama)\npython3 scripts/verify_pmids.py --parse-text references.txt --format-references refs_gbt.txt --citation-style gbt\n\n# Institutional mode: API key + polite pool\npython3 scripts/verify_pmids.py --source ./papers --verify-doi --ncbi-api-key $NCBI_API_KEY --mailto you@lab.org"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: pubmed-verifier\nlicense: MIT-0\nversion: 4.1.0\ndescription: >-\n  Reference checker for AI-fabricated citations: batch-verify PMIDs against\n  PubMed and catch the hallucination existence checks miss — a REAL PMID\n  pointing to a DIFFERENT paper. Five-state citation verification (correct /\n  mismatch / partial / invalid / unknown), citation-context parsing, Crossref\n  DOI cross-check, retraction detection (capped at partial), correct-PMID\n  suggestion, arXiv ID verification, plain-text reference list parsing,\n  formatted reference lists (GB/T 7714 / Vancouver / APA / AMA), OpenAlex DOI\n  fallback for non-Crossref registries, Chinese-reference title search,\n  CSV/JSON claims, HTML/JSON/text reports. Four data sources: NCBI, Europe PMC\n  fallback, Crossref, OpenAlex. Network failures are honestly reported as\n  unverified, never as \"not found\". Zero dependencies, runs fully local.\n  Triggers: verify PMIDs, check citations, citation audit, reference check,\n  PMID check, batch verify references, AI hallucination detection, verify DOI,\n  DOI check, citation formatting, reference formatter, GB/T 7714.\n---\n\n# PubMed Citation Verifier v4.1.0\n\nBatch verification of PMID citations via the PubMed E-utilities API. Not just\n\"does this PMID exist\" — **does this PMID point to the paper you claim?**\nZero dependencies, pure standard library, fully local.\n\n## When to use this skill\n\nInvoke it whenever citation truth matters:\n\n- \"Verify / check these PMIDs / this reference list\" (before submission or release)\n- Auditing citations in AI-generated text (fabricated or mismatched references)\n- Spot-checking a systematic review's bibliography\n- \"Does PMID 12345678 really say X?\" — point-of-doubt verification\n- Batch QA of a knowledge base's PMID/DOI citations\n- Automatic invocation: trigger this skill whenever the user asks to\n  \"verify/check/audit\" citations, PMIDs, DOIs, arXiv preprints or\n  references — no flags needed for a first pass; add `--claims-file`\n  when they supply the expected titles.\n\n**Trigger priority & tool choice** — explicit \"verify / check / audit\ncitations, references, PMIDs, DOIs\" requests invoke this skill first.\ncite-holmes is for deep research with machine-verified citations; when a\nrequest mixes research and verification, run the research first, then\nthis tool for the final reference audit.\n\n## The five-state verdict\n\n| Verdict | Meaning |\n|---------|---------|\n| ✅ Correct | PMID exists AND matches the claimed paper |\n| ⚠️ Mismatch | PMID exists but points to a **different** paper (the most common AI hallucination!) |\n| 🔶 Partial | Some metadata matches (e.g. author+journal but title differs) |\n| ❌ Invalid | PMID does not exist in PubMed |\n| ❓ Unknown | Not enough claimed metadata to cross-check — or both data sources unreachable (never misreported as invalid) |\n\n**Why existence checks are not enough:** a large share of fabricated\ncitations use REAL PMIDs that point to a different paper from the same\nyear/journal/field — in one of our o"},{"path":"examples/README.txt","content":"claims.sample.csv — reference format for --claims-file\n\n- pmid column: PubMed pipeline (full five-state verification)\n- rows with only arxiv_id (+title): arXiv pipeline (official API check)\n- doi column: DOI-splice cross-check against the PubMed-registered DOI\n  (PMID rows) / version-of-record pairing check (arXiv rows, v3.3.0) /\n  full five-state verdicts (doi-only rows, v3.4.0)\n- title: required for verdicts beyond existence-check\n- expected verdicts (as of v3.4.0): the 31018962 row carries a\n  deliberately inexact title → partial (title mismatch, everything else\n  matches — demonstrates the partial rung, not a data error); the\n  1706.03762 row → correct; the doi-only 10.1371 row → correct (its\n  second author \"Other\" is a synthetic placeholder; v3.8.0's\n  mis-attribution cap means a wrong author set would drop it to partial); the 24476887 row\n  (STAP cells) is RETRACTED on purpose → RETRACTED flag with a partial\n  cap — it demonstrates retraction detection, not a data error\n\nRun:  python3 scripts/verify_pmids.py --claims-file examples/claims.sample.csv\n\nLint first (offline, v3.5.0): python3 scripts/verify_pmids.py --lint-claims examples/claims.sample.csv\n\nBibliography (v3.6.0): refs.sample.bib audits a .bib file directly —\npython3 scripts/verify_pmids.py --bibliography examples/refs.sample.bib\n\nRIS (v3.7.0): refs.sample.ris audits Zotero/EndNote/Mendeley exports —\npython3 scripts/verify_pmids.py --bibliography examples/refs.sample.ris"},{"path":"README.md","content":"# PubMed Citation Verifier 🔬\n\n[![GitHub Stars](https://img.shields.io/github/stars/docsor1212/pubmed-verifier?style=social&label=Star)](https://github.com/docsor1212/pubmed-verifier/stargazers)\n\nBatch-verify PMID citations against PubMed API. Built for researchers, medical writers, and evidence-based medicine teams.\n\n## 简介（中文）\n\n**PubMed 文献引用批量验证工具**——把引用列表交给它，逐条核验 PMID/DOI/arXiv ID 是否真实存在、是否指向你声称的那篇论文（AI 幻觉引用最常见的形态就是「真 PMID 指向别的论文」）。五态判定（正确/不匹配/部分匹配/无效/待确认），撤稿检测全覆盖，网络故障如实标注「未判定」、绝不误报「不存在」。零依赖、纯本地运行。\n\n- **中文文献**：GB/T 7714 格式参考文献可解析（`……标题[J]. 刊名, 年… PMID: xxx`）；`--parse-text` 粘贴即核验整张文献表——行内 PMID/DOI 精确路由，中文条目经 OpenAlex 标题检索判定（无编号 ID 也能核）\n- **参考文献格式化**：`--format-references --citation-style gbt`（另支持 Vancouver/APA/AMA），先验证后排版，只排已验真的条目\n- **四数据源**：NCBI PubMed、Europe PMC 自动兜底、Crossref、OpenAlex（DataCite/Zenodo 等非 Crossref 注册的 DOI 也可判定）\n- **投稿就绪判定 + 审计底稿**：每份报告自带证据链，`--export-audit` 输出可复放的 JSON 工作底稿\n\n## Why?\n\nAcademic projects routinely contain hundreds of PMID citations. Manual verification is tedious and error-prone. During our own 225-reference audit, we found 3 invalid PMIDs and 6 cross-domain mismatches — errors that would have undermined the entire project.\n\n## Features\n\n- **Batch verification** — Scan entire project directories, extract all PMIDs, verify against PubMed in one run\n- **Mismatch detection** — Five-state verdicts catch a REAL PMID pointing to a DIFFERENT paper (the most common AI hallucination)\n- **Metadata validation** — Title, authors, journal and date fuzzy-compared against your claims (DOI cross-checked via Crossref with `--verify-doi`)\n- **Dual sources** — NCBI E-utilities primary, Europe PMC automatic fallback when NCBI is unreachable\n- **Network hardening** — Optional NCBI API key (3x faster batches), Crossref polite pool, 429 Retry-After backoff, UA rotation, per-host circuit breaker\n- **Content matching** — Keyword overlap scoring flags potentially irrelevant citations\n- **Three citation types in one audit** — PMIDs, DOIs and arXiv IDs (preprints) verified in a single scan; arXiv IDs checked against the official API (nonexistent = fabrication signal); claims rows keyed by DOI get full verdicts, and preprints surface their published version of record\n- **Claims lint** — `--lint-claims file.csv` validates a claims file offline (ID shapes, missing titles, unknown columns, duplicates) before any verification run\n- **Bibliography audit** — `--bibliography refs.bib` / `refs.ris` verifies BibTeX and RIS files directly (Zotero/EndNote/Mendeley exports; entries route by PMID > DOI > arXiv with full cross-checks); round-trips with `--export-bibtex` / `--export-ris`\n- **Plain-text reference list parsing** — `--parse-text refs.txt` verifies a reference list copied straight out of a manuscript draft: numbered entries (`[1]`, `1.`) split cleanly, inline PMID/DOI/arXiv IDs route exactly, and a title-only entry resolves via PubMed title search (honestly labeled `resolved_by: title_search`, never a mismatch source)\n- **Formatted reference list"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b6sjwend7cwwmhf7mdg1wzx82dfbf\",\n  \"slug\": \"pubmed-verifier\",\n  \"version\": \"4.1.0\",\n  \"publishedAt\": 1791564939147\n}"},{"path":"references/api_examples.md","content":"# PubMed E-utilities API Quick Reference\n\n## esummary — Article Metadata\n\n```bash\n# Single PMID\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=31018962&retmode=json\"\n\n# Batch (comma-separated)\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=31018962,22213727&retmode=json\"\n```\n\nResponse fields: `title`, `authors[].name`, `source` (journal), `pubdate`, `volume`, `pages`, `elocationid` (DOI).\n\nInvalid PMID → `{\"result\": {\"pmid\": {\"error\": \"cannot get document summary\"}}}` (the exact string has varied over time — the verifier only checks for the presence of an `error` key).\n\n## esearch — Find Articles by Query\n\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=Gattorno+classification+autoinflammatory&retmode=json&retmax=5\"\n```\n\nReturns `esearchresult.idlist` → array of PMIDs.\n\n## efetch — Full Abstracts\n\n```bash\n# efetch supports TEXT and XML only — retmode=json is NOT valid for efetch\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=37635643&rettype=abstract&retmode=text\"\n```\n\n## Deep Research Workflow (Systematic PubMed Search)\n\nWhen the user needs comprehensive medical literature research beyond simple PMID verification:\n\n### Step 1: esearch → keyword search → get PMIDs\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=<keywords>&retmax=10&sort=relevance\" | grep -oP '<Id>\\K\\d+'\n```\n\n### Step 2: esummary → metadata summaries (fast, compact)\n```bash\nPMIDS=\"pmid1,pmid2,pmid3\"\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=$PMIDS&retmode=json\" > /tmp/pubmed-results.json\npython3 -c \"\nimport json\ndata = json.load(open('/tmp/pubmed-results.json'))\nfor uid, art in data.get('result', {}).items():\n    if uid == 'uids': continue\n    authors = ', '.join([a.get('name','') for a in art.get('authors',[])[:6]])\n    print(f'PMID {uid}: {art.get(\\\"title\\\",\\\"\\\")}')\n    print(f'  {art.get(\\\"fulljournalname\\\",\\\"\\\")} {art.get(\\\"pubdate\\\",\\\"\\\")};{art.get(\\\"volume\\\",\\\"\\\")}:{art.get(\\\"pages\\\",\\\"\\\")}')\n    print(f'  DOI: {art.get(\\\"elocationid\\\",\\\"\\\")}')\n    print()\n\"\n```\n\n### Step 3: efetch → full abstracts for key papers\n```bash\ncurl -s \"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=$PMIDS&rettype=abstract&retmode=text\"\n```\n\n### Pitfalls\n- **efetch has no JSON mode** — use `retmode=text` (or XML) only\n- **esearch may return unrelated results** — always cross-check with esummary titles\n- **Rate limit**: 3 req/s without API key. Add `sleep 0.5` between batches\n- **Large author lists**: esummary `authors` array can be 50+ — truncate to first 6 for display\n\n## Rate Limits\n\n- Without API key: 3 requests/second\n- With API key (`&api_key=YOUR_KEY`): 10 requests/second\n- API key obtained from NCBI Settings page\n- Etiquette params: `&tool=pubmed-verifier&email=you@lab.org` (added automatically with `--mailto`)\n\n## Europe PMC (fallback source,"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2231,"uniquenessScore":40,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-09T20:21:50.745Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-09T20:21:50.745Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-09T23:50:28.317Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}