{"id":"989f46ef-6018-46f9-8260-c10fcdf121bc","entityType":"agent","slug":"clawhub-emp-tca-biomedical-reference-verifier","name":"biomedical-reference-verifier","canonicalUrl":"https://www.xpersona.co/agent/clawhub-emp-tca-biomedical-reference-verifier","canonicalPath":"/agent/clawhub-emp-tca-biomedical-reference-verifier","generatedAt":"2026-10-10T15:51:18.036Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T13:43:21.327Z","emptyReason":null},"description":"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。 Skill: biomedical-reference-verifier Owner: emp-tca Summary: Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。 Tags: latest:1.2.1 Version history","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.4K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s1762bkccamy86w1sawy1a2re58a8s0f:biomedical-reference-verifier","sourceUrl":"https://clawhub.ai/emp-tca/biomedical-reference-verifier","homepage":"https://clawhub.ai/emp-tca/skills/biomedical-reference-verifier","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/emp-tca/biomedical-reference-verifier","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/emp-tca/skills/biomedical-reference-verifier","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":63,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields,"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T13:43:21.327Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T13:43:21.327Z","emptyReason":null},"stars":null,"forks":null,"downloads":1406,"packageName":null,"latestVersion":"1.2.1","tractionLabel":"1.4K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T13:43:21.326Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T13:43:21.327Z","lastCrawledAt":"2026-10-10T13:43:21.326Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T13:43:21.326Z","lastVerifiedAt":null,"highlights":[{"version":"1.2.1","createdAt":"2026-09-14T13:20:09.429Z","changelog":"Privacy fix: stop reading USER_EMAIL and CLAWDBOT_EMAIL; omit contact email unless explicitly supplied. Add selectable English HTML reports with localized eligibility messages while preserving source citations and evidence. Add a capability-to-implementation map and privacy/language regression tests. Validation: 51 offline tests passed; English report layout, filtering and reset verified in a browser.","fileCount":22,"zipByteSize":80894},{"version":"1.2.0","createdAt":"2026-09-14T03:26:22.280Z","changelog":"Correct version numbering: 1.1.3 to 1.2.0. Compact HTML reports, stronger verification safeguards, four data-integrity fixes, and MIT-0 license. Supersedes the mistakenly numbered 2.1.0 release.","fileCount":20,"zipByteSize":70632},{"version":"1.1.3","createdAt":"2026-07-13T15:07:05.510Z","changelog":"Adds GitHub source provenance and links this ClawHub release to the public biomedical-reference-verifier repository. Core verification behavior is unchanged. 新增 GitHub 开源来源信息，并将 ClawHub 版本关联到公开 biomedical-reference-verifier 仓库。核心核验功能未改变。","fileCount":13,"zipByteSize":46290},{"version":"1.1.2","createdAt":"2026-07-11T17:34:15.695Z","changelog":"Adds Fast, Balanced, and Strict modes; bounded retries and provider-specific scheduling; PubMed batch-failure isolation; early DOI recovery through Crossref, OpenAlex, and PubMed; lossless audit/index conversion; artifact-tolerant and field-level result reuse; runtime metrics and request-event logging; modularized components; and expanded offline and real-interface regression coverage. Severe references remain report-only, and source files are never overwritten. 新增快速、均衡、严格三种模式；加入有限重试、分接口调度、PubMed批次失败隔离及多来源DOI提前恢复；支持审计与索引无损互转、产物容错、字段级复用、运行指标与请求事件记录；完成模块化整理并扩展离线及真实接口测试。严重错误仍仅报告，原文件不会被覆盖。","fileCount":13,"zipByteSize":46299},{"version":"1.1.1","createdAt":"2026-07-11T17:32:02.117Z","changelog":"- Updated skill description to be more concise and to include Chinese for improved clarity and accessibility. - Removed the file `skill-card.md`. - No changes to the core workflow, commands, functionality, or execution process. - Other content, usage rules, pipeline choices, and execution instructions remain unchanged.","fileCount":13,"zipByteSize":46181},{"version":"1.1.0","createdAt":"2026-07-11T16:18:05.839Z","changelog":"Biomedical Reference Verifier 1.1.0 introduces new execution modes and artifact management: - Added Fast, Balanced, and Strict execution modes for tunable verification speed and strictness. - Introduced artifact index format and round-trip conversion tools for result reuse. - Enhanced prior-result handling with robust validation, error skipping, and strict comparison rules. - Updated command-line interface to support new modes, artifact conversion, and result reuse options. - Improved audit policy and documentation in SKILL.md and references/verification_policy.md. - Modularized core logic into new verifier_* scripts for improved maintainability.","fileCount":13,"zipByteSize":46563},{"version":"1.0.0","createdAt":"2026-07-10T15:42:38.120Z","changelog":"Initial public release.","fileCount":6,"zipByteSize":36182}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s1762bkccamy86w1sawy1a2re58a8s0f:biomedical-reference-verifier","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T15:51:18.033Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-emp-tca-biomedical-reference-verifier/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T13:43:21.327Z","emptyReason":null},"readme":"Skill: biomedical-reference-verifier\n\nOwner: emp-tca\n\nSummary: Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。\n\nTags: latest:1.2.1\n\nVersion history:\n\nv1.2.1 | 2026-09-14T13:20:09.429Z | user\n\nPrivacy fix: stop reading USER_EMAIL and CLAWDBOT_EMAIL; omit contact email unless explicitly supplied. Add selectable English HTML reports with localized eligibility messages while preserving source citations and evidence. Add a capability-to-implementation map and privacy/language regression tests. Validation: 51 offline tests passed; English report layout, filtering and reset verified in a browser.\n\nv1.2.0 | 2026-09-14T03:26:22.280Z | user\n\nCorrect version numbering: 1.1.3 to 1.2.0. Compact HTML reports, stronger verification safeguards, four data-integrity fixes, and MIT-0 license. Supersedes the mistakenly numbered 2.1.0 release.\n\nv1.1.3 | 2026-07-13T15:07:05.510Z | user\n\nAdds GitHub source provenance and links this ClawHub release to the public biomedical-reference-verifier repository. Core verification behavior is unchanged.\n\n新增 GitHub 开源来源信息，并将 ClawHub 版本关联到公开 biomedical-reference-verifier 仓库。核心核验功能未改变。\n\nv1.1.2 | 2026-07-11T17:34:15.695Z | user\n\nAdds Fast, Balanced, and Strict modes; bounded retries and provider-specific scheduling; PubMed batch-failure isolation; early DOI recovery through Crossref, OpenAlex, and PubMed; lossless audit/index conversion; artifact-tolerant and field-level result reuse; runtime metrics and request-event logging; modularized components; and expanded offline and real-interface regression coverage. Severe references remain report-only, and source files are never overwritten.\n\n新增快速、均衡、严格三种模式；加入有限重试、分接口调度、PubMed批次失败隔离及多来源DOI提前恢复；支持审计与索引无损互转、产物容错、字段级复用、运行指标与请求事件记录；完成模块化整理并扩展离线及真实接口测试。严重错误仍仅报告，原文件不会被覆盖。\n\nv1.1.1 | 2026-07-11T17:32:02.117Z | auto\n\n- Updated skill description to be more concise and to include Chinese for improved clarity and accessibility.\n- Removed the file `skill-card.md`.\n- No changes to the core workflow, commands, functionality, or execution process.\n- Other content, usage rules, pipeline choices, and execution instructions remain unchanged.\n\nv1.1.0 | 2026-07-11T16:18:05.839Z | auto\n\nBiomedical Reference Verifier 1.1.0 introduces new execution modes and artifact management:\n\n- Added Fast, Balanced, and Strict execution modes for tunable verification speed and strictness.\n- Introduced artifact index format and round-trip conversion tools for result reuse.\n- Enhanced prior-result handling with robust validation, error skipping, and strict comparison rules.\n- Updated command-line interface to support new modes, artifact conversion, and result reuse options.\n- Improved audit policy and documentation in SKILL.md and references/verification_policy.md.\n- Modularized core logic into new verifier_* scripts for improved maintainability.\n\nv1.0.0 | 2026-07-10T15:42:38.120Z | user\n\nInitial public release.\n\nArchive index:\n\nArchive v1.2.1: 22 files, 80894 bytes\n\nFiles: assets/reference-audit-report-template.en.html (16935b), assets/reference-audit-report-template.html (16829b), CHANGELOG.md (1354b), LICENSE (906b), README.md (1576b), references/verification_policy.md (14486b), scripts/convert_reference_artifact.py (1437b), scripts/generate_html_report.py (5795b), scripts/self_test_verify_references.py (13008b), scripts/test_privacy_language.py (5606b), scripts/test_release_boundaries.py (7395b), scripts/test_verification_boundaries.py (9328b), scripts/verifier_models.py (2673b), scripts/verifier_network.py (2739b), scripts/verifier_policy.py (849b), scripts/verifier_prior_results.py (7706b), scripts/verifier_recovery.py (859b), scripts/verifier_runtime.py (1604b), scripts/verify_references.py (131637b), skill-card.md (2194b), SKILL.md (15841b), _meta.json (148b)\n\nFile v1.2.1:SKILL.md\n\n---\nname: biomedical-reference-verifier\nversion: 1.2.1\ndescription: \"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。\"\nmetadata:\n  openclaw:\n    requires:\n      bins:\n        - python3\n    envVars:\n      - name: NCBI_API_KEY\n        required: false\n        description: Optional NCBI API key for higher PubMed E-utilities rate limits.\n      - name: OPENALEX_API_KEY\n        required: false\n        description: Optional OpenAlex API key for authenticated metadata lookups.\n---\n\n# Biomedical Reference Verifier\n\nUse this skill only for reference-list authenticity checking, AI-caused reference cleanup, and citation-format normalization. Do not use it to judge whether papers are relevant, suitable, or correctly used in the manuscript body unless the user explicitly asks for reference authenticity plus body cleanup.\n\n## Core Rule\n\nAlways reduce the task to this sequence:\n\n1. Build or read `biomedical-reference-verifier.records.v1`.\n2. Choose the lightest pipeline.\n3. Run the script.\n4. Report result files and retained process files.\n5. Ask before expensive manual recovery or process-file deletion.\n\nLoad `references/verification_policy.md` before classifying authenticity errors, changing severe items, or explaining network/query policy.\n\n## Pipeline Choice\n\n- **Format only**: use when the user only asks to convert or unify citation style. Run `--pipeline format-only`. Do not query Crossref/PubMed/OpenAlex.\n- **Verify**: use when the user asks whether references are real, fake, wrong, AI-generated, DOI/PMID-mismatched, or metadata-corrupted. Run the default `verify` pipeline.\n- **Verify then format**: use when the user asks both to check authenticity and standardize style. Run `verify` with the requested `--citation-style`; the fixed copy is generated after verification.\n\n## Execution Mode\n\n- **Fast** (`--mode fast`): use for quick scans and large routine lists. Stop after primary DOI evidence when available.\n- **Balanced** (`--mode balanced`): default for ordinary authenticity checks. Auxiliary evidence lines receive a short 2.5-second grace period after primary DOI lookup completes (and supplied PMID lookup when enabled).\n- **Strict** (`--mode strict`): use for final or pre-submission checks. Wait for every enabled evidence line to complete or fail explicitly.\n\nDo not invent a local confidence score. For DOI-bearing references, resolve the DOI and directly compare returned title, authors, journal, year, DOI, and PMID with source fields. Keep the existing classification policy. Severe items are reported and retained; never auto-delete them.\n\nDo not run AI-assisted per-paper web searching by default. Stop after batch verification and title recovery, then ask the user whether to continue.\n\n## Machine Input\n\nPrefer `biomedical-reference-verifier.records.v1` JSON/JSONL. If the source is free text, convert it to records first using only values present in the user source. Do not fill missing source fields from Crossref, PubMed, OpenAlex, memory, or plausible guesses.\n\nMinimal record:\n\n```json\n{\n  \"schema\": \"biomedical-reference-verifier.records.v1\",\n  \"records\": [\n    {\n      \"index\": 1,\n      \"source\": {\n        \"original_text\": \"exact source reference\",\n        \"title\": \"title from source, or empty string\",\n        \"authors\": [\"First Author\"],\n        \"year\": \"2024\",\n        \"journal\": \"Journal from source\",\n        \"identifiers\": {\"doi\": \"10.xxxx/example\", \"pmid\": \"\", \"urls\": []},\n        \"source_lines\": [12],\n        \"context\": \"\"\n      }\n    }\n  ]\n}\n```\n\nUse `--input-mode records` for standardized JSON/JSONL. Use Markdown worksheet and `doi-context` modes only for compatibility.\n\n## Commands\n\nRun from this skill directory:\n\n```bash\npython3 scripts/verify_references.py refs.md --output-dir /tmp/reference-audit --citation-style ama\n```\n\nCommon variants:\n\n```bash\npython3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --mode fast\npython3 scripts/verify_references.py records.json --input-mode records --mode strict\npython3 scripts/verify_references.py records.json --input-mode records --reuse-results previous/reference-audit.json\npython3 scripts/verify_references.py records.json --input-mode records --write-index\npython3 scripts/convert_reference_artifact.py previous/reference-audit.json --to index --output reference-index.json\npython3 scripts/convert_reference_artifact.py reference-index.json --to audit --output restored-reference-audit.json\npython3 scripts/verify_references.py refs.md --pubmed-mode off --openalex-mode off\npython3 scripts/verify_references.py refs.md --doi-output append\npython3 scripts/verify_references.py refs.md --keep-process-json\npython3 scripts/generate_html_report.py previous/reference-audit.json --output reference-audit-report.html\npython3 scripts/verify_references.py refs.md --cleanup-process-files all\npython3 scripts/verify_references.py refs.md --cleanup-process-files normalized_input,extracted_references\n```\n\n## Output Contract\n\nResult files:\n\n- `reference-audit-summary.md`: concise chat-ready summary.\n- `reference-audit-detail.md`: detailed report with evidence links.\n- `reference-audit-report.html`: self-contained interactive report generated from the fixed template in `assets/reference-audit-report-template.html`.\n- `references.auto-fixed.md` or `document.auto-fixed.md`: fixed/formatted copy; never overwrites the source.\n\nThe HTML report is a default result file for both verification and format-only runs. Do not regenerate its page structure, CSS, or JavaScript in the task output. The script injects structured report data into the template marker and escapes script-breaking characters. Keep result cards in the input `results` order; never sort them by category, severity, title, or index.\n\nThe HTML report uses a compact, white, single-page list. All records, original/output citations, field differences, issues and evidence links are visible without expanding cards or changing pages. Preserve input order. Provide text search, category/status/field filters, reset, copy of filtered results with eligibility warnings, and print. Use small semantic tags rather than large decorative panels. Never equate title similarity with a confidence percentage.\n\nDisplay groups:\n- `correct`: `verified`.\n- `auto_fixed`: `minor_fix`, legacy `minor_format_error`.\n- `blocked`: identifier-only, metadata conflicts, fabrication, parser errors and unresolved items. Label **不可用**.\n- `unchecked`: `formatted_only`; formatting never proves authenticity.\n\nEach result separately records `verification_level`, `usable`, `repair_state`, and `field_differences`. Only verified/strongly matched, safely repaired results from the verification pipeline are eligible. A reference that cannot be verified must not be used. Keep the existing fabrication classifications, but distinguish an unavailable query from evidence that a query actually completed. Never imply that a timeout proves fabrication.\n\nProcess-file lifecycle:\n- Retain `reference-normalized-records.json`, the source extraction suitable for reuse and parser review.\n- Automatically clean newly generated `reference-normalized-input.md` and `references.extracted.md`: both can be rebuilt from the retained records. `--cleanup-process-files none` explicitly retains all generated process files.\n- Write `reference-audit.json` only with `--keep-process-json`; write `reference-index.json` only with `--write-index`. These retain reusable evidence and runtime information.\n- Remove task-created disposable scratch scripts and redundant previews after validation; never remove maintained skill scripts, user inputs, result files, or prior-run artifacts as generic cleanup.\n- After every round, report cleanup and ask which remaining process files to keep. Keep them if the user does not reply. A prior explicit keep/delete instruction takes precedence.\n\nThere is no hidden persistent cache. `--reuse-results` accepts either `reference-audit.json` or schema `biomedical-reference-verifier.index.v1`. The converter supports audit-to-index and index-to-audit round trips; the `results` array must remain identical after a round trip. In a conversation, ask about reuse once when such an artifact is available; do not repeatedly ask after the user decides.\n\nUser-edited prior artifacts are handled conservatively: malformed JSON is rejected with one concise error; invalid individual rows are skipped and rechecked; valid rows remain reusable. Do not guess repairs for damaged structured fields.\n\nPrior-result reuse requires all stored source fields to agree (including PMID, authors, journal and publication details), current policy version, compatible pipeline/mode/channels, a valid canonical record and a check date within 30 days. Failed, unresolved, incomplete and old artifacts are rechecked. Punctuation-only changes may be normalized. Reuse evidence, then regenerate the output for the requested citation style and DOI settings; never reuse stale formatted text. Temporary network failures receive one bounded retry; permanent 4xx failures do not. PubMed DOI batches split only when a batch fails. Early DOI recovery uses Crossref first, then OpenAlex and PubMed outside Fast mode.\n\n`--write-index` optionally writes `reference-index.json`; it is off by default. Network request events and phase timings are stored in `reference-audit.json` when `--keep-process-json` is used.\n\nNever delete user-requested result files. Delete only process files generated in the current run, according to the lifecycle above or an explicit user selection. Never delete an existing audit merely because this run did not request JSON.\n\n## Closeout\n\nAfter running:\n\n1. Paste or summarize `reference-audit-summary.md`.\n2. Link result files and retained process files.\n3. State which process files were cleaned or skipped.\n4. Ask whether to delete retained process files:\n   - A. delete all retained process files\n   - B. keep selected process files\n   - C. keep all process files for now\n5. If severe items exist, ask whether to continue AI-assisted recovery, delete, keep with warning notes, or stop with the report.\n6. Ask whether recovered DOI values should be appended when DOI output is optional.\n\n## Formatting and repair boundaries\n\nUse `formatted_only` for formatting without external queries. A missing or unreliable title leaves the original intact. Never invent an author. Preserve source DOI values; only newly recovered DOI values are subject to the append choice. Preserve volume, issue and page/article numbers in every supported style. The built-in renderer is a plain-text journal-article formatter, not a complete CSL engine: books, datasets and exact publisher typography need a dedicated renderer. Do not promise full style compliance for unsupported document types.\n\nPrefer author objects with `family`/`given`, or explicit `family, given` strings. Preserve surname-first initials and corporate names; retain ambiguous names rather than guessing. Compare authors, title, journal, date and supplied identifiers. Only an explicit provider abbreviation establishes journal equivalence; a different journal is a conflict. A one-year date difference is a reviewable difference unless date evidence explains it. Do not rewrite an identifier-only or partial/conflicting item automatically. Store field-level before/after values and evidence source.\n\nThis skill does not require another design skill to maintain its report template.\n\n## Real-list regression boundaries\n\n- Treat an explicit trailing `et al.` as an omitted author suffix. Compare every named author, in order, against the canonical prefix; do not require the abbreviated and complete author lists to have the same length. Preserve real prefix disagreements.\n- Commas in a scientific title are not evidence that it is an author list. Only apply the comma heuristic when each segment has author-name syntax.\n- Preserve PubMed `CollectiveName` and explicit family/given components. Normalize typographic apostrophes and diacritics for author comparison; do not infer missing initials or compound surnames.\n- Treat abbreviated page ranges and repeated electronic page endpoints as equivalent; distinct article numbers remain conflicts.\n- If a primary provider returns a recognized placeholder title such as `OUP accepted manuscript`, compare other enabled records for exactly the same DOI. Use a matching title/author/year record and disclose the fallback; do not replace the DOI. A placeholder alone cannot establish hijacking.\n- Retain same-identifier journal aliases and explicit publication years from provider evidence. Preserve a valid source year when it matches a known publication year. Save `provider_records` in structured results for reproducible diagnosis.\n\n## Implementation map and review scope\n\nThese files implement the advertised workflow; inspect them together rather than inferring missing functionality from one chunk:\n\n| Capability | Implementation | Offline verification |\n| --- | --- | --- |\n| Input normalization, Crossref/PubMed/OpenAlex lookup and field comparison | `scripts/verify_references.py` | `scripts/test_verification_boundaries.py` |\n| Formatting, output safety and query-failure handling | `scripts/verify_references.py` | `scripts/test_release_boundaries.py` |\n| HTML payload escaping and report generation | `scripts/generate_html_report.py` | `scripts/self_test_verify_references.py`, `scripts/test_privacy_language.py` |\n| Chinese and English report layout and filtering | `assets/reference-audit-report-template.html`, `assets/reference-audit-report-template.en.html` | `scripts/test_privacy_language.py` |\n| Evidence reuse and artifact conversion | `scripts/verifier_prior_results.py`, `scripts/convert_reference_artifact.py` | `scripts/self_test_verify_references.py` |\n\nThe skill executes local Python, reads the selected reference input, writes reports, and queries the documented bibliographic services. Formatting-only is offline. Verification sends DOI/PMID and, for title recovery, reference titles; it does not need manuscript paragraphs. Tests use subprocess argument lists to exercise local scripts. HTML templates are bundled and have no external script/font dependencies. This implementation map aids inspection; it is not a security certification.\n\n## Contact email and report language\n\nDo not read `USER_EMAIL` or `CLAWDBOT_EMAIL`. No contact email is included by default. Only pass `--email` when the user explicitly wants a contact email sent to the enabled bibliographic providers. The two optional provider API keys retain their existing purpose.\n\nChoose the HTML language from the user's request or conversation language: `--report-language en` for English and `--report-language zh` for Chinese. The CLI default remains Chinese for compatibility. Both templates show the complete report with the same filters, eligibility rules, copy and print functions. Preserve source citations and provider evidence in their original language; selecting English translates the interface and generated eligibility messages, not the cited research.\n\n```bash\npython3 scripts/verify_references.py records.json --input-mode records --report-language en\npython3 scripts/generate_html_report.py reference-audit.json --output report.en.html --language en\n```\n\nFile v1.2.1:README.md\n\n# Biomedical Reference Verifier\n\nVerify and normalize biomedical and life-science reference lists, with a focus on AI-caused citation errors.\n\nThe skill checks DOI, PMID, title, authors, journal, year, and other bibliographic metadata through Crossref, PubMed, and OpenAlex. It can also normalize references to AMA, APA, Vancouver, or GB/T 7714 style without overwriting the source document.\n\n## Highlights\n\n- Fast, Balanced, and Strict verification modes\n- DOI-first batch verification and early DOI recovery\n- Crossref, PubMed, and OpenAlex evidence\n- Bounded retries and PubMed batch-failure isolation\n- Detection of identifier hijacking and shifted identifiers\n- Reusable audit indexes with lossless audit/index conversion\n- Tolerant handling of user-edited prior artifacts\n- Runtime metrics and request-event logging\n- No hidden persistent cache\n- Severe references are reported, never automatically deleted\n\n## Requirements\n\n- Python 3\n- Standard library only\n- Optional `NCBI_API_KEY` for higher PubMed limits\n- Optional `OPENALEX_API_KEY`\n\n## Quick start\n\n```bash\npython3 scripts/verify_references.py refs.md --mode balanced --output-dir reference-audit\n```\n\nFormat without external verification:\n\n```bash\npython3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\n```\n\nRun the offline regression suite:\n\n```bash\npython3 scripts/self_test_verify_references.py\n```\n\nSee [SKILL.md](SKILL.md) for the complete workflow and [verification_policy.md](references/verification_policy.md) for classification and evidence rules.\n\n## License\n\nMIT-0\n\nFile v1.2.1:_meta.json\n\n{\n  \"ownerId\": \"kn748sfft0aksxdrw9332r67w58a99n2\",\n  \"slug\": \"biomedical-reference-verifier\",\n  \"version\": \"1.2.1\",\n  \"publishedAt\": 1789392009429\n}\n\nFile v1.2.1:references/verification_policy.md\n\n# Verification Policy\n\n## Evidence hierarchy\n\n1. Crossref is the fastest and strongest default primary source for journal DOI, title, authors, journal, publisher, and year.\n2. PubMed is the second evidence line for biomedical records, PMID mapping, and wrong PMID/title-pair detection.\n3. OpenAlex is the third default evidence line for fast DOI/PMID external-ID corroboration and title recovery when Crossref is weak.\n4. A DOI match from Crossref is the primary batch verification result. PubMed and OpenAlex matches can corroborate it, but should not block reporting.\n5. Semantic Scholar, Europe PMC, bioRxiv/medRxiv, DataCite, and OpenCitations are backup channels, not default full-batch channels.\n6. Search snippets, AI summaries, formatted APA/GB/T entries, and plausible journal names are not proof.\n7. `biomedical-reference-verifier.records.v1` is the source-of-truth input layer. Source fields must come from the user's text; database-returned fields must stay in evidence/report/output objects.\n\n## Error types\n\n- `verified`: DOI/PMID or title search resolves to a canonical record and title, year, and first-author evidence agree.\n- `verified_identifier_only`: DOI/PMID resolves to a canonical record, but the source title is missing or unreliable, so only the identifier is verified.\n- `parser_error`: the input row is malformed or the parsed/source title is an author line, table header, metadata field, or other non-title text; stop before treating the row as a bibliographic conflict.\n- `formatted_only`: formatting executed, authenticity not checked; not eligible for use on that basis.\n- `minor_fix` / `minor_format_error`: the paper is real and metadata agree; only DOI casing, punctuation, URL style, initials, journal abbreviation, or minor year formatting needs correction.\n- `partial_attribute_corruption`: the paper is real, but one or more attributes are corrupted, such as author, title, journal, year, volume/pages, DOI, or PMID.\n- `identifier_hijacking`: DOI, PMID, or URL is real, but it points to a different paper than the reference text.\n- `shifted_identifier`: DOI or PMID appears to belong to a nearby reference, commonly because one entry's DOI was attached to the previous or next citation.\n- `semantic_hallucination`: the paper exists, but a surrounding claim or summary is not supported by that paper.\n- `placeholder_generation`: the entry contains AI-like placeholder traces: generic title, invented DOI suffix, fake pages/volume, missing journal fields, or overly tidy but unverifiable metadata.\n- `total_fabrication`: title, authors, journal, year, DOI/PMID, and searches fail to identify any canonical record.\n- `unresolved`: automatic checks are insufficient; stop and ask before expensive manual recovery.\n\n## Matching thresholds\n\n- Title similarity >= 0.90: strong match.\n- Title similarity 0.86-0.89: acceptable match if year and first author agree.\n- Title similarity 0.78-0.85: possible match; mark as partial unless corroborated by DOI/PMID.\n- Title similarity < 0.70 for a supplied DOI/PMID: identifier hijacking unless adjacent-entry checks prove shifted identifier.\n- A one-year difference may be online-first drift, but remains a field difference unless source dates establish equivalence. Missing year/author evidence cannot satisfy full verification.\n- Author disagreement prevents automatic approval. Compare available authors and report differences. A differing journal is a conflict unless an explicit provider abbreviation matches.\n\n## Short-circuit rules\n\n- Build `biomedical-reference-verifier.records.v1` first. Do not let remote lookup run against raw mixed notes when the source can be converted to explicit source objects.\n- Required machine fields are `index`, `source.original_text`, `source.title`, `source.authors`, `source.year`, `source.journal`, `source.identifiers.doi`, `source.identifiers.pmid`, `source.identifiers.urls`, `source.source_lines`, and `source.context`.\n- During machine-record construction, AI may only extract values present in the source text. It must not use Crossref/PubMed/OpenAlex results, memory, or plausible completions to fill source fields.\n- Build `reference-normalized-records.json` before remote lookup. `reference-normalized-input.md` is a human-readable view of the same records.\n- If `source.title` is missing, too short, an author list, a table/header label, a DOI/PMID line, or metadata placeholder, mark the title unreliable.\n- If a DOI/PMID resolves but title is unreliable, mark `verified_identifier_only`; do not mark `identifier_hijacking`.\n- If a row has no reliable title and no resolvable DOI/PMID, mark `parser_error` and stop.\n- Use DOI verification first: `GET https://api.crossref.org/works/{doi}`. This is the primary source of DOI/title/author/journal/year truth.\n- Start enabled evidence lines together. Fast stops after primary evidence, Balanced allows a 2.5-second auxiliary grace period after primary DOI lookup and enabled supplied-PMID lookup, and Strict waits for every enabled line to complete or explicitly fail. Network timeout, rate limiting, not-found, parsing failure, mode skip, and expired budget must remain distinguishable.\n- If PubMed and OpenAlex are disabled with `--pubmed-mode off --openalex-mode off`, do not apply the grace timer to Crossref; Crossref DOI verification must complete because it is the primary authority.\n- OpenAlex checks use `GET https://api.openalex.org/works/doi:{doi}` or external IDs such as `pmid:{pmid}` for PMID, and `GET https://api.openalex.org/works?search={title}&per_page=3` for DOI-missing title recovery.\n- If an entry has no DOI, use title recovery in this order: Crossref title query first, OpenAlex title query only if the Crossref match is weak, PubMed title query only if biomedical corroboration is still needed. Record recovered DOI values in the report and ask before adding them to formatted citations.\n- Run quality-checked DOI-missing title recovery before full identifier verification, deduplicated and with provider-specific bounded concurrency. A recovered DOI may drive verification but must not silently modify the user source.\n- Do not use title search to silently override an existing DOI result. Existing DOI conflicts should be reported as conflicts, hijacking, or shifted identifiers.\n- Run near official limits without exceeding them: Crossref polite requests use no more than 10 requests/second and 3 concurrent requests; PubMed E-utilities uses no more than 3 requests/second without an API key and 10 requests/second with an NCBI API key; OpenAlex is budget/cost governed, so keep default concurrency conservative unless an API key and live limit headers justify more.\n- Use one bounded retry only for timeout, temporary connection failure, HTTP 429, or HTTP 5xx. Respect `Retry-After` within a bounded wait. Do not retry permanent 4xx responses. If a PubMed DOI batch fails, split it recursively until the failing DOI is isolated.\n- For large DOI-heavy documents, Crossref parallel DOI lookup is the primary full-batch path. PubMed and OpenAlex are corroboration channels, not reasons to block report generation.\n- If DOI/PMID resolves and title similarity is < 0.70, stop trusting that identifier for the current entry and mark `identifier_hijacking`.\n- Apply the < 0.70 hijacking rule only when the source title passed input quality checks.\n- Before finalizing `identifier_hijacking`, compare the resolved title against entries within two positions. If adjacent similarity is >= 0.86, mark `shifted_identifier`.\n- If several consecutive identifiers are shifted by the same offset, report a possible column/line offset instead of repairing each item manually.\n- For entries with DOI, run DOI lookup and report the DOI result; do not automatically title-search to replace the DOI.\n- For entries without DOI, run at most one Crossref title recovery query. If title recovery fails, mark `unresolved`, `placeholder_generation`, or `total_fabrication`; do not keep prompting the model to guess.\n- Do not run title recovery when `source_title` failed input quality checks. Rebuild the worksheet row first.\n- Use backup channels conditionally: Semantic Scholar only for title/abstract/citation corroboration after default lines are inconclusive; Europe PMC for biomedical PMCID/full-text evidence; bioRxiv/medRxiv for `10.1101/...` and published-DOI mapping; DataCite for non-journal dataset/software/report DOIs; OpenCitations for citation-network existence checks.\n- Parse every channel response by code into one canonical record shape: `source`, `title`, `authors`, `journal`, `year`, `doi`, `pmid`, `url`, and `score`. AI must not manually interpret structured API responses when parser adapters are available.\n- If the severe-error rate is high, stop after batch recovery and ask the user whether to continue AI-assisted per-paper search.\n\n## Network risk policy\n\n- Default DOI/PMID verification is low risk and should run without asking the user each time. The payload is limited to public identifiers required for the skill's purpose.\n- Default requests must not include `original_text`, manuscript paragraphs, abstracts, local evidence notes, or unpublished claims.\n- DOI-missing recovery may send only a short, quality-checked `source.title`.\n- Any deep search that sends title plus abstract, surrounding context, or manuscript-derived claims is higher risk and must be separated from default batch verification.\n- If Codex sandbox approval blocks external lookup, continue with local records normalization and report that external evidence was not executed.\n\n## Channel parser contract\n\n- Crossref `/works/{doi}` returns a `message` object; parse `title`, `author`, `container-title`, `issued`/published dates, `DOI`, and `URL`.\n- PubMed EFetch returns XML; parse `ArticleTitle`, `Author`, `Journal`, `PubDate`, `ArticleId IdType=\"doi\"`, and `PMID`.\n- OpenAlex works return work objects; parse `display_name`/`title`, `authorships`, `primary_location.source.display_name`, `publication_year`, `doi`, and `ids.pmid`.\n- Europe PMC search results return result objects; parse `title`, `authorString`, `journalTitle`, `pubYear`, `doi`, `pmid`, and `pmcid`.\n- Semantic Scholar Graph results return paper objects; parse `title`, `authors`, `venue`/`journal`, `year`, `externalIds.DOI`, `externalIds.PubMed`, and `url`.\n- DataCite DOI records return JSON:API objects; parse `data.attributes.titles`, `creators`, `publisher`/`container`, `publicationYear`, `doi`, and `url`.\n- bioRxiv/medRxiv API results return collection items; parse `title`, `authors`, `date`, `doi`, and `published_doi`.\n\n## Auto-fix policy\n\nAuto-fix only when canonical evidence is strong:\n\n- normalize DOI casing and links\n- strip trailing DOI punctuation\n- normalize DOI links in reports\n- record missing DOI candidates from canonical metadata and ask before appending them\n- include PMID only as secondary metadata when DOI-backed evidence exists\n- standardize journal title or abbreviation\n- normalize author initials and year formatting\n- report shifted DOI/PMID as a candidate relationship; retain the original until the proposed move is reviewed\n\nDo not auto-delete or silently replace severe items. For identifier-only, shifted, hijacked, fabricated, unresolved, and any partial/conflicting records, retain the original and report. Only verified/minor results may be automatically reformatted. Keep per-field before/after evidence.\n\nDo not create hidden persistent caches. Reuse prior results only from an explicitly supplied structured verifier artifact, and re-query new, changed, or incomplete records.\n\n## Format policy\n\n- Detect citation style before repair: APA, GB/T 7714, Vancouver/AMA, free-text, or mixed.\n- Preserve original format in the automatic fixed copy when possible.\n- If the list is mixed, recommend standardizing to the majority format after authenticity checks.\n- Formatting correctness is never authenticity evidence.\n\n## Body-vs-bibliography audit\n\nOnly execute this section when the user explicitly requests manuscript-body checking or cleanup. Bibliography verification alone does not authorize changing body claims.\n\n- Scan body text for narrative citations like `Author et al. (2023)` and parenthetical citations like `(Author, 2023)`.\n- Compare author/year/title semantics, not just numbered reference markers.\n- After deleting or flagging a fabricated reference, remove or flag body claims that depend on it.\n- Report uncertain claims separately instead of silently preserving them.\n\n## Eligibility, reuse and cleanup\n\n- The user requires a positive-use gate: references that cannot be verified are not usable. Identifier existence alone is insufficient. Formatting-only results are not authenticity passes. Preserve detailed reasons alongside this gate.\n- Results record verification coverage, eligibility, repair state and field differences independently. HTML groups are navigation aids, not replacements for these fields.\n- Reuse only complete compatible evidence checked within 30 days under the current policy, with all source fields unchanged. Stronger modes and additional channels require fresh checks. Regenerate formatted output under the current options.\n- Clean current-run redundant parser views/extracted copies by default; retain normalized source records and requested evidence artifacts. Honor `--cleanup-process-files none`. Ask after each round about retaining the remaining process files; no answer means keep. Never erase prior evidence just because an output option was omitted.\n\n## Real-reference matching details\n\nTrailing `et al.` denotes author truncation, not a named author. Compare the named prefix after conservative typographic normalization. Do not approve a different second or third author. Parse PubMed collective authors. A title containing commas is valid unless the comma-separated pieces actually have author-name syntax.\n\nRecognized publisher placeholder titles are inadequate metadata. A matching record from an enabled channel for the same DOI may supply the bibliographic title, with the fallback recorded. Without such evidence, retain an unresolved result; do not declare hijacking from the placeholder.\n\nPage forms such as `E5503-12` and `E5503-E5512`, or `e136-e136` and `e136`, are equivalent. Distinct electronic article numbers are not normalized into one another. Same-DOI journal aliases and explicit provider publication-year values are evidence of legitimate variants. Retain provider-specific records in the audit JSON.\n\nFile v1.2.1:CHANGELOG.md\n\n# Changelog\n\n## 1.2.1 — 2026-09-14\n\n- Remove ambient email reads and omit contact email from requests unless explicitly supplied.\n- Add selectable English HTML report template and localized generated eligibility messages.\n- Document implementation locations for capability review; add privacy and language regression tests.\n\n## 1.2.0 — 2026-09-14\n\nMinor release following public version 1.1.3. SKILL.md and the CLI both identify this release as 1.2.0.\n\n- Reject duplicate reference indices before verification or writing outputs, preventing silent loss or duplication of results.\n- Reject output paths that alias the source, an explicitly reused artifact, or another output, including symlinks and hardlinks.\n- Keep failed title queries distinguishable from completed searches with no match. Failed queries leave unresolved references unusable; successful empty searches retain the existing classification policy.\n- Do not cache failed title searches as successful empty results. Detect malformed search responses and incomplete PubMed record retrieval.\n- Include formatted-only entries in the Markdown detail report.\n- Update the verification policy version to 2026-09-14.1 so older evidence is rechecked.\n- Add 10 offline regression tests for these boundaries.\n- Add the MIT-0 license and exclude Python caches and macOS metadata from publication.\n\nFile v1.2.1:skill-card.md\n\n## Description:\n\nVerify and normalize biomedical and life-science reference lists, with a focus on AI-caused citation errors.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[emp-tca](https://clawhub.ai/user/emp-tca)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nExternal users and developers use this skill to check biomedical and life-science reference lists for DOI, PMID, title, author, journal, year, and metadata inconsistencies, then produce reviewable reports and normalized citations.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Verification runs may contact Crossref, PubMed/NCBI, and OpenAlex with public identifiers and sometimes citation titles.\n\nMitigation: Use format-only mode or disable providers for confidential drafts or unpublished research topics; provide contact email or API keys only when intentionally sending them to providers.\n\nRisk: Formatting-only output does not establish reference authenticity.\n\nMitigation: Use the verification pipeline before treating references as usable, and review the generated reports and evidence links for unresolved or blocked records.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/emp-tca/skills/biomedical-reference-verifier)\n- [README](README.md)\n- [Verification Policy](references/verification_policy.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, HTML, JSON, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown summaries, self-contained HTML report, JSON audit/index artifacts, fixed reference Markdown, and command guidance]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Writes local report files and never overwrites the source document; format-only mode is offline, while verification may contact Crossref, PubMed/NCBI, and OpenAlex.]\n\n## Skill Version(s):\n\n1.2.1 (source: frontmatter, changelog, release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.2.1:LICENSE\n\nMIT No Attribution\n\nCopyright (c) 2026 EMP-TCA\n\nPermission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the \"Software\"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.\n\nArchive v1.2.0: 20 files, 70632 bytes\n\nFiles: assets/reference-audit-report-template.html (16829b), CHANGELOG.md (1040b), LICENSE (906b), README.md (1576b), references/verification_policy.md (14486b), scripts/convert_reference_artifact.py (1437b), scripts/generate_html_report.py (4541b), scripts/self_test_verify_references.py (13008b), scripts/test_release_boundaries.py (7395b), scripts/test_verification_boundaries.py (9328b), scripts/verifier_models.py (2673b), scripts/verifier_network.py (2739b), scripts/verifier_policy.py (849b), scripts/verifier_prior_results.py (7706b), scripts/verifier_recovery.py (859b), scripts/verifier_runtime.py (1604b), scripts/verify_references.py (131223b), skill-card.md (2255b), SKILL.md (13309b), _meta.json (148b)\n\nFile v1.2.0:SKILL.md\n\n---\nname: biomedical-reference-verifier\nversion: 1.2.0\ndescription: \"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。\"\nmetadata:\n  openclaw:\n    requires:\n      bins:\n        - python3\n    envVars:\n      - name: NCBI_API_KEY\n        required: false\n        description: Optional NCBI API key for higher PubMed E-utilities rate limits.\n      - name: OPENALEX_API_KEY\n        required: false\n        description: Optional OpenAlex API key for authenticated metadata lookups.\n---\n\n# Biomedical Reference Verifier\n\nUse this skill only for reference-list authenticity checking, AI-caused reference cleanup, and citation-format normalization. Do not use it to judge whether papers are relevant, suitable, or correctly used in the manuscript body unless the user explicitly asks for reference authenticity plus body cleanup.\n\n## Core Rule\n\nAlways reduce the task to this sequence:\n\n1. Build or read `biomedical-reference-verifier.records.v1`.\n2. Choose the lightest pipeline.\n3. Run the script.\n4. Report result files and retained process files.\n5. Ask before expensive manual recovery or process-file deletion.\n\nLoad `references/verification_policy.md` before classifying authenticity errors, changing severe items, or explaining network/query policy.\n\n## Pipeline Choice\n\n- **Format only**: use when the user only asks to convert or unify citation style. Run `--pipeline format-only`. Do not query Crossref/PubMed/OpenAlex.\n- **Verify**: use when the user asks whether references are real, fake, wrong, AI-generated, DOI/PMID-mismatched, or metadata-corrupted. Run the default `verify` pipeline.\n- **Verify then format**: use when the user asks both to check authenticity and standardize style. Run `verify` with the requested `--citation-style`; the fixed copy is generated after verification.\n\n## Execution Mode\n\n- **Fast** (`--mode fast`): use for quick scans and large routine lists. Stop after primary DOI evidence when available.\n- **Balanced** (`--mode balanced`): default for ordinary authenticity checks. Auxiliary evidence lines receive a short 2.5-second grace period after primary DOI lookup completes (and supplied PMID lookup when enabled).\n- **Strict** (`--mode strict`): use for final or pre-submission checks. Wait for every enabled evidence line to complete or fail explicitly.\n\nDo not invent a local confidence score. For DOI-bearing references, resolve the DOI and directly compare returned title, authors, journal, year, DOI, and PMID with source fields. Keep the existing classification policy. Severe items are reported and retained; never auto-delete them.\n\nDo not run AI-assisted per-paper web searching by default. Stop after batch verification and title recovery, then ask the user whether to continue.\n\n## Machine Input\n\nPrefer `biomedical-reference-verifier.records.v1` JSON/JSONL. If the source is free text, convert it to records first using only values present in the user source. Do not fill missing source fields from Crossref, PubMed, OpenAlex, memory, or plausible guesses.\n\nMinimal record:\n\n```json\n{\n  \"schema\": \"biomedical-reference-verifier.records.v1\",\n  \"records\": [\n    {\n      \"index\": 1,\n      \"source\": {\n        \"original_text\": \"exact source reference\",\n        \"title\": \"title from source, or empty string\",\n        \"authors\": [\"First Author\"],\n        \"year\": \"2024\",\n        \"journal\": \"Journal from source\",\n        \"identifiers\": {\"doi\": \"10.xxxx/example\", \"pmid\": \"\", \"urls\": []},\n        \"source_lines\": [12],\n        \"context\": \"\"\n      }\n    }\n  ]\n}\n```\n\nUse `--input-mode records` for standardized JSON/JSONL. Use Markdown worksheet and `doi-context` modes only for compatibility.\n\n## Commands\n\nRun from this skill directory:\n\n```bash\npython3 scripts/verify_references.py refs.md --output-dir /tmp/reference-audit --citation-style ama\n```\n\nCommon variants:\n\n```bash\npython3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --mode fast\npython3 scripts/verify_references.py records.json --input-mode records --mode strict\npython3 scripts/verify_references.py records.json --input-mode records --reuse-results previous/reference-audit.json\npython3 scripts/verify_references.py records.json --input-mode records --write-index\npython3 scripts/convert_reference_artifact.py previous/reference-audit.json --to index --output reference-index.json\npython3 scripts/convert_reference_artifact.py reference-index.json --to audit --output restored-reference-audit.json\npython3 scripts/verify_references.py refs.md --pubmed-mode off --openalex-mode off\npython3 scripts/verify_references.py refs.md --doi-output append\npython3 scripts/verify_references.py refs.md --keep-process-json\npython3 scripts/generate_html_report.py previous/reference-audit.json --output reference-audit-report.html\npython3 scripts/verify_references.py refs.md --cleanup-process-files all\npython3 scripts/verify_references.py refs.md --cleanup-process-files normalized_input,extracted_references\n```\n\n## Output Contract\n\nResult files:\n\n- `reference-audit-summary.md`: concise chat-ready summary.\n- `reference-audit-detail.md`: detailed report with evidence links.\n- `reference-audit-report.html`: self-contained interactive report generated from the fixed template in `assets/reference-audit-report-template.html`.\n- `references.auto-fixed.md` or `document.auto-fixed.md`: fixed/formatted copy; never overwrites the source.\n\nThe HTML report is a default result file for both verification and format-only runs. Do not regenerate its page structure, CSS, or JavaScript in the task output. The script injects structured report data into the template marker and escapes script-breaking characters. Keep result cards in the input `results` order; never sort them by category, severity, title, or index.\n\nThe HTML report uses a compact, white, single-page list. All records, original/output citations, field differences, issues and evidence links are visible without expanding cards or changing pages. Preserve input order. Provide text search, category/status/field filters, reset, copy of filtered results with eligibility warnings, and print. Use small semantic tags rather than large decorative panels. Never equate title similarity with a confidence percentage.\n\nDisplay groups:\n- `correct`: `verified`.\n- `auto_fixed`: `minor_fix`, legacy `minor_format_error`.\n- `blocked`: identifier-only, metadata conflicts, fabrication, parser errors and unresolved items. Label **不可用**.\n- `unchecked`: `formatted_only`; formatting never proves authenticity.\n\nEach result separately records `verification_level`, `usable`, `repair_state`, and `field_differences`. Only verified/strongly matched, safely repaired results from the verification pipeline are eligible. A reference that cannot be verified must not be used. Keep the existing fabrication classifications, but distinguish an unavailable query from evidence that a query actually completed. Never imply that a timeout proves fabrication.\n\nProcess-file lifecycle:\n- Retain `reference-normalized-records.json`, the source extraction suitable for reuse and parser review.\n- Automatically clean newly generated `reference-normalized-input.md` and `references.extracted.md`: both can be rebuilt from the retained records. `--cleanup-process-files none` explicitly retains all generated process files.\n- Write `reference-audit.json` only with `--keep-process-json`; write `reference-index.json` only with `--write-index`. These retain reusable evidence and runtime information.\n- Remove task-created disposable scratch scripts and redundant previews after validation; never remove maintained skill scripts, user inputs, result files, or prior-run artifacts as generic cleanup.\n- After every round, report cleanup and ask which remaining process files to keep. Keep them if the user does not reply. A prior explicit keep/delete instruction takes precedence.\n\nThere is no hidden persistent cache. `--reuse-results` accepts either `reference-audit.json` or schema `biomedical-reference-verifier.index.v1`. The converter supports audit-to-index and index-to-audit round trips; the `results` array must remain identical after a round trip. In a conversation, ask about reuse once when such an artifact is available; do not repeatedly ask after the user decides.\n\nUser-edited prior artifacts are handled conservatively: malformed JSON is rejected with one concise error; invalid individual rows are skipped and rechecked; valid rows remain reusable. Do not guess repairs for damaged structured fields.\n\nPrior-result reuse requires all stored source fields to agree (including PMID, authors, journal and publication details), current policy version, compatible pipeline/mode/channels, a valid canonical record and a check date within 30 days. Failed, unresolved, incomplete and old artifacts are rechecked. Punctuation-only changes may be normalized. Reuse evidence, then regenerate the output for the requested citation style and DOI settings; never reuse stale formatted text. Temporary network failures receive one bounded retry; permanent 4xx failures do not. PubMed DOI batches split only when a batch fails. Early DOI recovery uses Crossref first, then OpenAlex and PubMed outside Fast mode.\n\n`--write-index` optionally writes `reference-index.json`; it is off by default. Network request events and phase timings are stored in `reference-audit.json` when `--keep-process-json` is used.\n\nNever delete user-requested result files. Delete only process files generated in the current run, according to the lifecycle above or an explicit user selection. Never delete an existing audit merely because this run did not request JSON.\n\n## Closeout\n\nAfter running:\n\n1. Paste or summarize `reference-audit-summary.md`.\n2. Link result files and retained process files.\n3. State which process files were cleaned or skipped.\n4. Ask whether to delete retained process files:\n   - A. delete all retained process files\n   - B. keep selected process files\n   - C. keep all process files for now\n5. If severe items exist, ask whether to continue AI-assisted recovery, delete, keep with warning notes, or stop with the report.\n6. Ask whether recovered DOI values should be appended when DOI output is optional.\n\n## Formatting and repair boundaries\n\nUse `formatted_only` for formatting without external queries. A missing or unreliable title leaves the original intact. Never invent an author. Preserve source DOI values; only newly recovered DOI values are subject to the append choice. Preserve volume, issue and page/article numbers in every supported style. The built-in renderer is a plain-text journal-article formatter, not a complete CSL engine: books, datasets and exact publisher typography need a dedicated renderer. Do not promise full style compliance for unsupported document types.\n\nPrefer author objects with `family`/`given`, or explicit `family, given` strings. Preserve surname-first initials and corporate names; retain ambiguous names rather than guessing. Compare authors, title, journal, date and supplied identifiers. Only an explicit provider abbreviation establishes journal equivalence; a different journal is a conflict. A one-year date difference is a reviewable difference unless date evidence explains it. Do not rewrite an identifier-only or partial/conflicting item automatically. Store field-level before/after values and evidence source.\n\nThis skill does not require another design skill to maintain its report template.\n\n## Real-list regression boundaries\n\n- Treat an explicit trailing `et al.` as an omitted author suffix. Compare every named author, in order, against the canonical prefix; do not require the abbreviated and complete author lists to have the same length. Preserve real prefix disagreements.\n- Commas in a scientific title are not evidence that it is an author list. Only apply the comma heuristic when each segment has author-name syntax.\n- Preserve PubMed `CollectiveName` and explicit family/given components. Normalize typographic apostrophes and diacritics for author comparison; do not infer missing initials or compound surnames.\n- Treat abbreviated page ranges and repeated electronic page endpoints as equivalent; distinct article numbers remain conflicts.\n- If a primary provider returns a recognized placeholder title such as `OUP accepted manuscript`, compare other enabled records for exactly the same DOI. Use a matching title/author/year record and disclose the fallback; do not replace the DOI. A placeholder alone cannot establish hijacking.\n- Retain same-identifier journal aliases and explicit publication years from provider evidence. Preserve a valid source year when it matches a known publication year. Save `provider_records` in structured results for reproducible diagnosis.\n\nFile v1.2.0:README.md\n\n# Biomedical Reference Verifier\n\nVerify and normalize biomedical and life-science reference lists, with a focus on AI-caused citation errors.\n\nThe skill checks DOI, PMID, title, authors, journal, year, and other bibliographic metadata through Crossref, PubMed, and OpenAlex. It can also normalize references to AMA, APA, Vancouver, or GB/T 7714 style without overwriting the source document.\n\n## Highlights\n\n- Fast, Balanced, and Strict verification modes\n- DOI-first batch verification and early DOI recovery\n- Crossref, PubMed, and OpenAlex evidence\n- Bounded retries and PubMed batch-failure isolation\n- Detection of identifier hijacking and shifted identifiers\n- Reusable audit indexes with lossless audit/index conversion\n- Tolerant handling of user-edited prior artifacts\n- Runtime metrics and request-event logging\n- No hidden persistent cache\n- Severe references are reported, never automatically deleted\n\n## Requirements\n\n- Python 3\n- Standard library only\n- Optional `NCBI_API_KEY` for higher PubMed limits\n- Optional `OPENALEX_API_KEY`\n\n## Quick start\n\n```bash\npython3 scripts/verify_references.py refs.md --mode balanced --output-dir reference-audit\n```\n\nFormat without external verification:\n\n```bash\npython3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\n```\n\nRun the offline regression suite:\n\n```bash\npython3 scripts/self_test_verify_references.py\n```\n\nSee [SKILL.md](SKILL.md) for the complete workflow and [verification_policy.md](references/verification_policy.md) for classification and evidence rules.\n\n## License\n\nMIT-0\n\nFile v1.2.0:_meta.json\n\n{\n  \"ownerId\": \"kn748sfft0aksxdrw9332r67w58a99n2\",\n  \"slug\": \"biomedical-reference-verifier\",\n  \"version\": \"1.2.0\",\n  \"publishedAt\": 1789356382280\n}\n\nFile v1.2.0:references/verification_policy.md\n\n# Verification Policy\n\n## Evidence hierarchy\n\n1. Crossref is the fastest and strongest default primary source for journal DOI, title, authors, journal, publisher, and year.\n2. PubMed is the second evidence line for biomedical records, PMID mapping, and wrong PMID/title-pair detection.\n3. OpenAlex is the third default evidence line for fast DOI/PMID external-ID corroboration and title recovery when Crossref is weak.\n4. A DOI match from Crossref is the primary batch verification result. PubMed and OpenAlex matches can corroborate it, but should not block reporting.\n5. Semantic Scholar, Europe PMC, bioRxiv/medRxiv, DataCite, and OpenCitations are backup channels, not default full-batch channels.\n6. Search snippets, AI summaries, formatted APA/GB/T entries, and plausible journal names are not proof.\n7. `biomedical-reference-verifier.records.v1` is the source-of-truth input layer. Source fields must come from the user's text; database-returned fields must stay in evidence/report/output objects.\n\n## Error types\n\n- `verified`: DOI/PMID or title search resolves to a canonical record and title, year, and first-author evidence agree.\n- `verified_identifier_only`: DOI/PMID resolves to a canonical record, but the source title is missing or unreliable, so only the identifier is verified.\n- `parser_error`: the input row is malformed or the parsed/source title is an author line, table header, metadata field, or other non-title text; stop before treating the row as a bibliographic conflict.\n- `formatted_only`: formatting executed, authenticity not checked; not eligible for use on that basis.\n- `minor_fix` / `minor_format_error`: the paper is real and metadata agree; only DOI casing, punctuation, URL style, initials, journal abbreviation, or minor year formatting needs correction.\n- `partial_attribute_corruption`: the paper is real, but one or more attributes are corrupted, such as author, title, journal, year, volume/pages, DOI, or PMID.\n- `identifier_hijacking`: DOI, PMID, or URL is real, but it points to a different paper than the reference text.\n- `shifted_identifier`: DOI or PMID appears to belong to a nearby reference, commonly because one entry's DOI was attached to the previous or next citation.\n- `semantic_hallucination`: the paper exists, but a surrounding claim or summary is not supported by that paper.\n- `placeholder_generation`: the entry contains AI-like placeholder traces: generic title, invented DOI suffix, fake pages/volume, missing journal fields, or overly tidy but unverifiable metadata.\n- `total_fabrication`: title, authors, journal, year, DOI/PMID, and searches fail to identify any canonical record.\n- `unresolved`: automatic checks are insufficient; stop and ask before expensive manual recovery.\n\n## Matching thresholds\n\n- Title similarity >= 0.90: strong match.\n- Title similarity 0.86-0.89: acceptable match if year and first author agree.\n- Title similarity 0.78-0.85: possible match; mark as partial unless corroborated by DOI/PMID.\n- Title similarity < 0.70 for a supplied DOI/PMID: identifier hijacking unless adjacent-entry checks prove shifted identifier.\n- A one-year difference may be online-first drift, but remains a field difference unless source dates establish equivalence. Missing year/author evidence cannot satisfy full verification.\n- Author disagreement prevents automatic approval. Compare available authors and report differences. A differing journal is a conflict unless an explicit provider abbreviation matches.\n\n## Short-circuit rules\n\n- Build `biomedical-reference-verifier.records.v1` first. Do not let remote lookup run against raw mixed notes when the source can be converted to explicit source objects.\n- Required machine fields are `index`, `source.original_text`, `source.title`, `source.authors`, `source.year`, `source.journal`, `source.identifiers.doi`, `source.identifiers.pmid`, `source.identifiers.urls`, `source.source_lines`, and `source.context`.\n- During machine-record construction, AI may only extract values present in the source text. It must not use Crossref/PubMed/OpenAlex results, memory, or plausible completions to fill source fields.\n- Build `reference-normalized-records.json` before remote lookup. `reference-normalized-input.md` is a human-readable view of the same records.\n- If `source.title` is missing, too short, an author list, a table/header label, a DOI/PMID line, or metadata placeholder, mark the title unreliable.\n- If a DOI/PMID resolves but title is unreliable, mark `verified_identifier_only`; do not mark `identifier_hijacking`.\n- If a row has no reliable title and no resolvable DOI/PMID, mark `parser_error` and stop.\n- Use DOI verification first: `GET https://api.crossref.org/works/{doi}`. This is the primary source of DOI/title/author/journal/year truth.\n- Start enabled evidence lines together. Fast stops after primary evidence, Balanced allows a 2.5-second auxiliary grace period after primary DOI lookup and enabled supplied-PMID lookup, and Strict waits for every enabled line to complete or explicitly fail. Network timeout, rate limiting, not-found, parsing failure, mode skip, and expired budget must remain distinguishable.\n- If PubMed and OpenAlex are disabled with `--pubmed-mode off --openalex-mode off`, do not apply the grace timer to Crossref; Crossref DOI verification must complete because it is the primary authority.\n- OpenAlex checks use `GET https://api.openalex.org/works/doi:{doi}` or external IDs such as `pmid:{pmid}` for PMID, and `GET https://api.openalex.org/works?search={title}&per_page=3` for DOI-missing title recovery.\n- If an entry has no DOI, use title recovery in this order: Crossref title query first, OpenAlex title query only if the Crossref match is weak, PubMed title query only if biomedical corroboration is still needed. Record recovered DOI values in the report and ask before adding them to formatted citations.\n- Run quality-checked DOI-missing title recovery before full identifier verification, deduplicated and with provider-specific bounded concurrency. A recovered DOI may drive verification but must not silently modify the user source.\n- Do not use title search to silently override an existing DOI result. Existing DOI conflicts should be reported as conflicts, hijacking, or shifted identifiers.\n- Run near official limits without exceeding them: Crossref polite requests use no more than 10 requests/second and 3 concurrent requests; PubMed E-utilities uses no more than 3 requests/second without an API key and 10 requests/second with an NCBI API key; OpenAlex is budget/cost governed, so keep default concurrency conservative unless an API key and live limit headers justify more.\n- Use one bounded retry only for timeout, temporary connection failure, HTTP 429, or HTTP 5xx. Respect `Retry-After` within a bounded wait. Do not retry permanent 4xx responses. If a PubMed DOI batch fails, split it recursively until the failing DOI is isolated.\n- For large DOI-heavy documents, Crossref parallel DOI lookup is the primary full-batch path. PubMed and OpenAlex are corroboration channels, not reasons to block report generation.\n- If DOI/PMID resolves and title similarity is < 0.70, stop trusting that identifier for the current entry and mark `identifier_hijacking`.\n- Apply the < 0.70 hijacking rule only when the source title passed input quality checks.\n- Before finalizing `identifier_hijacking`, compare the resolved title against entries within two positions. If adjacent similarity is >= 0.86, mark `shifted_identifier`.\n- If several consecutive identifiers are shifted by the same offset, report a possible column/line offset instead of repairing each item manually.\n- For entries with DOI, run DOI lookup and report the DOI result; do not automatically title-search to replace the DOI.\n- For entries without DOI, run at most one Crossref title recovery query. If title recovery fails, mark `unresolved`, `placeholder_generation`, or `total_fabrication`; do not keep prompting the model to guess.\n- Do not run title recovery when `source_title` failed input quality checks. Rebuild the worksheet row first.\n- Use backup channels conditionally: Semantic Scholar only for title/abstract/citation corroboration after default lines are inconclusive; Europe PMC for biomedical PMCID/full-text evidence; bioRxiv/medRxiv for `10.1101/...` and published-DOI mapping; DataCite for non-journal dataset/software/report DOIs; OpenCitations for citation-network existence checks.\n- Parse every channel response by code into one canonical record shape: `source`, `title`, `authors`, `journal`, `year`, `doi`, `pmid`, `url`, and `score`. AI must not manually interpret structured API responses when parser adapters are available.\n- If the severe-error rate is high, stop after batch recovery and ask the user whether to continue AI-assisted per-paper search.\n\n## Network risk policy\n\n- Default DOI/PMID verification is low risk and should run without asking the user each time. The payload is limited to public identifiers required for the skill's purpose.\n- Default requests must not include `original_text`, manuscript paragraphs, abstracts, local evidence notes, or unpublished claims.\n- DOI-missing recovery may send only a short, quality-checked `source.title`.\n- Any deep search that sends title plus abstract, surrounding context, or manuscript-derived claims is higher risk and must be separated from default batch verification.\n- If Codex sandbox approval blocks external lookup, continue with local records normalization and report that external evidence was not executed.\n\n## Channel parser contract\n\n- Crossref `/works/{doi}` returns a `message` object; parse `title`, `author`, `container-title`, `issued`/published dates, `DOI`, and `URL`.\n- PubMed EFetch returns XML; parse `ArticleTitle`, `Author`, `Journal`, `PubDate`, `ArticleId IdType=\"doi\"`, and `PMID`.\n- OpenAlex works return work objects; parse `display_name`/`title`, `authorships`, `primary_location.source.display_name`, `publication_year`, `doi`, and `ids.pmid`.\n- Europe PMC search results return result objects; parse `title`, `authorString`, `journalTitle`, `pubYear`, `doi`, `pmid`, and `pmcid`.\n- Semantic Scholar Graph results return paper objects; parse `title`, `authors`, `venue`/`journal`, `year`, `externalIds.DOI`, `externalIds.PubMed`, and `url`.\n- DataCite DOI records return JSON:API objects; parse `data.attributes.titles`, `creators`, `publisher`/`container`, `publicationYear`, `doi`, and `url`.\n- bioRxiv/medRxiv API results return collection items; parse `title`, `authors`, `date`, `doi`, and `published_doi`.\n\n## Auto-fix policy\n\nAuto-fix only when canonical evidence is strong:\n\n- normalize DOI casing and links\n- strip trailing DOI punctuation\n- normalize DOI links in reports\n- record missing DOI candidates from canonical metadata and ask before appending them\n- include PMID only as secondary metadata when DOI-backed evidence exists\n- standardize journal title or abbreviation\n- normalize author initials and year formatting\n- report shifted DOI/PMID as a candidate relationship; retain the original until the proposed move is reviewed\n\nDo not auto-delete or silently replace severe items. For identifier-only, shifted, hijacked, fabricated, unresolved, and any partial/conflicting records, retain the original and report. Only verified/minor results may be automatically reformatted. Keep per-field before/after evidence.\n\nDo not create hidden persistent caches. Reuse prior results only from an explicitly supplied structured verifier artifact, and re-query new, changed, or incomplete records.\n\n## Format policy\n\n- Detect citation style before repair: APA, GB/T 7714, Vancouver/AMA, free-text, or mixed.\n- Preserve original format in the automatic fixed copy when possible.\n- If the list is mixed, recommend standardizing to the majority format after authenticity checks.\n- Formatting correctness is never authenticity evidence.\n\n## Body-vs-bibliography audit\n\nOnly execute this section when the user explicitly requests manuscript-body checking or cleanup. Bibliography verification alone does not authorize changing body claims.\n\n- Scan body text for narrative citations like `Author et al. (2023)` and parenthetical citations like `(Author, 2023)`.\n- Compare author/year/title semantics, not just numbered reference markers.\n- After deleting or flagging a fabricated reference, remove or flag body claims that depend on it.\n- Report uncertain claims separately instead of silently preserving them.\n\n## Eligibility, reuse and cleanup\n\n- The user requires a positive-use gate: references that cannot be verified are not usable. Identifier existence alone is insufficient. Formatting-only results are not authenticity passes. Preserve detailed reasons alongside this gate.\n- Results record verification coverage, eligibility, repair state and field differences independently. HTML groups are navigation aids, not replacements for these fields.\n- Reuse only complete compatible evidence checked within 30 days under the current policy, with all source fields unchanged. Stronger modes and additional channels require fresh checks. Regenerate formatted output under the current options.\n- Clean current-run redundant parser views/extracted copies by default; retain normalized source records and requested evidence artifacts. Honor `--cleanup-process-files none`. Ask after each round about retaining the remaining process files; no answer means keep. Never erase prior evidence just because an output option was omitted.\n\n## Real-reference matching details\n\nTrailing `et al.` denotes author truncation, not a named author. Compare the named prefix after conservative typographic normalization. Do not approve a different second or third author. Parse PubMed collective authors. A title containing commas is valid unless the comma-separated pieces actually have author-name syntax.\n\nRecognized publisher placeholder titles are inadequate metadata. A matching record from an enabled channel for the same DOI may supply the bibliographic title, with the fallback recorded. Without such evidence, retain an unresolved result; do not declare hijacking from the placeholder.\n\nPage forms such as `E5503-12` and `E5503-E5512`, or `e136-e136` and `e136`, are equivalent. Distinct electronic article numbers are not normalized into one another. Same-DOI journal aliases and explicit provider publication-year values are evidence of legitimate variants. Retain provider-specific records in the audit JSON.\n\nFile v1.2.0:CHANGELOG.md\n\n# Changelog\n\n## 1.2.0 — 2026-09-14\n\nMinor release following public version 1.1.3. SKILL.md and the CLI both identify this release as 1.2.0.\n\n- Reject duplicate reference indices before verification or writing outputs, preventing silent loss or duplication of results.\n- Reject output paths that alias the source, an explicitly reused artifact, or another output, including symlinks and hardlinks.\n- Keep failed title queries distinguishable from completed searches with no match. Failed queries leave unresolved references unusable; successful empty searches retain the existing classification policy.\n- Do not cache failed title searches as successful empty results. Detect malformed search responses and incomplete PubMed record retrieval.\n- Include formatted-only entries in the Markdown detail report.\n- Update the verification policy version to 2026-09-14.1 so older evidence is rechecked.\n- Add 10 offline regression tests for these boundaries.\n- Add the MIT-0 license and exclude Python caches and macOS metadata from publication.\n\nFile v1.2.0:skill-card.md\n\n## Description:\n\nVerifies and normalizes biomedical and life-science reference lists by checking identifiers and bibliographic metadata, preserving evidence, and producing offline audit reports.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[emp-tca](https://clawhub.ai/user/emp-tca)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, researchers, medical writers, and reviewers use this skill to check whether biomedical references are authentic, identify corrupted or fabricated citation fields, and normalize citation style without overwriting the source document.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Ordinary verification can attach an ambient email address to third-party Crossref, PubMed, and OpenAlex reference queries.\n\nMitigation: Pass an explicit anonymous --email value unless a real contact address is intended.\n\nRisk: Confidential reference lists may disclose citation titles or identifiers to external metadata providers during verification.\n\nMitigation: Use format-only mode for local citation cleanup, or disable providers with the documented off flags when external verification is not appropriate.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/emp-tca/skills/biomedical-reference-verifier)\n- [README](README.md)\n- [Verification Policy](references/verification_policy.md)\n\n## Skill Output:\n\n**Output Type(s):** [Markdown, HTML, JSON, Text, Shell commands, Guidance]\n\n**Output Format:** [Markdown summaries and detail reports, a self-contained HTML audit report, JSON audit/index artifacts, fixed reference text files, and shell command guidance.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [The skill preserves source evidence, does not overwrite source documents, and can run format-only or with selected external metadata providers disabled.]\n\n## Skill Version(s):\n\n1.2.0 (source: frontmatter, changelog, release evidence; released 2026-09-14)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.2.0:LICENSE\n\nMIT No Attribution\n\nCopyright (c) 2026 EMP-TCA\n\nPermission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the \"Software\"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.\n\nArchive v1.1.3: 13 files, 46290 bytes\n\nFiles: references/verification_policy.md (11873b), scripts/convert_reference_artifact.py (1437b), scripts/self_test_verify_references.py (10332b), scripts/verifier_models.py (2044b), scripts/verifier_network.py (2739b), scripts/verifier_policy.py (849b), scripts/verifier_prior_results.py (5650b), scripts/verifier_recovery.py (859b), scripts/verifier_runtime.py (1604b), scripts/verify_references.py (115079b), skill-card.md (2624b), SKILL.md (7936b), _meta.json (148b)\n\nFile v1.1.3:SKILL.md\n\n---\nname: biomedical-reference-verifier\ndescription: \"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Makes your reference bibliography more reliable and convincing.. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。\"\nversion: 1.1.3\nmetadata:\n  openclaw:\n    requires:\n      bins:\n        - python3\n    envVars:\n      - name: NCBI_API_KEY\n        required: false\n        description: Optional NCBI API key for higher PubMed E-utilities rate limits.\n      - name: OPENALEX_API_KEY\n        required: false\n        description: Optional OpenAlex API key for authenticated metadata lookups.\n---\n\n# Biomedical Reference Verifier\n\nUse this skill only for reference-list authenticity checking, AI-caused reference cleanup, and citation-format normalization. Do not use it to judge whether papers are relevant, suitable, or correctly used in the manuscript body unless the user explicitly asks for reference authenticity plus body cleanup.\n\n## Core Rule\n\nAlways reduce the task to this sequence:\n\n1. Build or read `biomedical-reference-verifier.records.v1`.\n2. Choose the lightest pipeline.\n3. Run the script.\n4. Report result files and retained process files.\n5. Ask before expensive manual recovery or process-file deletion.\n\nLoad `references/verification_policy.md` before classifying authenticity errors, changing severe items, or explaining network/query policy.\n\n## Pipeline Choice\n\n- **Format only**: use when the user only asks to convert or unify citation style. Run `--pipeline format-only`. Do not query Crossref/PubMed/OpenAlex.\n- **Verify**: use when the user asks whether references are real, fake, wrong, AI-generated, DOI/PMID-mismatched, or metadata-corrupted. Run the default `verify` pipeline.\n- **Verify then format**: use when the user asks both to check authenticity and standardize style. Run `verify` with the requested `--citation-style`; the fixed copy is generated after verification.\n\n## Execution Mode\n\n- **Fast** (`--mode fast`): use for quick scans and large routine lists. Stop after primary DOI evidence when available.\n- **Balanced** (`--mode balanced`): default for ordinary authenticity checks. Auxiliary evidence lines receive a short 2.5-second grace period after the first line completes.\n- **Strict** (`--mode strict`): use for final or pre-submission checks. Wait for every enabled evidence line to complete or fail explicitly.\n\nDo not invent a local confidence score. For DOI-bearing references, resolve the DOI and directly compare returned title, authors, journal, year, DOI, and PMID with source fields. Keep the existing classification policy. Severe items are reported and retained; never auto-delete them.\n\nDo not run AI-assisted per-paper web searching by default. Stop after batch verification and title recovery, then ask the user whether to continue.\n\n## Machine Input\n\nPrefer `biomedical-reference-verifier.records.v1` JSON/JSONL. If the source is free text, convert it to records first using only values present in the user source. Do not fill missing source fields from Crossref, PubMed, OpenAlex, memory, or plausible guesses.\n\nMinimal record:\n\n```json\n{\n  \"schema\": \"biomedical-reference-verifier.records.v1\",\n  \"records\": [\n    {\n      \"index\": 1,\n      \"source\": {\n        \"original_text\": \"exact source reference\",\n        \"title\": \"title from source, or empty string\",\n        \"authors\": [\"First Author\"],\n        \"year\": \"2024\",\n        \"journal\": \"Journal from source\",\n        \"identifiers\": {\"doi\": \"10.xxxx/example\", \"pmid\": \"\", \"urls\": []},\n        \"source_lines\": [12],\n        \"context\": \"\"\n      }\n    }\n  ]\n}\n```\n\nUse `--input-mode records` for standardized JSON/JSONL. Use Markdown worksheet and `doi-context` modes only for compatibility.\n\n## Commands\n\nRun from this skill directory:\n\n```bash\npython3 scripts/verify_references.py refs.md --output-dir /tmp/reference-audit --citation-style ama\n```\n\nCommon variants:\n\n```bash\npython3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --mode fast\npython3 scripts/verify_references.py records.json --input-mode records --mode strict\npython3 scripts/verify_references.py records.json --input-mode records --reuse-results previous/reference-audit.json\npython3 scripts/verify_references.py records.json --input-mode records --write-index\npython3 scripts/convert_reference_artifact.py previous/reference-audit.json --to index --output reference-index.json\npython3 scripts/convert_reference_artifact.py reference-index.json --to audit --output restored-reference-audit.json\npython3 scripts/verify_references.py refs.md --pubmed-mode off --openalex-mode off\npython3 scripts/verify_references.py refs.md --doi-output append\npython3 scripts/verify_references.py refs.md --keep-process-json\npython3 scripts/verify_references.py refs.md --cleanup-process-files all\npython3 scripts/verify_references.py refs.md --cleanup-process-files normalized_input,extracted_references\n```\n\n## Output Contract\n\nResult files:\n\n- `reference-audit-summary.md`: concise chat-ready summary.\n- `reference-audit-detail.md`: detailed report with evidence links.\n- `references.auto-fixed.md` or `document.auto-fixed.md`: fixed/formatted copy; never overwrites the source.\n\nProcess files retained by default:\n\n- `reference-normalized-records.json`: machine-readable records.\n- `reference-normalized-input.md`: human-readable parser inspection table.\n- `references.extracted.md`: raw references extracted before verification.\n\nProcess files skipped by default:\n\n- `reference-audit.json`: write only with `--keep-process-json`.\n\nThere is no hidden persistent cache. `--reuse-results` accepts either `reference-audit.json` or schema `biomedical-reference-verifier.index.v1`. The converter supports audit-to-index and index-to-audit round trips; the `results` array must remain identical after a round trip. In a conversation, ask about reuse once when such an artifact is available; do not repeatedly ask after the user decides.\n\nUser-edited prior artifacts are handled conservatively: malformed JSON is rejected with one concise error; invalid individual rows are skipped and rechecked; valid rows remain reusable. Do not guess repairs for damaged structured fields.\n\nPrior-result reuse tolerates punctuation changes only when strict bibliographic fields still agree: DOI+title+year, or title+first author+year+journal. Temporary network failures receive one bounded retry; permanent 4xx failures do not. PubMed DOI batches split only when a batch fails. Early DOI recovery uses Crossref first, then OpenAlex and PubMed outside Fast mode.\n\n`--write-index` optionally writes `reference-index.json`; it is off by default. Network request events and phase timings are stored in `reference-audit.json` when `--keep-process-json` is used.\n\nNever delete user-requested result files. Delete only process files generated by the script in the selected output directory, and only when the user chooses deletion or `--cleanup-process-files` is explicit.\n\n## Closeout\n\nAfter running:\n\n1. Paste or summarize `reference-audit-summary.md`.\n2. Link result files and retained process files.\n3. State which process files were cleaned or skipped.\n4. Ask whether to delete retained process files:\n   - A. delete all retained process files\n   - B. keep selected process files\n   - C. keep all process files for now\n5. If severe items exist, ask whether to continue AI-assisted recovery, delete, keep with warning notes, or stop with the report.\n6. Ask whether recovered DOI values should be appended when DOI output is optional.\n\nFile v1.1.3:_meta.json\n\n{\n  \"ownerId\": \"kn748sfft0aksxdrw9332r67w58a99n2\",\n  \"slug\": \"biomedical-reference-verifier\",\n  \"version\": \"1.1.3\",\n  \"publishedAt\": 1783955225510\n}\n\nFile v1.1.3:references/verification_policy.md\n\n# Verification Policy\n\n## Evidence hierarchy\n\n1. Crossref is the fastest and strongest default primary source for journal DOI, title, authors, journal, publisher, and year.\n2. PubMed is the second evidence line for biomedical records, PMID mapping, and wrong PMID/title-pair detection.\n3. OpenAlex is the third default evidence line for fast DOI/PMID external-ID corroboration and title recovery when Crossref is weak.\n4. A DOI match from Crossref is the primary batch verification result. PubMed and OpenAlex matches can corroborate it, but should not block reporting.\n5. Semantic Scholar, Europe PMC, bioRxiv/medRxiv, DataCite, and OpenCitations are backup channels, not default full-batch channels.\n6. Search snippets, AI summaries, formatted APA/GB/T entries, and plausible journal names are not proof.\n7. `biomedical-reference-verifier.records.v1` is the source-of-truth input layer. Source fields must come from the user's text; database-returned fields must stay in evidence/report/output objects.\n\n## Error types\n\n- `verified`: DOI/PMID or title search resolves to a canonical record and title, year, and first-author evidence agree.\n- `verified_identifier_only`: DOI/PMID resolves to a canonical record, but the source title is missing or unreliable, so only the identifier is verified.\n- `parser_error`: the input row is malformed or the parsed/source title is an author line, table header, metadata field, or other non-title text; stop before treating the row as a bibliographic conflict.\n- `minor_fix` / `minor_format_error`: the paper is real and metadata agree; only DOI casing, punctuation, URL style, initials, journal abbreviation, or minor year formatting needs correction.\n- `partial_attribute_corruption`: the paper is real, but one or more attributes are corrupted, such as author, title, journal, year, volume/pages, DOI, or PMID.\n- `identifier_hijacking`: DOI, PMID, or URL is real, but it points to a different paper than the reference text.\n- `shifted_identifier`: DOI or PMID appears to belong to a nearby reference, commonly because one entry's DOI was attached to the previous or next citation.\n- `semantic_hallucination`: the paper exists, but a surrounding claim or summary is not supported by that paper.\n- `placeholder_generation`: the entry contains AI-like placeholder traces: generic title, invented DOI suffix, fake pages/volume, missing journal fields, or overly tidy but unverifiable metadata.\n- `total_fabrication`: title, authors, journal, year, DOI/PMID, and searches fail to identify any canonical record.\n- `unresolved`: automatic checks are insufficient; stop and ask before expensive manual recovery.\n\n## Matching thresholds\n\n- Title similarity >= 0.90: strong match.\n- Title similarity 0.86-0.89: acceptable match if year and first author agree.\n- Title similarity 0.78-0.85: possible match; mark as partial unless corroborated by DOI/PMID.\n- Title similarity < 0.70 for a supplied DOI/PMID: identifier hijacking unless adjacent-entry checks prove shifted identifier.\n- Year difference of 0-1 year can be online-first drift. Larger differences are corruption unless the source explains it.\n- First-author disagreement is a warning; combine it with title/year/journal evidence before deciding severity.\n\n## Short-circuit rules\n\n- Build `biomedical-reference-verifier.records.v1` first. Do not let remote lookup run against raw mixed notes when the source can be converted to explicit source objects.\n- Required machine fields are `index`, `source.original_text`, `source.title`, `source.authors`, `source.year`, `source.journal`, `source.identifiers.doi`, `source.identifiers.pmid`, `source.identifiers.urls`, `source.source_lines`, and `source.context`.\n- During machine-record construction, AI may only extract values present in the source text. It must not use Crossref/PubMed/OpenAlex results, memory, or plausible completions to fill source fields.\n- Build `reference-normalized-records.json` before remote lookup. `reference-normalized-input.md` is a human-readable view of the same records.\n- If `source.title` is missing, too short, an author list, a table/header label, a DOI/PMID line, or metadata placeholder, mark the title unreliable.\n- If a DOI/PMID resolves but title is unreliable, mark `verified_identifier_only`; do not mark `identifier_hijacking`.\n- If a row has no reliable title and no resolvable DOI/PMID, mark `parser_error` and stop.\n- Use DOI verification first: `GET https://api.crossref.org/works/{doi}`. This is the primary source of DOI/title/author/journal/year truth.\n- Start enabled evidence lines together. Fast stops after primary evidence, Balanced allows a 2.5-second auxiliary grace period, and Strict waits for every enabled line to complete or explicitly fail. Network timeout, rate limiting, not-found, parsing failure, mode skip, and expired budget must remain distinguishable.\n- If PubMed and OpenAlex are disabled with `--pubmed-mode off --openalex-mode off`, do not apply the grace timer to Crossref; Crossref DOI verification must complete because it is the primary authority.\n- OpenAlex checks use `GET https://api.openalex.org/works/doi:{doi}` or external IDs such as `pmid:{pmid}` for PMID, and `GET https://api.openalex.org/works?search={title}&per_page=3` for DOI-missing title recovery.\n- If an entry has no DOI, use title recovery in this order: Crossref title query first, OpenAlex title query only if the Crossref match is weak, PubMed title query only if biomedical corroboration is still needed. Record recovered DOI values in the report and ask before adding them to formatted citations.\n- Run quality-checked DOI-missing title recovery before full identifier verification, deduplicated and with provider-specific bounded concurrency. A recovered DOI may drive verification but must not silently modify the user source.\n- Do not use title search to silently override an existing DOI result. Existing DOI conflicts should be reported as conflicts, hijacking, or shifted identifiers.\n- Run near official limits without exceeding them: Crossref polite requests use no more than 10 requests/second and 3 concurrent requests; PubMed E-utilities uses no more than 3 requests/second without an API key and 10 requests/second with an NCBI API key; OpenAlex is budget/cost governed, so keep default concurrency conservative unless an API key and live limit headers justify more.\n- Use one bounded retry only for timeout, temporary connection failure, HTTP 429, or HTTP 5xx. Respect `Retry-After` within a bounded wait. Do not retry permanent 4xx responses. If a PubMed DOI batch fails, split it recursively until the failing DOI is isolated.\n- For large DOI-heavy documents, Crossref parallel DOI lookup is the primary full-batch path. PubMed and OpenAlex are corroboration channels, not reasons to block report generation.\n- If DOI/PMID resolves and title similarity is < 0.70, stop trusting that identifier for the current entry and mark `identifier_hijacking`.\n- Apply the < 0.70 hijacking rule only when the source title passed input quality checks.\n- Before finalizing `identifier_hijacking`, compare the resolved title against entries within two positions. If adjacent similarity is >= 0.86, mark `shifted_identifier`.\n- If several consecutive identifiers are shifted by the same offset, report a possible column/line offset instead of repairing each item manually.\n- For entries with DOI, run DOI lookup and report the DOI result; do not automatically title-search to replace the DOI.\n- For entries without DOI, run at most one Crossref title recovery query. If title recovery fails, mark `unresolved`, `placeholder_generation`, or `total_fabrication`; do not keep prompting the model to guess.\n- Do not run title recovery when `source_title` failed input quality checks. Rebuild the worksheet row first.\n- Use backup channels conditionally: Semantic Scholar only for title/abstract/citation corroboration after default lines are inconclusive; Europe PMC for biomedical PMCID/full-text evidence; bioRxiv/medRxiv for `10.1101/...` and published-DOI mapping; DataCite for non-journal dataset/software/report DOIs; OpenCitations for citation-network existence checks.\n- Parse every channel response by code into one canonical record shape: `source`, `title`, `authors`, `journal`, `year`, `doi`, `pmid`, `url`, and `score`. AI must not manually interpret structured API responses when parser adapters are available.\n- If the severe-error rate is high, stop after batch recovery and ask the user whether to continue AI-assisted per-paper search.\n\n## Network risk policy\n\n- Default DOI/PMID verification is low risk and should run without asking the user each time. The payload is limited to public identifiers required for the skill's purpose.\n- Default requests must not include `original_text`, manuscript paragraphs, abstracts, local evidence notes, or unpublished claims.\n- DOI-missing recovery may send only a short, quality-checked `source.title`.\n- Any deep search that sends title plus abstract, surrounding context, or manuscript-derived claims is higher risk and must be separated from default batch verification.\n- If Codex sandbox approval blocks external lookup, continue with local records normalization and report that external evidence was not executed.\n\n## Channel parser contract\n\n- Crossref `/works/{doi}` returns a `message` object; parse `title`, `author`, `container-title`, `issued`/published dates, `DOI`, and `URL`.\n- PubMed EFetch returns XML; parse `ArticleTitle`, `Author`, `Journal`, `PubDate`, `ArticleId IdType=\"doi\"`, and `PMID`.\n- OpenAlex works return work objects; parse `display_name`/`title`, `authorships`, `primary_location.source.display_name`, `publication_year`, `doi`, and `ids.pmid`.\n- Europe PMC search results return result objects; parse `title`, `authorString`, `journalTitle`, `pubYear`, `doi`, `pmid`, and `pmcid`.\n- Semantic Scholar Graph results return paper objects; parse `title`, `authors`, `venue`/`journal`, `year`, `externalIds.DOI`, `externalIds.PubMed`, and `url`.\n- DataCite DOI records return JSON:API objects; parse `data.attributes.titles`, `creators`, `publisher`/`container`, `publicationYear`, `doi`, and `url`.\n- bioRxiv/medRxiv API results return collection items; parse `title`, `authors`, `date`, `doi`, and `published_doi`.\n\n## Auto-fix policy\n\nAuto-fix only when canonical evidence is strong:\n\n- normalize DOI casing and links\n- strip trailing DOI punctuation\n- normalize DOI links in reports\n- record missing DOI candidates from canonical metadata and ask before appending them\n- include PMID only as secondary metadata when DOI-backed evidence exists\n- standardize journal title or abbreviation\n- normalize author initials and year formatting\n- move a shifted DOI/PMID only when nearby title similarity is strong\n\nDo not auto-delete or silently replace severe items. For `identifier_hijacking`, `total_fabrication`, and low-confidence `partial_attribute_corruption`, report and ask the user.\n\nDo not create hidden persistent caches. Reuse prior results only from an explicitly supplied structured verifier artifact, and re-query new, changed, or incomplete records.\n\n## Format policy\n\n- Detect citation style before repair: APA, GB/T 7714, Vancouver/AMA, free-text, or mixed.\n- Preserve original format in the automatic fixed copy when possible.\n- If the list is mixed, recommend standardizing to the majority format after authenticity checks.\n- Formatting correctness is never authenticity evidence.\n\n## Body-vs-bibliography audit\n\n- Scan body text for narrative citations like `Author et al. (2023)` and parenthetical citations like `(Author, 2023)`.\n- Compare author/year/title semantics, not just numbered reference markers.\n- After deleting or flagging a fabricated reference, remove or flag body claims that depend on it.\n- Report uncertain claims separately instead of silently preserving them.\n\nFile v1.1.3:skill-card.md\n\n## Description: <br>\nVerify or normalize biomedical and life-science reference lists when the task is about AI-caused reference errors, making bibliographies more reliable and convincing. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[emp-tca](https://clawhub.ai/user/emp-tca) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nResearchers, manuscript authors, editors, and developers use this skill to verify biomedical reference lists for fabricated or corrupted citations, normalize citation formats, and produce audit reports plus auto-fixed bibliography files. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Verification can send citation identifiers and sometimes short reference titles to public citation services such as Crossref, PubMed/NCBI, and OpenAlex. <br>\nMitigation: Use format-only mode for private documents, or disable verification channels that should not receive citation metadata. <br>\nRisk: Optional API keys or email settings may be sent to external services during authenticated or polite metadata lookups. <br>\nMitigation: Avoid setting NCBI_API_KEY, OPENALEX_API_KEY, or USER_EMAIL unless higher rate limits or authenticated lookups are needed. <br>\nRisk: Audit classifications and auto-fixed bibliographies can still contain mistakes or unresolved severe citation issues. <br>\nMitigation: Review generated reports before using the fixed bibliography, and keep severe or unresolved items for manual confirmation. <br>\n\n\n## Reference(s): <br>\n- [Verification Policy](references/verification_policy.md) <br>\n- [Biomedical Reference Verifier on ClawHub](https://clawhub.ai/emp-tca/skills/biomedical-reference-verifier) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] <br>\n**Output Format:** [Markdown reports, JSON or JSONL records, fixed bibliography files, and shell command guidance] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Writes reference-audit-summary.md, reference-audit-detail.md, references.auto-fixed.md or document.auto-fixed.md, and optional process JSON/index artifacts without overwriting the source.] <br>\n\n## Skill Version(s): <br>\n1.1.3 (source: frontmatter and server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.1.2: 13 files, 46299 bytes\n\nFiles: references/verification_policy.md (11873b), scripts/convert_reference_artifact.py (1437b), scripts/self_test_verify_references.py (10332b), scripts/verifier_models.py (2044b), scripts/verifier_network.py (2739b), scripts/verifier_policy.py (849b), scripts/verifier_prior_results.py (5650b), scripts/verifier_recovery.py (859b), scripts/verifier_runtime.py (1604b), scripts/verify_references.py (115079b), skill-card.md (2614b), SKILL.md (7936b), _meta.json (148b)\n\nFile v1.1.2:SKILL.md\n\n---\nname: biomedical-reference-verifier\ndescription: \"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Makes your reference bibliography more reliable and convincing.. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。\"\nversion: 1.1.2\nmetadata:\n  openclaw:\n    requires:\n      bins:\n        - python3\n    envVars:\n      - name: NCBI_API_KEY\n        required: false\n        description: Optional NCBI API key for higher PubMed E-utilities rate limits.\n      - name: OPENALEX_API_KEY\n        required: false\n        description: Optional OpenAlex API key for authenticated metadata lookups.\n---\n\n# Biomedical Reference Verifier\n\nUse this skill only for reference-list authenticity checking, AI-caused reference cleanup, and citation-format normalization. Do not use it to judge whether papers are relevant, suitable, or correctly used in the manuscript body unless the user explicitly asks for reference authenticity plus body cleanup.\n\n## Core Rule\n\nAlways reduce the task to this sequence:\n\n1. Build or read `biomedical-reference-verifier.records.v1`.\n2. Choose the lightest pipeline.\n3. Run the script.\n4. Report result files and retained process files.\n5. Ask before expensive manual recovery or process-file deletion.\n\nLoad `references/verification_policy.md` before classifying authenticity errors, changing severe items, or explaining network/query policy.\n\n## Pipeline Choice\n\n- **Format only**: use when the user only asks to convert or unify citation style. Run `--pipeline format-only`. Do not query Crossref/PubMed/OpenAlex.\n- **Verify**: use when the user asks whether references are real, fake, wrong, AI-generated, DOI/PMID-mismatched, or metadata-corrupted. Run the default `verify` pipeline.\n- **Verify then format**: use when the user asks both to check authenticity and standardize style. Run `verify` with the requested `--citation-style`; the fixed copy is generated after verification.\n\n## Execution Mode\n\n- **Fast** (`--mode fast`): use for quick scans and large routine lists. Stop after primary DOI evidence when available.\n- **Balanced** (`--mode balanced`): default for ordinary authenticity checks. Auxiliary evidence lines receive a short 2.5-second grace period after the first line completes.\n- **Strict** (`--mode strict`): use for final or pre-submission checks. Wait for every enabled evidence line to complete or fail explicitly.\n\nDo not invent a local confidence score. For DOI-bearing references, resolve the DOI and directly compare returned title, authors, journal, year, DOI, and PMID with source fields. Keep the existing classification policy. Severe items are reported and retained; never auto-delete them.\n\nDo not run AI-assisted per-paper web searching by default. Stop after batch verification and title recovery, then ask the user whether to continue.\n\n## Machine Input\n\nPrefer `biomedical-reference-verifier.records.v1` JSON/JSONL. If the source is free text, convert it to records first using only values present in the user source. Do not fill missing source fields from Crossref, PubMed, OpenAlex, memory, or plausible guesses.\n\nMinimal record:\n\n```json\n{\n  \"schema\": \"biomedical-reference-verifier.records.v1\",\n  \"records\": [\n    {\n      \"index\": 1,\n      \"source\": {\n        \"original_text\": \"exact source reference\",\n        \"title\": \"title from source, or empty string\",\n        \"authors\": [\"First Author\"],\n        \"year\": \"2024\",\n        \"journal\": \"Journal from source\",\n        \"identifiers\": {\"doi\": \"10.xxxx/example\", \"pmid\": \"\", \"urls\": []},\n        \"source_lines\": [12],\n        \"context\": \"\"\n      }\n    }\n  ]\n}\n```\n\nUse `--input-mode records` for standardized JSON/JSONL. Use Markdown worksheet and `doi-context` modes only for compatibility.\n\n## Commands\n\nRun from this skill directory:\n\n```bash\npython3 scripts/verify_references.py refs.md --output-dir /tmp/reference-audit --citation-style ama\n```\n\nCommon variants:\n\n```bash\npython3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --mode fast\npython3 scripts/verify_references.py records.json --input-mode records --mode strict\npython3 scripts/verify_references.py records.json --input-mode records --reuse-results previous/reference-audit.json\npython3 scripts/verify_references.py records.json --input-mode records --write-index\npython3 scripts/convert_reference_artifact.py previous/reference-audit.json --to index --output reference-index.json\npython3 scripts/convert_reference_artifact.py reference-index.json --to audit --output restored-reference-audit.json\npython3 scripts/verify_references.py refs.md --pubmed-mode off --openalex-mode off\npython3 scripts/verify_references.py refs.md --doi-output append\npython3 scripts/verify_references.py refs.md --keep-process-json\npython3 scripts/verify_references.py refs.md --cleanup-process-files all\npython3 scripts/verify_references.py refs.md --cleanup-process-files normalized_input,extracted_references\n```\n\n## Output Contract\n\nResult files:\n\n- `reference-audit-summary.md`: concise chat-ready summary.\n- `reference-audit-detail.md`: detailed report with evidence links.\n- `references.auto-fixed.md` or `document.auto-fixed.md`: fixed/formatted copy; never overwrites the source.\n\nProcess files retained by default:\n\n- `reference-normalized-records.json`: machine-readable records.\n- `reference-normalized-input.md`: human-readable parser inspection table.\n- `references.extracted.md`: raw references extracted before verification.\n\nProcess files skipped by default:\n\n- `reference-audit.json`: write only with `--keep-process-json`.\n\nThere is no hidden persistent cache. `--reuse-results` accepts either `reference-audit.json` or schema `biomedical-reference-verifier.index.v1`. The converter supports audit-to-index and index-to-audit round trips; the `results` array must remain identical after a round trip. In a conversation, ask about reuse once when such an artifact is available; do not repeatedly ask after the user decides.\n\nUser-edited prior artifacts are handled conservatively: malformed JSON is rejected with one concise error; invalid individual rows are skipped and rechecked; valid rows remain reusable. Do not guess repairs for damaged structured fields.\n\nPrior-result reuse tolerates punctuation changes only when strict bibliographic fields still agree: DOI+title+year, or title+first author+year+journal. Temporary network failures receive one bounded retry; permanent 4xx failures do not. PubMed DOI batches split only when a batch fails. Early DOI recovery uses Crossref first, then OpenAlex and PubMed outside Fast mode.\n\n`--write-index` optionally writes `reference-index.json`; it is off by default. Network request events and phase timings are stored in `reference-audit.json` when `--keep-process-json` is used.\n\nNever delete user-requested result files. Delete only process files generated by the script in the selected output directory, and only when the user chooses deletion or `--cleanup-process-files` is explicit.\n\n## Closeout\n\nAfter running:\n\n1. Paste or summarize `reference-audit-summary.md`.\n2. Link result files and retained process files.\n3. State which process files were cleaned or skipped.\n4. Ask whether to delete retained process files:\n   - A. delete all retained process files\n   - B. keep selected process files\n   - C. keep all process files for now\n5. If severe items exist, ask whether to continue AI-assisted recovery, delete, keep with warning notes, or stop with the report.\n6. Ask whether recovered DOI values should be appended when DOI output is optional.\n\nFile v1.1.2:_meta.json\n\n{\n  \"ownerId\": \"kn748sfft0aksxdrw9332r67w58a99n2\",\n  \"slug\": \"biomedical-reference-verifier\",\n  \"version\": \"1.1.2\",\n  \"publishedAt\": 1783791255695\n}\n\nFile v1.1.2:references/verification_policy.md\n\n# Verification Policy\n\n## Evidence hierarchy\n\n1. Crossref is the fastest and strongest default primary source for journal DOI, title, authors, journal, publisher, and year.\n2. PubMed is the second evidence line for biomedical records, PMID mapping, and wrong PMID/title-pair detection.\n3. OpenAlex is the third default evidence line for fast DOI/PMID external-ID corroboration and title recovery when Crossref is weak.\n4. A DOI match from Crossref is the primary batch verification result. PubMed and OpenAlex matches can corroborate it, but should not block reporting.\n5. Semantic Scholar, Europe PMC, bioRxiv/medRxiv, DataCite, and OpenCitations are backup channels, not default full-batch channels.\n6. Search snippets, AI summaries, formatted APA/GB/T entries, and plausible journal names are not proof.\n7. `biomedical-reference-verifier.records.v1` is the source-of-truth input layer. Source fields must come from the user's text; database-returned fields must stay in evidence/report/output objects.\n\n## Error types\n\n- `verified`: DOI/PMID or title search resolves to a canonical record and title, year, and first-author evidence agree.\n- `verified_identifier_only`: DOI/PMID resolves to a canonical record, but the source title is missing or unreliable, so only the identifier is verified.\n- `parser_error`: the input row is malformed or the parsed/source title is an author line, table header, metadata field, or other non-title text; stop before treating the row as a bibliographic conflict.\n- `minor_fix` / `minor_format_error`: the paper is real and metadata agree; only DOI casing, punctuation, URL style, initials, journal abbreviation, or minor year formatting needs correction.\n- `partial_attribute_corruption`: the paper is real, but one or more attributes are corrupted, such as author, title, journal, year, volume/pages, DOI, or PMID.\n- `identifier_hijacking`: DOI, PMID, or URL is real, but it points to a different paper than the reference text.\n- `shifted_identifier`: DOI or PMID appears to belong to a nearby reference, commonly because one entry's DOI was attached to the previous or next citation.\n- `semantic_hallucination`: the paper exists, but a surrounding claim or summary is not supported by that paper.\n- `placeholder_generation`: the entry contains AI-like placeholder traces: generic title, invented DOI suffix, fake pages/volume, missing journal fields, or overly tidy but unverifiable metadata.\n- `total_fabrication`: title, authors, journal, year, DOI/PMID, and searches fail to identify any canonical record.\n- `unresolved`: automatic checks are insufficient; stop and ask before expensive manual recovery.\n\n## Matching thresholds\n\n- Title similarity >= 0.90: strong match.\n- Title similarity 0.86-0.89: acceptable match if year and first author agree.\n- Title similarity 0.78-0.85: possible match; mark as partial unless corroborated by DOI/PMID.\n- Title similarity < 0.70 for a supplied DOI/PMID: identifier hijacking unless adjacent-entry checks prove shifted identifier.\n- Year difference of 0-1 year can be online-first drift. Larger differences are corruption unless the source explains it.\n- First-author disagreement is a warning; combine it with title/year/journal evidence before deciding severity.\n\n## Short-circuit rules\n\n- Build `biomedical-reference-verifier.records.v1` first. Do not let remote lookup run against raw mixed notes when the source can be converted to explicit source objects.\n- Required machine fields are `index`, `source.original_text`, `source.title`, `source.authors`, `source.year`, `source.journal`, `source.identifiers.doi`, `source.identifiers.pmid`, `source.identifiers.urls`, `source.source_lines`, and `source.context`.\n- During machine-record construction, AI may only extract values present in the source text. It must not use Crossref/PubMed/OpenAlex results, memory, or plausible completions to fill source fields.\n- Build `reference-normalized-records.json` before remote lookup. `reference-normalized-input.md` is a human-readable view of the same records.\n- If `source.title` is missing, too short, an author list, a table/header label, a DOI/PMID line, or metadata placeholder, mark the title unreliable.\n- If a DOI/PMID resolves but title is unreliable, mark `verified_identifier_only`; do not mark `identifier_hijacking`.\n- If a row has no reliable title and no resolvable DOI/PMID, mark `parser_error` and stop.\n- Use DOI verification first: `GET https://api.crossref.org/works/{doi}`. This is the primary source of DOI/title/author/journal/year truth.\n- Start enabled evidence lines together. Fast stops after primary evidence, Balanced allows a 2.5-second auxiliary grace period, and Strict waits for every enabled line to complete or explicitly fail. Network timeout, rate limiting, not-found, parsing failure, mode skip, and expired budget must remain distinguishable.\n- If PubMed and OpenAlex are disabled with `--pubmed-mode off --openalex-mode off`, do not apply the grace timer to Crossref; Crossref DOI verification must complete because it is the primary authority.\n- OpenAlex checks use `GET https://api.openalex.org/works/doi:{doi}` or external IDs such as `pmid:{pmid}` for PMID, and `GET https://api.openalex.org/works?search={title}&per_page=3` for DOI-missing title recovery.\n- If an entry has no DOI, use title recovery in this order: Crossref title query first, OpenAlex title query only if the Crossref match is weak, PubMed title query only if biomedical corroboration is still needed. Record recovered DOI values in the report and ask before adding them to formatted citations.\n- Run quality-checked DOI-missing title recovery before full identifier verification, deduplicated and with provider-specific bounded concurrency. A recovered DOI may drive verification but must not silently modify the user source.\n- Do not use title search to silently override an existing DOI result. Existing DOI conflicts should be reported as conflicts, hijacking, or shifted identifiers.\n- Run near official limits without exceeding them: Crossref polite requests use no more than 10 requests/second and 3 concurrent requests; PubMed E-utilities uses no more than 3 requests/second without an API key and 10 requests/second with an NCBI API key; OpenAlex is budget/cost governed, so keep default concurrency conservative unless an API key and live limit headers justify more.\n- Use one bounded retry only for timeout, temporary connection failure, HTTP 429, or HTTP 5xx. Respect `Retry-After` within a bounded wait. Do not retry permanent 4xx responses. If a PubMed DOI batch fails, split it recursively until the failing DOI is isolated.\n- For large DOI-heavy documents, Crossref parallel DOI lookup is the primary full-batch path. PubMed and OpenAlex are corroboration channels, not reasons to block report generation.\n- If DOI/PMID resolves and title similarity is < 0.70, stop trusting that identifier for the current entry and mark `identifier_hijacking`.\n- Apply the < 0.70 hijacking rule only when the source title passed input quality checks.\n- Before finalizing `identifier_hijacking`, compare the resolved title against entries within two positions. If adjacent similarity is >= 0.86, mark `shifted_identifier`.\n- If several consecutive identifiers are shifted by the same offset, report a possible column/line offset instead of repairing each item manually.\n- For entries with DOI, run DOI lookup and report the DOI result; do not automatically title-search to replace the DOI.\n- For entries without DOI, run at most one Crossref title recovery query. If title recovery fails, mark `unresolved`, `placeholder_generation`, or `total_fabrication`; do not keep prompting the model to guess.\n- Do not run title recovery when `source_title` failed input quality checks. Rebuild the worksheet row first.\n- Use backup channels conditionally: Semantic Scholar only for title/abstract/citation corroboration after default lines are inconclusive; Europe PMC for biomedical PMCID/full-text evidence; bioRxiv/medRxiv for `10.1101/...` and published-DOI mapping; DataCite for non-journal dataset/software/report DOIs; OpenCitations for citation-network existence checks.\n- Parse every channel response by code into one canonical record shape: `source`, `title`, `authors`, `journal`, `year`, `doi`, `pmid`, `url`, and `score`. AI must not manually interpret structured API responses when parser adapters are available.\n- If the severe-error rate is high, stop after batch recovery and ask the user whether to continue AI-assisted per-paper search.\n\n## Network risk policy\n\n- Default DOI/PMID verification is low risk and should run without asking the user each time. The payload is limited to public identifiers required for the skill's purpose.\n- Default requests must not include `original_text`, manuscript paragraphs, abstracts, local evidence notes, or unpublished claims.\n- DOI-missing recovery may send only a short, quality-checked `source.title`.\n- Any deep search that sends title plus abstract, surrounding context, or manuscript-derived claims is higher risk and must be separated from default batch verification.\n- If Codex sandbox approval blocks external lookup, continue with local records normalization and report that external evidence was not executed.\n\n## Channel parser contract\n\n- Crossref `/works/{doi}` returns a `message` object; parse `title`, `author`, `container-title`, `issued`/published dates, `DOI`, and `URL`.\n- PubMed EFetch returns XML; parse `ArticleTitle`, `Author`, `Journal`, `PubDate`, `ArticleId IdType=\"doi\"`, and `PMID`.\n- OpenAlex works return work objects; parse `display_name`/`title`, `authorships`, `primary_location.source.display_name`, `publication_year`, `doi`, and `ids.pmid`.\n- Europe PMC search results return result objects; parse `title`, `authorString`, `journalTitle`, `pubYear`, `doi`, `pmid`, and `pmcid`.\n- Semantic Scholar Graph results return paper objects; parse `title`, `authors`, `venue`/`journal`, `year`, `externalIds.DOI`, `externalIds.PubMed`, and `url`.\n- DataCite DOI records return JSON:API objects; parse `data.attributes.titles`, `creators`, `publisher`/`container`, `publicationYear`, `doi`, and `url`.\n- bioRxiv/medRxiv API results return collection items; parse `title`, `authors`, `date`, `doi`, and `published_doi`.\n\n## Auto-fix policy\n\nAuto-fix only when canonical evidence is strong:\n\n- normalize DOI casing and links\n- strip trailing DOI punctuation\n- normalize DOI links in reports\n- record missing DOI candidates from canonical metadata and ask before appending them\n- include PMID only as secondary metadata when DOI-backed evidence exists\n- standardize journal title or abbreviation\n- normalize author initials and year formatting\n- move a shifted DOI/PMID only when nearby title similarity is strong\n\nDo not auto-delete or silently replace severe items. For `identifier_hijacking`, `total_fabrication`, and low-confidence `partial_attribute_corruption`, report and ask the user.\n\nDo not create hidden persistent caches. Reuse prior results only from an explicitly supplied structured verifier artifact, and re-query new, changed, or incomplete records.\n\n## Format policy\n\n- Detect citation style before repair: APA, GB/T 7714, Vancouver/AMA, free-text, or mixed.\n- Preserve original format in the automatic fixed copy when possible.\n- If the list is mixed, recommend standardizing to the majority format after authenticity checks.\n- Formatting correctness is never authenticity evidence.\n\n## Body-vs-bibliography audit\n\n- Scan body text for narrative citations like `Author et al. (2023)` and parenthetical citations like `(Author, 2023)`.\n- Compare author/year/title semantics, not just numbered reference markers.\n- After deleting or flagging a fabricated reference, remove or flag body claims that depend on it.\n- Report uncertain claims separately instead of silently preserving them.\n\nFile v1.1.2:skill-card.md\n\n## Description: <br>\nVerifies and normalizes biomedical and life-science reference lists to help identify AI-caused citation errors and standardize bibliography formats. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[emp-tca](https://clawhub.ai/user/emp-tca) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers, researchers, and manuscript reviewers use this skill to check whether biomedical references are authentic, detect DOI/PMID or metadata corruption, and generate normalized reference outputs without overwriting the source. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Verification can send citation identifiers and selected title queries to third-party scholarly metadata services, which may reveal sensitive research topics. <br>\nMitigation: Use format-only mode or disable optional providers for confidential manuscripts, and avoid providing API credentials unless those lookups should be associated with the account. <br>\nRisk: Automatic fixed outputs may look publication-ready even when severe reference issues remain unresolved. <br>\nMitigation: Review the summary and detail reports before reuse; keep severe items as warnings, remove them, or run follow-up recovery only after an explicit user decision. <br>\nRisk: Network failures, rate limits, or disabled providers can leave some references unresolved. <br>\nMitigation: Use strict mode for final checks and treat unresolved entries as needing manual review rather than as verified references. <br>\n\n\n## Reference(s): <br>\n- [Verification Policy](references/verification_policy.md) <br>\n- [ClawHub Skill Page](https://clawhub.ai/emp-tca/skills/biomedical-reference-verifier) <br>\n- [DOI Resolver](https://doi.org/) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, JSON, files, shell commands, guidance] <br>\n**Output Format:** [Markdown reports, fixed reference-list files, retained process files, and optional JSON audit or index artifacts.] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Source files are not overwritten; severe reference issues remain report-only unless the user chooses a follow-up action.] <br>\n\n## Skill Version(s): <br>\n1.1.2 (source: frontmatter and server release) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.1.1: 13 files, 46181 bytes\n\nFiles: references/verification_policy.md (11873b), scripts/convert_reference_artifact.py (1437b), scripts/self_test_verify_references.py (10332b), scripts/verifier_models.py (2044b), scripts/verifier_network.py (2739b), scripts/verifier_policy.py (849b), scripts/verifier_prior_results.py (5650b), scripts/verifier_recovery.py (859b), scripts/verifier_runtime.py (1604b), scripts/verify_references.py (115079b), skill-card.md (2318b), SKILL.md (7936b), _meta.json (148b)\n\nFile v1.1.1:SKILL.md\n\n---\nname: biomedical-reference-verifier\ndescription: \"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Makes your reference bibliography more reliable and convincing.. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。\"\nversion: 1.1.1\nmetadata:\n  openclaw:\n    requires:\n      bins:\n        - python3\n    envVars:\n      - name: NCBI_API_KEY\n        required: false\n        description: Optional NCBI API key for higher PubMed E-utilities rate limits.\n      - name: OPENALEX_API_KEY\n        required: false\n        description: Optional OpenAlex API key for authenticated metadata lookups.\n---\n\n# Biomedical Reference Verifier\n\nUse this skill only for reference-list authenticity checking, AI-caused reference cleanup, and citation-format normalization. Do not use it to judge whether papers are relevant, suitable, or correctly used in the manuscript body unless the user explicitly asks for reference authenticity plus body cleanup.\n\n## Core Rule\n\nAlways reduce the task to this sequence:\n\n1. Build or read `biomedical-reference-verifier.records.v1`.\n2. Choose the lightest pipeline.\n3. Run the script.\n4. Report result files and retained process files.\n5. Ask before expensive manual recovery or process-file deletion.\n\nLoad `references/verification_policy.md` before classifying authenticity errors, changing severe items, or explaining network/query policy.\n\n## Pipeline Choice\n\n- **Format only**: use when the user only asks to convert or unify citation style. Run `--pipeline format-only`. Do not query Crossref/PubMed/OpenAlex.\n- **Verify**: use when the user asks whether references are real, fake, wrong, AI-generated, DOI/PMID-mismatched, or metadata-corrupted. Run the default `verify` pipeline.\n- **Verify then format**: use when the user asks both to check authenticity and standardize style. Run `verify` with the requested `--citation-style`; the fixed copy is generated after verification.\n\n## Execution Mode\n\n- **Fast** (`--mode fast`): use for quick scans and large routine lists. Stop after primary DOI evidence when available.\n- **Balanced** (`--mode balanced`): default for ordinary authenticity checks. Auxiliary evidence lines receive a short 2.5-second grace period after the first line completes.\n- **Strict** (`--mode strict`): use for final or pre-submission checks. Wait for every enabled evidence line to complete or fail explicitly.\n\nDo not invent a local confidence score. For DOI-bearing references, resolve the DOI and directly compare returned title, authors, journal, year, DOI, and PMID with source fields. Keep the existing classification policy. Severe items are reported and retained; never auto-delete them.\n\nDo not run AI-assisted per-paper web searching by default. Stop after batch verification and title recovery, then ask the user whether to continue.\n\n## Machine Input\n\nPrefer `biomedical-reference-verifier.records.v1` JSON/JSONL. If the source is free text, convert it to records first using only values present in the user source. Do not fill missing source fields from Crossref, PubMed, OpenAlex, memory, or plausible guesses.\n\nMinimal record:\n\n```json\n{\n  \"schema\": \"biomedical-reference-verifier.records.v1\",\n  \"records\": [\n    {\n      \"index\": 1,\n      \"source\": {\n        \"original_text\": \"exact source reference\",\n        \"title\": \"title from source, or empty string\",\n        \"authors\": [\"First Author\"],\n        \"year\": \"2024\",\n        \"journal\": \"Journal from source\",\n        \"identifiers\": {\"doi\": \"10.xxxx/example\", \"pmid\": \"\", \"urls\": []},\n        \"source_lines\": [12],\n        \"context\": \"\"\n      }\n    }\n  ]\n}\n```\n\nUse `--input-mode records` for standardized JSON/JSONL. Use Markdown worksheet and `doi-context` modes only for compatibility.\n\n## Commands\n\nRun from this skill directory:\n\n```bash\npython3 scripts/verify_references.py refs.md --output-dir /tmp/reference-audit --citation-style ama\n```\n\nCommon variants:\n\n```bash\npython3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --mode fast\npython3 scripts/verify_references.py records.json --input-mode records --mode strict\npython3 scripts/verify_references.py records.json --input-mode records --reuse-results previous/reference-audit.json\npython3 scripts/verify_references.py records.json --input-mode records --write-index\npython3 scripts/convert_reference_artifact.py previous/reference-audit.json --to index --output reference-index.json\npython3 scripts/convert_reference_artifact.py reference-index.json --to audit --output restored-reference-audit.json\npython3 scripts/verify_references.py refs.md --pubmed-mode off --openalex-mode off\npython3 scripts/verify_references.py refs.md --doi-output append\npython3 scripts/verify_references.py refs.md --keep-process-json\npython3 scripts/verify_references.py refs.md --cleanup-process-files all\npython3 scripts/verify_references.py refs.md --cleanup-process-files normalized_input,extracted_references\n```\n\n## Output Contract\n\nResult files:\n\n- `reference-audit-summary.md`: concise chat-ready summary.\n- `reference-audit-detail.md`: detailed report with evidence links.\n- `references.auto-fixed.md` or `document.auto-fixed.md`: fixed/formatted copy; never overwrites the source.\n\nProcess files retained by default:\n\n- `reference-normalized-records.json`: machine-readable records.\n- `reference-normalized-input.md`: human-readable parser inspection table.\n- `references.extracted.md`: raw references extracted before verification.\n\nProcess files skipped by default:\n\n- `reference-audit.json`: write only with `--keep-process-json`.\n\nThere is no hidden persistent cache. `--reuse-results` accepts either `reference-audit.json` or schema `biomedical-reference-verifier.index.v1`. The converter supports audit-to-index and index-to-audit round trips; the `results` array must remain identical after a round trip. In a conversation, ask about reuse once when such an artifact is available; do not repeatedly ask after the user decides.\n\nUser-edited prior artifacts are handled conservatively: malformed JSON is rejected with one concise error; invalid individual rows are skipped and rechecked; valid rows remain reusable. Do not guess repairs for damaged structured fields.\n\nPrior-result reuse tolerates punctuation changes only when strict bibliographic fields still agree: DOI+title+year, or title+first author+year+journal. Temporary network failures receive one bounded retry; permanent 4xx failures do not. PubMed DOI batches split only when a batch fails. Early DOI recovery uses Crossref first, then OpenAlex and PubMed outside Fast mode.\n\n`--write-index` optionally writes `reference-index.json`; it is off by default. Network request events and phase timings are stored in `reference-audit.json` when `--keep-process-json` is used.\n\nNever delete user-requested result files. Delete only process files generated by the script in the selected output directory, and only when the user chooses deletion or `--cleanup-process-files` is explicit.\n\n## Closeout\n\nAfter running:\n\n1. Paste or summarize `reference-audit-summary.md`.\n2. Link result files and retained process files.\n3. State which process files were cleaned or skipped.\n4. Ask whether to delete retained process files:\n   - A. delete all retained process files\n   - B. keep selected process files\n   - C. keep all process files for now\n5. If severe items exist, ask whether to continue AI-assisted recovery, delete, keep with warning notes, or stop with the report.\n6. Ask whether recovered DOI values should be appended when DOI output is optional.\n\nFile v1.1.1:_meta.json\n\n{\n  \"ownerId\": \"kn748sfft0aksxdrw9332r67w58a99n2\",\n  \"slug\": \"biomedical-reference-verifier\",\n  \"version\": \"1.1.1\",\n  \"publishedAt\": 1783791122117\n}\n\nFile v1.1.1:references/verification_policy.md\n\n# Verification Policy\n\n## Evidence hierarchy\n\n1. Crossref is the fastest and strongest default primary source for journal DOI, title, authors, journal, publisher, and year.\n2. PubMed is the second evidence line for biomedical records, PMID mapping, and wrong PMID/title-pair detection.\n3. OpenAlex is the third default evidence line for fast DOI/PMID external-ID corroboration and title recovery when Crossref is weak.\n4. A DOI match from Crossref is the primary batch verification result. PubMed and OpenAlex matches can corroborate it, but should not block reporting.\n5. Semantic Scholar, Europe PMC, bioRxiv/medRxiv, DataCite, and OpenCitations are backup channels, not default full-batch channels.\n6. Search snippets, AI summaries, formatted APA/GB/T entries, and plausible journal names are not proof.\n7. `biomedical-reference-verifier.records.v1` is the source-of-truth input layer. Source fields must come from the user's text; database-returned fields must stay in evidence/report/output objects.\n\n## Error types\n\n- `verified`: DOI/PMID or title search resolves to a canonical record and title, year, and first-author evidence agree.\n- `verified_identifier_only`: DOI/PMID resolves to a canonical record, but the source title is missing or unreliable, so only the identifier is verified.\n- `parser_error`: the input row is malformed or the parsed/source title is an author line, table header, metadata field, or other non-title text; stop before treating the row as a bibliographic conflict.\n- `minor_fix` / `minor_format_error`: the paper is real and metadata agree; only DOI casing, punctuation, URL style, initials, journal abbreviation, or minor year formatting needs correction.\n- `partial_attribute_corruption`: the paper is real, but one or more attributes are corrupted, such as author, title, journal, year, volume/pages, DOI, or PMID.\n- `identifier_hijacking`: DOI, PMID, or URL is real, but it points to a different paper than the reference text.\n- `shifted_identifier`: DOI or PMID appears to belong to a nearby reference, commonly because one entry's DOI was attached to the previous or next citation.\n- `semantic_hallucination`: the paper exists, but a surrounding claim or summary is not supported by that paper.\n- `placeholder_generation`: the entry contains AI-like placeholder traces: generic title, invented DOI suffix, fake pages/volume, missing journal fields, or overly tidy but unverifiable metadata.\n- `total_fabrication`: title, authors, journal, year, DOI/PMID, and searches fail to identify any canonical record.\n- `unresolved`: automatic checks are insufficient; stop and ask before expensive manual recovery.\n\n## Matching thresholds\n\n- Title similarity >= 0.90: strong match.\n- Title similarity 0.86-0.89: acceptable match if year and first author agree.\n- Title similarity 0.78-0.85: possible match; mark as partial unless corroborated by DOI/PMID.\n- Title similarity < 0.70 for a supplied DOI/PMID: identifier hijacking unless adjacent-entry checks prove shifted identifier.\n- Year difference of 0-1 year can be online-first drift. Larger differences are corruption unless the source explains it.\n- First-author disagreement is a warning; combine it with title/year/journal evidence before deciding severity.\n\n## Short-circuit rules\n\n- Build `biomedical-reference-verifier.records.v1` first. Do not let remote lookup run against raw mixed notes when the source can be converted to explicit source objects.\n- Required machine fields are `index`, `source.original_text`, `source.title`, `source.authors`, `source.year`, `source.journal`, `source.identifiers.doi`, `source.identifiers.pmid`, `source.identifiers.urls`, `source.source_lines`, and `source.context`.\n- During machine-record construction, AI may only extract values present in the source text. It must not use Crossref/PubMed/OpenAlex results, memory, or plausible completions to fill source fields.\n- Build `reference-normalized-records.json` before remote lookup. `reference-normalized-input.md` is a human-readable view of the same records.\n- If `source.title` is missing, too short, an author list, a table/header label, a DOI/PMID line, or metadata placeholder, mark the title unreliable.\n- If a DOI/PMID resolves but title is unreliable, mark `verified_identifier_only`; do not mark `identifier_hijacking`.\n- If a row has no reliable title and no resolvable DOI/PMID, mark `parser_error` and stop.\n- Use DOI verification first: `GET https://api.crossref.org/works/{doi}`. This is the primary source of DOI/title/author/journal/year truth.\n- Start enabled evidence lines together. Fast stops after primary evidence, Balanced allows a 2.5-second auxiliary grace period, and Strict waits for every enabled line to complete or explicitly fail. Network timeout, rate limiting, not-found, parsing failure, mode skip, and expired budget must remain distinguishable.\n- If PubMed and OpenAlex are disabled with `--pubmed-mode off --openalex-mode off`, do not apply the grace timer to Crossref; Crossref DOI verification must complete because it is the primary authority.\n- OpenAlex checks use `GET https://api.openalex.org/works/doi:{doi}` or external IDs such as `pmid:{pmid}` for PMID, and `GET https://api.openalex.org/works?search={title}&per_page=3` for DOI-missing title recovery.\n- If an entry has no DOI, use title recovery in this order: Crossref title query first, OpenAlex title query only if the Crossref match is weak, PubMed title query only if biomedical corroboration is still needed. Record recovered DOI values in the report and ask before adding them to formatted citations.\n- Run quality-checked DOI-missing title recovery before full identifier verification, deduplicated and with provider-specific bounded concurrency. A recovered DOI may drive verification but must not silently modify the user source.\n- Do not use title search to silently override an existing DOI result. Existing DOI conflicts should be reported as conflicts, hijacking, or shifted identifiers.\n- Run near official limits without exceeding them: Crossref polite requests use no more than 10 requests/second and 3 concurrent requests; PubMed E-utilities uses no more than 3 requests/second without an API key and 10 requests/second with an NCBI API key; OpenAlex is budget/cost governed, so keep default concurrency conservative unless an API key and live limit headers justify more.\n- Use one bounded retry only for timeout, temporary connection failure, HTTP 429, or HTTP 5xx. Respect `Retry-After` within a bounded wait. Do not retry permanent 4xx responses. If a PubMed DOI batch fails, split it recursively until the failing DOI is isolated.\n- For large DOI-heavy documents, Crossref parallel DOI lookup is the primary full-batch path. PubMed and OpenAlex are corroboration channels, not reasons to block report generation.\n- If DOI/PMID resolves and title similarity is < 0.70, stop trusting that identifier for the current entry and mark `identifier_hijacking`.\n- Apply the < 0.70 hijacking rule only when the source title passed input quality checks.\n- Before finalizing `identifier_hijacking`, compare the resolved title against entries within two positions. If adjacent similarity is >= 0.86, mark `shifted_identifier`.\n- If several consecutive identifiers are shifted by the same offset, report a possible column/line offset instead of repairing each item manually.\n- For entries with DOI, run DOI lookup and report the DOI result; do not automatically title-search to replace the DOI.\n- For entries without DOI, run at most one Crossref title recovery query. If title recovery fails, mark `unresolved`, `placeholder_generation`, or `total_fabrication`; do not keep prompting the model to guess.\n- Do not run title recovery when `source_title` failed input quality checks. Rebuild the worksheet row first.\n- Use backup channels conditionally: Semantic Scholar only for title/abstract/citation corroboration after default lines are inconclusive; Europe PMC for biomedical PMCID/full-text evidence; bioRxiv/medRxiv for `10.1101/...` and published-DOI mapping; DataCite for non-journal dataset/software/report DOIs; OpenCitations for citation-network existence checks.\n- Parse every channel response by code into one canonical record shape: `source`, `title`, `authors`, `journal`, `year`, `doi`, `pmid`, `url`, and `score`. AI must not manually interpret structured API responses when parser adapters are available.\n- If the severe-error rate is high, stop after batch recovery and ask the user whether to continue AI-assisted per-paper search.\n\n## Network risk policy\n\n- Default DOI/PMID verification is low risk and should run without asking the user each time. The payload is limited to public identifiers required for the skill's purpose.\n- Default requests must not include `original_text`, manuscript paragraphs, abstracts, local evidence notes, or unpublished claims.\n- DOI-missing recovery may send only a short, quality-checked `source.title`.\n- Any deep search that sends title plus abstract, surrounding context, or manuscript-derived claims is higher risk and must be separated from default batch verification.\n- If Codex sandbox approval blocks external lookup, continue with local records normalization and report that external evidence was not executed.\n\n## Channel parser contract\n\n- Crossref `/works/{doi}` returns a `message` object; parse `title`, `author`, `container-title`, `issued`/published dates, `DOI`, and `URL`.\n- PubMed EFetch returns XML; parse `ArticleTitle`, `Author`, `Journal`, `PubDate`, `ArticleId IdType=\"doi\"`, and `PMID`.\n- OpenAlex works return work objects; parse `display_name`/`title`, `authorships`, `primary_location.source.display_name`, `publication_year`, `doi`, and `ids.pmid`.\n- Europe PMC search results return result objects; parse `title`, `authorString`, `journalTitle`, `pubYear`, `doi`, `pmid`, and `pmcid`.\n- Semantic Scholar Graph results return paper objects; parse `title`, `authors`, `venue`/`journal`, `year`, `externalIds.DOI`, `externalIds.PubMed`, and `url`.\n- DataCite DOI records return JSON:API objects; parse `data.attributes.titles`, `creators`, `publisher`/`container`, `publicationYear`, `doi`, and `url`.\n- bioRxiv/medRxiv API results return collection items; parse `title`, `authors`, `date`, `doi`, and `published_doi`.\n\n## Auto-fix policy\n\nAuto-fix only when canonical evidence is strong:\n\n- normalize DOI casing and links\n- strip trailing DOI punctuation\n- normalize DOI links in reports\n- record missing DOI candidates from canonical metadata and ask before appending them\n- include PMID only as secondary metadata when DOI-backed evidence exists\n- standardize journal title or abbreviation\n- normalize author initials and year formatting\n- move a shifted DOI/PMID only when nearby title similarity is strong\n\nDo not auto-delete or silently replace severe items. For `identifier_hijacking`, `total_fabrication`, and low-confidence `partial_attribute_corruption`, report and ask the user.\n\nDo not create hidden persistent caches. Reuse prior results only from an explicitly supplied structured verifier artifact, and re-query new, changed, or incomplete records.\n\n## Format policy\n\n- Detect citation style before repair: APA, GB/T 7714, Vancouver/AMA, free-text, or mixed.\n- Preserve original format in the automatic fixed copy when possible.\n- If the list is mixed, recommend standardizing to the majority format after authenticity checks.\n- Formatting correctness is never authenticity evidence.\n\n## Body-vs-bibliography audit\n\n- Scan body text for narrative citations like `Author et al. (2023)` and parenthetical citations like `(Author, 2023)`.\n- Compare author/year/title semantics, not just numbered reference markers.\n- After deleting or flagging a fabricated reference, remove or flag body claims that depend on it.\n- Report uncertain claims separately instead of silently preserving them.\n\nFile v1.1.1:skill-card.md\n\n## Description: <br>\nVerifies or normalizes biomedical and life-science reference lists for AI-caused reference errors and citation-format cleanup. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[emp-tca](https://clawhub.ai/user/emp-tca) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers, researchers, editors, and reviewers use this skill to check whether biomedical references are real, mismatched, AI-generated, or metadata-corrupted, and to normalize reference lists into requested citation styles. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Verification mode may send citation identifiers, selected titles, and optional contact values to Crossref, PubMed/NCBI, and OpenAlex. <br>\nMitigation: Use the skill for public or non-sensitive bibliographies, or choose format-only mode and provider-off flags when network exposure is not acceptable. <br>\nRisk: Confidential manuscripts, grant materials, or unpublished study titles may contain sensitive research context. <br>\nMitigation: Avoid default verification for sensitive material unless external metadata lookup is approved; keep source fields limited to citation data needed for the check. <br>\n\n\n## Reference(s): <br>\n- [ClawHub Skill Page](https://clawhub.ai/emp-tca/skills/biomedical-reference-verifier) <br>\n- [Verification Policy](references/verification_policy.md) <br>\n- [DOI Resolver](https://doi.org/) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Markdown, JSON, Code, Shell commands, Guidance] <br>\n**Output Format:** [Markdown reports, fixed reference files, structured JSON or JSONL records, and shell commands] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Writes summary and detail reports plus a fixed/formatted copy without overwriting the source; optional process JSON and reusable index files can be retained.] <br>\n\n## Skill Version(s): <br>\n1.1.1 (source: frontmatter and server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.1.0: 13 files, 46563 bytes\n\nFiles: references/verification_policy.md (11873b), scripts/convert_reference_artifact.py (1437b), scripts/self_test_verify_references.py (10332b), scripts/verifier_models.py (2044b), scripts/verifier_network.py (2739b), scripts/verifier_policy.py (849b), scripts/verifier_prior_results.py (5650b), scripts/verifier_recovery.py (859b), scripts/verifier_runtime.py (1604b), scripts/verify_references.py (115079b), skill-card.md (2741b), SKILL.md (8476b), _meta.json (148b)\n\nFile v1.1.0:SKILL.md\n\n---\nname: biomedical-reference-verifier\ndescription: \"Verify or normalize biomedical and life-science reference lists only when the task is about AI-caused reference errors: whether references really exist, whether DOI/PMID/title/authors/journal/year are wrong, whether a bibliography has false or corrupted entries, or whether references should be reformatted into APA/AMA/GB/T 7714/Vancouver without judging topical relevance. Trigger examples: 帮我核验这些参考文献是不是真实存在; 检查这批 references 里有没有 AI 编造的文献; 这些文献可能是 AI 生成的，请查一下哪些是假的; 请验证这些 APA 引用的 DOI、作者、题名是否对应; 帮我审查参考文献真实性; 检查这些 AMA references 的 DOI 和期刊年份是否正确; 请把这些参考文献改成 AMA 格式; 帮我把参考文献统一成 APA. Do not trigger for general article review, manuscript logic, literature relevance, citation support, or whether papers are suitable for the article.\"\nversion: 1.1.0\nmetadata:\n  openclaw:\n    requires:\n      bins:\n        - python3\n    envVars:\n      - name: NCBI_API_KEY\n        required: false\n        description: Optional NCBI API key for higher PubMed E-utilities rate limits.\n      - name: OPENALEX_API_KEY\n        required: false\n        description: Optional OpenAlex API key for authenticated metadata lookups.\n---\n\n# Biomedical Reference Verifier\n\nUse this skill only for reference-list authenticity checking, AI-caused reference cleanup, and citation-format normalization. Do not use it to judge whether papers are relevant, suitable, or correctly used in the manuscript body unless the user explicitly asks for reference authenticity plus body cleanup.\n\n## Core Rule\n\nAlways reduce the task to this sequence:\n\n1. Build or read `biomedical-reference-verifier.records.v1`.\n2. Choose the lightest pipeline.\n3. Run the script.\n4. Report result files and retained process files.\n5. Ask before expensive manual recovery or process-file deletion.\n\nLoad `references/verification_policy.md` before classifying authenticity errors, changing severe items, or explaining network/query policy.\n\n## Pipeline Choice\n\n- **Format only**: use when the user only asks to convert or unify citation style. Run `--pipeline format-only`. Do not query Crossref/PubMed/OpenAlex.\n- **Verify**: use when the user asks whether references are real, fake, wrong, AI-generated, DOI/PMID-mismatched, or metadata-corrupted. Run the default `verify` pipeline.\n- **Verify then format**: use when the user asks both to check authenticity and standardize style. Run `verify` with the requested `--citation-style`; the fixed copy is generated after verification.\n\n## Execution Mode\n\n- **Fast** (`--mode fast`): use for quick scans and large routine lists. Stop after primary DOI evidence when available.\n- **Balanced** (`--mode balanced`): default for ordinary authenticity checks. Auxiliary evidence lines receive a short 2.5-second grace period after the first line completes.\n- **Strict** (`--mode strict`): use for final or pre-submission checks. Wait for every enabled evidence line to complete or fail explicitly.\n\nDo not invent a local confidence score. For DOI-bearing references, resolve the DOI and directly compare returned title, authors, journal, year, DOI, and PMID with source fields. Keep th\n\nArchive v1.0.0: 6 files, 36182 bytes\n\nFiles: references/verification_policy.md (11180b), scripts/self_test_verify_references.py (4540b), scripts/verify_references.py (109254b), skill-card.md (2592b), SKILL.md (5924b), _meta.json (148b)","readmeExcerpt":"Skill: biomedical-reference-verifier Owner: emp-tca Summary: Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。 Tags: latest:1.2.1 Version history","codeSnippets":[],"executableExamples":[{"language":"json","snippet":"{\n  \"schema\": \"biomedical-reference-verifier.records.v1\",\n  \"records\": [\n    {\n      \"index\": 1,\n      \"source\": {\n        \"original_text\": \"exact source reference\",\n        \"title\": \"title from source, or empty string\",\n        \"authors\": [\"First Author\"],\n        \"year\": \"2024\",\n        \"journal\": \"Journal from source\",\n        \"identifiers\": {\"doi\": \"10.xxxx/example\", \"pmid\": \"\", \"urls\": []},\n        \"source_lines\": [12],\n        \"context\": \"\"\n      }\n    }\n  ]\n}"},{"language":"bash","snippet":"python3 scripts/verify_references.py refs.md --output-dir /tmp/reference-audit --citation-style ama"},{"language":"bash","snippet":"python3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --citation-style ama\npython3 scripts/verify_references.py records.json --input-mode records --mode fast\npython3 scripts/verify_references.py records.json --input-mode records --mode strict\npython3 scripts/verify_references.py records.json --input-mode records --reuse-results previous/reference-audit.json\npython3 scripts/verify_references.py records.json --input-mode records --write-index\npython3 scripts/convert_reference_artifact.py previous/reference-audit.json --to index --output reference-index.json\npython3 scripts/convert_reference_artifact.py reference-index.json --to audit --output restored-reference-audit.json\npython3 scripts/verify_references.py refs.md --pubmed-mode off --openalex-mode off\npython3 scripts/verify_references.py refs.md --doi-output append\npython3 scripts/verify_references.py refs.md --keep-process-json\npython3 scripts/generate_html_report.py previous/reference-audit.json --output reference-audit-report.html\npython3 scripts/verify_references.py refs.md --cleanup-process-files all\npython3 scripts/verify_references.py refs.md --cleanup-process-files normalized_input,extracted_references"},{"language":"bash","snippet":"python3 scripts/verify_references.py records.json --input-mode records --report-language en\npython3 scripts/generate_html_report.py reference-audit.json --output report.en.html --language en"},{"language":"bash","snippet":"python3 scripts/verify_references.py refs.md --mode balanced --output-dir reference-audit"},{"language":"bash","snippet":"python3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: biomedical-reference-verifier\nversion: 1.2.1\ndescription: \"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。\"\nmetadata:\n  openclaw:\n    requires:\n      bins:\n        - python3\n    envVars:\n      - name: NCBI_API_KEY\n        required: false\n        description: Optional NCBI API key for higher PubMed E-utilities rate limits.\n      - name: OPENALEX_API_KEY\n        required: false\n        description: Optional OpenAlex API key for authenticated metadata lookups.\n---\n\n# Biomedical Reference Verifier\n\nUse this skill only for reference-list authenticity checking, AI-caused reference cleanup, and citation-format normalization. Do not use it to judge whether papers are relevant, suitable, or correctly used in the manuscript body unless the user explicitly asks for reference authenticity plus body cleanup.\n\n## Core Rule\n\nAlways reduce the task to this sequence:\n\n1. Build or read `biomedical-reference-verifier.records.v1`.\n2. Choose the lightest pipeline.\n3. Run the script.\n4. Report result files and retained process files.\n5. Ask before expensive manual recovery or process-file deletion.\n\nLoad `references/verification_policy.md` before classifying authenticity errors, changing severe items, or explaining network/query policy.\n\n## Pipeline Choice\n\n- **Format only**: use when the user only asks to convert or unify citation style. Run `--pipeline format-only`. Do not query Crossref/PubMed/OpenAlex.\n- **Verify**: use when the user asks whether references are real, fake, wrong, AI-generated, DOI/PMID-mismatched, or metadata-corrupted. Run the default `verify` pipeline.\n- **Verify then format**: use when the user asks both to check authenticity and standardize style. Run `verify` with the requested `--citation-style`; the fixed copy is generated after verification.\n\n## Execution Mode\n\n- **Fast** (`--mode fast`): use for quick scans and large routine lists. Stop after primary DOI evidence when available.\n- **Balanced** (`--mode balanced`): default for ordinary authenticity checks. Auxiliary evidence lines receive a short 2.5-second grace period after primary DOI lookup completes (and supplied PMID lookup when enabled).\n- **Strict** (`--mode strict`): use for final or pre-submission checks. Wait for every enabled evidence line to complete or fail explicitly.\n\nDo not invent a local confidence score. For DOI-bearing references, resolve the DOI and directly compare returned title, authors, journal, year, DOI, and PMID with source fields. Keep the existing classification policy. Severe items are reported and retained; never auto-delete them.\n\nDo not run AI-assisted per-paper web searching by default. Stop after batch verification and title recovery, then ask the user whether to continue."},{"path":"README.md","content":"# Biomedical Reference Verifier\n\nVerify and normalize biomedical and life-science reference lists, with a focus on AI-caused citation errors.\n\nThe skill checks DOI, PMID, title, authors, journal, year, and other bibliographic metadata through Crossref, PubMed, and OpenAlex. It can also normalize references to AMA, APA, Vancouver, or GB/T 7714 style without overwriting the source document.\n\n## Highlights\n\n- Fast, Balanced, and Strict verification modes\n- DOI-first batch verification and early DOI recovery\n- Crossref, PubMed, and OpenAlex evidence\n- Bounded retries and PubMed batch-failure isolation\n- Detection of identifier hijacking and shifted identifiers\n- Reusable audit indexes with lossless audit/index conversion\n- Tolerant handling of user-edited prior artifacts\n- Runtime metrics and request-event logging\n- No hidden persistent cache\n- Severe references are reported, never automatically deleted\n\n## Requirements\n\n- Python 3\n- Standard library only\n- Optional `NCBI_API_KEY` for higher PubMed limits\n- Optional `OPENALEX_API_KEY`\n\n## Quick start\n\n```bash\npython3 scripts/verify_references.py refs.md --mode balanced --output-dir reference-audit\n```\n\nFormat without external verification:\n\n```bash\npython3 scripts/verify_references.py refs.md --pipeline format-only --citation-style ama\n```\n\nRun the offline regression suite:\n\n```bash\npython3 scripts/self_test_verify_references.py\n```\n\nSee [SKILL.md](SKILL.md) for the complete workflow and [verification_policy.md](references/verification_policy.md) for classification and evidence rules.\n\n## License\n\nMIT-0"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn748sfft0aksxdrw9332r67w58a99n2\",\n  \"slug\": \"biomedical-reference-verifier\",\n  \"version\": \"1.2.1\",\n  \"publishedAt\": 1789392009429\n}"},{"path":"references/verification_policy.md","content":"# Verification Policy\n\n## Evidence hierarchy\n\n1. Crossref is the fastest and strongest default primary source for journal DOI, title, authors, journal, publisher, and year.\n2. PubMed is the second evidence line for biomedical records, PMID mapping, and wrong PMID/title-pair detection.\n3. OpenAlex is the third default evidence line for fast DOI/PMID external-ID corroboration and title recovery when Crossref is weak.\n4. A DOI match from Crossref is the primary batch verification result. PubMed and OpenAlex matches can corroborate it, but should not block reporting.\n5. Semantic Scholar, Europe PMC, bioRxiv/medRxiv, DataCite, and OpenCitations are backup channels, not default full-batch channels.\n6. Search snippets, AI summaries, formatted APA/GB/T entries, and plausible journal names are not proof.\n7. `biomedical-reference-verifier.records.v1` is the source-of-truth input layer. Source fields must come from the user's text; database-returned fields must stay in evidence/report/output objects.\n\n## Error types\n\n- `verified`: DOI/PMID or title search resolves to a canonical record and title, year, and first-author evidence agree.\n- `verified_identifier_only`: DOI/PMID resolves to a canonical record, but the source title is missing or unreliable, so only the identifier is verified.\n- `parser_error`: the input row is malformed or the parsed/source title is an author line, table header, metadata field, or other non-title text; stop before treating the row as a bibliographic conflict.\n- `formatted_only`: formatting executed, authenticity not checked; not eligible for use on that basis.\n- `minor_fix` / `minor_format_error`: the paper is real and metadata agree; only DOI casing, punctuation, URL style, initials, journal abbreviation, or minor year formatting needs correction.\n- `partial_attribute_corruption`: the paper is real, but one or more attributes are corrupted, such as author, title, journal, year, volume/pages, DOI, or PMID.\n- `identifier_hijacking`: DOI, PMID, or URL is real, but it points to a different paper than the reference text.\n- `shifted_identifier`: DOI or PMID appears to belong to a nearby reference, commonly because one entry's DOI was attached to the previous or next citation.\n- `semantic_hallucination`: the paper exists, but a surrounding claim or summary is not supported by that paper.\n- `placeholder_generation`: the entry contains AI-like placeholder traces: generic title, invented DOI suffix, fake pages/volume, missing journal fields, or overly tidy but unverifiable metadata.\n- `total_fabrication`: title, authors, journal, year, DOI/PMID, and searches fail to identify any canonical record.\n- `unresolved`: automatic checks are insufficient; stop and ask before expensive manual recovery.\n\n## Matching thresholds\n\n- Title similarity >= 0.90: strong match.\n- Title similarity 0.86-0.89: acceptable match if year and first author agree.\n- Title similarity 0.78-0.85: possible match; mark as partial unless corroborated by DOI/PMID.\n- Title si"},{"path":"CHANGELOG.md","content":"# Changelog\n\n## 1.2.1 — 2026-09-14\n\n- Remove ambient email reads and omit contact email from requests unless explicitly supplied.\n- Add selectable English HTML report template and localized generated eligibility messages.\n- Document implementation locations for capability review; add privacy and language regression tests.\n\n## 1.2.0 — 2026-09-14\n\nMinor release following public version 1.1.3. SKILL.md and the CLI both identify this release as 1.2.0.\n\n- Reject duplicate reference indices before verification or writing outputs, preventing silent loss or duplication of results.\n- Reject output paths that alias the source, an explicitly reused artifact, or another output, including symlinks and hardlinks.\n- Keep failed title queries distinguishable from completed searches with no match. Failed queries leave unresolved references unusable; successful empty searches retain the existing classification policy.\n- Do not cache failed title searches as successful empty results. Detect malformed search responses and incomplete PubMed record retrieval.\n- Include formatted-only entries in the Markdown detail report.\n- Update the verification policy version to 2026-09-14.1 so older evidence is rechecked.\n- Add 10 offline regression tests for these boundaries.\n- Add the MIT-0 license and exclude Python caches and macOS metadata from publication."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。 Skill: biomedical-reference-verifier Owner: emp-tca Summary: Verify or normalize biomedical and life-science reference lists, when the task is about AI-caused reference errors. Checks identifiers and bibliographic fields, preserves source evidence, and produces a searchable offline report. 生物医学/生命科学参考文献真实性验证skill，可以对参考文献列表（引文列表）进行多轮核查和错误修复，附带引文格式整理功能，能统一规范化所有引文为AMA、APA、GB/T 7714等格式。 Tags: latest:1.2.1 Version history","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":1895,"uniquenessScore":45,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T13:43:21.327Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T13:43:21.327Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T15:51:18.036Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}