{"id":"f2513f35-cee0-4730-93ee-886014cd0189","entityType":"agent","slug":"clawhub-docsor1212-paper-polisher-pro","name":"Paper Polisher Pro — AI Detector & Academic Polishing","canonicalUrl":"https://www.xpersona.co/agent/clawhub-docsor1212-paper-polisher-pro","canonicalPath":"/agent/clawhub-docsor1212-paper-polisher-pro","generatedAt":"2026-10-10T07:42:05.088Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T01:31:08.446Z","emptyReason":null},"description":"AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell), metaphor audit, quality report, AIGC compliance label check (China 2025-09 labeling rules), paragraph-level attribution, journal precheck, sentence-level rewrite suggestions (locates and advises, never auto-rewrites), plus `--batch DIR` for thesis-scale batch rewriting (per-file AI-rate scores directory-wide). Bilingual CN/EN, 100% local, zero upload, zero credentials; bundled unit-test suite + AST-based zero-network self-verification. v3 delivers a recalibrated multi-layer rule engine (11 core layers + discourse/smoothness heuristics) + token-spectrum layer + length-routed fusion + optional supervised Qwen3-0.6B ONNX layer (AUROC 1.0 on held-out test) + LLM fingerprint attribution (GLM/DeepSeek/Qwen/Kimi/MiniMax/GPT/Claude/Gemini) + freshness pipeline. Base-engine numbers reproduce from the bundled held-out evaluation; supervised columns are author-side measurements (model not bundled). Skill: Paper Polisher Pro — AI Detector & Academic Polishing Owner: docsor1212 Summary: AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell), metaphor audit, quality report, AIGC compliance label check (China 2025-09 labeling rules), paragraph-level attribution, journal precheck, sentence-level rewrite suggestions (locates and advises, never auto-rewrites), plus --batch DIR","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.8K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s17dagtwyk21qs6vpz98bzcrh1853t29:paper-polisher-pro","sourceUrl":"https://clawhub.ai/docsor1212/paper-polisher-pro","homepage":"https://clawhub.ai/docsor1212/skills/paper-polisher-pro","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/docsor1212/paper-polisher-pro","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/docsor1212/skills/paper-polisher-pro","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":65,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell), metaphor audit, quality report, AIGC compliance label check (C"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T01:31:08.446Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T01:31:08.446Z","emptyReason":null},"stars":null,"forks":null,"downloads":1805,"packageName":null,"latestVersion":"5.1.0","tractionLabel":"1.8K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T01:31:08.446Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T01:31:08.446Z","lastCrawledAt":"2026-10-10T01:31:08.446Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T01:31:08.446Z","lastVerifiedAt":null,"highlights":[{"version":"5.1.0","createdAt":"2026-10-09T16:59:52.833Z","changelog":"v5.1.0: rewrite-closure & batch-report release — rewrite-effect regression check (pp_rewrite_check.py: original vs revised engine-source comparison — document score/risk-band migration, paragraph-level aligned deltas, feature-type clears vs remaining, edit extent; relative reference under this engine's criteria only; pp.py rewrite-check + pp_api.rewrite_check); batch HTML summary report (pp_batch_report.py: renders --batch --csv into a single self-contained HTML — cards, histogram, per-file table; zero deps, no engine re-run); batch-CSV concurrency semantics and degraded_notice reading guide documented in FAQ. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。","fileCount":102,"zipByteSize":756120},{"version":"5.0.0","createdAt":"2026-10-08T18:57:11.895Z","changelog":"v5.0.0: verification & rewrite-suggestions release — sentence-level rewrite suggestions (pp_fix_suggest.py: which sentences, why, how to improve — locate & strategy, never auto-rewrite; wired into pp_workflow and pp.py fix); bundled unit-test suite (tests/, 44 stdlib cases, pp.py test); AST-structured zero-network self-verification (pp_verify.py); literary-narrative register notice (disclosure only, no scoring change); requirements.txt; pp.py quickstart zero-model demo. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。","fileCount":98,"zipByteSize":741262},{"version":"4.9.0","createdAt":"2026-10-08T04:32:19.691Z","changelog":"v4.9.0: mixed-document calibration release — paragraph-level hi/med/lo thresholds calibrated on the controlled mixed benchmark (333 paragraphs with ground truth; best operating point precision 0.60 at 63% coverage — honestly below the automatic-verdict bar, ranking aid only; references/para_thresholds.json ships with the package and surfaces in mixed_document assessments); unified entry scripts/pp.py routes all eleven subcommands; batch CSV writes are atomic; gate layers retry once before fallback. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。","fileCount":84,"zipByteSize":712810},{"version":"4.8.0","createdAt":"2026-10-07T04:32:14.303Z","changelog":"v4.8.0: measurement & output-contract integrity (independent third-party test round, 13 findings verified and addressed) — run_eval cache keys now bind engine mode (rules-only evaluation no longer replays supervised cached scores; reproducible fingerprint-and-mode-bound: rules 0.8985 old-gen / 0.7149 current-gen on the bundled sample); --profile journal --format json emits pure JSON with a journal_precheck field and both statistics explained; risk_bands (active thresholds + calibration tier) surfaced in every report; mixed_signal denoised (both extremes >=25% of paragraphs); single-file >5MB warning; --batch directory requirement in the error message; honest disclosures for real-world medical FPR and fingerprint attribution limits in the FAQ; corpus scope (n=1,304 bundled vs n=5,251 full) and terminology count (2,308 loaded) corrected. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。","fileCount":81,"zipByteSize":705522},{"version":"4.7.0","createdAt":"2026-10-06T05:06:51.523Z","changelog":"v4.6.0: onboarding release — new scripts/pp_workflow.py end-to-end self-check command (AI-rate detection + paragraph attribution + 4-layer gate + terminology + translation-smell + style + quality report + AIGC label check in one pass; Markdown report + full JSON; also pp_api.workflow()); TL;DR quick-start layer in both docs; historical release notes moved to CHANGELOG.md; consolidated anti-patterns section; EN FAQ aligned with ZH; superseded eval artifacts archived; SDK/setup errors carry recovery hints. Zero new dependencies. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。","fileCount":79,"zipByteSize":680761},{"version":"4.6.0","createdAt":"2026-10-05T04:30:46.544Z","changelog":"v4.6.0: onboarding release — new scripts/pp_workflow.py end-to-end self-check command (AI-rate detection + paragraph attribution + 4-layer gate + terminology + translation-smell + style + quality report + AIGC label check in one pass; Markdown report + full JSON; also pp_api.workflow()); TL;DR quick-start layer in both docs; historical release notes moved to CHANGELOG.md; consolidated anti-patterns section; EN FAQ aligned with ZH; superseded eval artifacts archived; SDK/setup errors carry recovery hints. Zero new dependencies.","fileCount":75,"zipByteSize":521386},{"version":"4.5.0","createdAt":"2026-10-03T19:44:24.822Z","changelog":"v4.5.0: programmable-interface release — new scripts/pp_api.py Python SDK (in-process detect_text with CLI parity; gate_text/term_report/smell_report/style_report/quality_report_file/attribution/model_fingerprint/doctor_summary; zero network, stdlib-only, JSON dicts, exceptions over silent failures) and scripts/pp_setup.py one-command supervised model installation (md5 verified against the author-signed fingerprint registry, unknown weights rejected, inference canary, --check status). Zero new dependencies.","fileCount":73,"zipByteSize":507910},{"version":"4.4.0","createdAt":"2026-10-03T04:43:20.000Z","changelog":"v4.4.0: measurement-integrity release — headline numbers re-measured fingerprint-bound: pre-2026 held-out AUROC 0.9998, current-gen 0.9400, human FPR@med 2.3% (supersedes the 0.9022/0.6542 artifacts: stale score-cache replay + unverified model lineage); run_eval cache keys bind the model md5 (model_fp in result JSONs); new eval/check_leak.py contamination audit halted the v36 retrain (285 eval-set samples had leaked into training; numbers void; supervised model stays v35 md5 2631df3d388b); ONNX export variable-length self-test; PP_ORT_THREADS; pp_doctor model fingerprint. Zero new dependencies.","fileCount":70,"zipByteSize":498051}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17dagtwyk21qs6vpz98bzcrh1853t29:paper-polisher-pro","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T07:42:05.082Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-docsor1212-paper-polisher-pro/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T01:31:08.446Z","emptyReason":null},"readme":"Skill: Paper Polisher Pro — AI Detector & Academic Polishing\n\nOwner: docsor1212\n\nSummary: AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell), metaphor audit, quality report, AIGC compliance label check (China 2025-09 labeling rules), paragraph-level attribution, journal precheck, sentence-level rewrite suggestions (locates and advises, never auto-rewrites), plus `--batch DIR` for thesis-scale batch rewriting (per-file AI-rate scores directory-wide). Bilingual CN/EN, 100% local, zero upload, zero credentials; bundled unit-test suite + AST-based zero-network self-verification. v3 delivers a recalibrated multi-layer rule engine (11 core layers + discourse/smoothness heuristics) + token-spectrum layer + length-routed fusion + optional supervised Qwen3-0.6B ONNX layer (AUROC 1.0 on held-out test) + LLM fingerprint attribution (GLM/DeepSeek/Qwen/Kimi/MiniMax/GPT/Claude/Gemini) + freshness pipeline. Base-engine numbers reproduce from the bundled held-out evaluation; supervised columns are author-side measurements (model not bundled).\n\nTags: latest:5.1.0\n\nVersion history:\n\nv5.1.0 | 2026-10-09T16:59:52.833Z | user\n\nv5.1.0: rewrite-closure & batch-report release — rewrite-effect regression check (pp_rewrite_check.py: original vs revised engine-source comparison — document score/risk-band migration, paragraph-level aligned deltas, feature-type clears vs remaining, edit extent; relative reference under this engine's criteria only; pp.py rewrite-check + pp_api.rewrite_check); batch HTML summary report (pp_batch_report.py: renders --batch --csv into a single self-contained HTML — cards, histogram, per-file table; zero deps, no engine re-run); batch-CSV concurrency semantics and degraded_notice reading guide documented in FAQ. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。\n\nv5.0.0 | 2026-10-08T18:57:11.895Z | user\n\nv5.0.0: verification & rewrite-suggestions release — sentence-level rewrite suggestions (pp_fix_suggest.py: which sentences, why, how to improve — locate & strategy, never auto-rewrite; wired into pp_workflow and pp.py fix); bundled unit-test suite (tests/, 44 stdlib cases, pp.py test); AST-structured zero-network self-verification (pp_verify.py); literary-narrative register notice (disclosure only, no scoring change); requirements.txt; pp.py quickstart zero-model demo. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。\n\nv4.9.0 | 2026-10-08T04:32:19.691Z | user\n\nv4.9.0: mixed-document calibration release — paragraph-level hi/med/lo thresholds calibrated on the controlled mixed benchmark (333 paragraphs with ground truth; best operating point precision 0.60 at 63% coverage — honestly below the automatic-verdict bar, ranking aid only; references/para_thresholds.json ships with the package and surfaces in mixed_document assessments); unified entry scripts/pp.py routes all eleven subcommands; batch CSV writes are atomic; gate layers retry once before fallback. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。\n\nv4.8.0 | 2026-10-07T04:32:14.303Z | user\n\nv4.8.0: measurement & output-contract integrity (independent third-party test round, 13 findings verified and addressed) — run_eval cache keys now bind engine mode (rules-only evaluation no longer replays supervised cached scores; reproducible fingerprint-and-mode-bound: rules 0.8985 old-gen / 0.7149 current-gen on the bundled sample); --profile journal --format json emits pure JSON with a journal_precheck field and both statistics explained; risk_bands (active thresholds + calibration tier) surfaced in every report; mixed_signal denoised (both extremes >=25% of paragraphs); single-file >5MB warning; --batch directory requirement in the error message; honest disclosures for real-world medical FPR and fingerprint attribution limits in the FAQ; corpus scope (n=1,304 bundled vs n=5,251 full) and terminology count (2,308 loaded) corrected. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。\n\nv4.7.0 | 2026-10-06T05:06:51.523Z | user\n\nv4.6.0: onboarding release — new scripts/pp_workflow.py end-to-end self-check command (AI-rate detection + paragraph attribution + 4-layer gate + terminology + translation-smell + style + quality report + AIGC label check in one pass; Markdown report + full JSON; also pp_api.workflow()); TL;DR quick-start layer in both docs; historical release notes moved to CHANGELOG.md; consolidated anti-patterns section; EN FAQ aligned with ZH; superseded eval artifacts archived; SDK/setup errors carry recovery hints. Zero new dependencies. 觉得有用的话，欢迎 Star（GitHub）或收藏（SkillHub 技能页）。\n\nv4.6.0 | 2026-10-05T04:30:46.544Z | user\n\nv4.6.0: onboarding release — new scripts/pp_workflow.py end-to-end self-check command (AI-rate detection + paragraph attribution + 4-layer gate + terminology + translation-smell + style + quality report + AIGC label check in one pass; Markdown report + full JSON; also pp_api.workflow()); TL;DR quick-start layer in both docs; historical release notes moved to CHANGELOG.md; consolidated anti-patterns section; EN FAQ aligned with ZH; superseded eval artifacts archived; SDK/setup errors carry recovery hints. Zero new dependencies.\n\nv4.5.0 | 2026-10-03T19:44:24.822Z | user\n\nv4.5.0: programmable-interface release — new scripts/pp_api.py Python SDK (in-process detect_text with CLI parity; gate_text/term_report/smell_report/style_report/quality_report_file/attribution/model_fingerprint/doctor_summary; zero network, stdlib-only, JSON dicts, exceptions over silent failures) and scripts/pp_setup.py one-command supervised model installation (md5 verified against the author-signed fingerprint registry, unknown weights rejected, inference canary, --check status). Zero new dependencies.\n\nv4.4.0 | 2026-10-03T04:43:20.000Z | user\n\nv4.4.0: measurement-integrity release — headline numbers re-measured fingerprint-bound: pre-2026 held-out AUROC 0.9998, current-gen 0.9400, human FPR@med 2.3% (supersedes the 0.9022/0.6542 artifacts: stale score-cache replay + unverified model lineage); run_eval cache keys bind the model md5 (model_fp in result JSONs); new eval/check_leak.py contamination audit halted the v36 retrain (285 eval-set samples had leaked into training; numbers void; supervised model stays v35 md5 2631df3d388b); ONNX export variable-length self-test; PP_ORT_THREADS; pp_doctor model fingerprint. Zero new dependencies.\n\nv4.3.0 | 2026-10-01T04:30:37.897Z | user\n\nv4.3.0: generation-split evaluation infrastructure — new 443-doc current-gen eval set (312 fresh AI from 9 families + 131 held-out humans); first quantified generation-gap numbers (pre-2026 corpus AUROC 0.9022 vs current-gen 0.6542); run_eval --layer import fix; spectrum-v2 and L13 current-gen negative results data-closed; supervised retraining on fresh samples planned.\n\nv4.2.0 | 2026-09-30T04:45:37.184Z | user\n\nv4.2.0: batch recursion (--recursive walks subdirectories, relative paths); GitHub README landing page (description, quick start, family table, compliance statement); gate layer-3 distribution verification on 40 held-out documents.\n\nv4.1.0 | 2026-09-29T04:00:36.499Z | user\n\nv4.1.0: smoothness layer (L13 surprisal-variation) fused into the score at 0.06 weight across all length bands — held-out A/B shows AUROC bit-identical to baseline (0.9022) and human FP unchanged; batch mode gains --csv per-file output; paragraph report carries mixed/register/encoding disclosures consistent with JSON.\n\nv4.0.0 | 2026-09-28T13:43:09.119Z | user\n\nv4.0.0: paper-workflow family referral loop (SkillHub display-name entries with the arXiv hallucinated-citation policy hook); task-word-root coverage in description (academic writing / polish / batch rewriting guidance / terminology); family-section cleanup. First release under the major-version numbering rule.\n\nv3.12.0 | 2026-09-28T04:45:31.933Z | user\n\nv3.12.0: translation-smell layer revived in deai_gate (schema mismatch fix — the layer had been silently neutral; it now genuinely contributes to the fused score); integrity notice on every report output; paragraph-count consistency; negative results recorded (spectrum v2 blend no-op, L13 mid-length declined on data).\n\nv3.11.0 | 2026-09-27T04:15:24.676Z | user\n\nv3.11.0: batch detection (--batch DIR — per-file scores + aggregate stats, deterministic, >5MB skipped) built for thesis-scale self-review; paragraph report now carries the academic-integrity notice; tier-2 n-gram negative result honestly recorded (285-sample mining yields only topic noise after guards). 100% local, zero new dependencies.\n\nv3.10.0 | 2026-09-26T04:15:30.001Z | user\n\nv3.9.0: honest-boundary guardrails — mixed-register signal (mixed_signal + per-paragraph guidance), register hint, encoding warning, gate layer-divergence disclosure. Includes v3.7.0 docs upgrade (centralized FAQ, script cheat sheet, clearer edge-case errors). 100% local, zero new dependencies.\n\nv3.9.0 | 2026-09-25T04:30:32.078Z | user\n\nv3.9.0: honest-boundary guardrails — mixed-register signal (mixed_signal + per-paragraph guidance), register hint, encoding warning, gate layer-divergence disclosure. Includes v3.7.0 docs upgrade (centralized FAQ, script cheat sheet, clearer edge-case errors). 100% local, zero new dependencies.\n\nv3.8.0 | 2026-09-24T12:05:07.785Z | user\n\nv3.8.0: honest-boundary guardrails — mixed-register signal (mixed_signal + per-paragraph guidance), register hint, encoding warning, gate layer-divergence disclosure. Includes v3.7.0 docs upgrade (centralized FAQ, script cheat sheet, clearer edge-case errors). 100% local, zero new dependencies.\n\nv3.7.0 | 2026-09-23T04:00:24.843Z | user\n\nv3.7.0: evaluation-driven docs upgrade — centralized FAQ section (9 answers), per-script cheat-sheet table, clearer style_distance edge-case errors. Targets TRACE C/R dimension feedback.\n\nv3.6.0 | 2026-09-20T20:10:31.511Z | user\n\nv3.6.0: academic-integrity guardrails (explicit integrity_notice in every report, new Academic Integrity section; positioning = author self-review and disclosure compliance, not detector evasion); English docs hardened to quality-framed wording per platform policy.\n\nv3.5.0 | 2026-09-19T19:47:05.268Z | user\n\nv3.5.0: supervised v3 line first public release (base engine AUROC 0.9187 held-out; optional supervised layer up to 1.0). New: degraded-mode disclosure, <100-char no-verdict iron law enforced, pp_doctor self-check, Chinese-Windows UTF-8 hardening, deai_gate fallback loop, honest bilingual docs, Paper Toolbox family section.\n\nv1.0.1 | 2026-04-30T05:37:34.616Z | user\n\nEnglish SKILL.md for ClawHub (was incorrectly in Chinese), streamlined descriptions, SEO keyword optimization, 3 Xiaohongshu promo posts\n\nv1.0.0 | 2026-04-29T04:18:04.952Z | user\n\nInitial release: AI detection engine (6-layer, 300+ rules, 14 model fingerprints), terminology checker (2255 terms), n-gram similarity analyzer, quality report generator. Supports Chinese and English academic papers.\n\nArchive index:\n\nArchive v5.1.0: 102 files, 756120 bytes\n\nFiles: CHANGELOG.md (39029b), data/terminology.json (319057b), eval/attack_gen.py (4568b), eval/build_mixed_bench.py (4412b), eval/calibrate_mixed_para.py (5053b), eval/check_leak.py (4872b), eval/corpus_builder.py (6397b), eval/gap_eval.py (4832b), eval/measure_mixed_para.py (5057b), eval/release_smoke.py (43797b), eval/results/archive/accept_v380_repro.json (1762b), eval/results/archive/baseline_v2.json (1785b), eval/results/archive/layer_surprisal_v3100_surprisal.json (130b), eval/results/archive/phase1_debt.json (1779b), eval/results/archive/phase2_test.json (1850b), eval/results/archive/v3100_check.json (1756b), eval/results/archive/v3100_fresh.json (1756b), eval/results/archive/v350_release.json (1852b), eval/results/archive/v36_gen2026.STALE-cache-poisoned.json (1342b), eval/results/archive/v36_oldgen.STALE-cache-poisoned.json (1755b), eval/results/gen2026.json (1099b), eval/results/gen2026v2.json (1340b), eval/results/gen2026v2spec.json (1344b), eval/results/layer_surprisal_gen2026.json (125b), eval/results/leak_audit_20261003.json (1538b), eval/results/mixed_bench_20261005.json (342819b), eval/results/mixed_para_20261005.json (10879b), eval/results/score_cache.json (543942b), eval/results/v3_final_test.json (1864b), eval/results/v3_fp_test.json (1863b), eval/results/v3_local_test.json (1864b), eval/results/v35ctl_gen2026.json (1375b), eval/results/v35ctl_oldgen.json (1740b), eval/results/v36fix_oldgen.json (1754b), eval/results/v370_fresh.json (1755b), eval/results/v370_fresh2.json (1756b), eval/results/v370_release.json (1757b), eval/results/v370_release2.json (1758b), eval/results/v380_merge.json (1755b), eval/results/v390_final.json (1755b), eval/results/v410_l13fusion.json (1759b), eval/results/v480_rules_gen2026.json (1387b), eval/results/v480_rules_oldgen.json (1799b), eval/results/v500check.json (1736b), eval/results/v500rules.json (1791b), eval/run_eval.py (12579b), LICENSE.md (918b), README.md (3252b), references/ai_patterns_en.json (13858b), references/ai_patterns_zh.json (7537b), references/article-review-workflow.md (3200b), references/fusion_config.json (1083b), references/model_fingerprints.json (11185b), references/para_thresholds.json (1767b), references/sentence_patterns_zh.json (10269b), references/supervised_models.json (636b), references/synonyms_general.json (58842b), references/token_spectrum_v2.json (184509b), references/token_spectrum_zh.json (194373b), requirements.txt (771b), scripts/ai_detector.py (55750b), scripts/aigc_label_check.py (7205b), scripts/build_spectrum.py (3157b), scripts/calibrate_v3.py (5587b), scripts/deai_gate.py (8482b), scripts/fingerprint_miner.py (5681b), scripts/freshness_refresh.py (3792b), scripts/layers_lm.py (9648b), scripts/layers_surface.py (8671b), scripts/ngram_similarity.py (7846b), scripts/paragraph_report.py (5949b), scripts/pattern_recalibrator.py (6619b), scripts/perplexity.py (12178b), scripts/pp_api.py (10872b), scripts/pp_batch_report.py (8753b), scripts/pp_doctor.py (11133b), scripts/pp_fix_suggest.py (15987b), scripts/pp_rewrite_check.py (11598b), scripts/pp_setup.py (6026b), scripts/pp_split.py (681b)\n\nFile v5.1.0:SKILL.md\n\n---\nname: paper-polisher\nversion: 5.1.0\nauthor: DoctorQ Lab\ndescription: >-\n  AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell),\n  metaphor audit, quality report, AIGC compliance label check (China 2025-09\n  labeling rules), paragraph-level attribution, journal precheck, sentence-level\n  rewrite suggestions (locates and advises, never auto-rewrites), plus `--batch DIR`\n  for thesis-scale batch rewriting (per-file AI-rate scores directory-wide).\n  Bilingual CN/EN, 100% local, zero upload, zero credentials; bundled unit-test suite\n  + AST-based zero-network self-verification.\n  v3 delivers a recalibrated multi-layer\n  rule engine (11 core layers + discourse/smoothness heuristics) + token-spectrum layer + length-routed fusion + optional\n  supervised Qwen3-0.6B ONNX layer (AUROC 1.0 on held-out test) + LLM\n  fingerprint attribution (GLM/DeepSeek/Qwen/Kimi/MiniMax/GPT/Claude/Gemini) +\n  freshness pipeline. Base-engine numbers reproduce from the bundled held-out\n  evaluation; supervised columns are author-side measurements (model not bundled).\ntags: [ai-detection, deai, academic-writing, paraphrase, paper-polish]\n---\n\n# Paper Polisher Pro v3\n\nAI writing detection (AI-rate self-check for authors) · academic polishing guidance · terminology standardization · translation-smell check · quality report · AIGC compliance label check · paragraph-level attribution · journal precheck.\n100% local, zero upload, zero credentials, pure standard library (optional onnxruntime enhancement layer).\n\n> ## ⛔ Iron laws\n> 1. **Only reproducible numbers.** Every metric comes from the held-out (test split) evaluation in `eval/run_eval.py`; unsupported claims like \"100% detection rate / F1 98.3%\" from older docs have been removed.\n> 2. **No verdict on short text.** Texts under 100 characters get `risk=unknown` (community lesson: short-text false positives are uncontrollable).\n> 3. **Fingerprints attribute, never score.** (Measured 2026-08-15: injecting fingerprints into the detector doubled human false positives.)\n> 4. **Calibration/evaluation separation.** Spectrum, weights and thresholds are built on the calib half only; the test half is reserved for final evaluation (an in-sample AUROC of 0.9972 collapsed to a real 0.9187 once split).\n\n## TL;DR\n\n- **What**: 100% local AI-rate self-check + academic polishing toolkit for Chinese academic text (optional supervised model for best accuracy; English gets advisory rules-only scores).\n- **30-second start**: `python3 scripts/pp.py quickstart` (zero-model, zero-file demo) · `python3 scripts/pp.py detect draft.txt --format json` · sentence-level rewrite suggestions: `python3 scripts/pp.py fix draft.txt` · full self-check report: `python3 scripts/pp.py workflow draft.txt` · environment: `python3 scripts/pp.py doctor` (one entry routes all subcommands)\n- **Measured** (held-out, fingerprint-bound md5 2631df3d388b): AUROC 0.9998 pre-2026 / 0.9400 current-generation; human FPR@medium 2.3%.\n- **Know the limits**: texts <100 chars get `risk=unknown` by design · medical text in degraded mode is over-scored · authors' self-check only — never for evading institutional AI detection.\n- **Where to look next**: capability boundary matrix below · end-to-end example in § Quick start · FAQ near the end · full history in `CHANGELOG.md`.\n\n## Academic integrity\n\nThis tool is for **authors self-reviewing and improving their own writing quality** — clearer sentences, consistent terminology, natural style. It is not designed to evade institutional AI-detection systems, and it must not be used to misrepresent AI-generated work as human-written. Follow your institution's AI-use and disclosure policies; the bundled `aigc_label_check.py` exists to help you **comply** with disclosure and labeling rules (e.g., China's 2025-09 labeling measures) — to declare AI assistance properly, not to hide it. Every AI-risk report (`ai_detector.py` / `deai_gate.py`) carries an explicit `integrity_notice` to this effect.\n\n## Measured performance (C-ReD + DetectRL-ZH, held-out test half, n=5,251)\n\n> Corpus scope note (v4.8.0): the v3.0-era table below was measured on the **full** held-out test half (n=5,251). The **bundled** sample corpus is a subset — its test half is n=1,304; reproducible per-corpus numbers: rules-only (PP_NO_SUP) 0.8985 old-gen / 0.7149 current-gen, fused 0.9998 / 0.9400 (`eval/results/v480_rules_*.json`, `v35ctl_*.json`; cache keys bind model fingerprint AND engine mode).\n\n| Metric | v2.0 baseline | v3.0 rules+spectrum | v3.1 +supervised | **v3.4 supervised + edit-regression v2** |\n|---|---|---|---|---|\n| AUROC (test half) | 0.7046 | 0.9187 | 0.9997 | **1.0** |\n| TPR@FPR5% | 30.4% | 49.0% | 99.95% | **100%** |\n| TPR@FPR1% | 16.7% | 24.9% | 99.88% | **100%** |\n| Human FPR @calibrated p99 | not measured | not measured | 3.56% (30/844) | **0.71% (6/844)** |\n| Paraphrase/mixed-attack AUROC | 0.64 | 0.89 | 1.0 (in-corpus) | **1.0** |\n| Attack \"AI-assisted\" recall | — | — | 71.1% | **86.6%** |\n| OOD plain-narrative/film recall | — | — | 1/6 | **5/6 supervised-only · 6/6 local fusion** |\n\n> **v4.4.0 fingerprint-bound re-measurement** of the shipping supervised model (md5 `2631df3d388b`): AUROC **0.9998** (test half, n_base=927), TPR@FPR1% 99.4%, human FPR@medium 2.3% — `eval/results/v35ctl_oldgen.json`. Columns above are preserved as version-era records (earlier model lineage; binaries were not fingerprinted before v4.4.0).\n\n**Which column applies to you?** The base package runs the **v3.0 rules+spectrum engine** (0.9187 AUROC column, measured on the full held-out corpus; pre-4.4.0 archived baselines predate fingerprint binding — every eval result since v4.4.0 carries the deployed model's md5 as `model_fp`, current bound numbers in `eval/results/v35ctl_*.json`). The two right-hand columns require the optional local supervised model (see below). The engine tells you honestly which mode you are in: every report carries `degraded_mode` / `degraded_notice` when the supervised layer is absent or skipped.\n\n### Capability boundary matrix (read before trusting any detector)\n\n| Scenario | Behavior |\n|---|---|\n| Chinese academic prose, full stack | Best case (AUROC 0.9998 held-out, human FPR@medium 2.3% — v4.4.0 fingerprint-bound) |\n| Base package without model | Rules+spectrum (0.9187); **medical register over-scored** (rules-only human FPR @medium: ~59% medical vs ~2% general) → trust only @high verdicts on medical text |\n| English text | Language gating skips the Chinese-trained supervised layer by design; rules-only English skeleton, advisory only |\n| Mixed human+AI documents | Document-level AUROC 0.52-0.54 (inherent averaging limitation); paragraph-level AUROC 0.69 with **calibrated best operating point P=0.60 at 63% coverage** (`references/para_thresholds.json`) — below the automatic-verdict bar; use `pp_workflow.py`/`paragraph_report.py` rankings for human review only |\n| Edit-extent regression head | ρ=0.540 — reported as metadata, never used in verdicts |\n| **Current-generation models (2026-09 sampling)** | **AUROC 0.9400** (v4.4.0 fingerprint-bound re-measurement, 443-doc current-gen eval set: 9 families incl. K3/K2.7/Qwen3.7-3.8/DS-V4/V4.1/GLM-5.3/M3) vs 0.9998 pre-2026 held-out — a modest verified gap. The earlier 0.6542-vs-0.9022 figure was a measurement artifact (stale score-cache replay + unverified model lineage); both classes are structurally prevented since v4.4.0 (`model_fp` in every result JSON)\n| Colloquial / oral-register text | The style layer is calibrated on academic prose; treat style scores as advisory outside that register |\n\n## Safety and behavior statement\n\n- **100% local**: every feature runs on-device. The codebase makes zero network calls — no network client libraries, no network utilities. Verify structurally: `python3 scripts/pp_verify.py` (AST-level scan; exit 0 = zero network calls, all scripts compile). Legacy text grep `grep -rEin \"urllib|requests|socket|import http\" scripts/` may show a few URL *strings* in report footers — those are data in string constants, not network code; the AST verifier distinguishes the two.\n- **No upload, no credentials**: reads and transmits no credentials, keys, or personal data; the only environment variable, `PP_NO_SUP`, is a local behavior toggle.\n- **No persistence**: creates no scheduled tasks, autostart entries, or system config changes; temp files (inter-layer JSON, probe text) are deleted after use.\n- **No remote code**: loads no remote models or scripts; the optional supervised model is placed by the user at a local path.\n- **Data boundary**: reads/writes only user-specified files, the system temp dir, and its own package data directories (calibration/freshness artifacts); reports go only where the user points them.\n- **Academic integrity**: see the section above — for author self-review and quality improvement with policy-compliant disclosure; not for evading detection.\n\n## What's new in v5.1.0\n\n- **Rewrite-effect regression check (`scripts/pp_rewrite_check.py`, also `pp.py rewrite-check`)**: the loop-closer after the rewrite suggestions — *how much did your revision actually change?* Engine-source comparison of the original vs the revised draft: document-level score and risk-band migration, paragraph-level difflib-aligned per-paragraph deltas, feature-type counts cleared vs remaining (same seven types as `pp_fix_suggest`), and edit extent (char ratio + replaced-paragraph rate). Relative reference under this engine's criteria only — never an institutional verdict; ships with `--json`, `--demo`, and `pp_api.rewrite_check()` for programmatic use.\n- **Batch HTML summary report (`scripts/pp_batch_report.py`, also `pp.py batch-report`)**: `--batch --csv` now renders into a single self-contained HTML — totals/mean/risk-band cards, a score histogram, and a per-file table sorted by score with ERROR/unknown rows surfaced. Zero dependencies, no engine re-run (renders the existing CSV).\n- **Docs**: the batch CSV concurrency semantics (atomic temp+rename since v4.9.0 — concurrent batch runs cannot clobber each other's CSV) and a reading guide for `degraded_notice` are now stated explicitly (see FAQ).\n\n## What's new in v5.0.0\n\n- **Sentence-level rewrite suggestions (`scripts/pp_fix_suggest.py`, also `pp.py fix`)**: the natural next question after a score — *which sentences, why, and how to improve them*. Each flagged sentence lists its concrete features (AI clichés, filler phrases, template patterns, vague qualifiers, connective openers, dash/colon habits, uniform rhythm) with a per-type rewrite strategy. Guidance only: it locates and suggests, never auto-rewrites — the editing decision stays with the author. Ships with `--json` for programmatic use and a built-in two-sample demo (`--demo`).\n- **Bundled unit-test suite (`tests/`, `pp.py test`)**: 44 stdlib-unittest cases covering the iron laws (short text/empty/GBK), report field contracts, JSON purity, the gate, the workflow Markdown layout, data files, and the new tools — runnable in seconds without the model, so anyone can verify behavior on their own machine.\n- **Structured zero-network self-verification (`scripts/pp_verify.py`, also `pp.py verify`)**: replaces the old grep advice with an AST-level scan of every script — catches network imports/calls and curl/wget-style subprocess commands, while URL *strings* in report footers are correctly treated as data. Exit 0 = clean; `--json` for pipelines.\n- **Register awareness 2.0**: literary-narrative texts (dialogue quotes + time progression + inner monologue cues) now get a dedicated register notice explaining that this register sits outside the academic calibration domain — measured literary classics can reach high band in this engine — so the result is not mistaken for AI evidence. Disclosure only; no scoring change.\n- **`requirements.txt`** ships with the package: core = zero third-party dependencies; the two optional supervised-layer deps (onnxruntime/numpy) are declared and commented.\n- **`pp.py quickstart`**: zero-model, zero-file one-command demo (detect → fix → doctor) on a built-in sample.\n- **Freshness visible in `pp_doctor`**: the doctor now reports the latest held-out evaluation record alongside the fingerprint-registry coverage row.\n\n## What's new in v4.9.0\n\n- **Mixed-document calibration (closing the v4.7.0 backlog)**: paragraph-level hi/med/lo thresholds are now calibrated on the controlled mixed benchmark (333 paragraphs with ground truth; `eval/calibrate_mixed_para.py` → `references/para_thresholds.json`). Honest result: the best operating point (≥50) reaches precision 0.60 at 63% AI-paragraph coverage — below the automatic-verdict bar, so paragraph attribution remains a ranking aid for human review; the calibrated numbers and their scope ship in `references/para_thresholds.json` and surface in every `mixed_document` assessment.\n- **Unified entry (`scripts/pp.py`)**: one command routes all eleven subcommands (detect/gate/workflow/term/smell/style/quality/aigc/paragraph/setup/doctor) — no more script-navigation cost.\n- **Reliability**: batch CSV writes are now atomic (temp+rename — concurrent batch runs no longer clobber the same CSV); gate layers retry once on crash/timeout before falling back.\n\n## Anti-patterns (avoid these)\n\n- **Don't feed <100 chars** and expect a verdict — `risk=unknown` is by design (short-text false positives are uncontrollable); 300+ chars recommended.\n- **Don't trust degraded-mode scores on medical text** — rules-only over-scores medical register (~59% human FPR @medium); install the supervised model or trust only `@high`.\n- **Don't treat scores as CNKI/Wanfang equivalents** — thresholds are calibrated on our own held-out corpus; self-check only.\n- **Don't substitute or re-quantize the model file** — measured probability drift; only author-signed fingerprints pass `pp_setup.py`.\n- **Don't use it to evade institutional AI detection** — the integrity notice ships on every report; disclose per your institution's policy.\n- **Don't run batch on >5 MB files** — skipped by design; split first.\n\n## Architecture (v3)\n\n```\nai_detector.py            Main engine: 8 rule layers (125 recalibrated patterns, markdown caps,\n                          EN openers, paragraph-level language) + length-routed fusion\n + layers_surface.py      L9 surface stats L10 token-spectrum (9,955-token delta spectrum)\n                          L11 chain-of-thought features\n + ai_detector L12        discourse-structure heuristics (v3.7.0: hook/reversal/slogan/engagement)\n + fusion_config.json     Weights & thresholds (calib-half grid search + human p95/p99)\n + model_fingerprints.json v4 fingerprint registry (13 families incl. GLM-5.3 & Kimi K-series self-sampled; attribution only)\n + layers_lm.py           Optional supervised layer (local ONNX + pure-Python Qwen tokenizer;\n                          PP_NO_SUP=1 falls back to rules)\nparagraph_report.py       Paragraph-level attribution HTML (pattern×spectrum 50/50 fusion)\naigc_label_check.py       AIGC compliance labels (China labeling rules 2025-09: metadata/C2PA/explicit)\nfingerprint_miner.py      Fingerprint mining (new model drop → sample → mine → register)\npattern_recalibrator.py   Data-driven pattern recalibration (human-hit filtering)\nbuild_spectrum.py / calibrate_v3.py   Spectrum build / weight calibration\nfreshness_refresh.py         Monthly freshness pipeline (sample → rebuild → calibrate → regression)\npp_doctor.py              Environment self-check (v3.5; v5.0.0 adds latest-eval-record row)\npp_verify.py              AST-level structured zero-network self-verification (v5.0.0)\npp_fix_suggest.py         Sentence-level rewrite suggestions (v5.0.0: locate + strategy, no auto-rewrite)\ntests/                    Bundled unit-test suite, `python3 -m unittest discover -s tests -t .` (v5.0.0)\nrequirements.txt          Dependency declaration: core zero-dep; optional supervised-layer extras (v5.0.0)\neval/                     corpus_builder / attack_gen / run_eval (AUROC, TPR@FPR, per-model, attack decay)\n```\n\n## Quick start\n\n```bash\n# Zero-model, zero-file one-command demo (v5.0.0)\npython scripts/pp.py quickstart\n# AI writing detection (probability + layered evidence + fingerprint attribution)\npython scripts/ai_detector.py draft.txt --format json\n# Sentence-level rewrite suggestions (v5.0.0: which sentences, why, how to improve)\npython scripts/pp_fix_suggest.py draft.txt --top 10\n# Rewrite-effect regression check (v5.1.0: original vs revised, engine-source comparison)\npython scripts/pp_rewrite_check.py draft_original.txt draft_revised.txt\n# Batch CSV -> self-contained HTML summary (v5.1.0)\npython scripts/pp_batch_report.py scores.csv -o report.html\n# Journal precheck (suspected-AIGC ratio vs the 20-25% reference line, non-interchangeable disclaimer)\npython scripts/ai_detector.py draft.txt --profile journal\n# Paragraph-level attribution (locate human/AI collaboration)\npython scripts/paragraph_report.py draft.txt --output report.html\n# AIGC compliance label check (docx/pdf/png/txt)\npython scripts/aigc_label_check.py manuscript.docx figures/*.png\n# Terminology / translation smell / 4-layer gate (same as v2)\npython scripts/term_check.py draft.txt --auto-fix\npython scripts/translation_smell_check.py draft.txt\npython scripts/deai_gate.py draft.txt\n# Environment self-check\npython scripts/pp_doctor.py\n# Structured zero-network self-verification + bundled unit tests (v5.0.0)\npython scripts/pp_verify.py\npython -m unittest discover -s tests -t .          # or: python scripts/pp.py test\n# Held-out regression (mandatory after any engine change)\npython eval/run_eval.py --split test --tag mytag\n```\n\n### Optional supervised layer (recommended, v3.2+)\n\n```bash\npip install onnxruntime regex          # the two optional dependencies\n# Place the two model files exactly as shipped by the authors:\n#   ~/.cache/paper-polisher/qwen3-detector/model.int8.onnx\n#   ~/.cache/paper-polisher/qwen3-detector/tokenizer.json\npython scripts/layers_lm.py            # self-test: supervised_available: true\n# ai_detector.py fuses automatically afterwards (0.9*supervised + 0.1*rules);\n# PP_NO_SUP=1 temporarily falls back to rules-only.\n# ⚠️ Do not substitute other exports or quantizations — measured probability drift; use exactly these files.\n```\n\n### Python API (programmatic use)\n\n```python\nimport sys; sys.path.insert(0, \"<skill>/scripts\")\nfrom pp_api import detect_text, gate_text, doctor_summary\nr = detect_text(\"中文学术文本，建议 300 字以上。\" * 10, lang=\"zh\")\nprint(r[\"overall_ai_score\"], r[\"overall_risk\"], r[\"degraded_mode\"])\n```\n\n`detect_text` runs in-process (no subprocess) and returns the same JSON structure as the CLI. Every function returns JSON-able dicts and raises on bad input — no silent failures. Zero network, stdlib-only.\n\n### One-command supervised setup\n\n```bash\npython3 scripts/pp_setup.py --model <author-signed model.onnx>   # verify md5 -> install -> inference canary\npython3 scripts/pp_setup.py --check                              # current installation status\n```\n\nOnly author-signed fingerprints (`references/supervised_models.json`) are accepted; unknown weights are rejected before anything is touched. Nothing is downloaded — the model always comes from the authors' channel as a local file.\n\n### End-to-end workflow (one command)\n\n```bash\npython3 scripts/pp_workflow.py draft.txt      # writes draft.workflow.md + draft.workflow.json\n```\n\nRuns the full self-check in one pass — AI-rate detection (fused engine), paragraph-level\nattribution with hi/med/lo classification, the 4-layer gate, terminology, translation-smell,\nstyle, quality report and AIGC label self-check — and produces a single readable Markdown\nreport plus machine-readable JSON. Programmatic: `from pp_api import workflow`.\n\n## FAQ\n\n**Q: Why no risk verdict for texts under 100 characters?**\nShort-text false positives are uncontrollable (a few sentences carry no style distribution). The tool returns `risk=unknown` by design; submit 100+ chars (300+ recommended).\n\n**Q: Why is my medical text scored high?**\nYou are most likely in degraded mode (optional supervised model not installed). Rules-only scoring systematically over-scores medical register (held-out human FPR at @medium: ~59% medical vs ~2% general). For medical text trust only @high verdicts, or install the supervised layer (next question).\n\n**Q: How do I install the supervised model and confirm it works?**\none command: `python3 scripts/pp_setup.py --model <author-signed model.onnx>` — it verifies the md5 against the signed registry, installs, and runs an inference canary (GREEN = active; `degraded_mode=false` in reports confirms it). Manual placement of the two files at `~/.cache/paper-polisher/qwen3-detector/` still works and `python scripts/pp_doctor.py` remains the full check.\n\n**Q: What do degraded_mode / degraded_notice mean?**\nEngine-mode disclosure: true = rules+spectrum fallback, reason in the notice (model missing / PP_NO_SUP=1 / English language gating). See the capability boundary matrix.\n\n**Q: Is this score interchangeable with CNKI/Wanfang official checks?**\nNo. Thresholds are calibrated on our own held-out corpus and are not interchangeable with any institutional detector; self-check only (stated in journal-profile output too).\n\n**Q: A deai_gate layer shows \"解析失败\" (parse failure) — what now?**\nThat layer falls back to a neutral 50; other layers and the verdict are unaffected. Usually a subprocess timeout or odd input encoding; retry once, then run `pp_doctor.py`.\n\n**Q: What is the edit-extent estimate?**\nA supervised-layer regression head estimating how much the text was AI-edited (0-1). Limited discriminative power (ρ=0.54) — report metadata only, never used in verdicts.\n\n**Q: Is English supported?**\nPartially: the supervised layer is Chinese-trained, so English skips fusion by design and gets rules-only skeleton scoring, advisory only (stated in the report).\n\n**Q: What about documents that mix human and AI writing?**\nWatch the mixed-register signal (`mixed_signal=true`): document-level scores are diluted by human paragraphs or pushed up by AI ones — unreliable either way. Run `paragraph_report.py` for per-paragraph attribution and work paragraph by paragraph.\n\n**Q: Why does a real, human-written journal paper still score medium/high?**\nTwo measured reasons: distribution shift (our held-out corpus differs from real journal PDF→text, which carries layout noise) and register calibration. Paragraph attribution is the actionable signal — use it to locate suspect passages for human review; the document-level score is a triage hint, not a verdict. The journal precheck (distribution口径) offers a second view and may disagree with the main score by design.\n\n**Q: The fingerprint attribution says GPT-4o but my text is from another model?**\nAttribution is heuristic (top-n candidates, never scored) and may misattribute — known case: GLM-generated text has been attributed elsewhere. Treat family hints as weak evidence; the detection score and paragraph attribution are the substantive outputs.\n\n**Q: How do I use the AIGC label check?**\n`python scripts/aigc_label_check.py manuscript.docx figures/*.png` — checks metadata / C2PA watermark / explicit declaration (China 2025-09 labeling rules). Exit 0 = labeled, 1 = unlabeled; both are normal runs.\n\n**Q: How can I verify the \"100% local / zero upload\" claim myself?**\nRun `python3 scripts/pp_verify.py` — an AST-level structural scan of every script. It flags network imports/calls and curl/wget-style subprocess commands, while URL strings in report footers are correctly treated as data (the old grep advice could not tell the two apart). Exit 0 = zero network calls. For behavior-level checks, the bundled test suite (`python -m unittest discover -s tests -t .`) exercises the iron laws and report contracts on your own machine.\n\n**Q: Can several batch runs write to the same CSV safely?**\nYes — since v4.9.0 batch CSV writes are atomic (write to a temp file, then rename into place). Two concurrent batch runs pointing at the same CSV cannot interleave or clobber each other's rows; the file always contains one complete run's output.\n\n**Q: How should I read `degraded_notice`?**\nIt tells you which full-mode ingredient was skipped and why: supervised model missing, `PP_NO_SUP=1`, or English language gating. Consequences differ — degraded rules-only mode systematically over-scores medical register (see the boundary matrix), while English gating only means the Chinese-trained supervised layer does not apply. The notice names the reason so you can decide whether to install the model, unset the toggle, or treat scores as advisory.\n\n**Q: What do the rewrite suggestions do — do they change my text?**\nNo. `pp_fix_suggest.py` (`pp.py fix`) only locates sentences carrying improvable features and explains the improvement strategy per feature type (e.g., replace an AI cliché with a concrete claim, split a template sentence). It never edits your file; the editing decision and execution stay with the author. `--demo` shows a built-in two-sample walkthrough.\n\n### Script cheat sheet\n\n| Script | Purpose | Key flags | Output |\n|---|---|---|---|\n| ai_detector.py | Main AI-writing detector | `--lang auto\\|zh\\|en` `--format json\\|text\\|summary` `--profile journal` `--batch DIR` | Score + paragraph detail + fingerprints (JSON incl. degraded_mode/integrity_notice) |\n| pp_fix_suggest.py | Sentence-level rewrite suggestions (v5.0.0) | `--top N` `--json` `--demo` | Per-sentence features + rewrite strategies (locates and suggests, never auto-rewrites) |\n| pp_doctor.py | Environment self-check | `--json` | Data/deps/model probes; exit 0 = green |\n| pp_rewrite_check.py | Rewrite-effect regression check (v5.1.0) | `original revised` `--json` `--demo` | Score/risk-band migration + paragraph deltas + feature-type clears/remaining |\n| pp_batch_report.py | Batch CSV → HTML summary (v5.1.0) | `csv_file` `-o out.html` | Totals/mean/band cards + histogram + per-file table |\n| pp_verify.py | Structured zero-network self-verification (v5.0.0) | `--json` | AST-level scan verdict; exit 0 = zero network calls |\n| quality_report.py | Overall quality report | `--json` | Readability/quality dimensions for the manuscript |\n| deai_gate.py | 4-layer fused gate | `--json` | composite score + verdict band (<35 pass / 35-55 review / ≥55 suspect) |\n| paragraph_report.py | Paragraph attribution | `--output report.html` | HTML report |\n| term_check.py | Terminology (2,308 terms) | `--auto-fix` `--output` | Standardization rate + fixed file |\n| translation_smell_check.py | Translation-smell scan | `--json` | Hits + blind-spot terms |\n| style_distance.py | Stylometry (human-likeness) | `--json` | style_score + verdict (advisory outside academic register) |\n| aigc_label_check.py | AIGC compliance labels | files: docx/pdf/png/txt | Per-file label verdict |\n\n## Trigger words (Chinese)\n\n`润色论文`查AI率` `论文AI率` `AIGC检测` `AIGC率` `GPT检测` `查AI写作` `论文润色` `改写论文` `AI论文检测` `学术写作助手` `AI写作检测` `毕业论文润色` `学位论文降重` `SCI论文编辑` `手稿润色` `AI写作评分` `AI改写检测` `文风对标顶刊` `这篇文章像不像AI`\n\n## Related tools\n\n- **cn-med-oa** — free Chinese medical literature (OA) download & citation metadata\n- **pubmed-verifier** — verify PMID/DOI references before submission\n- **cite-holmes** — deep research with machine-verified citations\n- **academic-figures** — publication-ready scientific figures in one command\n- **doc-holmes** — layout-preserving PDF translation\n- **paper-rewriter** — same-source de-AI rewriting companion (full rewrite pipeline)\n\nDocs & site: **docsor.cn**\n\n## Fingerprint freshness (against \"detectors lag one generation\")\n\nCoverage as of 2026-09-26: kimi-k3 & kimi-k2.7 registered (OpenCode Go fresh sampling, attribution-verified); qwen3.8 / deepseek-v4 / deepseek-v4.1 / minimax-m3 sampled — mining produced no family-distinctive low-FP patterns, honestly unregistered; glm-5.3 refreshed (no new patterns). Next: deepseek-v4.1 & minimax-m3 with larger corpora.\n\nOn a new-model release day: `python scripts/fingerprint_miner.py --corpus <new_samples.jsonl> --model <family> --apply`\nMonthly full pass: `python scripts/freshness_refresh.py` (schedule it with your own system timer, e.g. monthly; the script never creates or modifies system schedules). Compare adjacent `eval/results/freshness_*.json`; investigate if AUROC drops by more than 3 percentage points.\n\n## Version history (condensed)\n\n- **v5.1.0 (2026-10-10)** — rewrite-closure & batch-report release: rewrite-effect regression check (`pp_rewrite_check.py`, engine-source original-vs-revised comparison); batch HTML summary report (`pp_batch_report.py`); batch-CSV concurrency and degraded_notice reading guide documented.\n- **v5.0.0 (2026-10-09)** — verification & rewrite-suggestions release: bundled unit-test suite (44 stdlib cases, `pp.py test`); AST-level zero-network self-verification (`pp_verify.py`); sentence-level rewrite suggestion engine (`pp_fix_suggest.py` — locate + strategy, never auto-rewrite, also wired into `pp_workflow` and `pp.py fix`); literary-narrative register notice (disclosure only, no scoring change); `requirements.txt`; `pp.py quickstart`; doctor freshness row.\n- **v4.9.0 (2026-10-08)** — mixed-document calibration release: paragraph thresholds calibrated on the 333-paragraph benchmark (best operating point P=0.60/coverage 63% — honestly below auto-verdict bar; ranking aid only); unified `pp.py` entry; atomic batch CSV; gate layer retry.\n- **v4.8.0 (2026-10-07)** — measurement & output-contract integrity (independent test round, 13 findings addressed): eval cache keys bind engine mode (rules-only reproducible: 0.8985/0.7149 bundled); journal+json purity; risk_bands surfaced; mixed_signal denoised; 5MB/batch CLI guards; real-world FPR & attribution limits disclosed in FAQ; corpus/terminology counting precision.\n- **v4.7.0 (2026-10-06)** — mixed-document special: controlled per-paragraph benchmark quantifies paragraph-level AUROC 0.69 (document-level 0.52-0.54); pp_workflow gains mixed_document assessment + prominent mixed flagging; description Trigger-on routing words; README China mirror link.\n- **v4.6.0 (2026-10-05)** — onboarding release: end-to-end workflow command (pp_workflow.py / pp_api.workflow); TL;DR layer; historical notes moved to CHANGELOG.md; consolidated anti-patterns section; eval archive cleanup; recovery hints in SDK errors.\n- **v4.5.0 (2026-10-04)** — programmable-interface release: pp_api.py SDK (in-process detect + 8 helpers, zero network, JSON dicts) and pp_setup.py one-command model installation with the author-signed fingerprint registry; closing the two biggest usability friction points reported with the 4.4.x interface (programmatic integration and model setup).\n- **v4.4.0 (2026-10-03)** — measurement-integrity release: verified re-baseline (held-out 0.9998 / current-gen 0.9400 / human FPR@med 2.3%, superseding 0.9022/0.6542 artifacts); eval cache keys bind model md5; contamination audit (eval/check_leak.py) halts the v36 retrain (285 eval-set samples had leaked into training); ONNX export self-test; PP_ORT_THREADS; pp_doctor fingerprint; smoke degradation checks made mode-aware.\n- **v4.3.0 (2026-10-01)** — generation-split eval infrastructure; first quantified generation-gap numbers (0.9022 vs 0.6542); spectrum/L13 current-gen negative results recorded.\n- **v4.2.0 (2026-09-30)** — batch recursion; GitHub README landing page; gate layer-3 distribution verification.\n- **v4.1.0 (2026-09-29)** — smoothness layer fused into the score (A/B-verified zero regression); batch CSV; paragraph report disclosures.\n- **v4.0.0 (2026-09-28)** — family referral loop; task-word-root coverage; family-section cleanup.\n- **v3.12.0 (2026-09-28)** — translation-smell layer revived (schema fix); integrity notice on every report; paragraph-count consistency; spectrum-v2 & L13-mid negative results recorded.\n- v3.11.0 (2026-09-27) — batch detection (`--batch DIR`); paragraph report integrity notice; tier-2 n-gram negative result recorded.\n- v3.10.0 (2026-09-26) — fingerprint freshness phase 3: kimi-k2.7 registered (attribution-verified); contaminated candidates rolled back per quality gate.\n- **v3.9.0 (2026-09-25)** — discourse smoothness disclosure (L13 surprisal-variation, standalone AUROC 0.8256 held-out; NOT fused per iron law); layer-evaluation mode in run_eval (`--layer`).\n- v3.8.0 (2026-09-24) — mixed-register signal, register hint, encoding warning, gate layer-divergence disclosure; safety & behavior statement; qwen3.8/v4 fingerprint mining (honestly unregistered).\n- v3.7.0 (2026-09-23) — discourse-structure heuristic layer L12 (8 groups, density-scaled cap 30; LES-20260923-021 blind-spot fix; AUROC 0.9022 unchanged); kimi-k3 fingerprint registered (OpenCode Go sampling); centralized FAQ; script cheat sheet; documentation wording cleanup.\n- v3.6.0 (2026-09-21) — academic-integrity guardrails (integrity_notice + section); CH-side wording cleanup.\n- **v3.5.0 (2026-09-20)** — degraded-mode disclosure (engine mode + medical-register warning with held-out numbers); iron law #2 enforced (<100 chars → risk=unknown, quality_report shows \"cannot judge\" instead of misleading green); `pp_doctor.py` self-check; `deai_gate.py` usage guard; honest dual-language docs rebuild.\n- v3.4.3 — markdown table-separator rows filtered from paragraph scoring (6/8 flagged rows in real MD manuscripts were false positives).\n- v3.4.2 — fixed CJK double-count in language detection (Chinese journal PDFs misrouted to EN rules); degenerate PDF hard-line-break paragraph rebuilding (747→19 segments); paragraph-level language routing dead code fixed.\n- v3.4.1 — language gating: English text skips the Chinese-trained supervised layer (measured EN OOD p_ai=0.9996 → EN false positives 91.7→17.2). Rule-editor experiment: negative result, honestly abandoned.\n- v3.4.0 — edit-extent regression v2 (1,620 pairs, token-level distance, two-stage training): human FPR@p99 1.66%→0.71%; attack \"AI-assisted\" recall 86.6%.\n- v3.2/v3.3 — supervised layer v3.2 (4-dim head, local ONNX fp16, pure-Python Qwen tokenizer); OOD blind spots honestly recorded then closed (film-register recall 1/6→5/6, GLM-5.3 probe 9/9).\n- v3.1 — Qwen3-0.6B LoRA supervised layer (AUROC 0.9997 held-out).\n- v3.0 — eval-driven rebuild: recalibrated pattern library (693→125 patterns, 568 dead/inverted signals removed), token-spectrum layer, length-routed fusion, calib/test leak-proof split, fingerprint registry v4, paragraph attribution, AIGC label check, journal precheck, freshness pipeline, honest docs. AUROC 0.7046→0.9187.\n- v2.0.x — 9-layer rule engine + terminology library (baseline column above; non-reproducible claims removed).\n\nFile v5.1.0:README.md\n\n# Paper Polisher Pro — 论文降AI润色工具 · AI率检测\n\n[![GitHub Stars](https://img.shields.io/github/stars/docsor1212/paper-polisher-pro?style=social&label=Star)](https://github.com/docsor1212/paper-polisher-pro)\n\nAI 痕迹检测（AI率）· 去AI化改写建议 · 句子级改写建议（哪几句像AI、怎么改）· 术语标准化 · 翻译腔检查 · 质量报告 · AIGC 合规标识检查 · 段落级归因 · 期刊口径预检。\n\n**100% 本地运行，零上传，零凭证**——论文数据不出本机。零网络承诺可用包内 `pp_verify.py`（AST 结构化扫描）自行验证，行为契约可用随包 `tests/` 单元测试套件（44 用例）在自己机器上复跑。\n\n## 这是什么\n\n面向学术写作者的 AI 痕迹自查工具：概率化输出（非二元判定）、分层证据、指纹归因（GLM / DeepSeek / Qwen / Kimi / MiniMax / GPT / Claude / Gemini），全部指标可由随包留出集评测复现。\n\n- 基础引擎（纯规则+词频谱）：留出集 AUROC 0.9187\n- 可选监督层（本地 Qwen3-0.6B ONNX）：AUROC 1.0（作者侧实测）\n- 短文本不出判定（<100 字，误报铁律）\n\n## 快速开始\n\n```bash\n# 零模型零文件一键体验\npython scripts/pp.py quickstart\n\n# AI 痕迹检测（AI率）\npython scripts/ai_detector.py draft.txt --format json\n\n# 句子级改写建议（定位+策略，不代改）\npython scripts/pp_fix_suggest.py draft.txt --top 10\n\n# 四层融合门禁\npython scripts/deai_gate.py draft.txt\n\n# 批量检测一个目录\npython scripts/ai_detector.py --batch ./drafts --csv scores.csv\n\n# 环境自检\npython scripts/pp_doctor.py\n\n# Python 编程接口（零网络，import 即用）\npython -c \"import sys; sys.path.insert(0,'scripts'); from pp_api import detect_text; \\\nprint(detect_text(open('draft.txt').read())['overall_ai_score'])\"\n\n# 监督层一键装模（作者签发模型文件，指纹校验+推理自检）\npython scripts/pp_setup.py --model <作者签发模型.onnx>\n```\n\n## 安装\n\n克隆本仓库后直接使用，纯 Python 标准库即可运行（可选 onnxruntime 增强监督层，依赖声明见 `requirements.txt`）。\n\n```bash\ngit clone https://github.com/docsor1212/paper-polisher-pro\ncd paper-polisher-pro\npython scripts/ai_detector.py your_draft.txt --format summary\n```\n\n**China mirror (ModelScope 魔搭)**: <https://modelscope.cn/skills/Docsor/paper-polisher-pro> — if you find this skill useful, a like there helps others find it.\n\n## 论文工作流家族\n\n写作是一条链，每环有专用工具（均在本账号下）：\n\n| 工具 | 用途 |\n|---|---|\n| **paper-polisher-pro**（本仓库） | AI率检测·润色·降重·质量报告 |\n| **paper-rewriter** | 论文降AI改写执行 |\n| **pubmed-verifier** | PMID/DOI 引用核验 |\n| **cite-holmes** | 深度调研 × 引用自证 |\n| **cn-med-oa** | 中文医学文献 OA 下载 |\n| **academic-figures** | 出版级科研图表 |\n| **doc-holmes** | PDF 精准翻译 |\n\n文档站：[docsor.cn](https://docsor.cn)\n\n## 合规声明\n\n本工具供作者自查与写作质量改进，**不用于规避机构的 AIGC 检测**；请遵循所在机构的 AI 使用与披露政策（标识合规可用包内 aigc_label_check 自查）。许可：MIT-0。\n\nFile v5.1.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b6sjwend7cwwmhf7mdg1wzx82dfbf\",\n  \"slug\": \"paper-polisher-pro\",\n  \"version\": \"5.1.0\",\n  \"publishedAt\": 1791565192833\n}\n\nFile v5.1.0:references/ai_patterns_en.json\n\n{\n  \"version\": \"3.1.0\",\n  \"description\": \"English AI writing pattern library\",\n  \"categories\": {\n    \"filler_phrases\": {\n      \"weight\": 3,\n      \"description\": \"High-frequency AI filler phrases (strong signal)\",\n      \"patterns\": [\n        \"it is worth noting that\",\n        \"it is important to note that\",\n        \"it should be noted that\",\n        \"it is worth mentioning that\",\n        \"it is crucial to understand\",\n        \"it is essential to recognize\",\n        \"it is imperative to\",\n        \"it is paramount to\",\n        \"in conclusion\",\n        \"to summarize\",\n        \"in summary\",\n        \"to sum up\",\n        \"all in all\",\n        \"at the end of the day\",\n        \"when all is said and done\",\n        \"needless to say\",\n        \"it goes without saying\",\n        \"as a matter of fact\",\n        \"in today's world\",\n        \"in this day and age\",\n        \"in the modern world\",\n        \"in recent years\",\n        \"with the development of\",\n        \"with the advancement of\",\n        \"in the era of\",\n        \"in the age of\",\n        \"plays a crucial role\",\n        \"plays a vital role\",\n        \"plays an important role\",\n        \"plays a significant role\",\n        \"is of paramount importance\",\n        \"is of great significance\",\n        \"has gained significant attention\",\n        \"has attracted considerable attention\",\n        \"has been widely studied\",\n        \"has been extensively investigated\",\n        \"a growing body of evidence\",\n        \"an increasing number of\",\n        \"a wide range of\",\n        \"a variety of\",\n        \"a plethora of\",\n        \"a myriad of\",\n        \"delve into\",\n        \"shed light on\",\n        \"pave the way for\",\n        \"open new avenues\",\n        \"bridge the gap\",\n        \"furthermore\",\n        \"moreover\",\n        \"additionally\",\n        \"consequently\",\n        \"nevertheless\",\n        \"nonetheless\",\n        \"subsequently\",\n        \"notably\",\n        \"specifically\",\n        \"fundamentally\",\n        \"intriguingly\",\n        \"notably\",\n        \"remarkably\",\n        \"underscores the importance\",\n        \"highlights the significance\",\n        \"serves as a testament\",\n        \"poised to\",\n        \"tailored to\",\n        \"groundbreaking\",\n        \"pivotal\",\n        \"indispensable\",\n        \"multifaceted\",\n        \"nuanced\",\n        \"comprehensive\",\n        \"robust\",\n        \"state-of-the-art\",\n        \"cutting-edge\"\n      ]\n    },\n    \"conclusion_patterns\": {\n      \"weight\": 4,\n      \"description\": \"AI-typical conclusion patterns\",\n      \"patterns\": [\n        \"in conclusion,?.{0,30}(important|significant|crucial)\",\n        \"this study (demonstrates|shows|reveals|suggests) that\",\n        \"these findings (suggest|indicate|demonstrate|highlight)\",\n        \"future research (should|could|may|might) (focus|explore|investigate|examine)\",\n        \"the results of this study (provide|offer|contribute)\",\n        \"this research (contributes|adds|provides) (to|a|new)\",\n        \"taken together,?.{0,20}(suggest|indicate|demonstrate)\",\n        \"overall,?.{0,20}(demonstrate|show|suggest|highlight)\"\n      ]\n    },\n    \"transition_overuse\": {\n      \"weight\": 2,\n      \"description\": \"AI-overused transition words\",\n      \"patterns\": [\n        \"furthermore\",\n        \"moreover\",\n        \"additionally\",\n        \"in addition\",\n        \"consequently\",\n        \"therefore\",\n        \"thus\",\n        \"hence\",\n        \"nevertheless\",\n        \"nonetheless\",\n        \"however\",\n        \"on the other hand\",\n        \"in contrast\",\n        \"similarly\",\n        \"likewise\",\n        \"correspondingly\",\n        \"subsequently\",\n        \"meanwhile\",\n        \"in turn\",\n        \"as a result\"\n      ]\n    },\n    \"sentence_structure\": {\n      \"weight\": 2,\n      \"description\": \"AI-typical sentence structures\",\n      \"patterns\": [\n        \"not only .{2,40} but also\",\n        \"both .{2,30} and .{2,30}\",\n        \"either .{2,20} or .{2,20}\",\n        \"neither .{2,20} nor .{2,20}\",\n        \"by .{2,20}(ing)?,? .{0,10}(enable|allow|facilitate|promote)\",\n        \"through .{2,30},? .{0,10}(achieve|realize|accomplish)\",\n        \"with .{2,20}(development|advancement|emergence)\",\n        \"despite .{2,30},? .{0,10}(remain|continue|persist)\"\n      ]\n    },\n    \"vague_expressions\": {\n      \"weight\": 2,\n      \"description\": \"AI-typical vague/hedging expressions\",\n      \"patterns\": [\n        \"to a certain extent\",\n        \"to some degree\",\n        \"in some ways\",\n        \"in a sense\",\n        \"generally speaking\",\n        \"broadly speaking\",\n        \"for the most part\",\n        \"by and large\",\n        \"in general\",\n        \"it can be argued that\",\n        \"one could argue that\",\n        \"it seems that\",\n        \"it appears that\",\n        \"there is evidence to suggest\",\n        \"it has been suggested that\",\n        \"it is widely believed\",\n        \"it is generally accepted\"\n      ]\n    },\n    \"redundant_expressions\": {\n      \"weight\": 1,\n      \"description\": \"AI redundant expressions\",\n      \"patterns\": [\n        \"conduct(ed)? an in-depth (analysis|study|investigation|examination)\",\n        \"carry out a comprehensive (review|analysis|study)\",\n        \"make(s)? a significant contribution\",\n        \"perform(ed)? a detailed (analysis|examination)\",\n        \"provide(s)? valuable insights into\",\n        \"offer(s)? a novel perspective on\",\n        \"present(s)? a comprehensive overview\"\n      ]\n    },\n    \"rlhf_alignment\": {\n      \"description\": \"RLHF alignment patterns: excessive hedging, over-polite, disclaimer-like, balanced-to-a-fault\",\n      \"weight\": 3,\n      \"patterns\": [\n        \"it is worth noting that\",\n        \"it should be noted that\",\n        \"it is important to emphasize\",\n        \"it bears mentioning\",\n        \"it is worth mentioning\",\n        \"needless to say\",\n        \"importantly,\",\n        \"notably,\",\n        \"of note,\",\n        \"significantly,\",\n        \"however, it is important to remember\",\n        \"however, this does not mean\",\n        \"this should not be taken to mean\",\n        \"it is crucial to keep in mind\",\n        \"it is essential to note\",\n        \"it is also important to consider\",\n        \"one should also bear in mind\",\n        \"it is equally important\",\n        \"on the one hand\",\n        \"on the other hand\",\n        \"while it is true that\",\n        \"having said that\",\n        \"at the same time,\",\n        \"that being said,\",\n        \"in fairness,\",\n        \"to be fair,\",\n        \"furthermore, it is\",\n        \"moreover, it is\",\n        \"in addition to the above\",\n        \"not only.*but also\",\n        \"beyond that,\",\n        \"even more importantly\",\n        \"above all,\",\n        \"most importantly,\",\n        \"plays a crucial role\",\n        \"is of paramount importance\",\n        \"cannot be overstated\",\n        \"holds significant promise\",\n        \"has far-reaching implications\",\n        \"represents a significant\",\n        \"serves as a testament\",\n        \"stands as a testament\",\n        \"paves the way for\",\n        \"opens up new avenues\",\n        \"sheds light on\",\n        \"I hope this helps\",\n        \"I hope you find this\",\n        \"feel free to\",\n        \"please don't hesitate\",\n        \"if you have any questions\",\n        \"as always,\",\n        \"happy to help\",\n        \"in many ways\",\n        \"to a certain extent\",\n        \"to some degree\",\n        \"in a sense\",\n        \"arguably,\",\n        \"one could argue\",\n        \"it could be argued\",\n        \"there is a case to be made\",\n        \"the truth is more nuanced\",\n        \"the reality is more complex\"\n      ]\n    },\n    \"academic_formal\": {\n      \"weight\": 3,\n      \"description\": \"AI academic formal language\",\n      \"patterns\": [\n        \"characterized by\",\n        \"approach to\",\n        \"management of\",\n        \"including but not limited\",\n        \"plays an? (?:crucial|essential|pivotal|important|key) role\",\n        \"it is (?:essential|crucial|important|imperative) (?:to|that)\",\n        \"significantly impact\",\n        \"significantly improv\",\n        \"a comprehensive (?:approach|understanding|management|review)\",\n        \"in terms of\",\n        \"the management of\",\n        \"a structured approach\",\n        \"it is (?:a|an) (?:chronic|autoimmune|rare|common|serious)\",\n        \"is (?:defined|characterized|diagnosed) (?:by|as)\",\n        \"(?:primary|secondary|tertiary) (?:cause|prevention|treatment)\",\n        \"the (?:primary|main|key|critical|essential) (?:goal|objective|aim|target)\",\n        \"this (?:condition|disease|disorder|syndrome)\",\n        \"(?:early|timely|prompt) (?:diagnosis|treatment|intervention|recognition)\",\n        \"(?:treatment|management|therapy) (?:options|approaches|strategies|modalities)\",\n        \"clinical (?:presentation|manifestation|features|characteristics)\",\n        \"have been (?:associated|linked|shown|demonstrated) (?:to|with)\",\n        \"is (?:associated|linked) (?:with|to) an? increased\",\n        \"(?:prognosis|outcome) (?:depends|varies|is)\",\n        \"(?:regular|routine|periodic|close) (?:monitoring|follow-up|surveillance)\",\n        \"significantly improves?\",\n        \"highlighting the need\",\n        \"necessitat(?:es?|ing)\",\n        \"are critical (?:to|for)\",\n        \"is essential (?:to|for|in)\",\n        \"In refractory cases\",\n        \"may be considered\",\n        \"remain(?:s)? high-risk\",\n        \"Early diagnosis and (?:treatment|intervention)\",\n        \"despite (?:treatment|therapy|this)\",\n        \"further research (?:into|is needed|on)\",\n        \"optimizing (?:maternal|patient|clinical|treatment)\",\n        \"which significantly\",\n        \"Close.*?monitoring is\",\n        \"adverse outcomes\",\n        \"multidisciplinary (?:care|approach|team)\",\n        \"targeted therapies\",\n        \"pregnancy outcomes\",\n        \"clinical criteria\",\n        \"pathophysiology involves\",\n        \"cornerstone (?:disease|treatment|therapy)\",\n        \"primarily used (?:to|for)\",\n        \"acts (?:by|through|via)\",\n        \"widely used (?:for|in|to)\",\n        \"is commonly associated\",\n        \"mistakenly target\",\n        \"is an? (?:important|essential|useful) (?:tool|marker|indicator)\",\n        \"(?:screening|diagnostic) (?:tool|test|method|approach)\",\n        \"should be (?:individualized|tailored|guided|determined)\",\n        \"(?:reduce|minimize|prevent) (?:false|misdiagnosis|error)\",\n        \"combined (?:with|use|approach)\",\n        \"depending on (?:the |individual |patient )\",\n        \"under (?:medical|physician|clinical) (?:guidance|supervision)\",\n        \"(?:oral|subcutaneous|intravenous) (?:administration|dose|route)\",\n        \"starting dose\",\n        \"maximum.*?dose\",\n        \"monitor.*?(?:regularly|closely|periodically)\",\n        \"avoid.*?(?:concomitant|combination|with)\",\n        \"(?:contraindicated|not recommended) (?:in|during|for)\",\n        \"live (?:vaccine|vaccination)\",\n        \"bone marrow suppression\",\n        \"hepatotoxicity\",\n        \"are (?:auto|commonly|frequently|often|typically) associated\",\n        \"mistakenly (?:target|attack)\",\n        \"doesn't always (?:confirm|indicate|mean)\",\n        \"leading to (?:tissue|organ|cell|systemic)\",\n        \"trigger (?:inflammation|immune|autoimmune)\",\n        \"A positive.*?(?:test|result|finding)\",\n        \"(?:low|high) (?:levels|titers|doses)\",\n        \"can occur in (?:healthy|normal)\",\n        \"specific.*?(?:patterns|markers|criteria|findings)\",\n        \"(?:help|assist|aid) (?:diagnos|monitor|detect|identif)\",\n        \"as well as (?:its|their|the)\",\n        \"own (?:cells|tissues|body|immune)\",\n        \"systemic (?:symptoms|manifestations|features)\",\n        \"(?:80|100|120|150|200) words\",\n        \"remains (?:challenging|the|a|first-line)\",\n        \"has revolutionized\",\n        \"hinges on\",\n        \"first-line therapy\",\n        \"steroid-sparing\",\n        \"refractory (?:disease|cases|myositis)\",\n        \"(?:quickly|rapidly) followed by\",\n        \"personalized (?:based|approach|treatment)\",\n        \"demand(?:s)? (?:a )?(?:comprehensive|multidisciplinary)\",\n        \"balanced? with\",\n        \"efficacy with tox\",\n        \"minimize cumulative\",\n        \"preserv(?:e|ation) (?:renal|organ|function)\",\n        \"guides? biopsy\",\n        \"paradigm shifts?\",\n        \"increasingly prioritized\",\n        \"subclassification\",\n        \"subtype\",\n        \"remain(?:s)? (?:the|a|challenging|first-line|important)\",\n        \"balanc(?:e|ing) efficacy\",\n        \"carry (?:notable|significant) (?:safety|risk)\",\n        \"Regulatory agencies\",\n        \"Major adverse events include\",\n        \"black.?box warnings?\",\n        \"increasingly used for\",\n        \"broad immunomodulation\",\n        \"prompting.*?warning\",\n        \"susceptibility to (?:serious|opportunistic)\",\n        \"underlying risk factors\",\n        \"particularly (?:lymphoma|infection|malignancy)\",\n        \"have been observed\",\n        \"higher rates of\",\n        \"safety warnings\",\n        \"notable safety\",\n        \"Thromboembolic events\",\n        \"deep vein thrombosis\",\n        \"such as.*?are effective\",\n        \"including.*?and.*?are\",\n        \"arise from\",\n        \"broader (?:clinical|therapeutic|safety)\",\n        \"Evolving.*?Paradigms\",\n        \"multi-factorial interplay\",\n        \"pathogenic hallmark\",\n        \"has undergone a profound transformation\",\n        \"transformative.*?advancement\",\n        \"pivotal breakthrough\",\n        \"The landscape of.*?has undergone\",\n        \"Elucidat\",\n        \"armamentarium\",\n        \"fundamentally driven\",\n        \"tailored therapies\",\n        \"are emerging\",\n        \"comprise a spectrum of\",\n        \"share a pathogenic\",\n        \"simple.*?model to a complex\",\n        \"recognition of the alternative\",\n        \"historically replaced\",\n        \"largest.*?ever conducted\",\n        \"Despite advancements in\",\n        \"has transitioned from\"\n      ]\n    },\n    \"en_markdown\": {\n      \"weight\": 8,\n      \"description\": \"Markdown formatting (AI output indicator)\",\n      \"patterns\": [\n        \"\\\\*\\\\*.*?\\\\*\\\\*\",\n        \"^#{1,4}\\\\s\",\n        \"\\\\n\\\\d+\\\\.\\\\s\"\n      ]\n    }\n  }\n}\n\nFile v5.1.0:references/ai_patterns_zh.json\n\n{\n  \"version\": \"4.0.0\",\n  \"description\": \"v3数据驱动重校准版 (基线语料实测FPR过滤)\",\n  \"recalibration\": {\n    \"date\": \"2026-08-15\",\n    \"corpus\": \"eval/corpora_small/eval_zh.jsonl\",\n    \"humans\": 245,\n    \"ais\": 1645,\n    \"dropped\": 568,\n    \"demoted\": 5\n  },\n  \"categories\": {\n    \"weak_signals\": {\n      \"weight\": 1,\n      \"description\": \"v3重校准: 人类语料命中3%+且lift<2的弱信号, 仅组合计分(单独命中不足以判AI)\",\n      \"patterns\": [\n        \"然后\",\n        \"因此\",\n        \"再次\",\n        \"似乎\",\n        \"先.*再.*\"\n      ]\n    },\n    \"filler_phrases\": {\n      \"weight\": 3,\n      \"description\": \"各模型通用AI高频填充短语\",\n      \"patterns\": [\n        \"值得注意的是\",\n        \"综上所述\",\n        \"总而言之\",\n        \"具体而言\",\n        \"深入探讨\",\n        \"深入理解\",\n        \"系统梳理\",\n        \"深入分析\",\n        \"旨在探讨\",\n        \"具有重要意义\",\n        \"提供了新的视角\",\n        \"奠定了基础\",\n        \"本文旨在\",\n        \"本研究旨在\",\n        \"本研究通过\",\n        \"本研究为\",\n        \"与此同时\",\n        \"在此基础上\",\n        \"尤其是\",\n        \"特别是\",\n        \"不容忽视\",\n        \"至关重要\",\n        \"关键作用\",\n        \"重要意义\",\n        \"不可忽视\"\n      ]\n    },\n    \"ai_cliches\": {\n      \"weight\": 4,\n      \"description\": \"AI标志性套话和陈词滥调（最强信号）\",\n      \"patterns\": [\n        \"令人印象深刻\",\n        \"令人满意\",\n        \"新思路\",\n        \"日益受到\",\n        \"不断完善\",\n        \"展现出.*潜力\",\n        \"展现出.*价值\",\n        \"为.*提供了新\",\n        \"值得关注\",\n        \"引起了广泛关注\",\n        \"为.{2,15}(铺平|奠定|打开)了.{2,10}(道路|基础|大门|新方向)\"\n      ]\n    },\n    \"conclusion_patterns\": {\n      \"weight\": 4,\n      \"description\": \"AI标志性结尾模式\",\n      \"patterns\": [\n        \"本研究(为|对).{0,30}(提供了|具有一定的)\",\n        \"研究结果.{0,20}(表明|显示|证实)\"\n      ]\n    },\n    \"transition_words\": {\n      \"weight\": 2,\n      \"description\": \"AI过度使用的过渡词\",\n      \"patterns\": [\n        \"此外\",\n        \"同时\",\n        \"然而\",\n        \"与此同时\",\n        \"另一方面\",\n        \"进一步地\",\n        \"总的来说\",\n        \"首先\",\n        \"其次\"\n      ]\n    },\n    \"sentence_structure\": {\n      \"weight\": 2,\n      \"description\": \"AI典型句式特征\",\n      \"patterns\": [\n        \"不仅.{2,20}而且.{2,30}\",\n        \"不仅.{2,20}还.{2,30}\",\n        \"既.{2,15}又.{2,15}\",\n        \"一方面.{5,40}另一方面.{5,40}\",\n        \"通过.{2,30}实现了.{2,30}\",\n        \"通过.{2,30}为.{2,30}(提供了|奠定了)\",\n        \"随着.{2,20}的(发展|深入|推进|不断|广泛)\",\n        \"对于.*来说\",\n        \"对.*而言\",\n        \"作为.*的.*之一\",\n        \"在.*中(扮演|发挥|起到)\"\n      ]\n    },\n    \"vague_expressions\": {\n      \"weight\": 2,\n      \"description\": \"AI倾向使用的模糊表达\",\n      \"patterns\": [\n        \"在一定程度上\",\n        \"整体而言\",\n        \"可以说\",\n        \"总体而言\",\n        \"总体来说\"\n      ]\n    },\n    \"redundant_expressions\": {\n      \"weight\": 1,\n      \"description\": \"AI冗余表达\",\n      \"patterns\": [\n        \"进行(了)?(深入|详细|全面|系统)的?(分析|研究|探讨|阐述)\",\n        \"对.{2,15}进行了.{2,15}(分析|研究|探讨)\",\n        \"进行了.*分析\",\n        \"进行了.*研究\",\n        \"进行了.*探讨\"\n      ]\n    },\n    \"deepseek_fingerprint\": {\n      \"weight\": 4,\n      \"description\": \"DeepSeek模型指纹：教科书体+数据驱动+编号结构（最强信号）\",\n      \"patterns\": [\n        \"大规模.*研究\"\n      ]\n    },\n    \"glm_fingerprint\": {\n      \"weight\": 3,\n      \"description\": \"GLM/智谱模型指纹：四平八稳综述体\",\n      \"patterns\": [\n        \"关注的焦点\",\n        \"确保.*安全\",\n        \"以确保.*安全\"\n      ]\n    },\n    \"qwen_fingerprint\": {\n      \"weight\": 3,\n      \"description\": \"Qwen/通义模型指纹：编号体+标准论文格式\",\n      \"patterns\": [\n        \"存在.*的风险\",\n        \"存在.*的风险\"\n      ]\n    },\n    \"kimi_fingerprint\": {\n      \"weight\": 3,\n      \"description\": \"Kimi/月之暗面指纹：口语+专业混搭+破折号+伪随性\",\n      \"patterns\": [\n        \"实际.*问题\",\n        \"这个问题\",\n        \"这个.*问题\"\n      ]\n    },\n    \"minimax_fingerprint\": {\n      \"weight\": 4,\n      \"description\": \"MiniMax指纹：文艺比喻+排比+情感色彩（最强信号）\",\n      \"patterns\": [\n        \"在.*与.*之间\"\n      ]\n    },\n    \"colloquial_in_academic\": {\n      \"weight\": 3,\n      \"description\": \"口语化表达侵入学术文本（Kimi/MiniMax等模型的伪口语风格）\",\n      \"patterns\": [\n        \"实际.*问题\",\n        \"核心问题\",\n        \"关键在于\"\n      ]\n    },\n    \"step_fingerprint\": {\n      \"weight\": 3,\n      \"description\": \"Step/阶跃指纹：西式中译+'让我们'式表达\",\n      \"patterns\": [\n        \"让我们.*\"\n      ]\n    },\n    \"rlhf_alignment\": {\n      \"weight\": 3,\n      \"description\": \"RLHF对齐特征：过度礼貌、回避性表达、总结性语句、免责声明等\",\n      \"patterns\": [\n        \"值得一提的是\",\n        \"在一定程度上\",\n        \"总的来说\",\n        \"总体而言\",\n        \"更重要的是\",\n        \"在此基础上\",\n        \"具有重要意义\",\n        \"具有重要的\"\n      ]\n    },\n    \"medical_templates\": {\n      \"weight\": 4,\n      \"description\": \"医学学术模板化表达：AI在医学综述/研究论文中高频使用的固定句式和短语\",\n      \"patterns\": [\n        \"在.*方面\",\n        \"在.*方面，\",\n        \"被认为是\",\n        \"关键环节\",\n        \"本研究旨在探讨\",\n        \"独立危险因素\"\n      ]\n    },\n    \"structural_markers\": {\n      \"weight\": 3,\n      \"description\": \"AI写作的结构性标记\",\n      \"patterns\": [\n        \"具有重要意义\",\n        \"发挥着\",\n        \"治疗策略\",\n        \"个体化治疗\",\n        \"分为.*?型\",\n        \"需.*?长期\",\n        \"诊断.*?基于\",\n        \"旨在.*?控制\",\n        \"旨在.*?保护\",\n        \"改善.*?预后\",\n        \"个体化\",\n        \"近年来.*?降低\",\n        \"优化.*治疗\",\n        \"优化.*方案\",\n        \"精准.*干预\"\n      ]\n    },\n    \"markdown_format\": {\n      \"weight\": 2,\n      \"description\": \"Markdown格式（AI输出特征）\",\n      \"patterns\": [\n        \"\\\\n\\\\d+\\\\.\\\\s\"\n      ]\n    },\n    \"translationese\": {\n      \"weight\": 3,\n      \"description\": \"翻译腔：英文句法直译痕迹，非自然中文医学表述（被动堆叠/连词链/定语堆叠/名词化）\",\n      \"patterns\": [\n        \"被视为\",\n        \"被(证明|发现|证实)\",\n        \"扮演.{0,4}角色\",\n        \"发挥着.{0,8}作用\",\n        \"对.{2,20}进行(了)?(详细|系统|全面|深入|进一步)\",\n        \"使得.{2,30}\",\n        \"从而(导致|使|促进|降低|提高|增加|减少)\",\n        \"从.{2,10}的(角度|视角|立场)来看\"\n      ]\n    },\n    \"fingerprint_kimi_k3\": {\n      \"weight\": 3,\n      \"description\": \"v4挖掘: kimi-k3 (mined 2026-09-23 via opencode-go)\",\n      \"patterns\": [\n        \"实践中\",\n        \"体会到\"\n      ]\n    },\n    \"fingerprint_kimi_k2_7\": {\n      \"weight\": 3,\n      \"description\": \"v4挖掘: kimi-k2.7 (mined 2026-09-26 via opencode-go)\",\n      \"patterns\": [\n        \"实践中\"\n      ]\n    }\n  }\n}\n\nFile v5.1.0:references/article-review-workflow.md\n\n# Article Review Workflow (Terminology + AI Detection + Metaphor Audit)\n\n> Three-task parallel review for medical/scientific articles before publication.\n\n## Workflow Sequence\n\n### Phase 1: Parallel Tool Runs (do these simultaneously)\n\n```bash\n# Task A: Terminology check\npython3 scripts/term_check.py <article.md>\n\n# Task B: AI detection with JSON output\npython3 scripts/ai_detector.py <article.md> --lang auto --format json --output /tmp/ai_report.json\n```\n\n### Phase 2: Metaphor/Analogy Audit (manual, by agent)\n\nExtract ALL metaphors, analogies, and figurative language from the article. For each:\n\n| Field | What to assess |\n|-------|---------------|\n| Location | Chapter/section + line number |\n| Quoted text | The exact metaphor |\n| Scientific accuracy | Does the analogy correctly represent the underlying biology/medicine? |\n| Appropriateness | Is it suitable for the target audience? (e.g., peer physicians vs. lay public) |\n| Verdict | ✅ Keep / ⚠️ Modify / ❌ Remove |\n\n#### Metaphor Quality Criteria\n\n1. **Accuracy first**: The analogy must not distort the science. Example: \"alarmin = smoke alarm\" is accurate because alarmins signal danger without being the danger itself.\n2. **Consistency**: Maintain one imagery system per concept thread. Mixing metaphors (e.g., \"car engine\" → \"fire\" → \"gasoline\") breaks coherence. Fix by picking one system and sticking with it.\n3. **No over-explanation**: State the metaphor once. Don't add \"you can't just pretend...\" follow-up sentences that explain the metaphor to death.\n4. **Audience match**: For peer physicians, metaphors should illuminate mechanism, not dumb it down. Avoid overly literary/poetic phrasing in clinical sections.\n\n### Phase 3: Present Findings (concise!)\n\nPresent a single consolidated table to the user. Do NOT write a long essay.\n- **Keep it actionable**: table format with verdict + specific fix\n- **Get confirmation before patching**: list all proposed changes, ask once, execute all\n- **Don't make the user ask \"进度?\"** — if the report is ready, say so immediately\n\n### Phase 4: Execute Patches\n\nApply all confirmed changes, then re-run tools to verify improvement.\n\n## AI Detection: What's Worth Fixing\n\nNot every 50+ paragraph needs rewriting. Prioritize:\n\n1. **Model fingerprints** (DeepSeek/GLM/Qwen specific patterns) — highest priority, always fix\n2. **Classic AI openers** (\"值得注意的是\", \"综上所述\", \"在此基础上\") — easy wins\n3. **RLHF alignment patterns** — fix if concentrated in one section\n4. **Markdown bold patterns** — ignore, these are formatting not AI tells\n\n## Metaphor Audit: Common Pitfalls\n\n| Pitfall | Example | Fix |\n|---------|---------|-----|\n| Mixed imagery systems | \"油门卡死\" then \"火上浇油\" | Pick one system (engine → \"第二台发动机\") |\n| Over-explained metaphor | \"like X. You can't just pretend Y isn't happening.\" | Delete the explanation sentence |\n| Literary excess in clinical text | \"面纱\" + \"狰狞的脸\" | Simplify to direct language (\"伪装\") |\n| Chains too long | A→B→C→D extended metaphor | Cut to A→B, max |\n| Scientifically inaccurate | Any analogy that misrepresents pathophysiology | Rewrite or remove |\n\nFile v5.1.0:references/fusion_config.json\n\n{\n \"version\": \"1.0\",\n \"calibrated\": \"2026-08-15\",\n \"overall_auroc\": 0.9787,\n \"bands\": {\n  \"short\": {\n   \"weights\": {\n    \"para\": 0.4,\n    \"spectrum\": 0.54,\n    \"surface\": 0.0,\n    \"cot\": 0.0,\n    \"surprisal\": 0.06\n   },\n   \"band_auroc\": 0.9684,\n   \"n\": 1362\n  },\n  \"mid\": {\n   \"weights\": {\n    \"para\": 0.4,\n    \"spectrum\": 0.54,\n    \"surface\": 0.0,\n    \"cot\": 0.0,\n    \"surprisal\": 0.06\n   },\n   \"band_auroc\": 0.982,\n   \"n\": 2660\n  },\n  \"long\": {\n   \"weights\": {\n    \"para\": 0.4,\n    \"spectrum\": 0.54,\n    \"surface\": 0.0,\n    \"cot\": 0.0,\n    \"surprisal\": 0.06\n   },\n   \"band_auroc\": 0.9894,\n   \"n\": 733\n  }\n },\n \"thresholds\": {\n  \"medium\": 0.6368,\n  \"high\": 0.7486\n },\n \"v2\": {\n  \"w_supervised\": 0.9,\n  \"thresholds\": {\n   \"medium\": 0.3444,\n   \"high\": 0.5383\n  },\n  \"calibrated\": \"2026-08-16\",\n  \"model\": \"qwen3-0.6b-lora-v34-4dim-fp16-onnx\",\n  \"calib_sample\": {\n   \"human\": 260,\n   \"ai_recall_at_high\": 1.0\n  },\n  \"note\": \"fusion_v2 = v3.0_fused*(1-w)+supervised*w; 阈值=calib半人类blended分布p95/p99; v3.2模型分布较s43整体上移, 旧阈值(0.064/0.0814)已废弃\"\n }\n}\n\nFile v5.1.0:references/model_fingerprints.json\n\n{\n \"version\": \"4.0\",\n \"families\": {\n  \"glm-5.3\": {\n   \"versions\": [\n    \"GLM-5.3(2026-08)\",\n    \"GLM-5.2路由\"\n   ],\n   \"sampled_via\": \"self:GLM-5.3(agent)\",\n   \"patterns\": [\n    {\n     \"p\": \"真实世\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.0,\n     \"lift\": 22.22\n    },\n    {\n     \"p\": \"活动度\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.0,\n     \"lift\": 22.22\n    },\n    {\n     \"p\": \"和患者\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.0,\n     \"lift\": 22.22\n    },\n    {\n     \"p\": \"实世界\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.0,\n     \"lift\": 22.22\n    },\n    {\n     \"p\": \"二十年\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.002,\n     \"lift\": 19.01\n    },\n    {\n     \"p\": \"没有人\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.014,\n     \"lift\": 9.45\n    },\n    {\n     \"p\": \"的事情\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.015,\n     \"lift\": 8.82\n    },\n    {\n     \"p\": \"患者的\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.015,\n     \"lift\": 8.82\n    },\n    {\n     \"p\": \"在一起\",\n     \"w\": 4,\n     \"cov_model\": 0.222,\n     \"fpr_human\": 0.02,\n     \"lift\": 7.34\n    }\n   ],\n   \"sampled_docs\": 9,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"deepseek-r1\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"本研究\",\n     \"w\": 4,\n     \"cov_model\": 0.15,\n     \"fpr_human\": 0.022,\n     \"lift\": 4.71\n    }\n   ],\n   \"sampled_docs\": 379,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"deepseek-v3\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"提供了\",\n     \"w\": 4,\n     \"cov_model\": 0.197,\n     \"fpr_human\": 0.02,\n     \"lift\": 6.52\n    },\n    {\n     \"p\": \"本研究\",\n     \"w\": 4,\n     \"cov_model\": 0.189,\n     \"fpr_human\": 0.022,\n     \"lift\": 5.92\n    }\n   ],\n   \"sampled_docs\": 370,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"qwen-2.5\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"近年来\",\n     \"w\": 4,\n     \"cov_model\": 0.154,\n     \"fpr_human\": 0.014,\n     \"lift\": 6.53\n    },\n    {\n     \"p\": \"每个人\",\n     \"w\": 4,\n     \"cov_model\": 0.151,\n     \"fpr_human\": 0.015,\n     \"lift\": 5.99\n    },\n    {\n     \"p\": \"本研究\",\n     \"w\": 4,\n     \"cov_model\": 0.18,\n     \"fpr_human\": 0.022,\n     \"lift\": 5.62\n    },\n    {\n     \"p\": \"的重要\",\n     \"w\": 4,\n     \"cov_model\": 0.203,\n     \"fpr_human\": 0.03,\n     \"lift\": 5.03\n    },\n    {\n     \"p\": \"的故事\",\n     \"w\": 4,\n     \"cov_model\": 0.151,\n     \"fpr_human\": 0.02,\n     \"lift\": 4.99\n    },\n    {\n     \"p\": \"同时也\",\n     \"w\": 4,\n     \"cov_model\": 0.174,\n     \"fpr_human\": 0.025,\n     \"lift\": 4.94\n    },\n    {\n     \"p\": \"之间的\",\n     \"w\": 4,\n     \"cov_model\": 0.177,\n     \"fpr_human\": 0.039,\n     \"lift\": 3.62\n    },\n    {\n     \"p\": \"重要的\",\n     \"w\": 4,\n     \"cov_model\": 0.182,\n     \"fpr_human\": 0.044,\n     \"lift\": 3.38\n    },\n    {\n     \"p\": \"过程中\",\n     \"w\": 4,\n     \"cov_model\": 0.201,\n     \"fpr_human\": 0.051,\n     \"lift\": 3.3\n    },\n    {\n     \"p\": \"进一步\",\n     \"w\": 3,\n     \"cov_model\": 0.174,\n     \"fpr_human\": 0.049,\n     \"lift\": 2.96\n    }\n   ],\n   \"sampled_docs\": 384,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"qwen-3\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"有助于\",\n     \"w\": 4,\n     \"cov_model\": 0.155,\n     \"fpr_human\": 0.007,\n     \"lift\": 9.23\n    },\n    {\n     \"p\": \"真正的\",\n     \"w\": 4,\n     \"cov_model\": 0.157,\n     \"fpr_human\": 0.014,\n     \"lift\": 6.69\n    },\n    {\n     \"p\": \"近年来\",\n     \"w\": 4,\n     \"cov_model\": 0.152,\n     \"fpr_human\": 0.014,\n     \"lift\": 6.46\n    },\n    {\n     \"p\": \"本研究\",\n     \"w\": 4,\n     \"cov_model\": 0.152,\n     \"fpr_human\": 0.022,\n     \"lift\": 4.76\n    },\n    {\n     \"p\": \"是一种\",\n     \"w\": 4,\n     \"cov_model\": 0.16,\n     \"fpr_human\": 0.039,\n     \"lift\": 3.28\n    },\n    {\n     \"p\": \"进一步\",\n     \"w\": 3,\n     \"cov_model\": 0.171,\n     \"fpr_human\": 0.049,\n     \"lift\": 2.89\n    }\n   ],\n   \"sampled_docs\": 375,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"gpt-4o\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"未来的\",\n     \"w\": 4,\n     \"cov_model\": 0.201,\n     \"fpr_human\": 0.005,\n     \"lift\": 13.31\n    },\n    {\n     \"p\": \"不仅是\",\n     \"w\": 4,\n     \"cov_model\": 0.224,\n     \"fpr_human\": 0.008,\n     \"lift\": 12.14\n    },\n    {\n     \"p\": \"复杂的\",\n     \"w\": 4,\n     \"cov_model\": 0.172,\n     \"fpr_human\": 0.007,\n     \"lift\": 10.26\n    },\n    {\n     \"p\": \"不仅仅\",\n     \"w\": 4,\n     \"cov_model\": 0.154,\n     \"fpr_human\": 0.005,\n     \"lift\": 10.2\n    },\n    {\n     \"p\": \"有助于\",\n     \"w\": 4,\n     \"cov_model\": 0.167,\n     \"fpr_human\": 0.007,\n     \"lift\": 9.95\n    },\n    {\n     \"p\": \"提供了\",\n     \"w\": 4,\n     \"cov_model\": 0.266,\n     \"fpr_human\": 0.02,\n     \"lift\": 8.78\n    },\n    {\n     \"p\": \"生活中\",\n     \"w\": 4,\n     \"cov_model\": 0.18,\n     \"fpr_human\": 0.014,\n     \"lift\": 7.64\n    },\n    {\n     \"p\": \"的重要\",\n     \"w\": 4,\n     \"cov_model\": 0.234,\n     \"fpr_human\": 0.03,\n     \"lift\": 5.8\n    },\n    {\n     \"p\": \"尤其是\",\n     \"w\": 4,\n     \"cov_model\": 0.193,\n     \"fpr_human\": 0.024,\n     \"lift\": 5.73\n    },\n    {\n     \"p\": \"本研究\",\n     \"w\": 4,\n     \"cov_model\": 0.174,\n     \"fpr_human\": 0.022,\n     \"lift\": 5.46\n    },\n    {\n     \"p\": \"生活的\",\n     \"w\": 4,\n     \"cov_model\": 0.159,\n     \"fpr_human\": 0.024,\n     \"lift\": 4.72\n    },\n    {\n     \"p\": \"是一种\",\n     \"w\": 4,\n     \"cov_model\": 0.208,\n     \"fpr_human\": 0.039,\n     \"lift\": 4.26\n    }\n   ],\n   \"sampled_docs\": 384,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"gpt-3.5-turbo\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"研究旨\",\n     \"w\": 4,\n     \"cov_model\": 0.154,\n     \"fpr_human\": 0.002,\n     \"lift\": 13.18\n    },\n    {\n     \"p\": \"究旨在\",\n     \"w\": 4,\n     \"cov_model\": 0.154,\n     \"fpr_human\": 0.002,\n     \"lift\": 13.18\n    },\n    {\n     \"p\": \"提供了\",\n     \"w\": 4,\n     \"cov_model\": 0.193,\n     \"fpr_human\": 0.02,\n     \"lift\": 6.38\n    },\n    {\n     \"p\": \"本研究\",\n     \"w\": 4,\n     \"cov_model\": 0.178,\n     \"fpr_human\": 0.022,\n     \"lift\": 5.56\n    },\n    {\n     \"p\": \"让我们\",\n     \"w\": 4,\n     \"cov_model\": 0.191,\n     \"fpr_human\": 0.025,\n     \"lift\": 5.39\n    },\n    {\n     \"p\": \"是一种\",\n     \"w\": 4,\n     \"cov_model\": 0.175,\n     \"fpr_human\": 0.039,\n     \"lift\": 3.58\n    }\n   ],\n   \"sampled_docs\": 383,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"claude-3.5-haiku\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"不仅仅\",\n     \"w\": 4,\n     \"cov_model\": 0.236,\n     \"fpr_human\": 0.005,\n     \"lift\": 15.64\n    },\n    {\n     \"p\": \"仅仅是\",\n     \"w\": 4,\n     \"cov_model\": 0.196,\n     \"fpr_human\": 0.008,\n     \"lift\": 10.64\n    },\n    {\n     \"p\": \"提供了\",\n     \"w\": 4,\n     \"cov_model\": 0.165,\n     \"fpr_human\": 0.02,\n     \"lift\": 5.45\n    },\n    {\n     \"p\": \"本研究\",\n     \"w\": 4,\n     \"cov_model\": 0.173,\n     \"fpr_human\": 0.022,\n     \"lift\": 5.41\n    },\n    {\n     \"p\": \"的重要\",\n     \"w\": 4,\n     \"cov_model\": 0.181,\n     \"fpr_human\": 0.03,\n     \"lift\": 4.47\n    },\n    {\n     \"p\": \"重要的\",\n     \"w\": 4,\n     \"cov_model\": 0.212,\n     \"fpr_human\": 0.044,\n     \"lift\": 3.93\n    },\n    {\n     \"p\": \"是一种\",\n     \"w\": 4,\n     \"cov_model\": 0.162,\n     \"fpr_human\": 0.039,\n     \"lift\": 3.32\n    }\n   ],\n   \"sampled_docs\": 382,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"gemini-2.5-flash\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"本研究旨在\",\n     \"w\": 4,\n     \"cov_model\": 0.167,\n     \"fpr_human\": 0.002,\n     \"lift\": 14.26\n    },\n    {\n     \"p\": \"未来的\",\n     \"w\": 4,\n     \"cov_model\": 0.161,\n     \"fpr_human\": 0.005,\n     \"lift\": 10.7\n    },\n    {\n     \"p\": \"独特的\",\n     \"w\": 4,\n     \"cov_model\": 0.189,\n     \"fpr_human\": 0.008,\n     \"lift\": 10.22\n    },\n    {\n     \"p\": \"仅仅是\",\n     \"w\": 4,\n     \"cov_model\": 0.153,\n     \"fpr_human\": 0.008,\n     \"lift\": 8.29\n    },\n    {\n     \"p\": \"提供了\",\n     \"w\": 4,\n     \"cov_model\": 0.186,\n     \"fpr_human\": 0.02,\n     \"lift\": 6.14\n    },\n    {\n     \"p\": \"是一种\",\n     \"w\": 4,\n     \"cov_model\": 0.194,\n     \"fpr_human\": 0.039,\n     \"lift\": 3.97\n    },\n    {\n     \"p\": \"重要的\",\n     \"w\": 4,\n     \"cov_model\": 0.172,\n     \"fpr_human\": 0.044,\n     \"lift\": 3.19\n    },\n    {\n     \"p\": \"进一步\",\n     \"w\": 4,\n     \"cov_model\": 0.186,\n     \"fpr_human\": 0.049,\n     \"lift\": 3.15\n    }\n   ],\n   \"sampled_docs\": 366,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"doubao-1.5-pro\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"提供了\",\n     \"w\": 4,\n     \"cov_model\": 0.172,\n     \"fpr_human\": 0.02,\n     \"lift\": 5.68\n    },\n    {\n     \"p\": \"让我们\",\n     \"w\": 4,\n     \"cov_model\": 0.182,\n     \"fpr_human\": 0.025,\n     \"lift\": 5.15\n    },\n    {\n     \"p\": \"过程中\",\n     \"w\": 3,\n     \"cov_model\": 0.152,\n     \"fpr_human\": 0.051,\n     \"lift\": 2.51\n    }\n   ],\n   \"sampled_docs\": 407,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"glm\": {\n   \"versions\": [],\n   \"sampled_via\": \"\",\n   \"patterns\": [\n    {\n     \"p\": \"供了有\",\n     \"w\": 4,\n     \"cov_model\": 0.177,\n     \"fpr_human\": 0.0,\n     \"lift\": 17.67\n    },\n    {\n     \"p\": \"探讨了\",\n     \"w\": 4,\n     \"cov_model\": 0.228,\n     \"fpr_human\": 0.003,\n     \"lift\": 17.08\n    },\n    {\n     \"p\": \"本研究为\",\n     \"w\": 4,\n     \"cov_model\": 0.168,\n     \"fpr_human\": 0.0,\n     \"lift\": 16.81\n    },\n    {\n     \"p\": \"味无穷\",\n     \"w\": 4,\n     \"cov_model\": 0.159,\n     \"fpr_human\": 0.0,\n     \"lift\": 15.95\n    },\n    {\n     \"p\": \"回味无\",\n     \"w\": 4,\n     \"cov_model\": 0.159,\n     \"fpr_human\": 0.0,\n     \"lift\": 15.95\n    },\n    {\n     \"p\": \"揭示了\",\n     \"w\": 4,\n     \"cov_model\": 0.185,\n     \"fpr_human\": 0.002,\n     \"lift\": 15.86\n    },\n    {\n     \"p\": \"为我国\",\n     \"w\": 4,\n     \"cov_model\": 0.155,\n     \"fpr_human\": 0.0,\n     \"lift\": 15.52\n    },\n    {\n     \"p\": \"供了新\",\n     \"w\": 4,\n     \"cov_model\": 0.151,\n     \"fpr_human\": 0.0,\n     \"lift\": 15.09\n    },\n    {\n     \"p\": \"提的是\",\n     \"w\": 4,\n     \"cov_model\": 0.185,\n     \"fpr_human\": 0.003,\n     \"lift\": 13.85\n    },\n    {\n     \"p\": \"提供了\",\n     \"w\": 4,\n     \"cov_model\": 0.418,\n     \"fpr_human\": 0.02,\n     \"lift\": 13.81\n    },\n    {\n     \"p\": \"一提的\",\n     \"w\": 4,\n     \"cov_model\": 0.181,\n     \"fpr_human\": 0.003,\n     \"lift\": 13.53\n    },\n    {\n     \"p\": \"值得一提\",\n     \"w\": 4,\n     \"cov_model\": 0.181,\n     \"fpr_human\": 0.003,\n     \"lift\": 13.53\n    }\n   ],\n   \"sampled_docs\": 232,\n   \"mined_at\": \"2026-08-15\"\n  },\n  \"kimi-k3\": {\n   \"versions\": [\n    \"2026-09\"\n   ],\n   \"sampled_via\": \"opencode-go\",\n   \"patterns\": [\n    {\n     \"p\": \"实践中\",\n     \"w\": 4,\n     \"cov_model\": 0.2,\n     \"fpr_human\": 0.0,\n     \"lift\": 20.0\n    },\n    {\n     \"p\": \"体会到\",\n     \"w\": 4,\n     \"cov_model\": 0.2,\n     \"fpr_human\": 0.0,\n     \"lift\": 20.0\n    }\n   ],\n   \"sampled_docs\": 40,\n   \"mined_at\": \"2026-09-23\"\n  },\n  \"kimi-k2.7\": {\n   \"versions\": [\n    \"2026-09\"\n   ],\n   \"sampled_via\": \"opencode-go\",\n   \"patterns\": [\n    {\n     \"p\": \"实践中\",\n     \"w\": 4,\n     \"cov_model\": 0.176,\n     \"fpr_human\": 0.0,\n     \"lift\": 17.65\n    }\n   ],\n   \"sampled_docs\": 34,\n   \"mined_at\": \"2026-09-26\"\n  }\n }\n}\n\nFile v5.1.0:references/para_thresholds.json\n\n{\n \"version\": \"v4.9.0\",\n \"bench\": \"mixed_bench_20261005.json\",\n \"n_paragraphs\": 333,\n \"n_ai_paragraphs\": 134,\n \"recommended_high\": {\n  \"para_high\": 50.0,\n  \"precision\": 0.6,\n  \"ai_para_coverage\": 0.627,\n  \"f1\": 0.6131\n },\n \"review_band\": {\n  \"para_medium_floor\": 45.0,\n  \"ai_para_coverage_ge_floor\": 0.701\n },\n \"confusion_at_high\": {\n  \"tp\": 84,\n  \"fp\": 56,\n  \"fn\": 50,\n  \"tn\": 143\n },\n \"honest_conclusion\": \"高置信段落判定（≥50 分）精度 0.60，但 AI 段覆盖率仅 63%——段落归因的用途=「高置信段实锤 + 剩余段人工复核排序」，不是段落级自动判定。文档级 AUROC 0.52-0.54 的原理性局限不变。阈值校准于同基准（n=333 段，单参数），属校准集内选点。\",\n \"grid_top8\": [\n  {\n   \"para_high\": 50.0,\n   \"precision\": 0.6,\n   \"ai_para_coverage\": 0.627,\n   \"f1\": 0.6131\n  },\n  {\n   \"para_high\": 52.5,\n   \"precision\": 0.603,\n   \"ai_para_coverage\": 0.545,\n   \"f1\": 0.5725\n  },\n  {\n   \"para_high\": 55.0,\n   \"precision\": 0.63,\n   \"ai_para_coverage\": 0.507,\n   \"f1\": 0.562\n  },\n  {\n   \"para_high\": 57.5,\n   \"precision\": 0.567,\n   \"ai_para_coverage\": 0.381,\n   \"f1\": 0.4554\n  },\n  {\n   \"para_high\": 60.0,\n   \"precision\": 0.544,\n   \"ai_para_coverage\": 0.321,\n   \"f1\": 0.4038\n  },\n  {\n   \"para_high\": 62.5,\n   \"precision\": 0.515,\n   \"ai_para_coverage\": 0.254,\n   \"f1\": 0.34\n  },\n  {\n   \"para_high\": 65.0,\n   \"precision\": 0.533,\n   \"ai_para_coverage\": 0.239,\n   \"f1\": 0.3299\n  },\n  {\n   \"para_high\": 67.5,\n   \"precision\": 0.542,\n   \"ai_para_coverage\": 0.194,\n   \"f1\": 0.2857\n  }\n ],\n \"grid_domain\": [\n  50.0,\n  95.0\n ],\n \"grid_domain_note\": \"网格域 [50,95]；全局 F1 最优在 47.5（P=0.589/F1=0.637）——现选 50.0 为偏精度保守选点，honest_conclusion 的「≥50」限定据此。\"\n}\n\nFile v5.1.0:references/sentence_patterns_zh.json\n\n{\n  \"version\": \"1.0.0\",\n  \"description\": \"常见AI学术句式模板（中文），用于检测和改写建议\",\n  \"note\": \"仅覆盖高频可模板化的句式，不做语法分析\",\n  \"patterns\": [\n    {\n      \"id\": \"sp001\",\n      \"category\": \"研究目的\",\n      \"pattern\": \"本研究旨在探讨{topic}的{aspect}\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"关于{topic}的{aspect}，我们做了一个分析\",\n        \"{topic}的{aspect}到底如何？我们看了看数据\",\n        \"我们关注的问题是{topic}中{aspect}的情况\"\n      ]\n    },\n    {\n      \"id\": \"sp002\",\n      \"category\": \"研究目的\",\n      \"pattern\": \"本研究/本文的目的是{action}\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"我们做这项研究主要是想{action}\",\n        \"为了弄清楚{action}，我们设计了这项分析\",\n        \"这项工作围绕{action}展开\"\n      ]\n    },\n    {\n      \"id\": \"sp003\",\n      \"category\": \"方法描述\",\n      \"pattern\": \"回顾性分析了{time}期间收治的{n}例{disease}患者的临床资料\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"我们调取了{time}间{n}例{disease}患者的病历\",\n        \"{time}共有{n}例{disease}患者入组\",\n        \"从{time}的病历中筛选出{n}例{disease}进行分析\"\n      ]\n    },\n    {\n      \"id\": \"sp004\",\n      \"category\": \"结果描述\",\n      \"pattern\": \"结果显示{group1}与{group2}在{variable}方面存在显著差异\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"{group1}和{group2}的{variable}差别明显\",\n        \"比较两组的{variable}，差异有统计学意义\",\n        \"{variable}在两组间拉开了差距\"\n      ]\n    },\n    {\n      \"id\": \"sp005\",\n      \"category\": \"结果描述\",\n      \"pattern\": \"多因素{analysis}分析显示{factor1}、{factor2}和{factor3}是{outcome}的独立危险因素\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"把{factor1}、{factor2}、{factor3}放在一起看，发现它们各自独立影响{outcome}\",\n        \"{analysis}告诉我们，即使控制了其他变量，{factor1}等仍然与{outcome}相关\",\n        \"排除了混杂因素后，{factor1}、{factor2}、{factor3}依然显著\"\n      ]\n    },\n    {\n      \"id\": \"sp006\",\n      \"category\": \"结论\",\n      \"pattern\": \"结论：临床上应重视{topic}的评估，合理使用{treatment}\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"简单说就是：{topic}不能掉以轻心，{treatment}要用得当\",\n        \"临床工作中要关注{topic}，{treatment}方案因人而异\",\n        \"提醒临床同道注意{topic}，用好{treatment}\"\n      ]\n    },\n    {\n      \"id\": \"sp007\",\n      \"category\": \"过渡\",\n      \"pattern\": \"然而，其{aspect}日益引起临床关注\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"但{aspect}这块一直不太清楚\",\n        \"{aspect}是个不能回避的问题\",\n        \"临床中{aspect}的争议也不少\"\n      ]\n    },\n    {\n      \"id\": \"sp008\",\n      \"category\": \"综述\",\n      \"pattern\": \"本文综述了{topic}的发生机制、临床表现及处理策略，为临床实践提供参考\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"下面从机制、临床和治疗方案三个角度梳理{topic}\",\n        \"这篇综述聚焦{topic}，讨论病理机制、临床特点和怎么办\",\n        \"我们整理了{topic}相关的最新证据，供同行参考\"\n      ]\n    },\n    {\n      \"id\": \"sp009\",\n      \"category\": \"局限\",\n      \"pattern\": \"本研究仍存在一定局限性，包括{limit1}和{limit2}。未来需要开展更大规模的{study_type}以进一步验证上述结论\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"当然这项研究也有短板：{limit1}和{limit2}。后续需要更大样本的{study_type}来确认\",\n        \"不足之处在于{limit1}和{limit2}，期待后续有人做{study_type}来回答这些问题\",\n        \"客观地说，{limit1}是个问题，{limit2}也需要改进。下一步应该做{study_type}\"\n      ]\n    },\n    {\n      \"id\": \"sp010\",\n      \"category\": \"综述开头\",\n      \"pattern\": \"近年来，{topic}在{field}领域取得了突破性进展\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"过去几年{field}领域关于{topic}的进展很快\",\n        \"{field}圈子里{topic}的变化可以用翻天覆地来形容\",\n        \"{topic}是{field}这两年最热的话题之一\"\n      ]\n    },\n    {\n      \"id\": \"sp011\",\n      \"category\": \"方法\",\n      \"pattern\": \"所有患者均签署知情同意书并通过{committee}审批\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"研究方案经{committee}批准，患者知情同意\",\n        \"{committee}审批通过，纳入患者均签了知情同意\",\n        \"伦理方面：{committee}同意，知情同意书签了\"\n      ]\n    },\n    {\n      \"id\": \"sp012\",\n      \"category\": \"方法\",\n      \"pattern\": \"采用{software}进行数据分析，计量资料以{stat}表示\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"数据用{software}处理，连续变量报{stat}\",\n        \"统计软件是{software}，连续变量用{stat}描述\",\n        \"计量资料以{stat}呈现在{software}中分析\"\n      ]\n    },\n    {\n      \"id\": \"sp013\",\n      \"category\": \"方法\",\n      \"pattern\": \"符合纳入标准的患者被随机分为{group1}和{group2}\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"按纳入标准筛完后，患者被随机分到{group1}或{group2}\",\n        \"入组患者随机进入{group1}或{group2}\",\n        \"随机分组：{group1} vs {group2}\"\n      ]\n    },\n    {\n      \"id\": \"sp014\",\n      \"category\": \"讨论\",\n      \"pattern\": \"与既往研究结果一致，提示{finding}\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"这个发现跟之前的研究对得上，说明{finding}\",\n        \"前人也有类似的发现，{finding}这个观点更加站得住脚了\",\n        \"我们的结果印证了{finding}——多个团队都看到了这个现象\"\n      ]\n    },\n    {\n      \"id\": \"sp015\",\n      \"category\": \"讨论\",\n      \"pattern\": \"本研究的创新之处在于{innovation}\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"这项工作新在{innovation}\",\n        \"跟已有工作最大的不同是{innovation}\",\n        \"值得关注的亮点：{innovation}\"\n      ]\n    },\n    {\n      \"id\": \"sp016\",\n      \"category\": \"综述结尾\",\n      \"pattern\": \"综上所述，{summary}。未来需要{future}\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"总结来说：{summary}。下一步{future}\",\n        \"归拢一下：{summary}。展望一下：{future}\",\n        \"{summary}。至于{future}，留给后续研究\"\n      ]\n    },\n    {\n      \"id\": \"sp017\",\n      \"category\": \"结果\",\n      \"pattern\": \"在随访期间共有{n}例患者出现{event}\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"随访过程中{n}例发生了{event}\",\n        \"{n}例在随访中出现{event}\",\n        \"{event}发生在{n}例患者中\"\n      ]\n    },\n    {\n      \"id\": \"sp018\",\n      \"category\": \"结果\",\n      \"pattern\": \"{treatment}组与{control}组的{endpoint}比较，差异有统计学意义(P<{p})\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"{treatment}组的{endpoint}明显优于{control}(P<{p})\",\n        \"两组{endpoint}拉开了差距，{treatment}更好(P<{p})\",\n        \"{endpoint}方面，{treatment}显著胜出(P<{p})\"\n      ]\n    },\n    {\n      \"id\": \"sp019\",\n      \"category\": \"背景\",\n      \"pattern\": \"{disease}是一种常见的{type}疾病，其发病机制尚未完全阐明\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"{disease}这种{type}病很常见，但到底怎么发生的还没完全搞清楚\",\n        \"{disease}在{type}领域是老话题了，不过发病机制这块还有不少空白\",\n        \"说到{type}疾病，{disease}发病率不低，但病因至今不完全明确\"\n      ]\n    },\n    {\n      \"id\": \"sp020\",\n      \"category\": \"背景\",\n      \"pattern\": \"早期诊断和及时治疗对改善{disease}预后具有重要意义\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"{disease}越早发现越好治，这已经是共识了\",\n        \"{disease}预后好不好，很大程度上取决于什么时候开始治\",\n        \"抓早期、及时干预，{disease}的结局会好很多\"\n      ]\n    },\n    {\n      \"id\": \"sp021\",\n      \"category\": \"过渡\",\n      \"pattern\": \"值得注意的是，{finding}\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"有意思的是{finding}\",\n        \"一个意外的发现：{finding}\",\n        \"{finding}，这个值得单独说说\"\n      ]\n    },\n    {\n      \"id\": \"sp022\",\n      \"category\": \"方法\",\n      \"pattern\": \"根据{criteria}进行诊断\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"诊断标准参照{criteria}\",\n        \"按{criteria}来确诊\",\n        \"{criteria}是诊断依据\"\n      ]\n    },\n    {\n      \"id\": \"sp023\",\n      \"category\": \"讨论\",\n      \"pattern\": \"分析其原因可能与{factor}有关\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"推测跟{factor}脱不了干系\",\n        \"可能的原因之一是{factor}\",\n        \"{factor}可能是背后的推手\"\n      ]\n    },\n    {\n      \"id\": \"sp024\",\n      \"category\": \"安全\",\n      \"pattern\": \"{treatment}总体安全性良好，不良反应多为轻度\",\n      \"ai_frequency\": \"very_high\",\n      \"alternatives\": [\n        \"{treatment}安全可耐受，严重副反应很少见\",\n        \"用{treatment}的大多没什么大问题，偶尔有点轻微不适\",\n        \"{treatment}的安全性数据让人放心，严重不良事件基本没有\"\n      ]\n    },\n    {\n      \"id\": \"sp025\",\n      \"category\": \"统计\",\n      \"pattern\": \"以P<{threshold}为差异有统计学意义\",\n      \"ai_frequency\": \"high\",\n      \"alternatives\": [\n        \"显著性水平设为{threshold}\",\n        \"P值小于{threshold}视为有意义\",\n        \"统计检验水准α={threshold}\"\n      ]\n    }\n  ]\n}\n\nFile v5.1.0:references/supervised_models.json\n\n{\n  \"comment\": \"作者签发的监督层模型注册表——pp_setup.py 只安装这里登记的指纹（v4.5.0 起）。模型不随包分发（100% 本地契约），由作者渠道提供文件后经本命令校验安装。\",\n  \"active\": \"v35\",\n  \"models\": [\n    {\n      \"id\": \"v35\",\n      \"md5_12\": \"2631df3d388b\",\n      \"tokenizer_md5_12\": \"6423133b9cc1\",\n      \"n_labels\": 4,\n      \"maxlen\": 1024,\n      \"note\": \"2026-10-03 重导出验证件（export_fp16.py 变长自检通过）。指纹绑定评测：老代留出 AUROC 0.9998 / 当打代 0.9400 / 人类误报@med 2.3%（eval/results/v35ctl_*.json, v4.4.0）。\"\n    }\n  ]\n}\n\nFile v5.1.0:references/synonyms_general.json\n\n{\n  \"version\": \"1.1.0\",\n  \"description\": \"通用同义词库（降重用），按领域分组\",\n  \"note\": \"专业术语不在同义词库中，由term_check.py保护\",\n  \"general\": {\n    \"表明\": [\n      \"显示\",\n      \"证实\",\n      \"说明\",\n      \"揭示\",\n      \"呈现\",\n      \"反映出\"\n    ],\n    \"显示\": [\n      \"表明\",\n      \"呈现\",\n      \"体现\",\n      \"展示\",\n      \"证实\"\n    ],\n    \"证实\": [\n      \"验证\",\n      \"确认\",\n      \"证明\",\n      \"表明\",\n      \"印证\"\n    ],\n    \"说明\": [\n      \"表明\",\n      \"揭示\",\n      \"反映出\",\n      \"意味着\",\n      \"佐证\"\n    ],\n    \"重要\": [\n      \"关键\",\n      \"核心\",\n      \"主要\",\n      \"突出\",\n      \"显著\"\n    ],\n    \"显著\": [\n      \"明显\",\n      \"突出\",\n      \"大幅\",\n      \"可观\",\n      \"引人注目\"\n    ],\n    \"提高\": [\n      \"提升\",\n      \"增强\",\n      \"改善\",\n      \"促进\",\n      \"推进\"\n    ],\n    \"降低\": [\n      \"减少\",\n      \"下降\",\n      \"削弱\",\n      \"缩减\",\n      \"抑制\"\n    ],\n    \"增加\": [\n      \"增长\",\n      \"上升\",\n      \"增多\",\n      \"扩大\",\n      \"提升\"\n    ],\n    \"减少\": [\n      \"降低\",\n      \"缩减\",\n      \"削减\",\n      \"下降\",\n      \"缩小\"\n    ],\n    \"影响\": [\n      \"作用\",\n      \"效应\",\n      \"干预\",\n      \"冲击\",\n      \"波及\"\n    ],\n    \"相关\": [\n      \"关联\",\n      \"有关\",\n      \"联系\",\n      \"涉及\",\n      \"牵涉\"\n    ],\n    \"包括\": [\n      \"涵盖\",\n      \"包含\",\n      \"涉及\",\n      \"囊括\",\n      \"纳入\"\n    ],\n    \"发现\": [\n      \"观察到\",\n      \"检测到\",\n      \"识别出\",\n      \"注意到\",\n      \"获得\"\n    ],\n    \"使用\": [\n      \"采用\",\n      \"运用\",\n      \"借助\",\n      \"利用\",\n      \"通过\"\n    ],\n    \"方法\": [\n      \"手段\",\n      \"途径\",\n      \"策略\",\n      \"方式\",\n      \"技术\"\n    ],\n    \"结果\": [\n      \"结局\",\n      \"成效\",\n      \"产出\",\n      \"发现\",\n      \"所得\"\n    ],\n    \"分析\": [\n      \"剖析\",\n      \"研讨\",\n      \"审视\",\n      \"考察\",\n      \"探究\"\n    ],\n    \"研究\": [\n      \"考察\",\n      \"探索\",\n      \"调查\",\n      \"钻研\",\n      \"探究\"\n    ],\n    \"认为\": [\n      \"主张\",\n      \"提出\",\n      \"推断\",\n      \"倾向于\",\n      \"判断\"\n    ],\n    \"指出\": [\n      \"提出\",\n      \"强调\",\n      \"阐明\",\n      \"论述\",\n      \"提及\"\n    ],\n    \"探讨\": [\n      \"讨论\",\n      \"探究\",\n      \"剖析\",\n      \"审视\",\n      \"辨析\"\n    ],\n    \"评价\": [\n      \"评估\",\n      \"判定\",\n      \"衡量\",\n      \"鉴定\",\n      \"考量\"\n    ],\n    \"比较\": [\n      \"对比\",\n      \"对照\",\n      \"相较\",\n      \"类比\",\n      \"参照\"\n    ],\n    \"观察\": [\n      \"监测\",\n      \"跟踪\",\n      \"注视\",\n      \"留意\",\n      \"审视\"\n    ],\n    \"变化\": [\n      \"改变\",\n      \"变动\",\n      \"转变\",\n      \"波动\",\n      \"演变\"\n    ],\n    \"特点\": [\n      \"特征\",\n      \"特性\",\n      \"属性\",\n      \"特色\",\n      \"特质\"\n    ],\n    \"问题\": [\n      \"议题\",\n      \"难题\",\n      \"困境\",\n      \"挑战\",\n      \"瓶颈\"\n    ],\n    \"目的\": [\n      \"目标\",\n      \"意图\",\n      \"用意\",\n      \"导向\",\n      \"宗旨\"\n    ],\n    \"过程\": [\n      \"流程\",\n      \"历程\",\n      \"阶段\",\n      \"演进\",\n      \"经过\"\n    ],\n    \"基础\": [\n      \"根基\",\n      \"前提\",\n      \"依托\",\n      \"基石\",\n      \"底座\"\n    ],\n    \"条件\": [\n      \"前提\",\n      \"环境\",\n      \"背景\",\n      \"情境\",\n      \"因素\"\n    ],\n    \"领域\": [\n      \"范畴\",\n      \"范围\",\n      \"区间\",\n      \"方面\",\n      \"方向\"\n    ],\n    \"趋势\": [\n      \"走向\",\n      \"动向\",\n      \"态势\",\n      \"倾向\",\n      \"潮流\"\n    ],\n    \"效果\": [\n      \"成效\",\n      \"功效\",\n      \"作用\",\n      \"效用\",\n      \"成果\"\n    ],\n    \"需求\": [\n      \"需要\",\n      \"诉求\",\n      \"要求\",\n      \"渴望\",\n      \"诉求\"\n    ],\n    \"现状\": [\n      \"情形\",\n      \"态势\",\n      \"格局\",\n      \"面貌\",\n      \"状况\"\n    ],\n    \"原因\": [\n      \"缘由\",\n      \"成因\",\n      \"根源\",\n      \"起因\",\n      \"动因\"\n    ],\n    \"意义\": [\n      \"价值\",\n      \"重要性\",\n      \"必要性\",\n      \"作用\",\n      \"影响\"\n    ],\n    \"方面\": [\n      \"角度\",\n      \"层面\",\n      \"维度\",\n      \"方向\",\n      \"视角\"\n    ],\n    \"通过\": [\n      \"借助\",\n      \"依托\",\n      \"凭借\",\n      \"利用\",\n      \"经由\"\n    ],\n    \"针对\": [\n      \"面向\",\n      \"聚焦于\",\n      \"围绕\",\n      \"就\",\n      \"针对\"\n    ],\n    \"结合\": [\n      \"融合\",\n      \"整合\",\n      \"串联\",\n      \"联合\",\n      \"兼顾\"\n    ],\n    \"考虑\": [\n      \"考量\",\n      \"权衡\",\n      \"斟酌\",\n      \"顾及\",\n      \"统筹\"\n    ],\n    \"确定\": [\n      \"明确\",\n      \"认定\",\n      \"判定\",\n      \"核实\",\n      \"确认\"\n    ],\n    \"提供\": [\n      \"给予\",\n      \"供给\",\n      \"带来\",\n      \"贡献\",\n      \"输出\"\n    ],\n    \"保证\": [\n      \"确保\",\n      \"保障\",\n      \"维护\",\n      \"维持\",\n      \"巩固\"\n    ],\n    \"实现\": [\n      \"达成\",\n      \"完成\",\n      \"促成\",\n      \"落实\",\n      \"做到\"\n    ],\n    \"发展\": [\n      \"演进\",\n      \"演变\",\n      \"拓展\",\n      \"推进\",\n      \"深化\"\n    ],\n    \"应用\": [\n      \"运用\",\n      \"采用\",\n      \"使用\",\n      \"实践\",\n      \"部署\"\n    ],\n    \"提出\": [\n      \"构建\",\n      \"创立\",\n      \"设计\",\n      \"倡导\",\n      \"构思\"\n    ],\n    \"建立\": [\n      \"构建\",\n      \"创建\",\n      \"搭建\",\n      \"确立\",\n      \"组建\"\n    ],\n    \"完善\": [\n      \"优化\",\n      \"改进\",\n      \"健全\",\n      \"充实\",\n      \"改良\"\n    ],\n    \"优化\": [\n      \"改良\",\n      \"改进\",\n      \"提升\",\n      \"升级\",\n      \"精调\"\n    ],\n    \"验证\": [\n      \"检验\",\n      \"证实\",\n      \"确认\",\n      \"考核\",\n      \"校验\"\n    ],\n    \"探索\": [\n      \"发掘\",\n      \"寻求\",\n      \"尝试\",\n      \"摸索\",\n      \"开拓\"\n    ],\n    \"解决\": [\n      \"应对\",\n      \"处理\",\n      \"化解\",\n      \"攻克\",\n      \"应对\"\n    ],\n    \"开展\": [\n      \"实施\",\n      \"推进\",\n      \"进行\",\n      \"启动\",\n      \"组织\"\n    ],\n    \"调查\": [\n      \"调研\",\n      \"考察\",\n      \"摸底\",\n      \"排查\",\n      \"探访\"\n    ],\n    \"收集\": [\n      \"采集\",\n      \"获取\",\n      \"汇集\",\n      \"征询\",\n      \"调取\"\n    ],\n    \"筛选\": [\n      \"甄别\",\n      \"遴选\",\n      \"挑选\",\n      \"过滤\",\n      \"淘汰\"\n    ],\n    \"纳入\": [\n      \"收入\",\n      \"纳入研究\",\n      \"选入\",\n      \"招募\",\n      \"入组\"\n    ],\n    \"排除\": [\n      \"剔除\",\n      \"排除在外\",\n      \"不予纳入\",\n      \"筛除\",\n      \"舍去\"\n    ],\n    \"随访\": [\n      \"追踪\",\n      \"跟踪\",\n      \"定期复查\",\n      \"远期观察\",\n      \"追踪观察\"\n    ],\n    \"确诊\": [\n      \"明确诊断\",\n      \"确定诊断\",\n      \"诊断明确\",\n      \"确诊为\",\n      \"判定为\"\n    ],\n    \"治疗\": [\n      \"医治\",\n      \"施治\",\n      \"干预\",\n      \"处理\",\n      \"救治\"\n    ],\n    \"预防\": [\n      \"防治\",\n      \"防备\",\n      \"防范\",\n      \"预防性\",\n      \"规避\"\n    ],\n    \"评估\": [\n      \"判定\",\n      \"测量\",\n      \"衡量\",\n      \"考评\",\n      \"鉴定\"\n    ],\n    \"监测\": [\n      \"监控\",\n      \"追踪\",\n      \"动态观察\",\n      \"定期检查\",\n      \"密切观察\"\n    ],\n    \"记录\": [\n      \"记载\",\n      \"登记\",\n      \"记录在案\",\n      \"建档\",\n      \"载入\"\n    ],\n    \"统计\": [\n      \"计量\",\n      \"测算\",\n      \"统计分析\",\n      \"量化\",\n      \"清点\"\n    ],\n    \"报道\": [\n      \"报告\",\n      \"文献报道\",\n      \"已有报告\",\n      \"先前报道\",\n      \"见诸文献\"\n    ],\n    \"提示\": [\n      \"暗示\",\n      \"表明\",\n      \"意味着\",\n      \"反映\",\n      \"透露\"\n    ],\n    \"证明\": [\n      \"证实\",\n      \"确证\",\n      \"印证\",\n      \"佐证\",\n      \"支持\"\n    ],\n    \"推测\": [\n      \"猜测\",\n      \"推断\",\n      \"估计\",\n      \"料想\",\n      \"揣测\"\n    ],\n    \"强调\": [\n      \"着重指出\",\n      \"特别提出\",\n      \"突出\",\n      \"重申\",\n      \"申明\"\n    ],\n    \"质疑\": [\n      \"怀疑\",\n      \"商榷\",\n      \"存疑\",\n      \"提出异议\",\n      \"争议\"\n    ],\n    \"反对\": [\n      \"不认同\",\n      \"否定\",\n      \"批驳\",\n      \"驳斥\",\n      \"持异议\"\n    ],\n    \"支持\": [\n      \"赞同\",\n      \"拥护\",\n      \"背书\",\n      \"肯定\",\n      \"认同\"\n    ],\n    \"建议\": [\n      \"提议\",\n      \"推荐\",\n      \"倡导\",\n      \"主张\",\n      \"呼吁\"\n    ],\n    \"总结\": [\n      \"归纳\",\n      \"概括\",\n      \"总结性\",\n      \"小结\",\n      \"概述\"\n    ],\n    \"描述\": [\n      \"刻画\",\n      \"描绘\",\n      \"叙述\",\n      \"陈述\",\n      \"表征\"\n    ],\n    \"解释\": [\n      \"阐释\",\n      \"诠释\",\n      \"解读\",\n      \"阐明\",\n      \"析理\"\n    ],\n    \"讨论\": [\n      \"商榷\",\n      \"探讨\",\n      \"议论\",\n      \"研判\",\n      \"辨析\"\n    ],\n    \"改善\": [\n      \"好转\",\n      \"进步\",\n      \"改观\",\n      \"转好\",\n      \"缓解\"\n    ],\n    \"恶化\": [\n      \"加重\",\n      \"进展\",\n      \"变差\",\n      \"转劣\",\n      \"退步\"\n    ],\n    \"恢复\": [\n      \"复原\",\n      \"回归\",\n      \"好转\",\n      \"改善至\",\n      \"痊愈\"\n    ],\n    \"维持\": [\n      \"保持\",\n      \"延续\",\n      \"稳定于\",\n      \"控制在\",\n      \"保住\"\n    ],\n    \"波动\": [\n      \"起伏\",\n      \"震荡\",\n      \"不稳定\",\n      \"忽高忽低\",\n      \"变动\"\n    ],\n    \"升高\": [\n      \"上升\",\n      \"增高\",\n      \"攀升\",\n      \"上涨\",\n      \"上扬\"\n    ],\n    \"下降\": [\n      \"降低\",\n      \"回落\",\n      \"下跌\",\n      \"减少\",\n      \"递减\"\n    ],\n    \"实施\": [\n      \"执行\",\n      \"落实\",\n      \"开展\",\n      \"推行\",\n      \"付诸实施\"\n    ],\n    \"制定\": [\n      \"拟定\",\n      \"编制\",\n      \"规划\",\n      \"设计\",\n      \"出台\"\n    ],\n    \"修改\": [\n      \"修订\",\n      \"更正\",\n      \"调整\",\n      \"修正\",\n      \"改动\"\n    ],\n    \"删除\": [\n      \"删去\",\n      \"移除\",\n      \"剔除\",\n      \"裁剪\",\n      \"去掉\"\n    ],\n    \"替换\": [\n      \"替代\",\n      \"置换\",\n      \"更换\",\n      \"取代\",\n      \"代用\"\n    ],\n    \"调整\": [\n      \"调节\",\n      \"微调\",\n      \"修订\",\n      \"校准\",\n      \"矫正\"\n    ],\n    \"明显\": [\n      \"显著\",\n      \"突出\",\n      \"可观\",\n      \"引人注目\",\n      \"醒目\"\n    ],\n    \"轻微\": [\n      \"微弱\",\n      \"微小\",\n      \"轻微的\",\n      \"不显著\",\n      \"不甚明显\"\n    ],\n    \"严重\": [\n      \"重度\",\n      \"危重\",\n      \"沉重\",\n      \"严峻\",\n      \"剧烈\"\n    ],\n    \"复杂\": [\n      \"繁复\",\n      \"错综\",\n      \"多重\",\n      \"棘手\",\n      \"盘根错节\"\n    ],\n    \"简单\": [\n      \"简易\",\n      \"简便\",\n      \"单一\",\n      \"简洁\",\n      \"不复杂\"\n    ],\n    \"准确\": [\n      \"精确\",\n      \"精准\",\n      \"可靠\",\n      \"高精度\",\n      \"确当\"\n    ],\n    \"稳定\": [\n      \"平稳\",\n      \"可靠\",\n      \"恒定\",\n      \"波动小\",\n      \"均衡\"\n    ],\n    \"安全\": [\n      \"可靠\",\n      \"低风险\",\n      \"耐受性好\",\n      \"无严重不良反应\",\n      \"可控\"\n    ],\n    \"有效\": [\n      \"奏效\",\n      \"管用\",\n      \"起效\",\n      \"有确切疗效\",\n      \"达到预期\"\n    ],\n    \"常见\": [\n      \"多见\",\n      \"普遍\",\n      \"多发\",\n      \"频发\",\n      \"屡见不鲜\"\n    ],\n    \"罕见\": [\n      \"少见\",\n      \"稀有\",\n      \"不常见\",\n      \"罕见性\",\n      \"低发\"\n    ],\n    \"典型\": [\n      \"经典\",\n      \"代表性\",\n      \"标准\",\n      \"特征性\",\n      \"标志性\"\n    ],\n    \"特殊\": [\n      \"特定\",\n      \"独特\",\n      \"异常\",\n      \"非典型\",\n      \"特别\"\n    ],\n    \"必要\": [\n      \"必需\",\n      \"不可或缺\",\n      \"必不可少\",\n      \"必须\",\n      \"须要\"\n    ],\n    \"充分\": [\n      \"足够\",\n      \"充足\",\n      \"完备\",\n      \"详尽\",\n      \"充裕\"\n    ],\n    \"明确\": [\n      \"清晰\",\n      \"确切\",\n      \"清楚\",\n      \"明朗\",\n      \"不含糊\"\n    ],\n    \"合理\": [\n      \"得当\",\n      \"妥当\",\n      \"恰当\",\n      \"适宜\",\n      \"适切\"\n    ],\n    \"患者\": [\n      \"病人\",\n      \"受试者\",\n      \"病例\",\n      \"就诊者\",\n      \"研究对象\"\n    ],\n    \"症状\": [\n      \"临床表现\",\n      \"征象\",\n      \"主诉\",\n      \"不适\",\n      \"症候\"\n    ],\n    \"体征\": [\n      \"体格检查所见\",\n      \"查体发现\",\n      \"阳性体征\",\n      \"客观指标\",\n      \"物理征\"\n    ],\n    \"预后\": [\n      \"转归\",\n      \"结局\",\n      \"远期结果\",\n      \"随访结果\",\n      \"疾病走向\"\n    ],\n    \"并发症\": [\n      \"合并症\",\n      \"继发病变\",\n      \"不良反应\",\n      \"副反应\",\n      \"合并情况\"\n    ],\n    \"禁忌症\": [\n      \"禁忌\",\n      \"不适用\",\n      \"禁用\",\n      \"不推荐\",\n      \"避免使用\"\n    ],\n    \"适应症\": [\n      \"适用范围\",\n      \"适应证\",\n      \"适用人群\",\n      \"用药指征\",\n      \"推荐使用\"\n    ],\n    \"不良反应\": [\n      \"副作用\",\n      \"不良事件\",\n      \"药物反应\",\n      \"毒副反应\",\n      \"耐受性问题\"\n    ],\n    \"危险因素\": [\n      \"风险因素\",\n      \"致病因素\",\n      \"相关因素\",\n      \"诱因\",\n      \"高危因素\"\n    ],\n    \"对照组\": [\n      \"对照\",\n      \"对比组\",\n      \"参考组\",\n      \"常规治疗组\",\n      \"基准组\"\n    ],\n    \"治疗组\": [\n      \"干预组\",\n      \"实验组\",\n      \"用药组\",\n      \"试验组\",\n      \"研究组\"\n    ],\n    \"剂量\": [\n      \"用量\",\n      \"给药量\",\n      \"药量\",\n      \"用药剂量\",\n      \"治疗剂量\"\n    ],\n    \"疗程\": [\n      \"治疗周期\",\n      \"用药时长\",\n      \"治疗时间\",\n      \"干预时长\",\n      \"治疗阶段\"\n    ],\n    \"样本\": [\n      \"样本量\",\n      \"研究对象\",\n      \"受试者\",\n      \"数据集\",\n      \"采样\"\n    ],\n    \"数据\": [\n      \"资料\",\n      \"信息\",\n      \"数据集\",\n      \"数值\",\n      \"原始记录\"\n    ],\n    \"指标\": [\n      \"参数\",\n      \"变量\",\n      \"衡量标准\",\n      \"评价标准\",\n      \"观测指标\"\n    ],\n    \"标准\": [\n      \"准则\",\n      \"规范\",\n      \"尺度\",\n      \"判定标准\",\n      \"参照依据\"\n    ],\n    \"差异\": [\n      \"差别\",\n      \"差距\",\n      \"不同\",\n      \"分歧\",\n      \"距\"\n    ],\n    \"关联\": [\n      \"联系\",\n      \"关系\",\n      \"相关性\",\n      \"纽带\",\n      \"联结\"\n    ],\n    \"分布\": [\n      \"散布\",\n      \"频次分布\",\n      \"构成\",\n      \"占比\",\n      \"比例\"\n    ],\n    \"概率\": [\n      \"几率\",\n      \"可能性\",\n      \"风险\",\n      \"发生率\",\n      \"机率\"\n    ],\n    \"灵敏度\": [\n      \"敏感度\",\n      \"真阳性率\",\n      \"检出率\",\n      \"识别率\",\n      \"敏感性\"\n    ],\n    \"特异度\": [\n      \"特异性\",\n      \"真阴性率\",\n      \"排除率\",\n      \"鉴别力\",\n      \"准确排除率\"\n    ],\n    \"文献\": [\n      \"参考资料\",\n      \"已有研究\",\n      \"前人工作\",\n      \"既往报告\",\n      \"相关论文\"\n    ],\n    \"证据\": [\n      \"依据\",\n      \"佐证\",\n      \"凭证\",\n      \"数据支持\",\n      \"实证\"\n    ],\n    \"理论\": [\n      \"学说\",\n      \"理论框架\",\n      \"理论依据\",\n      \"理论基础\",\n      \"构想\"\n    ],\n    \"假说\": [\n      \"假设\",\n      \"推测\",\n      \"猜想\",\n      \"预设\",\n      \"设想\"\n    ],\n    \"创新\": [\n      \"新颖性\",\n      \"原创\",\n      \"突破\",\n      \"新发现\",\n      \"新意\"\n    ],\n    \"局限\": [\n      \"不足\",\n      \"缺陷\",\n      \"短板\",\n      \"限制\",\n      \"欠缺\"\n    ],\n    \"展望\": [\n      \"前景\",\n      \"未来方向\",\n      \"发展方向\",\n      \"下一步\",\n      \"后续工作\"\n    ],\n    \"相比之下\": [\n      \"与此不同\",\n      \"相反\",\n      \"比较而言\",\n      \"与之相比\",\n      \"反观\"\n    ],\n    \"与此同时\": [\n      \"同时\",\n      \"在此期间\",\n      \"同期\",\n      \"另一当面\",\n      \"并行地\"\n    ],\n    \"综上所述\": [\n      \"总之\",\n      \"综合来看\",\n      \"归纳起来\",\n      \"总而言之\",\n      \"整体而言\"\n    ],\n    \"值得注意的是\": [\n      \"需要关注的是\",\n      \"引人注目的是\",\n      \"特别指出\",\n      \"应当注意\",\n      \"不可忽视的是\"\n    ],\n    \"在一定程度上\": [\n      \"部分地\",\n      \"从某种意义上\",\n      \"某种程度上\",\n      \"一定范围内\",\n      \"相对而言\"\n    ],\n    \"众所周知\": [\n      \"公认地\",\n      \"普遍认为\",\n      \"常识是\",\n      \"已有共识\",\n      \"广为人知\"\n    ],\n    \"据我们所知\": [\n      \"据文献检索\",\n      \"据已有报道\",\n      \"据掌握资料\",\n      \"据目前了解\",\n      \"检索文献发现\"\n    ],\n    \"首次\": [\n      \"第一次\",\n      \"开创性地\",\n      \"首次报道\",\n      \"先驱性\",\n      \"前所未有地\"\n    ],\n    \"给药\": [\n      \"用药\",\n      \"施药\",\n      \"投药\",\n      \"给药方案\",\n      \"用药方案\"\n    ],\n    \"手术\": [\n      \"术式\",\n      \"外科干预\",\n      \"手术方式\",\n      \"手术治疗\",\n      \"操作\"\n    ],\n    \"活检\": [\n      \"组织活检\",\n      \"病理取样\",\n      \"穿刺\",\n      \"取材\",\n      \"标本采集\"\n    ],\n    \"穿刺\": [\n      \"穿刺术\",\n      \"穿刺活检\",\n      \"抽吸\",\n      \"穿刺取样\",\n      \"针吸\"\n    ],\n    \"输血\": [\n      \"血液输注\",\n      \"输注\",\n      \"血液制品\",\n      \"成分输血\",\n      \"输血治疗\"\n    ],\n    \"转诊\": [\n      \"转院\",\n      \"转介\",\n      \"转至\",\n      \"转送\",\n      \"转入\"\n    ],\n    \"出院\": [\n      \"出院\",\n      \"出院后\",\n      \"出院随访\",\n      \"离院\",\n      \"结束住院\"\n    ],\n    \"入院\": [\n      \"收治\",\n      \"收治入院\",\n      \"住院\",\n      \"入院治疗\",\n      \"入院诊断\"\n    ],\n    \"抢救\": [\n      \"急救\",\n      \"紧急处理\",\n      \"急诊处理\",\n      \"复苏\",\n      \"生命支持\"\n    ],\n    \"计算\": [\n      \"测算\",\n      \"核算\",\n      \"推算\",\n      \"统计得出\",\n      \"得出\"\n    ],\n    \"回归\": [\n      \"回归分析\",\n      \"拟合\",\n      \"建立模型\",\n      \"建模\",\n      \"多因素分析\"\n    ],\n    \"校正\": [\n      \"调整\",\n      \"控制变量\",\n      \"纳入协变量\",\n      \"分层分析\",\n      \"标准化\"\n    ],\n    \"分层\": [\n      \"分组\",\n      \"亚组\",\n      \"分层分析\",\n      \"按层次\",\n      \"分类讨论\"\n    ],\n    \"匹配\": [\n      \"配对\",\n      \"匹配分析\",\n      \"倾向性匹配\",\n      \"配比\",\n      \"均衡\"\n    ],\n    \"随机\": [\n      \"随机化\",\n      \"随机分组\",\n      \"随机分配\",\n      \"随机对照\",\n      \"盲法\"\n    ],\n    \"盲法\": [\n      \"双盲\",\n      \"单盲\",\n      \"开放标签\",\n      \"设盲\",\n      \"非盲\"\n    ],\n    \"中止\": [\n      \"终止\",\n      \"停止\",\n      \"中断\",\n      \"提前结束\",\n      \"叫停\"\n    ],\n    \"完成\": [\n      \"完成率\",\n      \"依从\",\n      \"按要求完成\",\n      \"顺利结束\",\n      \"执行完毕\"\n    ],\n    \"不显著\": [\n      \"无统计学意义\",\n      \"未达显著\",\n      \"无明显差异\",\n      \"差异不显著\",\n      \"无统计学差异\"\n    ],\n    \"独立\": [\n      \"独立性\",\n      \"独立预测\",\n      \"独立因素\",\n      \"自变量\",\n      \"单独\"\n    ],\n    \"线性\": [\n      \"线性关系\",\n      \"呈线性\",\n      \"正相关\",\n      \"负相关\",\n      \"剂量依赖\"\n    ],\n    \"非线性\": [\n      \"曲线关系\",\n      \"阈值效应\",\n      \"J型\",\n      \"U型\",\n      \"拐点\"\n    ],\n    \"发病\": [\n      \"起病\",\n      \"疾病发生\",\n      \"发病过程\",\n      \"疾病进展\",\n      \"新发\"\n    ],\n    \"缓解\": [\n      \"好转\",\n      \"症状消退\",\n      \"临床缓解\",\n      \"病情改善\",\n      \"部分缓解\"\n    ],\n    \"复发\": [\n      \"再发\",\n      \"反复\",\n      \"再次发作\",\n      \"病情反复\",\n      \"复发率\"\n    ],\n    \"死亡\": [\n      \"致死\",\n      \"死亡病例\",\n      \"病死率\",\n      \"死亡风险\",\n      \"致命\"\n    ],\n    \"生存\": [\n      \"存活\",\n      \"生存率\",\n      \"生存期\",\n      \"生存时间\",\n      \"无病生存\"\n    ],\n    \"治愈\": [\n      \"痊愈\",\n      \"根治\",\n      \"完全缓解\",\n      \"临床治愈\",\n      \"康复\"\n    ],\n    \"转移\": [\n      \"播散\",\n      \"远处转移\",\n      \"血行转移\",\n      \"淋巴转移\",\n      \"种植转移\"\n    ],\n    \"感染\": [\n      \"继发感染\",\n      \"合并感染\",\n      \"院内感染\",\n      \"机会性感染\",\n      \"侵袭性\"\n    ],\n    \"疗效\": [\n      \"治疗效果\",\n      \"有效性\",\n      \"治疗效果\",\n      \"临床获益\",\n      \"治疗应答\"\n    ],\n    \"毒性\": [\n      \"毒副作用\",\n      \"毒性反应\",\n      \"器官毒性\",\n      \"药物毒性\",\n      \"安全性\"\n    ],\n    \"耐受\": [\n      \"耐受性\",\n      \"耐受良好\",\n      \"可耐受\",\n      \"依从性好\",\n      \"可接受\"\n    ],\n    \"耐药\": [\n      \"抗药性\",\n      \"药物抵抗\",\n      \"继发耐药\",\n      \"原发耐药\",\n      \"获得性耐药\"\n    ],\n    \"联用\": [\n      \"联合用药\",\n      \"合并用药\",\n      \"联合方案\",\n      \"联合治疗\",\n      \"组合疗法\"\n    ],\n    \"单药\": [\n      \"单独用药\",\n      \"单药治疗\",\n      \"单一方案\",\n      \"单用\",\n      \"不联合\"\n    ],\n    \"影像\": [\n      \"影像学\",\n      \"影像检查\",\n      \"影像学表现\",\n      \"影像发现\",\n      \"放射学\"\n    ],\n    \"病理\": [\n      \"组织学\",\n      \"病理学\",\n      \"病理改变\",\n      \"病理结果\",\n      \"镜下所见\"\n    ],\n    \"实验室\": [\n      \"化验\",\n      \"检验\",\n      \"实验室检查\",\n      \"生化指标\",\n      \"血液学\"\n    ],\n    \"基因\": [\n      \"遗传学\",\n      \"基因检测\",\n      \"分子标志物\",\n      \"基因突变\",\n      \"基因组学\"\n    ],\n    \"回顾\": [\n      \"回顾性\",\n      \"回顾分析\",\n      \"历史数据\",\n      \"回顾研究\",\n      \"事后分析\"\n    ],\n    \"前瞻\": [\n      \"前瞻性\",\n      \"前瞻研究\",\n      \"纵向\",\n      \"队列\",\n      \"随访研究\"\n    ],\n    \"横断面\": [\n      \"断面研究\",\n      \"时点调查\",\n      \"现况调查\",\n      \"一次性\",\n      \"时间点\"\n    ],\n    \"多中心\": [\n      \"多中心研究\",\n      \"多家医院\",\n      \"多机构\",\n      \"跨中心\",\n      \"协作研究\"\n    ],\n    \"单中心\": [\n      \"单一机构\",\n      \"单家医院\",\n      \"本单位\",\n      \"院内\",\n      \"机构内\"\n    ],\n    \"双盲\": [\n      \"双盲设计\",\n      \"安慰剂对照\",\n      \"设盲\",\n      \"随机双盲\",\n      \"盲法评估\"\n    ],\n    \"安慰剂\": [\n      \"安慰剂对照\",\n      \"假治疗\",\n      \"空白对照\",\n      \"对照组干预\",\n      \"模拟治疗\"\n    ],\n    \"然而\": [\n      \"但\",\n      \"不过\",\n      \"尽管如此\",\n      \"值得注意的是\",\n      \"另一方面\"\n    ],\n    \"此外\": [\n      \"另外\",\n      \"同时\",\n      \"补充一点\",\n      \"不仅如此\",\n      \"除此之外\"\n    ],\n    \"因此\": [\n      \"所以\",\n      \"由此可见\",\n      \"基于此\",\n      \"有鉴于此\",\n      \"由此可见\"\n    ],\n    \"结论\": [\n      \"得出结论\",\n      \"总而言之\",\n      \"归纳为\",\n      \"研究结论\",\n      \"最终结论\"\n    ],\n    \"背景\": [\n      \"研究背景\",\n      \"问题的提出\",\n      \"起因\",\n      \"缘起\",\n      \"为什么做这项研究\"\n    ],\n    \"创新点\": [\n      \"亮点\",\n      \"新贡献\",\n      \"创新之处\",\n      \"突破\",\n      \"与既往研究不同之处\"\n    ],\n    \"显著改善\": [\n      \"明显好转\",\n      \"有了质的飞跃\",\n      \"进步很大\",\n      \"效果突出\",\n      \"大幅提升\"\n    ],\n    \"深入研究\": [\n      \"仔细分析\",\n      \"细致考察\",\n      \"深入调查\",\n      \"重点攻关\",\n      \"全面剖析\"\n    ],\n    \"广泛关注\": [\n      \"热议\",\n      \"引起重视\",\n      \"备受瞩目\",\n      \"学界关注\",\n      \"焦点话题\"\n    ],\n    \"重要意义\": [\n      \"价值很大\",\n      \"不可忽视\",\n      \"举足轻重\",\n      \"至关重要\",\n      \"影响深远\"\n    ],\n    \"亟待解决\": [\n      \"迫切需要\",\n      \"当务之急\",\n      \"迫在眉睫\",\n      \"亟待\",\n      \"亟需\"\n    ],\n    \"日益增长\": [\n      \"越来越多\",\n      \"不断增加\",\n      \"逐年上升\",\n      \"持续攀升\",\n      \"与日俱增\"\n    ],\n    \"发挥着重要作用\": [\n      \"占据核心地位\",\n      \"不可或缺\",\n      \"具有关键价值\",\n      \"至关重要\",\n      \"意义突出\"\n    ],\n    \"取得了显著进展\": [\n      \"进步很大\",\n      \"成效显著\",\n      \"有了突破\",\n      \"硕果累累\",\n      \"进展迅速\"\n    ],\n    \"为...提供了\": [\n      \"有助于\",\n      \"有助于推动\",\n      \"助力\",\n      \"赋能\",\n      \"给...带来\"\n    ],\n    \"在...方面\": [\n      \"从...角度看\",\n      \"就...而言\",\n      \"在...维度\",\n      \"在...层面\",\n      \"关于...\"\n    ],\n    \"本研究\": [\n      \"我们的研究\",\n      \"本项工作\",\n      \"这项调查\",\n      \"本次分析\",\n      \"我们的分析\"\n    ],\n    \"既往研究\": [\n      \"之前的研究\",\n      \"前人工作\",\n      \"已有文献\",\n      \"早期报道\",\n      \"先前的研究\"\n    ],\n    \"本研究首次\": [\n      \"我们是第一个\",\n      \"此前未见报道\",\n      \"据文献检索\",\n      \"首次发现\",\n      \"率先\"\n    ],\n    \"有一定的\": [\n      \"存在某种程度的\",\n      \"部分\",\n      \"某种意义上\",\n      \"在一定程度上\",\n      \"不乏\"\n    ],\n    \"一定程度上\": [\n      \"部分地\",\n      \"在某种范围内\",\n      \"相对地\",\n      \"有限度地\",\n      \"有条件地\"\n    ],\n    \"具有重要意义\": [\n      \"意义重大\",\n      \"非常关键\",\n      \"不可小觑\",\n      \"举足轻重\",\n      \"价值突出\"\n    ],\n    \"得到了验证\": [\n      \"证实了\",\n      \"得到了印证\",\n      \"确证了\",\n      \"支持了\",\n      \"佐证了\"\n    ],\n    \"尚不清楚\": [\n      \"还不确定\",\n      \"仍有疑问\",\n      \"有待阐明\",\n      \"不甚明了\",\n      \"尚无定论\"\n    ],\n    \"需要进一步\": [\n      \"还有待\",\n      \"仍需\",\n      \"后续需要\",\n      \"下一步应当\",\n      \"亟待\"\n    ],\n    \"越来越多的\": [\n      \"日益增多的\",\n      \"不断增长的\",\n      \"越来越多的\",\n      \"与日俱增的\",\n      \"逐渐增加的\"\n    ],\n    \"日益引起重视\": [\n      \"越来越受关注\",\n      \"备受瞩目\",\n      \"引发热议\",\n      \"引起广泛讨论\",\n      \"成为焦点\"\n    ],\n    \"口服\": [\n      \"口服给药\",\n      \"经口\",\n      \"经口服用\",\n      \"PO\",\n      \"口服途径\"\n    ],\n    \"静脉\": [\n      \"静脉注射\",\n      \"静注\",\n      \"静脉输注\",\n      \"IV\",\n      \"经静脉\"\n    ],\n    \"皮下\": [\n      \"皮下注射\",\n      \"皮下给药\",\n      \"SC\",\n      \"皮下途径\",\n      \"皮下埋植\"\n    ],\n    \"肌注\": [\n      \"肌肉注射\",\n      \"IM\",\n      \"肌内注射\",\n      \"注射给药\",\n      \"肌肉途径\"\n    ],\n    \"每日一次\": [\n      \"一天一次\",\n      \"qd\",\n      \"每日单次\",\n      \"每天一次\",\n      \"一日一次\"\n    ],\n    \"每日两次\": [\n      \"一天两次\",\n      \"bid\",\n      \"每日两次给药\",\n      \"每天两次\",\n      \"早晚各一次\"\n    ],\n    \"初步诊断\": [\n      \"印象诊断\",\n      \"入院诊断\",\n      \"初步印象\",\n      \"疑似诊断\",\n      \"暂定诊断\"\n    ],\n    \"鉴别诊断\": [\n      \"鉴别\",\n      \"鉴别要点\",\n      \"排除诊断\",\n      \"鉴别分析\",\n      \"需与...鉴别\"\n    ],\n    \"辅助检查\": [\n      \"实验室检查\",\n      \"检查结果\",\n      \"检验数据\",\n      \"化验结果\",\n      \"相关检查\"\n    ],\n    \"影像学检查\": [\n      \"影像\",\n      \"CT\",\n      \"MRI\",\n      \"X线\",\n      \"超声检查\"\n    ],\n    \"实验室检查\": [\n      \"化验\",\n      \"血常规\",\n      \"生化\",\n      \"免疫学检查\",\n      \"血清学\"\n    ],\n    \"一线治疗\": [\n      \"首选治疗\",\n      \"初始治疗\",\n      \"一线方案\",\n      \"标准治疗\",\n      \"常规治疗\"\n    ],\n    \"二线治疗\": [\n      \"替代治疗\",\n      \"备选方案\",\n      \"挽救治疗\",\n      \"二线方案\",\n      \"后续治疗\"\n    ],\n    \"维持治疗\": [\n      \"长期治疗\",\n      \"持续治疗\",\n      \"维持用药\",\n      \"巩固治疗\",\n      \"维持剂量\"\n    ],\n    \"诱导缓解\": [\n      \"初始治疗\",\n      \"诱导期\",\n      \"强化治疗\",\n      \"积极治疗\",\n      \"冲击治疗\"\n    ],\n    \"联合治疗\": [\n      \"组合治疗\",\n      \"联合方案\",\n      \"综合治疗\",\n      \"多药联合\",\n      \"序贯治疗\"\n    ],\n    \"对症治疗\": [\n      \"支持治疗\",\n      \"姑息治疗\",\n      \"对症处理\",\n      \"支持性治疗\",\n      \"缓解治疗\"\n    ],\n    \"主要终点\": [\n      \"主要结局\",\n      \"首要终点\",\n      \"primary endpoint\",\n      \"主要指标\",\n      \"核心结局\"\n    ],\n    \"次要终点\": [\n      \"次要结局\",\n      \"secondary endpoint\",\n      \"次要指标\",\n      \"附加结局\",\n      \"辅助指标\"\n    ],\n    \"无进展生存\": [\n      \"PFS\",\n      \"疾病未进展时间\",\n      \"无进展生存期\",\n      \"肿瘤控制时间\",\n      \"至进展时间\"\n    ],\n    \"总生存\": [\n      \"OS\",\n      \"整体生存\",\n      \"总生存期\",\n      \"存活时间\",\n      \"生存期\"\n    ],\n    \"无病生存\": [\n      \"DFS\",\n      \"无病生存期\",\n      \"无复发时间\",\n      \"治愈后随访时间\",\n      \"至复发时间\"\n    ],\n    \"置信区间\": [\n      \"CI\",\n      \"可信区间\",\n      \"置信水平\",\n      \"区间估计\",\n      \"95%CI\"\n    ],\n    \"比值比\": [\n      \"OR\",\n      \"优势比\",\n      \"odds ratio\",\n      \"机会比\",\n      \"相对比值\"\n    ],\n    \"风险比\": [\n      \"HR\",\n      \"hazard ratio\",\n      \"危害比\",\n      \"相对风险\",\n      \"风险系数\"\n    ],\n    \"相对风险\": [\n      \"RR\",\n      \"relative risk\",\n      \"风险比\",\n      \"相对危险度\",\n      \"相对危险性\"\n    ],\n    \"绝对风险\": [\n      \"AR\",\n      \"absolute risk\",\n      \"绝对危险度\",\n      \"事件发生率\",\n      \"实际风险\"\n    ],\n    \"P值\": [\n      \"显著性水平\",\n      \"P-value\",\n      \"统计学P值\",\n      \"概率值\",\n      \"α水平\"\n    ],\n    \"显著性\": [\n      \"统计学意义\",\n      \"统计显著\",\n      \"显著差异\",\n      \"有显著性\",\n      \"达到显著\"\n    ],\n    \"具有挑战性\": [\n      \"不容易\",\n      \"有难度\",\n      \"困难\",\n      \"问题复杂\",\n      \"棘手\"\n    ],\n    \"提供了新的视角\": [\n      \"看到了不同的角度\",\n      \"带来新想法\",\n      \"打开新思路\",\n      \"提供了线索\",\n      \"另辟蹊径\"\n    ],\n    \"填补了空白\": [\n      \"弥补了不足\",\n      \"弥补空白\",\n      \"补充了证据\",\n      \"丰富了文献\",\n      \"补上了缺口\"\n    ],\n    \"为...奠定基础\": [\n      \"给...打下基础\",\n      \"做了铺垫\",\n      \"奠定了\",\n      \"提供了前提\",\n      \"做出了奠基\"\n    ],\n    \"越来越受到\": [\n      \"日益得到\",\n      \"逐渐获得\",\n      \"不断获得\",\n      \"越来越获得\",\n      \"愈发得到\"\n    ],\n    \"取得了满意的效果\": [\n      \"效果不错\",\n      \"结果令人满意\",\n      \"达到了预期\",\n      \"表现良好\",\n      \"效果理想\"\n    ],\n    \"具有良好的\": [\n      \"表现出较好的\",\n      \"具备不错的\",\n      \"呈现出良好的\",\n      \"显示出优良的\",\n      \"展现了良好的\"\n    ],\n    \"在临床实践中\": [\n      \"实际工作中\",\n      \"临床工作中\",\n      \"日常诊疗中\",\n      \"实践中\",\n      \"在真实世界中\"\n    ],\n    \"作为一种\": [\n      \"作为一类\",\n      \"属于\",\n      \"被归类为\",\n      \"是一种\",\n      \"归入\"\n    ],\n    \"发挥着越来越重要的作用\": [\n      \"地位越来越突出\",\n      \"越来越不可或缺\",\n      \"作用越来越关键\",\n      \"日益重要\",\n      \"变得举足轻重\"\n    ],\n    \"in_summary\": [\n      \"to summarize\",\n      \"in brief\",\n      \"taken together\",\n      \"overall\",\n      \"collectively\"\n    ],\n    \"interestingly\": [\n      \"notably\",\n      \"of note\",\n      \"strikingly\",\n      \"remarkably\",\n      \"curiously\"\n    ],\n    \"consistent_with\": [\n      \"in line with\",\n      \"matching\",\n      \"concordant with\",\n      \"aligning with\",\n      \"supporting\"\n    ],\n    \"in_contrast_to\": [\n      \"unlike\",\n      \"differing from\",\n      \"as opposed to\",\n      \"compared with\",\n      \"versus\"\n    ],\n    \"despite_the_fact\": [\n      \"even though\",\n      \"although\",\n      \"while\",\n      \"notwithstanding\",\n      \"regardless of\"\n    ],\n    \"as_demonstrated_by\": [\n      \"as shown by\",\n      \"evidenced by\",\n      \"illustrated by\",\n      \"supported by\",\n      \"confirmed by\"\n    ],\n    \"the_present_study\": [\n      \"our study\",\n      \"this work\",\n      \"the current investigation\",\n      \"our analysis\",\n      \"this report\"\n    ],\n    \"previous_studies\": [\n      \"earlier work\",\n      \"prior research\",\n      \"past studies\",\n      \"existing literature\",\n      \"earlier reports\"\n    ],\n    \"免疫抑制\": [\n      \"免疫调节\",\n      \"免疫控制\",\n      \"免疫压制\",\n      \"免疫干预\"\n    ],\n    \"自身抗体\": [\n      \"自免抗体\",\n      \"自身免疫抗体\",\n      \"免疫标志物\",\n      \"血清学标志\"\n    ],\n    \"炎症\": [\n      \"炎性反应\",\n      \"炎症反应\",\n      \"炎性\",\n      \"炎症性\"\n    ],\n    \"免疫球蛋白\": [\n      \"Ig\",\n      \"抗体\",\n      \"丙种球蛋白\",\n      \"免疫球蛋白制剂\"\n    ],\n    \"补体\": [\n      \"补体系统\",\n      \"补体C3\",\n      \"补体C4\",\n      \"补体水平\"\n    ],\n    \"关节炎\": [\n      \"关节炎症\",\n      \"关节受累\",\n      \"关节病变\",\n      \"关节损害\"\n    ],\n    \"皮疹\": [\n      \"皮肤损害\",\n      \"皮肤表现\",\n      \"皮损\",\n      \"皮肤病变\"\n    ],\n    \"浆膜炎\": [\n      \"浆膜腔积液\",\n      \"胸腔积液\",\n      \"心包积液\",\n      \"腹水\"\n    ],\n    \"肾脏受累\": [\n      \"肾损害\",\n      \"肾炎\",\n      \"肾病\",\n      \"狼疮性肾炎\"\n    ],\n    \"神经系统受累\": [\n      \"神经精神\",\n      \"中枢受累\",\n      \"脑血管事件\",\n      \"癫痫发作\"\n    ],\n    \"血液系统\": [\n      \"血细胞减少\",\n      \"贫血\",\n      \"白细胞减少\",\n      \"血小板减少\"\n    ],\n    \"引起\": [\n      \"导致\",\n      \"诱发\",\n      \"引发\",\n      \"促发\",\n      \"触发\"\n    ],\n    \"防止\": [\n      \"预防\",\n      \"阻止\",\n      \"避免\",\n      \"防范\",\n      \"遏制\"\n    ],\n    \"促进\": [\n      \"推动\",\n      \"加快\",\n      \"加速\",\n      \"有助于\",\n      \"利好\"\n    ],\n    \"抑制\": [\n      \"压制\",\n      \"遏制\",\n      \"阻滞\",\n      \"减慢\",\n      \"减弱\"\n    ],\n    \"激活\": [\n      \"活化\",\n      \"启动\",\n      \"触发\",\n      \"唤起\",\n      \"激发\"\n    ],\n    \"调节\": [\n      \"调控\",\n      \"调整\",\n      \"调理\",\n      \"影响\",\n      \"干预\"\n    ],\n    \"表达\": [\n      \"呈现\",\n      \"显示\",\n      \"表现为\",\n      \"体现出\",\n      \"反映为\"\n    ],\n    \"参与\": [\n      \"介入\",\n      \"涉及\",\n      \"加入\",\n      \"牵涉\",\n      \"介入其中\"\n    ],\n    \"介导\": [\n      \"介导的\",\n      \"通过...途径\",\n      \"经由\",\n      \"以...为中介\",\n      \"介导机制\"\n    ],\n    \"机制\": [\n      \"机理\",\n      \"原理\",\n      \"作用机制\",\n      \"病理生理\",\n      \"通路\"\n    ],\n    \"通路\": [\n      \"信号通路\",\n      \"信号传导\",\n      \"级联反应\",\n      \"信号转导\",\n      \"路径\"\n    ],\n    \"靶点\": [\n      \"作用靶点\",\n      \"药物靶点\",\n      \"治疗靶点\",\n      \"靶标\",\n      \"目标分子\"\n    ],\n    \"标志物\": [\n      \"生物标志物\",\n      \"指标\",\n      \"预测因子\",\n      \"标记物\",\n      \"biomarker\"\n    ],\n    \"队列\": [\n      \"队列研究\",\n      \"人群\",\n      \"样本队列\",\n      \"研究队列\",\n      \"随访队列\"\n    ],\n    \"终点\": [\n      \"结局指标\",\n      \"研究终点\",\n      \"评价终点\",\n      \"观察终点\",\n      \"目标终点\"\n    ],\n    \"基线\": [\n      \"起始\",\n      \"基线水平\",\n      \"初始状态\",\n      \"治疗前\",\n      \"入组时\"\n    ],\n    \"安全性\": [\n      \"安全\",\n      \"不良反应\",\n      \"耐受性\",\n      \"副作用\",\n      \"药物不良事件\"\n    ],\n    \"有效性\": [\n      \"疗效\",\n      \"治疗效果\",\n      \"临床获益\",\n      \"有效性评价\",\n      \"治疗效果\"\n    ],\n    \"依从性\": [\n      \"遵医行为\",\n      \"服药依从\",\n      \"配合程度\",\n      \"遵从度\",\n      \"执行度\"\n    ],\n    \"持续\": [\n      \"持续性的\",\n      \"延续\",\n      \"不断\",\n      \"一直\",\n      \"持久\"\n    ],\n    \"反复\": [\n      \"反复性\",\n      \"多次\",\n      \"频繁\",\n      \"周期性\",\n      \"间歇性\"\n    ],\n    \"急性\": [\n      \"急骤\",\n      \"突发\",\n      \"急性发作\",\n      \"暴发性\",\n      \"来势凶猛\"\n    ],\n    \"慢性\": [\n      \"长期\",\n      \"迁延\",\n      \"慢性病程\",\n      \"持续性\",\n      \"长期性\"\n    ],\n    \"轻度\": [\n      \"轻微\",\n      \"I级\",\n      \"低度\",\n      \"不严重\",\n      \"轻微的\"\n    ],\n    \"中度\": [\n      \"中等\",\n      \"II级\",\n      \"中度严重\",\n      \"中等程度\",\n      \"有一定影响\"\n    ],\n    \"重度\": [\n      \"严重\",\n      \"III级\",\n      \"高度\",\n      \"危重\",\n      \"重症\"\n    ],\n    \"早期\": [\n      \"初期\",\n      \"早期阶段\",\n      \"发病初期\",\n      \"起始期\",\n      \"疾病早期\"\n    ],\n    \"晚期\": [\n      \"终末期\",\n      \"晚期\",\n      \"末期\",\n      \"进展期\",\n      \"晚期阶段\"\n    ],\n    \"首先\": [\n      \"第一\",\n      \"起初\",\n      \"先是\",\n      \"一开始\",\n      \"最先\"\n    ],\n    \"其次\": [\n      \"第二\",\n      \"接下来\",\n      \"然后\",\n      \"再者\",\n      \"除此之外\"\n    ],\n    \"最后\": [\n      \"最终\",\n      \"末了\",\n      \"归根结底\",\n      \"总的来看\",\n      \"归根到底\"\n    ],\n    \"同时\": [\n      \"与此并行\",\n      \"在...的同时\",\n      \"同步地\",\n      \"一边...一边\",\n      \"并行\"\n    ],\n    \"综上\": [\n      \"综合来看\",\n      \"总而言之\",\n      \"整体而言\",\n      \"归根到底\",\n      \"一句话\"\n    ],\n    \"换言之\": [\n      \"换句话说\",\n      \"也就是说\",\n      \"具体而言\",\n      \"简言之\",\n      \"说白了\"\n    ],\n    \"具体而言\": [\n      \"具体来说\",\n      \"详细地说\",\n      \"具体讲\",\n      \"准确说\",\n      \"细说\"\n    ],\n    \"不仅...而且\": [\n      \"既...又\",\n      \"除了...还\",\n      \"一方面...另一方面\",\n      \"不只是...同时\"\n    ],\n    \"一方面...另一方面\": [\n      \"既有...也有\",\n      \"既存在...也存在\",\n      \"从两个角度看\",\n      \"兼具\"\n    ],\n    \"既...又\": [\n      \"不但...而且\",\n      \"同时具备\",\n      \"兼具\",\n      \"并存\",\n      \"两者兼有\"\n    ],\n    \"根据\": [\n      \"依据\",\n      \"按照\",\n      \"基于\",\n      \"参照\",\n      \"按照\"\n    ],\n    \"基于\": [\n      \"根据\",\n      \"依据\",\n      \"立足于\",\n      \"建立在\",\n      \"以...为基础\"\n    ],\n    \"围绕\": [\n      \"聚焦于\",\n      \"围绕\",\n      \"以...为中心\",\n      \"针对\",\n      \"就...展开\"\n    ],\n    \"借助\": [\n      \"利用\",\n      \"通过\",\n      \"凭借\",\n      \"依托\",\n      \"仰仗\"\n    ],\n    \"利用\": [\n      \"使用\",\n      \"采用\",\n      \"运用\",\n      \"借助\",\n      \"应用\"\n    ],\n    \"采用\": [\n      \"使用\",\n      \"选取\",\n      \"选用\",\n      \"采纳\",\n      \"采取\"\n    ],\n    \"选取\": [\n      \"选择\",\n      \"挑出\",\n      \"选定\",\n      \"筛选出\",\n      \"甄选\"\n    ],\n    \"包含\": [\n      \"包括\",\n      \"含有\",\n      \"涵盖\",\n      \"涉及\",\n      \"囊括\"\n    ],\n    \"涉及\": [\n      \"涵盖\",\n      \"涉及\",\n      \"包括\",\n      \"牵涉\",\n      \"涵盖\"\n    ],\n    \"覆盖\": [\n      \"涵盖\",\n      \"涉及\",\n      \"包含\",\n      \"包括\",\n      \"囊括\"\n    ],\n    \"忽略\": [\n      \"忽视\",\n      \"未考虑\",\n      \"未纳入\",\n      \"遗漏\",\n      \"未涉及\"\n    ],\n    \"突出\": [\n      \"凸显\",\n      \"彰显\",\n      \"体现\",\n      \"显示出\",\n      \"尤为\"\n    ],\n    \"体现\": [\n      \"反映\",\n      \"表现\",\n      \"展示\",\n      \"呈现\",\n      \"彰显\"\n    ],\n    \"反映\": [\n      \"体现\",\n      \"展示\",\n      \"呈现\",\n      \"表现\",\n      \"说明\"\n    ],\n    \"significant\": [\n      \"notable\",\n      \"substantial\",\n      \"considerable\",\n      \"marked\",\n      \"pronounced\"\n    ],\n    \"demonstrate\": [\n      \"show\",\n      \"reveal\",\n      \"indicate\",\n      \"illustrate\",\n      \"exhibit\"\n    ],\n    \"investigate\": [\n      \"examine\",\n      \"explore\",\n      \"study\",\n      \"assess\",\n      \"evaluate\"\n    ],\n    \"subsequently\": [\n      \"then\",\n      \"next\",\n      \"afterward\",\n      \"later\",\n      \"following this\"\n    ],\n    \"approximately\": [\n      \"about\",\n      \"roughly\",\n      \"around\",\n      \"nearly\",\n      \"close to\"\n    ],\n    \"predominantly\": [\n      \"mainly\",\n      \"chiefly\",\n      \"primarily\",\n      \"largely\",\n      \"mostly\"\n    ],\n    \"comprehensive\": [\n      \"thorough\",\n      \"extensive\",\n      \"complete\",\n      \"in-depth\",\n      \"systematic\"\n    ],\n    \"efficacy\": [\n      \"effectiveness\",\n      \"performance\",\n      \"therapeutic effect\",\n      \"benefit\",\n      \"impact\"\n    ],\n    \"substantial\": [\n      \"considerable\",\n      \"significant\",\n      \"large\",\n      \"meaningful\",\n      \"notable\"\n    ],\n    \"methodology\": [\n      \"approach\",\n      \"methods\",\n      \"technique\",\n      \"procedure\",\n      \"protocol\"\n    ],\n    \"抗核抗体\": [\n      \"ANA\",\n      \"抗核抗体谱\",\n      \"核抗体\",\n      \"ANA阳性\"\n    ],\n    \"类风湿因子\": [\n      \"RF\",\n      \"类风湿因子阳性\",\n      \"RF阳性\",\n      \"乳胶凝集试验\"\n    ],\n    \"抗CCP抗体\": [\n      \"抗环瓜氨酸肽抗体\",\n      \"ACPA\",\n      \"抗瓜氨酸化蛋白抗体\"\n    ],\n    \"抗磷脂抗体\": [\n      \"aPL\",\n      \"抗心磷脂抗体\",\n      \"狼疮抗凝物\",\n      \"抗β2GP1\"\n    ],\n    \"血沉\": [\n      \"ESR\",\n      \"红细胞沉降率\",\n      \"沉降率\",\n      \"血沉加快\"\n    ],\n    \"C反应蛋白\": [\n      \"CRP\",\n      \"C-反应蛋白\",\n      \"超敏CRP\",\n      \"hs-CRP\"\n    ],\n    \"白细胞介素\": [\n      \"IL\",\n      \"白介素\",\n      \"细胞因子\",\n      \"炎症因子\"\n    ],\n    \"肿瘤坏死因子\": [\n      \"TNF\",\n      \"TNF-α\",\n      \"肿瘤坏死因子α\",\n      \"促炎因子\"\n    ],\n    \"干扰素\": [\n      \"IFN\",\n      \"干扰素α\",\n      \"干扰素β\",\n      \"干扰素γ\"\n    ],\n    \"T细胞\": [\n      \"T淋巴细胞\",\n      \"Treg\",\n      \"效应T细胞\",\n      \"辅助性T细胞\"\n    ],\n    \"B细胞\": [\n      \"B淋巴细胞\",\n      \"浆细胞\",\n      \"记忆B细胞\",\n      \"B细胞活化\"\n    ],\n    \"巨噬细胞\": [\n      \"Mφ\",\n      \"单核巨噬细胞\",\n      \"组织细胞\",\n      \"吞噬细胞\"\n    ],\n    \"中性粒细胞\": [\n      \"PMN\",\n      \"嗜中性粒细胞\",\n      \"粒细胞\",\n      \"多形核白细胞\"\n    ],\n    \"糖皮质激素\": [\n      \"激素\",\n      \"皮质类固醇\",\n      \"GC\",\n      \"泼尼松\",\n      \"甲泼尼龙\"\n    ],\n    \"泼尼松\": [\n      \"强的松\",\n      \"泼尼松龙\",\n      \"醋酸泼尼松\",\n      \"Prednisone\"\n    ],\n    \"甲氨蝶呤\": [\n      \"MTX\",\n      \"氨甲蝶呤\",\n      \"甲氨蝶呤钠\"\n    ],\n    \"环磷酰胺\": [\n      \"CTX\",\n      \"环磷酰胺冲击\",\n      \"CYC\"\n    ],\n    \"硫唑嘌呤\": [\n      \"AZA\",\n      \"依木兰\",\n      \"咪唑硫嘌呤\"\n    ],\n    \"羟氯喹\": [\n      \"HCQ\",\n      \"硫酸羟氯喹\",\n      \"氯喹\",\n      \"Plaquenil\"\n    ],\n    \"环孢素\": [\n      \"CsA\",\n      \"环孢素A\",\n      \"新山地明\",\n      \"CsA\"\n    ],\n    \"吗替麦考酚酯\": [\n      \"MMF\",\n      \"霉酚酸酯\",\n      \"骁悉\"\n    ],\n    \"他克莫司\": [\n      \"FK506\",\n      \"普乐可复\",\n      \"Tac\"\n    ],\n    \"利妥昔单抗\": [\n      \"RTX\",\n      \"美罗华\",\n      \"抗CD20单抗\"\n    ],\n    \"托珠单抗\": [\n      \"TCZ\",\n      \"雅美罗\",\n      \"抗IL-6受体单抗\"\n    ],\n    \"阿达木单抗\": [\n      \"ADA\",\n      \"修美乐\",\n      \"抗TNF单抗\"\n    ],\n    \"英夫利昔单抗\": [\n      \"IFX\",\n      \"类克\",\n      \"抗TNF-α单抗\"\n    ],\n    \"依那西普\": [\n      \"ETN\",\n      \"恩利\",\n      \"TNF受体融合蛋白\"\n    ],\n    \"JAK抑制剂\": [\n      \"托法替布\",\n      \"巴瑞替尼\",\n      \"乌帕替尼\",\n      \"JAKi\"\n    ],\n    \"心脏\": [\n      \"心肌\",\n      \"心包\",\n      \"心血管\",\n      \"心脏受累\"\n    ],\n    \"肺\": [\n      \"肺部\",\n      \"肺间质\",\n      \"胸膜\",\n      \"肺动脉高压\"\n    ],\n    \"肝脏\": [\n      \"肝功能\",\n      \"肝损害\",\n      \"药物性肝损伤\",\n      \"转氨酶升高\"\n    ],\n    \"肾脏\": [\n      \"肾功能\",\n      \"肾小球\",\n      \"肾小管\",\n      \"肾活检\"\n    ],\n    \"胃肠道\": [\n      \"消化系统\",\n      \"消化道\",\n      \"肠道\",\n      \"胃\"\n    ],\n    \"皮肤\": [\n      \"皮肤黏膜\",\n      \"皮损\",\n      \"蝶形红斑\",\n      \"光敏感\"\n    ],\n    \"眼部\": [\n      \"眼\",\n      \"视力\",\n      \"葡萄膜炎\",\n      \"视网膜\"\n    ],\n    \"关节\": [\n      \"关节腔\",\n      \"滑膜\",\n      \"软骨\",\n      \"骨破坏\"\n    ],\n    \"肌肉\": [\n      \"骨骼肌\",\n      \"肌力\",\n      \"肌痛\",\n      \"肌无力\"\n    ],\n    \"中枢神经\": [\n      \"脑\",\n      \"脑血管\",\n      \"认知\",\n      \"癫痫\"\n    ],\n    \"起初\": [\n      \"最初\",\n      \"一开始\",\n      \"起先\",\n      \"早先\",\n      \"先\"\n    ],\n    \"随后\": [\n      \"后来\",\n      \"接着\",\n      \"之后\",\n      \"紧接着\",\n      \"随之而来\"\n    ],\n    \"最终\": [\n      \"最后\",\n      \"到头来\",\n      \"终归\",\n      \"结局是\",\n      \"终究\"\n    ],\n    \"相对\": [\n      \"相比之下\",\n      \"比较而言\",\n      \"从...角度看\",\n      \"就...来说\",\n      \"相对而言\"\n    ],\n    \"绝对\": [\n      \"完全\",\n      \"百分之百\",\n      \"确凿\",\n      \"无疑\",\n      \"毫无悬念\"\n    ],\n    \"总体\": [\n      \"整体\",\n      \"大面上\",\n      \"从全局看\",\n      \"总体而言\",\n      \"宏观上\"\n    ],\n    \"部分\": [\n      \"一些\",\n      \"某些\",\n      \"有的\",\n      \"其中\",\n      \"一部分\"\n    ],\n    \"多数\": [\n      \"大多数\",\n      \"大部分\",\n      \"绝大多数\",\n      \"多半\",\n      \"过半\"\n    ],\n    \"少数\": [\n      \"少数\",\n      \"极少数\",\n      \"个别\",\n      \"寥寥\",\n      \"为数不多\"\n    ],\n    \"全部\": [\n      \"所有\",\n      \"无一例外\",\n      \"全部\",\n      \"统统\",\n      \"整体\"\n    ],\n    \"几乎所有\": [\n      \"绝大多数\",\n      \"几乎\",\n      \"差不多\",\n      \"基本\",\n      \"压倒性\"\n    ],\n    \"确认\": [\n      \"确认\",\n      \"确认了\",\n      \"得到确认\",\n      \"明确\",\n      \"认定\"\n    ],\n    \"否定\": [\n      \"否认\",\n      \"排除\",\n      \"不成立\",\n      \"推翻\",\n      \"不予支持\"\n    ],\n    \"同意\": [\n      \"赞同\",\n      \"认可\",\n      \"点头\",\n      \"赞成\",\n      \"支持\"\n    ],\n    \"接受\": [\n      \"采纳\",\n      \"认同\",\n      \"接纳\",\n      \"同意\",\n      \"予以接受\"\n    ],\n    \"拒绝\": [\n      \"不接受\",\n      \"婉拒\",\n      \"否决\",\n      \"驳回\",\n      \"不予采纳\"\n    ],\n    \"承认\": [\n      \"认可\",\n      \"承认了\",\n      \"坦承\",\n      \"确认\",\n      \"供认\"\n    ],\n    \"否认\": [\n      \"不承认\",\n      \"否认\",\n      \"驳斥\",\n      \"澄清\",\n      \"否定了\"\n    ],\n    \"优势\": [\n      \"优点\",\n      \"长处\",\n      \"有利之处\",\n      \"亮点\",\n      \"强项\"\n    ],\n    \"不足\": [\n      \"缺点\",\n      \"短板\",\n      \"局限\",\n      \"缺陷\",\n      \"欠缺\"\n    ],\n    \"方式\": [\n      \"方法\",\n      \"途径\",\n      \"手段\",\n      \"模式\",\n      \"办法\"\n    ],\n    \"范围\": [\n      \"范畴\",\n      \"领域\",\n      \"区间\",\n      \"覆盖面\",\n      \"涉及面\"\n    ],\n    \"程度\": [\n      \"水平\",\n      \"级别\",\n      \"强度\",\n      \"范围\",\n      \"量级\"\n    ],\n    \"频率\": [\n      \"频次\",\n      \"发生率\",\n      \"次数\",\n      \"出现频率\",\n      \"频度\"\n    ],\n    \"比例\": [\n      \"占比\",\n      \"百分比\",\n      \"比重\",\n      \"份额\",\n      \"构成比\"\n    ],\n    \"速度\": [\n      \"速率\",\n      \"效率\",\n      \"进程\",\n      \"节奏\",\n      \"进展速度\"\n    ],\n    \"质量\": [\n      \"品质\",\n      \"水准\",\n      \"水平\",\n      \"品级\",\n      \"优劣\"\n    ],\n    \"近期\": [\n      \"最近\",\n      \"近期内\",\n      \"这一段时间\",\n      \"近来\",\n      \"最近一段时间\"\n    ],\n    \"远期\": [\n      \"长期\",\n      \"远期\",\n      \"长期来看\",\n      \"长期随访\",\n      \"后续\"\n    ],\n    \"暂时\": [\n      \"临时\",\n      \"暂时性\",\n      \"一时\",\n      \"短期\",\n      \"暂时地\"\n    ],\n    \"永久\": [\n      \"永久性\",\n      \"长期\",\n      \"终身\",\n      \"不可逆\",\n      \"长久\"\n    ],\n    \"间断\": [\n      \"间歇性\",\n      \"断断续续\",\n      \"时有时无\",\n      \"周期性\",\n      \"反复\"\n    ],\n    \"迅速\": [\n      \"快速\",\n      \"很快\",\n      \"短期内\",\n      \"急剧\",\n      \"急剧地\"\n    ],\n    \"缓慢\": [\n      \"逐渐\",\n      \"缓慢地\",\n      \"渐进性\",\n      \"逐步\",\n      \"慢慢\"\n    ],\n    \"目前\": [\n      \"当前\",\n      \"现在\",\n      \"现阶段\",\n      \"当今\",\n      \"至今\"\n    ],\n    \"此前\": [\n      \"之前\",\n      \"此前\",\n      \"早前\",\n      \"先前\",\n      \"在这之前\"\n    ],\n    \"至今\": [\n      \"到目前为止\",\n      \"迄今\",\n      \"至今\",\n      \"迄今为止\",\n      \"截至目前\"\n    ],\n    \"预计\": [\n      \"预估\",\n      \"预计\",\n      \"预期\",\n      \"估算\",\n      \"测算\"\n    ],\n    \"预期\": [\n      \"预料\",\n      \"预判\",\n      \"展望\",\n      \"预计\",\n      \"设想\"\n    ],\n    \"实际\": [\n      \"实际中\",\n      \"现实\",\n      \"事实上\",\n      \"实践中\",\n      \"真实世界\"\n    ],\n    \"理论上\": [\n      \"从理论上说\",\n      \"理论上讲\",\n      \"学理上\",\n      \"从机理看\",\n      \"机制上\"\n    ],\n    \"实践中\": [\n      \"实际工作中\",\n      \"临床实践中\",\n      \"真实世界中\",\n      \"现实世界\",\n      \"日常工作中\"\n    ],\n    \"常规\": [\n      \"常规的\",\n      \"标准的\",\n      \"通用的\",\n      \"惯用的\",\n      \"一般的\"\n    ],\n    \"新型\": [\n      \"新一代\",\n      \"新型\",\n      \"新兴的\",\n      \"新一代的\",\n      \"新近出现的\"\n    ],\n    \"传统\": [\n      \"经典的\",\n      \"传统的\",\n      \"老一代\",\n      \"既往常用的\",\n      \"传统意义上的\"\n    ],\n    \"替代\": [\n      \"备选\",\n      \"替代方案\",\n      \"其他选择\",\n      \"另一选择\",\n      \"可供选择\"\n    ],\n    \"首选\": [\n      \"一线\",\n      \"优先选择\",\n      \"推荐首选\",\n      \"优选举荐\",\n      \"第一选择\"\n    ],\n    \"大量\": [\n      \"大量\",\n      \"众多\",\n      \"大量\",\n      \"可观的\",\n      \"相当的\"\n    ],\n    \"少量\": [\n      \"少量\",\n      \"少许\",\n      \"微量\",\n      \"一点点\",\n      \"些许\"\n    ],\n    \"大量研究\": [\n      \"多项研究\",\n      \"众多文献\",\n      \"大量证据\",\n      \"多项报道\",\n      \"丰富的研究\"\n    ],\n    \"少量研究\": [\n      \"少数研究\",\n      \"少量文献\",\n      \"有限的证据\",\n      \"初步证据\",\n      \"为数不多的报道\"\n    ],\n    \"一致\": [\n      \"一致地\",\n      \"统一的\",\n      \"一致的\",\n      \"高度一致\",\n      \"共识\"\n    ],\n    \"不一致\": [\n      \"存在分歧\",\n      \"有争议\",\n      \"尚有争议\",\n      \"存在不同观点\",\n      \"结果不一\"\n    ],\n    \"明显高于\": [\n      \"远高于\",\n      \"显著高于\",\n      \"大幅高于\",\n      \"显著超过\",\n      \"远远高于\"\n    ],\n    \"明显低于\": [\n      \"远低于\",\n      \"显著低于\",\n      \"大幅低于\",\n      \"远不及\",\n      \"显著不及\"\n    ],\n    \"显著高于\": [\n      \"远超\",\n      \"大幅超过\",\n      \"明显高于\",\n      \"远高于\",\n      \"显著超过\"\n    ],\n    \"显著低于\": [\n      \"远低于\",\n      \"大幅低于\",\n      \"显著低于\",\n      \"远远低于\",\n      \"不及\"\n    ],\n    \"临床表现\": [\n      \"临床特点\",\n      \"症状和体征\",\n      \"临床征象\",\n      \"就诊表现\",\n      \"临床特征\"\n    ],\n    \"实验室发现\": [\n      \"检验结果\",\n      \"化验发现\",\n      \"实验室指标\",\n      \"血清学发现\",\n      \"检验数据\"\n    ],\n    \"影像学发现\": [\n      \"影像表现\",\n      \"放射学所见\",\n      \"CT所见\",\n      \"MRI表现\",\n      \"影像学征象\"\n    ],\n    \"病理结果\": [\n      \"病理所见\",\n      \"组织学发现\",\n      \"镜下所见\",\n      \"病理诊断\",\n      \"活检结果\"\n    ],\n    \"随访结果\": [\n      \"远期结果\",\n      \"随访数据\",\n      \"追踪结果\",\n      \"远期随访\",\n      \"随访所见\"\n    ],\n    \"治疗反应\": [\n      \"治疗应答\",\n      \"治疗响应\",\n      \"治疗效果\",\n      \"临床获益\",\n      \"治疗结果\"\n    ],\n    \"药物相关\": [\n      \"药物引起的\",\n      \"药源性\",\n      \"药物所致\",\n      \"药物诱发的\",\n      \"与药物有关\"\n    ],\n    \"疾病相关\": [\n      \"疾病引起的\",\n      \"病变所致\",\n      \"疾病相关的\",\n      \"原发病相关\",\n      \"与疾病有关\"\n    ],\n    \"从...角度\": [\n      \"从...维度\",\n      \"就...而言\",\n      \"站在...立场\",\n      \"从...视角看\",\n      \"以...为切入点\"\n    ],\n    \"以...为\": [\n      \"将...作为\",\n      \"选择...为\",\n      \"用...做\",\n      \"拿...当\",\n      \"设定...为\"\n    ],\n    \"相对于\": [\n      \"相较于\",\n      \"与...相比\",\n      \"对比\",\n      \"参照\",\n      \"较之于\"\n    ],\n    \"不同于\": [\n      \"区别于\",\n      \"与...不同\",\n      \"异于\",\n      \"有异于\",\n      \"不同于\"\n    ],\n    \"类似于\": [\n      \"类似于\",\n      \"类似于\",\n      \"近似于\",\n      \"接近于\",\n      \"近似\"\n    ],\n    \"取决于\": [\n      \"依赖于\",\n      \"视...而定\",\n      \"受制于\",\n      \"关键在于\",\n      \"决定于\"\n    ],\n    \"独立于\": [\n      \"不依赖于\",\n      \"不受...影响\",\n      \"与...无关\",\n      \"不取决于\",\n      \"脱离于\"\n    ],\n    \"伴随着\": [\n      \"同时出现\",\n      \"伴有\",\n      \"合并\",\n      \"伴随着\",\n      \"并发\"\n    ],\n    \"倾向于\": [\n      \"更偏向于\",\n      \"倾向于\",\n      \"偏好\",\n      \"更多地\",\n      \"主要\"\n    ],\n    \"旨在\": [\n      \"目的在于\",\n      \"意在\",\n      \"为了\",\n      \"目标是\",\n      \"致力于\"\n    ],\n    \"非常\": [\n      \"极为\",\n      \"十分\",\n      \"相当\",\n      \"特别\",\n      \"尤其\"\n    ],\n    \"极其\": [\n      \"极为\",\n      \"非常\",\n      \"极度\",\n      \"特别\",\n      \"无比\"\n    ],\n    \"相当\": [\n      \"颇为\",\n      \"相当\",\n      \"比较\",\n      \"较为\",\n      \"十分\"\n    ],\n    \"较为\": [\n      \"比较\",\n      \"相对\",\n      \"较为\",\n      \"一定程度\",\n      \"尚算\"\n    ],\n    \"略微\": [\n      \"稍有\",\n      \"略\",\n      \"轻微地\",\n      \"有些\",\n      \"有一点\"\n    ],\n    \"完全\": [\n      \"彻底\",\n      \"完全地\",\n      \"百分之百\",\n      \"全然\",\n      \"彻底地\"\n    ],\n    \"几乎\": [\n      \"差不多\",\n      \"将近\",\n      \"接近\",\n      \"几近\",\n      \"近乎\"\n    ],\n    \"基本\": [\n      \"基本上\",\n      \"大体上\",\n      \"大致\",\n      \"基本上来说\",\n      \"总体上\"\n    ],\n    \"推进\": [\n      \"推动\",\n      \"促进\",\n      \"加速\",\n      \"驱动\",\n      \"推动\"\n    ],\n    \"推动\": [\n      \"促进\",\n      \"助推\",\n      \"驱动\",\n      \"助力\",\n      \"推进\"\n    ],\n    \"获取\": [\n      \"得到\",\n      \"取得\",\n      \"获得\",\n      \"采集\",\n      \"拿到\"\n    ],\n    \"产生\": [\n      \"引起\",\n      \"导致\",\n      \"带来\",\n      \"引发\",\n      \"造成\"\n    ],\n    \"造成\": [\n      \"导致\",\n      \"引起\",\n      \"带来\",\n      \"酿成\",\n      \"产生\"\n    ],\n    \"导致\": [\n      \"引起\",\n      \"造成\",\n      \"带来\",\n      \"引发\",\n      \"促发\"\n    ],\n    \"引发\": [\n      \"引起\",\n      \"触发\",\n      \"导致\",\n      \"诱发\",\n      \"造成\"\n    ],\n    \"带来\": [\n      \"产生\",\n      \"造成\",\n      \"引起\",\n      \"带来\",\n      \"导致\"\n    ],\n    \"转变\": [\n      \"转变\",\n      \"转化\",\n      \"演变\",\n      \"变化\",\n      \"转型\"\n    ],\n    \"演变\": [\n      \"演化\",\n      \"发展\",\n      \"变化\",\n      \"演进\",\n      \"变迁\"\n    ],\n    \"呈现\": [\n      \"表现出\",\n      \"呈现出\",\n      \"显示为\",\n      \"展现为\",\n      \"体现为\"\n    ],\n    \"展现\": [\n      \"展示\",\n      \"呈现\",\n      \"表现\",\n      \"体现\",\n      \"显露\"\n    ],\n    \"揭示\": [\n      \"揭露\",\n      \"揭示出\",\n      \"暴露出\",\n      \"展示\",\n      \"呈现\"\n    ],\n    \"阐明\": [\n      \"阐明\",\n      \"说清楚\",\n      \"解释清楚\",\n      \"论述\",\n      \"阐释\"\n    ],\n    \"论述\": [\n      \"论述\",\n      \"阐述\",\n      \"讨论\",\n      \"分析\",\n      \"论证\"\n    ],\n    \"论证\": [\n      \"论证\",\n      \"论说\",\n      \"说明\",\n      \"论证了\",\n      \"论证过程\"\n    ],\n    \"阐述\": [\n      \"阐释\",\n      \"详述\",\n      \"论述\",\n      \"说明\",\n      \"阐明\"\n    ],\n    \"详述\": [\n      \"详细说明\",\n      \"详述\",\n      \"深入阐述\",\n      \"详细讨论\",\n      \"展开论述\"\n    ]\n  },\n  \"academic_connectors\": {\n    \"however\": [\n      \"但\",\n      \"然而\",\n      \"不过\",\n      \"但须指出\",\n      \"反观\"\n    ],\n    \"therefore\": [\n      \"因此\",\n      \"由此\",\n      \"基于此\",\n      \"由此可见\",\n      \"这说明\"\n    ],\n    \"in_addition\": [\n      \"另外\",\n      \"同时\",\n      \"除此之外\",\n      \"再者\",\n      \"补充而言\"\n    ],\n    \"for_example\": [\n      \"例如\",\n      \"以...为例\",\n      \"具体来看\",\n      \"举例来说\",\n      \"以...为证\"\n    ],\n    \"in_conclusion\": [\n      \"综上\",\n      \"由上可知\",\n      \"归纳起来\",\n      \"总结来看\",\n      \"概括而言\"\n    ],\n    \"furthermore\": [\n      \"进一步地\",\n      \"更值得注意的是\",\n      \"不仅如此\",\n      \"深入来看\",\n      \"再者\"\n    ],\n    \"specifically\": [\n      \"具体而言\",\n      \"确切地说\",\n      \"详细来看\",\n      \"就...而言\",\n      \"针对性地看\"\n    ],\n    \"overall\": [\n      \"总体来看\",\n      \"从全局看\",\n      \"整体而言\",\n      \"综合来看\",\n      \"统观全局\"\n    ],\n    \"in_contrast\": [\n      \"conversely\",\n      \"on the other hand\",\n      \"whereas\",\n      \"while\",\n      \"alternatively\"\n    ],\n    \"notably\": [\n      \"importantly\",\n      \"significantly\",\n      \"remarkably\",\n      \"in particular\",\n      \"especially\"\n    ],\n    \"to_our_knowledge\": [\n      \"to the best of our knowledge\",\n      \"as far as we know\",\n      \"we are not aware of\",\n      \"based on available literature\"\n    ],\n    \"for_the_first_time\": [\n      \"novel\",\n      \"previously unreported\",\n      \"initial\",\n      \"pioneering\",\n      \"first report of\"\n    ],\n    \"remains_unclear\": [\n      \"is not well understood\",\n      \"has yet to be elucidated\",\n      \"is still debated\",\n      \"remains controversial\"\n    ],\n    \"increasing_evidence\": [\n      \"growing evidence\",\n      \"mounting data\",\n      \"accumulating evidence\",\n      \"emerging data\"\n    ],\n    \"it_should_be_noted\": [\n      \"it is important to note\",\n      \"we emphasize that\",\n      \"of note\",\n      \"critically\"\n    ],\n    \"further_studies\": [\n      \"future research\",\n      \"additional investigation\",\n      \"further investigation\",\n      \"subsequent studies\"\n    ]\n  }\n}\n\nArchive v5.0.0: 98 files, 741262 bytes\n\nFiles: CHANGELOG.md (37029b), data/terminology.json (319057b), eval/attack_gen.py (4568b), eval/build_mixed_bench.py (4412b), eval/calibrate_mixed_para.py (5053b), eval/check_leak.py (4872b), eval/corpus_builder.py (6397b), eval/gap_eval.py (4832b), eval/measure_mixed_para.py (5057b), eval/release_smoke.py (39918b), eval/results/archive/accept_v380_repro.json (1762b), eval/results/archive/baseline_v2.json (1785b), eval/results/archive/layer_surprisal_v3100_surprisal.json (130b), eval/results/archive/phase1_debt.json (1779b), eval/results/archive/phase2_test.json (1850b), eval/results/archive/v3100_check.json (1756b), eval/results/archive/v3100_fresh.json (1756b), eval/results/archive/v350_release.json (1852b), eval/results/archive/v36_gen2026.STALE-cache-poisoned.json (1342b), eval/results/archive/v36_oldgen.STALE-cache-poisoned.json (1755b), eval/results/gen2026.json (1099b), eval/results/gen2026v2.json (1340b), eval/results/gen2026v2spec.json (1344b), eval/results/layer_surprisal_gen2026.json (125b), eval/results/leak_audit_20261003.json (1538b), eval/results/mixed_bench_20261005.json (342819b), eval/results/mixed_para_20261005.json (10879b), eval/results/score_cache.json (543942b), eval/results/v3_final_test.json (1864b), eval/results/v3_fp_test.json (1863b), eval/results/v3_local_test.json (1864b), eval/results/v35ctl_gen2026.json (1375b), eval/results/v35ctl_oldgen.json (1740b), eval/results/v36fix_oldgen.json (1754b), eval/results/v370_fresh.json (1755b), eval/results/v370_fresh2.json (1756b), eval/results/v370_release.json (1757b), eval/results/v370_release2.json (1758b), eval/results/v380_merge.json (1755b), eval/results/v390_final.json (1755b), eval/results/v410_l13fusion.json (1759b), eval/results/v480_rules_gen2026.json (1387b), eval/results/v480_rules_oldgen.json (1799b), eval/results/v500check.json (1736b), eval/results/v500rules.json (1791b), eval/run_eval.py (12579b), LICENSE.md (918b), README.md (3252b), references/ai_patterns_en.json (13858b), references/ai_patterns_zh.json (7537b), references/article-review-workflow.md (3200b), references/fusion_config.json (1083b), references/model_fingerprints.json (11185b), references/para_thresholds.json (1767b), references/sentence_patterns_zh.json (10269b), references/supervised_models.json (636b), references/synonyms_general.json (58842b), references/token_spectrum_v2.json (184509b), references/token_spectrum_zh.json (194373b), requirements.txt (771b), scripts/ai_detector.py (55750b), scripts/aigc_label_check.py (7205b), scripts/build_spectrum.py (3157b), scripts/calibrate_v3.py (5587b), scripts/deai_gate.py (8482b), scripts/fingerprint_miner.py (5681b), scripts/freshness_refresh.py (3792b), scripts/layers_lm.py (9648b), scripts/layers_surface.py (8671b), scripts/ngram_similarity.py (7846b), scripts/paragraph_report.py (5949b), scripts/pattern_recalibrator.py (6619b), scripts/perplexity.py (12178b), scripts/pp_api.py (9756b), scripts/pp_doctor.py (11133b), scripts/pp_fix_suggest.py (15987b), scripts/pp_setup.py (6026b), scripts/pp_split.py (681b), scripts/pp_verify.py (8720b), scripts/pp_workflow.py (13040b)\n\nFile v5.0.0:SKILL.md\n\n---\nname: paper-polisher\nversion: 5.0.0\nauthor: DoctorQ Lab\ndescription: >-\n  AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell),\n  metaphor audit, quality report, AIGC compliance label check (China 2025-09\n  labeling rules), paragraph-level attribution, journal precheck, sentence-level\n  rewrite suggestions (locates and advises, never auto-rewrites), plus `--batch DIR`\n  for thesis-scale batch rewriting (per-file AI-rate scores directory-wide).\n  Bilingual CN/EN, 100% local, zero upload, zero credentials; bundled unit-test suite\n  + AST-based zero-network self-verification.\n  v3 delivers a recalibrated multi-layer\n  rule engine (11 core layers + discourse/smoothness heuristics) + token-spectrum layer + length-routed fusion + optional\n  supervised Qwen3-0.6B ONNX layer (AUROC 1.0 on held-out test) + LLM\n  fingerprint attribution (GLM/DeepSeek/Qwen/Kimi/MiniMax/GPT/Claude/Gemini) +\n  freshness pipeline. Base-engine numbers reproduce from the bundled held-out\n  evaluation; supervised columns are author-side measurements (model not bundled).\ntags: [ai-detection, deai, academic-writing, paraphrase, paper-polish]\n---\n\n# Paper Polisher Pro v3\n\nAI writing detection (AI-rate self-check for authors) · academic polishing guidance · terminology standardization · translation-smell check · quality report · AIGC compliance label check · paragraph-level attribution · journal precheck.\n100% local, zero upload, zero credentials, pure standard library (optional onnxruntime enhancement layer).\n\n> ## ⛔ Iron laws\n> 1. **Only reproducible numbers.** Every metric comes from the held-out (test split) evaluation in `eval/run_eval.py`; unsupported claims like \"100% detection rate / F1 98.3%\" from older docs have been removed.\n> 2. **No verdict on short text.** Texts under 100 characters get `risk=unknown` (community lesson: short-text false positives are uncontrollable).\n> 3. **Fingerprints attribute, never score.** (Measured 2026-08-15: injecting fingerprints into the detector doubled human false positives.)\n> 4. **Calibration/evaluation separation.** Spectrum, weights and thresholds are built on the calib half only; the test half is reserved for final evaluation (an in-sample AUROC of 0.9972 collapsed to a real 0.9187 once split).\n\n## TL;DR\n\n- **What**: 100% local AI-rate self-check + academic polishing toolkit for Chinese academic text (optional supervised model for best accuracy; English gets advisory rules-only scores).\n- **30-second start**: `python3 scripts/pp.py quickstart` (zero-model, zero-file demo) · `python3 scripts/pp.py detect draft.txt --format json` · sentence-level rewrite suggestions: `python3 scripts/pp.py fix draft.txt` · full self-check report: `python3 scripts/pp.py workflow draft.txt` · environment: `python3 scripts/pp.py doctor` (one entry routes all subcommands)\n- **Measured** (held-out, fingerprint-bound md5 2631df3d388b): AUROC 0.9998 pre-2026 / 0.9400 current-generation; human FPR@medium 2.3%.\n- **Know the limits**: texts <100 chars get `risk=unknown` by design · medical text in degraded mode is over-scored · authors' self-check only — never for evading institutional AI detection.\n- **Where to look next**: capability boundary matrix below · end-to-end example in § Quick start · FAQ near the end · full history in `CHANGELOG.md`.\n\n## Academic integrity\n\nThis tool is for **authors self-reviewing and improving their own writing quality** — clearer sentences, consistent terminology, natural style. It is not designed to evade institutional AI-detection systems, and it must not be used to misrepresent AI-generated work as human-written. Follow your institution's AI-use and disclosure policies; the bundled `aigc_label_check.py` exists to help you **comply** with disclosure and labeling rules (e.g., China's 2025-09 labeling measures) — to declare AI assistance properly, not to hide it. Every AI-risk report (`ai_detector.py` / `deai_gate.py`) carries an explicit `integrity_notice` to this effect.\n\n## Measured performance (C-ReD + DetectRL-ZH, held-out test half, n=5,251)\n\n> Corpus scope note (v4.8.0): the v3.0-era table below was measured on the **full** held-out test half (n=5,251). The **bundled** sample corpus is a subset — its test half is n=1,304; reproducible per-corpus numbers: rules-only (PP_NO_SUP) 0.8985 old-gen / 0.7149 current-gen, fused 0.9998 / 0.9400 (`eval/results/v480_rules_*.json`, `v35ctl_*.json`; cache keys bind model fingerprint AND engine mode).\n\n| Metric | v2.0 baseline | v3.0 rules+spectrum | v3.1 +supervised | **v3.4 supervised + edit-regression v2** |\n|---|---|---|---|---|\n| AUROC (test half) | 0.7046 | 0.9187 | 0.9997 | **1.0** |\n| TPR@FPR5% | 30.4% | 49.0% | 99.95% | **100%** |\n| TPR@FPR1% | 16.7% | 24.9% | 99.88% | **100%** |\n| Human FPR @calibrated p99 | not measured | not measured | 3.56% (30/844) | **0.71% (6/844)** |\n| Paraphrase/mixed-attack AUROC | 0.64 | 0.89 | 1.0 (in-corpus) | **1.0** |\n| Attack \"AI-assisted\" recall | — | — | 71.1% | **86.6%** |\n| OOD plain-narrative/film recall | — | — | 1/6 | **5/6 supervised-only · 6/6 local fusion** |\n\n> **v4.4.0 fingerprint-bound re-measurement** of the shipping supervised model (md5 `2631df3d388b`): AUROC **0.9998** (test half, n_base=927), TPR@FPR1% 99.4%, human FPR@medium 2.3% — `eval/results/v35ctl_oldgen.json`. Columns above are preserved as version-era records (earlier model lineage; binaries were not fingerprinted before v4.4.0).\n\n**Which column applies to you?** The base package runs the **v3.0 rules+spectrum engine** (0.9187 AUROC column, measured on the full held-out corpus; pre-4.4.0 archived baselines predate fingerprint binding — every eval result since v4.4.0 carries the deployed model's md5 as `model_fp`, current bound numbers in `eval/results/v35ctl_*.json`). The two right-hand columns require the optional local supervised model (see below). The engine tells you honestly which mode you are in: every report carries `degraded_mode` / `degraded_notice` when the supervised layer is absent or skipped.\n\n### Capability boundary matrix (read before trusting any detector)\n\n| Scenario | Behavior |\n|---|---|\n| Chinese academic prose, full stack | Best case (AUROC 0.9998 held-out, human FPR@medium 2.3% — v4.4.0 fingerprint-bound) |\n| Base package without model | Rules+spectrum (0.9187); **medical register over-scored** (rules-only human FPR @medium: ~59% medical vs ~2% general) → trust only @high verdicts on medical text |\n| English text | Language gating skips the Chinese-trained supervised layer by design; rules-only English skeleton, advisory only |\n| Mixed human+AI documents | Document-level AUROC 0.52-0.54 (inherent averaging limitation); paragraph-level AUROC 0.69 with **calibrated best operating point P=0.60 at 63% coverage** (`references/para_thresholds.json`) — below the automatic-verdict bar; use `pp_workflow.py`/`paragraph_report.py` rankings for human review only |\n| Edit-extent regression head | ρ=0.540 — reported as metadata, never used in verdicts |\n| **Current-generation models (2026-09 sampling)** | **AUROC 0.9400** (v4.4.0 fingerprint-bound re-measurement, 443-doc current-gen eval set: 9 families incl. K3/K2.7/Qwen3.7-3.8/DS-V4/V4.1/GLM-5.3/M3) vs 0.9998 pre-2026 held-out — a modest verified gap. The earlier 0.6542-vs-0.9022 figure was a measurement artifact (stale score-cache replay + unverified model lineage); both classes are structurally prevented since v4.4.0 (`model_fp` in every result JSON)\n| Colloquial / oral-register text | The style layer is calibrated on academic prose; treat style scores as advisory outside that register |\n\n## Safety and behavior statement\n\n- **100% local**: every feature runs on-device. The codebase makes zero network calls — no network client libraries, no network utilities. Verify structurally: `python3 scripts/pp_verify.py` (AST-level scan; exit 0 = zero network calls, all scripts compile). Legacy text grep `grep -rEin \"urllib|requests|socket|import http\" scripts/` may show a few URL *strings* in report footers — those are data in string constants, not network code; the AST verifier distinguishes the two.\n- **No upload, no credentials**: reads and transmits no credentials, keys, or personal data; the only environment variable, `PP_NO_SUP`, is a local behavior toggle.\n- **No persistence**: creates no scheduled tasks, autostart entries, or system config changes; temp files (inter-layer JSON, probe text) are deleted after use.\n- **No remote code**: loads no remote models or scripts; the optional supervised model is placed by the user at a local path.\n- **Data boundary**: reads/writes only user-specified files, the system temp dir, and its own package data directories (calibration/freshness artifacts); reports go only where the user points them.\n- **Academic integrity**: see the section above — for author self-review and quality improvement with policy-compliant disclosure; not for evading detection.\n\n## What's new in v5.0.0\n\n- **Sentence-level rewrite suggestions (`scripts/pp_fix_suggest.py`, also `pp.py fix`)**: the natural next question after a score — *which sentences, why, and how to improve them*. Each flagged sentence lists its concrete features (AI clichés, filler phrases, template patterns, vague qualifiers, connective openers, dash/colon habits, uniform rhythm) with a per-type rewrite strategy. Guidance only: it locates and suggests, never auto-rewrites — the editing decision stays with the author. Ships with `--json` for programmatic use and a built-in two-sample demo (`--demo`).\n- **Bundled unit-test suite (`tests/`, `pp.py test`)**: 44 stdlib-unittest cases covering the iron laws (short text/empty/GBK), report field contracts, JSON purity, the gate, the workflow Markdown layout, data files, and the new tools — runnable in seconds without the model, so anyone can verify behavior on their own machine.\n- **Structured zero-network self-verification (`scripts/pp_verify.py`, also `pp.py verify`)**: replaces the old grep advice with an AST-level scan of every script — catches network imports/calls and curl/wget-style subprocess commands, while URL *strings* in report footers are correctly treated as data. Exit 0 = clean; `--json` for pipelines.\n- **Register awareness 2.0**: literary-narrative texts (dialogue quotes + time progression + inner monologue cues) now get a dedicated register notice explaining that this register sits outside the academic calibration domain — measured literary classics can reach high band in this engine — so the result is not mistaken for AI evidence. Disclosure only; no scoring change.\n- **`requirements.txt`** ships with the package: core = zero third-party dependencies; the two optional supervised-layer deps (onnxruntime/numpy) are declared and commented.\n- **`pp.py quickstart`**: zero-model, zero-file one-command demo (detect → fix → doctor) on a built-in sample.\n- **Freshness visible in `pp_doctor`**: the doctor now reports the latest held-out evaluation record alongside the fingerprint-registry coverage row.\n\n## What's new in v4.9.0\n\n- **Mixed-document calibration (closing the v4.7.0 backlog)**: paragraph-level hi/med/lo thresholds are now calibrated on the controlled mixed benchmark (333 paragraphs with ground truth; `eval/calibrate_mixed_para.py` → `references/para_thresholds.json`). Honest result: the best operating point (≥50) reaches precision 0.60 at 63% AI-paragraph coverage — below the automatic-verdict bar, so paragraph attribution remains a ranking aid for human review; the calibrated numbers and their scope ship in `references/para_thresholds.json` and surface in every `mixed_document` assessment.\n- **Unified entry (`scripts/pp.py`)**: one command routes all eleven subcommands (detect/gate/workflow/term/smell/style/quality/aigc/paragraph/setup/doctor) — no more script-navigation cost.\n- **Reliability**: batch CSV writes are now atomic (temp+rename — concurrent batch runs no longer clobber the same CSV); gate layers retry once on crash/timeout before falling back.\n\n## Anti-patterns (avoid these)\n\n- **Don't feed <100 chars** and expect a verdict — `risk=unknown` is by design (short-text false positives are uncontrollable); 300+ chars recommended.\n- **Don't trust degraded-mode scores on medical text** — rules-only over-scores medical register (~59% human FPR @medium); install the supervised model or trust only `@high`.\n- **Don't treat scores as CNKI/Wanfang equivalents** — thresholds are calibrated on our own held-out corpus; self-check only.\n- **Don't substitute or re-quantize the model file** — measured probability drift; only author-signed fingerprints pass `pp_setup.py`.\n- **Don't use it to evade institutional AI detection** — the integrity notice ships on every report; disclose per your institution's policy.\n- **Don't run batch on >5 MB files** — skipped by design; split first.\n\n## Architecture (v3)\n\n```\nai_detector.py            Main engine: 8 rule layers (125 recalibrated patterns, markdown caps,\n                          EN openers, paragraph-level language) + length-routed fusion\n + layers_surface.py      L9 surface stats L10 token-spectrum (9,955-token delta spectrum)\n                          L11 chain-of-thought features\n + ai_detector L12        discourse-structure heuristics (v3.7.0: hook/reversal/slogan/engagement)\n + fusion_config.json     Weights & thresholds (calib-half grid search + human p95/p99)\n + model_fingerprints.json v4 fingerprint registry (13 families incl. GLM-5.3 & Kimi K-series self-sampled; attribution only)\n + layers_lm.py           Optional supervised layer (local ONNX + pure-Python Qwen tokenizer;\n                          PP_NO_SUP=1 falls back to rules)\nparagraph_report.py       Paragraph-level attribution HTML (pattern×spectrum 50/50 fusion)\naigc_label_check.py       AIGC compliance labels (China labeling rules 2025-09: metadata/C2PA/explicit)\nfingerprint_miner.py      Fingerprint mining (new model drop → sample → mine → register)\npattern_recalibrator.py   Data-driven pattern recalibration (human-hit filtering)\nbuild_spectrum.py / calibrate_v3.py   Spectrum build / weight calibration\nfreshness_refresh.py         Monthly freshness pipeline (sample → rebuild → calibrate → regression)\npp_doctor.py              Environment self-check (v3.5; v5.0.0 adds latest-eval-record row)\npp_verify.py              AST-level structured zero-network self-verification (v5.0.0)\npp_fix_suggest.py         Sentence-level rewrite suggestions (v5.0.0: locate + strategy, no auto-rewrite)\ntests/                    Bundled unit-test suite, `python3 -m unittest discover -s tests -t .` (v5.0.0)\nrequirements.txt          Dependency declaration: core zero-dep; optional supervised-layer extras (v5.0.0)\neval/                     corpus_builder / attack_gen / run_eval (AUROC, TPR@FPR, per-model, attack decay)\n```\n\n## Quick start\n\n```bash\n# Zero-model, zero-file one-command demo (v5.0.0)\npython scripts/pp.py quickstart\n# AI writing detection (probability + layered evidence + fingerprint attribution)\npython scripts/ai_detector.py draft.txt --format json\n# Sentence-level rewrite suggestions (v5.0.0: which sentences, why, how to improve)\npython scripts/pp_fix_suggest.py draft.txt --top 10\n# Journal precheck (suspected-AIGC ratio vs the 20-25% reference line, non-interchangeable disclaimer)\npython scripts/ai_detector.py draft.txt --profile journal\n# Paragraph-level attribution (locate human/AI collaboration)\npython scripts/paragraph_report.py draft.txt --output report.html\n# AIGC compliance label check (docx/pdf/png/txt)\npython scripts/aigc_label_check.py manuscript.docx figures/*.png\n# Terminology / translation smell / 4-layer gate (same as v2)\npython scripts/term_check.py draft.txt --auto-fix\npython scripts/translation_smell_check.py draft.txt\npython scripts/deai_gate.py draft.txt\n# Environment self-check\npython scripts/pp_doctor.py\n# Structured zero-network self-verification + bundled unit tests (v5.0.0)\npython scripts/pp_verify.py\npython -m unittest discover -s tests -t .          # or: python scripts/pp.py test\n# Held-out regression (mandatory after any engine change)\npython eval/run_eval.py --split test --tag mytag\n```\n\n### Optional supervised layer (recommended, v3.2+)\n\n```bash\npip install onnxruntime regex          # the two optional dependencies\n# Place the two model files exactly as shipped by the authors:\n#   ~/.cache/paper-polisher/qwen3-detector/model.int8.onnx\n#   ~/.cache/paper-polisher/qwen3-detector/tokenizer.json\npython scripts/layers_lm.py            # self-test: supervised_available: true\n# ai_detector.py fuses automatically aft\n\nArchive v4.9.0: 84 files, 712810 bytes\n\nFiles: CHANGELOG.md (31805b), data/terminology.json (319057b), eval/attack_gen.py (4568b), eval/build_mixed_bench.py (4412b), eval/calibrate_mixed_para.py (5053b), eval/check_leak.py (4872b), eval/corpus_builder.py (6397b), eval/gap_eval.py (4832b), eval/measure_mixed_para.py (5057b), eval/release_smoke.py (32153b), eval/results/archive/accept_v380_repro.json (1762b), eval/results/archive/baseline_v2.json (1785b), eval/results/archive/layer_surprisal_v3100_surprisal.json (130b), eval/results/archive/phase1_debt.json (1779b), eval/results/archive/phase2_test.json (1850b), eval/results/archive/v3100_check.json (1756b), eval/results/archive/v3100_fresh.json (1756b), eval/results/archive/v350_release.json (1852b), eval/results/archive/v36_gen2026.STALE-cache-poisoned.json (1342b), eval/results/archive/v36_oldgen.STALE-cache-poisoned.json (1755b), eval/results/gen2026.json (1099b), eval/results/gen2026v2.json (1340b), eval/results/gen2026v2spec.json (1344b), eval/results/layer_surprisal_gen2026.json (125b), eval/results/leak_audit_20261003.json (1538b), eval/results/mixed_bench_20261005.json (342819b), eval/results/mixed_para_20261005.json (10879b), eval/results/score_cache.json (543942b), eval/results/v3_final_test.json (1864b), eval/results/v3_fp_test.json (1863b), eval/results/v3_local_test.json (1864b), eval/results/v35ctl_gen2026.json (1375b), eval/results/v35ctl_oldgen.json (1740b), eval/results/v36fix_oldgen.json (1754b), eval/results/v370_fresh.json (1755b), eval/results/v370_fresh2.json (1756b), eval/results/v370_release.json (1757b), eval/results/v370_release2.json (1758b), eval/results/v380_merge.json (1755b), eval/results/v390_final.json (1755b), eval/results/v410_l13fusion.json (1759b), eval/results/v480_rules_gen2026.json (1387b), eval/results/v480_rules_oldgen.json (1799b), eval/run_eval.py (12579b), LICENSE.md (918b), README.md (2809b), references/ai_patterns_en.json (13858b), references/ai_patterns_zh.json (7537b), references/article-review-workflow.md (3200b), references/fusion_config.json (1083b), references/model_fingerprints.json (11185b), references/para_thresholds.json (1767b), references/sentence_patterns_zh.json (10269b), references/supervised_models.json (636b), references/syno...","readmeExcerpt":"Skill: Paper Polisher Pro — AI Detector & Academic Polishing Owner: docsor1212 Summary: AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell), metaphor audit, quality report, AIGC compliance label check (China 2025-09 labeling rules), paragraph-level attribution, journal precheck, sentence-level rewrite suggestions (locates and advises, never auto-rewrites), plus --batch DIR","codeSnippets":[],"executableExamples":[{"language":"text","snippet":"ai_detector.py            Main engine: 8 rule layers (125 recalibrated patterns, markdown caps,\n                          EN openers, paragraph-level language) + length-routed fusion\n + layers_surface.py      L9 surface stats L10 token-spectrum (9,955-token delta spectrum)\n                          L11 chain-of-thought features\n + ai_detector L12        discourse-structure heuristics (v3.7.0: hook/reversal/slogan/engagement)\n + fusion_config.json     Weights & thresholds (calib-half grid search + human p95/p99)\n + model_fingerprints.json v4 fingerprint registry (13 families incl. GLM-5.3 & Kimi K-series self-sampled; attribution only)\n + layers_lm.py           Optional supervised layer (local ONNX + pure-Python Qwen tokenizer;\n                          PP_NO_SUP=1 falls back to rules)\nparagraph_report.py       Paragraph-level attribution HTML (pattern×spectrum 50/50 fusion)\naigc_label_check.py       AIGC compliance labels (China labeling rules 2025-09: metadata/C2PA/explicit)\nfingerprint_miner.py      Fingerprint mining (new model drop → sample → mine → register)\npattern_recalibrator.py   Data-driven pattern recalibration (human-hit filtering)\nbuild_spectrum.py / calibrate_v3.py   Spectrum build / weight calibration\nfreshness_refresh.py         Monthly freshness pipeline (sample → rebuild → calibrate → regression)\npp_doctor.py              Environment self-check (v3.5; v5.0.0 adds latest-eval-record row)\npp_verify.py              AST-level structured zero-network self-verification (v5.0.0)\npp_fix_suggest.py         Sentence-level rewrite suggestions (v5.0.0: locate + strategy, no auto-rewrite)\ntests/                    Bundled unit-test suite, `python3 -m unittest discover -s tests -t .` (v5.0.0)\nrequirements.txt          Dependency declaration: core zero-dep; optional supervised-layer extras (v5.0.0)\neval/                     corpus_builder / attack_gen / run_eval (AUROC, TPR@FPR, per-model, attack decay)"},{"language":"bash","snippet":"# Zero-model, zero-file one-command demo (v5.0.0)\npython scripts/pp.py quickstart\n# AI writing detection (probability + layered evidence + fingerprint attribution)\npython scripts/ai_detector.py draft.txt --format json\n# Sentence-level rewrite suggestions (v5.0.0: which sentences, why, how to improve)\npython scripts/pp_fix_suggest.py draft.txt --top 10\n# Rewrite-effect regression check (v5.1.0: original vs revised, engine-source comparison)\npython scripts/pp_rewrite_check.py draft_original.txt draft_revised.txt\n# Batch CSV -> self-contained HTML summary (v5.1.0)\npython scripts/pp_batch_report.py scores.csv -o report.html\n# Journal precheck (suspected-AIGC ratio vs the 20-25% reference line, non-interchangeable disclaimer)\npython scripts/ai_detector.py draft.txt --profile journal\n# Paragraph-level attribution (locate human/AI collaboration)\npython scripts/paragraph_report.py draft.txt --output report.html\n# AIGC compliance label check (docx/pdf/png/txt)\npython scripts/aigc_label_check.py manuscript.docx figures/*.png\n# Terminology / translation smell / 4-layer gate (same as v2)\npython scripts/term_check.py draft.txt --auto-fix\npython scripts/translation_smell_check.py draft.txt\npython scripts/deai_gate.py draft.txt\n# Environment self-check\npython scripts/pp_doctor.py\n# Structured zero-network self-verification + bundled unit tests (v5.0.0)\npython scripts/pp_verify.py\npython -m unittest discover -s tests -t .          # or: python scripts/pp.py test\n# Held-out regression (mandatory after any engine change)\npython eval/run_eval.py --split test --tag mytag"},{"language":"bash","snippet":"pip install onnxruntime regex          # the two optional dependencies\n# Place the two model files exactly as shipped by the authors:\n#   ~/.cache/paper-polisher/qwen3-detector/model.int8.onnx\n#   ~/.cache/paper-polisher/qwen3-detector/tokenizer.json\npython scripts/layers_lm.py            # self-test: supervised_available: true\n# ai_detector.py fuses automatically afterwards (0.9*supervised + 0.1*rules);\n# PP_NO_SUP=1 temporarily falls back to rules-only.\n# ⚠️ Do not substitute other exports or quantizations — measured probability drift; use exactly these files."},{"language":"python","snippet":"import sys; sys.path.insert(0, \"<skill>/scripts\")\nfrom pp_api import detect_text, gate_text, doctor_summary\nr = detect_text(\"中文学术文本，建议 300 字以上。\" * 10, lang=\"zh\")\nprint(r[\"overall_ai_score\"], r[\"overall_risk\"], r[\"degraded_mode\"])"},{"language":"bash","snippet":"python3 scripts/pp_setup.py --model <author-signed model.onnx>   # verify md5 -> install -> inference canary\npython3 scripts/pp_setup.py --check                              # current installation status"},{"language":"bash","snippet":"python3 scripts/pp_workflow.py draft.txt      # writes draft.workflow.md + draft.workflow.json"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: paper-polisher\nversion: 5.1.0\nauthor: DoctorQ Lab\ndescription: >-\n  AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell),\n  metaphor audit, quality report, AIGC compliance label check (China 2025-09\n  labeling rules), paragraph-level attribution, journal precheck, sentence-level\n  rewrite suggestions (locates and advises, never auto-rewrites), plus `--batch DIR`\n  for thesis-scale batch rewriting (per-file AI-rate scores directory-wide).\n  Bilingual CN/EN, 100% local, zero upload, zero credentials; bundled unit-test suite\n  + AST-based zero-network self-verification.\n  v3 delivers a recalibrated multi-layer\n  rule engine (11 core layers + discourse/smoothness heuristics) + token-spectrum layer + length-routed fusion + optional\n  supervised Qwen3-0.6B ONNX layer (AUROC 1.0 on held-out test) + LLM\n  fingerprint attribution (GLM/DeepSeek/Qwen/Kimi/MiniMax/GPT/Claude/Gemini) +\n  freshness pipeline. Base-engine numbers reproduce from the bundled held-out\n  evaluation; supervised columns are author-side measurements (model not bundled).\ntags: [ai-detection, deai, academic-writing, paraphrase, paper-polish]\n---\n\n# Paper Polisher Pro v3\n\nAI writing detection (AI-rate self-check for authors) · academic polishing guidance · terminology standardization · translation-smell check · quality report · AIGC compliance label check · paragraph-level attribution · journal precheck.\n100% local, zero upload, zero credentials, pure standard library (optional onnxruntime enhancement layer).\n\n> ## ⛔ Iron laws\n> 1. **Only reproducible numbers.** Every metric comes from the held-out (test split) evaluation in `eval/run_eval.py`; unsupported claims like \"100% detection rate / F1 98.3%\" from older docs have been removed.\n> 2. **No verdict on short text.** Texts under 100 characters get `risk=unknown` (community lesson: short-text false positives are uncontrollable).\n> 3. **Fingerprints attribute, never score.** (Measured 2026-08-15: injecting fingerprints into the detector doubled human false positives.)\n> 4. **Calibration/evaluation separation.** Spectrum, weights and thresholds are built on the calib half only; the test half is reserved for final evaluation (an in-sample AUROC of 0.9972 collapsed to a real 0.9187 once split).\n\n## TL;DR\n\n- **What**: 100% local AI-rate self-check + academic polishing toolkit for Chinese academic text (optional supervised model for best accuracy; English gets advisory rules-only scores).\n- **30-second start**: `python3 scripts/pp.py quickstart` (zero-model, zero-file demo) · `python3 scripts/pp.py detect draft.txt --format json` · sentence-level rewrite suggestions: `python3 scripts/pp.py fix draft.txt` · full self-check report: `python3 scripts/pp.py workflow draft.txt` · environment: `python3 scripts/pp.py doctor` (one entry routes all subcommands)\n- **Measured** (held-out, fingerprint-bound md5 2631df3d388b): AUROC 0.9998 pre-2026 / 0.9400 current-generation; human FPR@medium 2.3%.\n- **"},{"path":"README.md","content":"# Paper Polisher Pro — 论文降AI润色工具 · AI率检测\n\n[![GitHub Stars](https://img.shields.io/github/stars/docsor1212/paper-polisher-pro?style=social&label=Star)](https://github.com/docsor1212/paper-polisher-pro)\n\nAI 痕迹检测（AI率）· 去AI化改写建议 · 句子级改写建议（哪几句像AI、怎么改）· 术语标准化 · 翻译腔检查 · 质量报告 · AIGC 合规标识检查 · 段落级归因 · 期刊口径预检。\n\n**100% 本地运行，零上传，零凭证**——论文数据不出本机。零网络承诺可用包内 `pp_verify.py`（AST 结构化扫描）自行验证，行为契约可用随包 `tests/` 单元测试套件（44 用例）在自己机器上复跑。\n\n## 这是什么\n\n面向学术写作者的 AI 痕迹自查工具：概率化输出（非二元判定）、分层证据、指纹归因（GLM / DeepSeek / Qwen / Kimi / MiniMax / GPT / Claude / Gemini），全部指标可由随包留出集评测复现。\n\n- 基础引擎（纯规则+词频谱）：留出集 AUROC 0.9187\n- 可选监督层（本地 Qwen3-0.6B ONNX）：AUROC 1.0（作者侧实测）\n- 短文本不出判定（<100 字，误报铁律）\n\n## 快速开始\n\n```bash\n# 零模型零文件一键体验\npython scripts/pp.py quickstart\n\n# AI 痕迹检测（AI率）\npython scripts/ai_detector.py draft.txt --format json\n\n# 句子级改写建议（定位+策略，不代改）\npython scripts/pp_fix_suggest.py draft.txt --top 10\n\n# 四层融合门禁\npython scripts/deai_gate.py draft.txt\n\n# 批量检测一个目录\npython scripts/ai_detector.py --batch ./drafts --csv scores.csv\n\n# 环境自检\npython scripts/pp_doctor.py\n\n# Python 编程接口（零网络，import 即用）\npython -c \"import sys; sys.path.insert(0,'scripts'); from pp_api import detect_text; \\\nprint(detect_text(open('draft.txt').read())['overall_ai_score'])\"\n\n# 监督层一键装模（作者签发模型文件，指纹校验+推理自检）\npython scripts/pp_setup.py --model <作者签发模型.onnx>\n```\n\n## 安装\n\n克隆本仓库后直接使用，纯 Python 标准库即可运行（可选 onnxruntime 增强监督层，依赖声明见 `requirements.txt`）。\n\n```bash\ngit clone https://github.com/docsor1212/paper-polisher-pro\ncd paper-polisher-pro\npython scripts/ai_detector.py your_draft.txt --format summary\n```\n\n**China mirror (ModelScope 魔搭)**: <https://modelscope.cn/skills/Docsor/paper-polisher-pro> — if you find this skill useful, a like there helps others find it.\n\n## 论文工作流家族\n\n写作是一条链，每环有专用工具（均在本账号下）：\n\n| 工具 | 用途 |\n|---|---|\n| **paper-polisher-pro**（本仓库） | AI率检测·润色·降重·质量报告 |\n| **paper-rewriter** | 论文降AI改写执行 |\n| **pubmed-verifier** | PMID/DOI 引用核验 |\n| **cite-holmes** | 深度调研 × 引用自证 |\n| **cn-med-oa** | 中文医学文献 OA 下载 |\n| **academic-figures** | 出版级科研图表 |\n| **doc-holmes** | PDF 精准翻译 |\n\n文档站：[docsor.cn](https://docsor.cn)\n\n## 合规声明\n\n本工具供作者自查与写作质量改进，**不用于规避机构的 AIGC 检测**；请遵循所在机构的 AI 使用与披露政策（标识合规可用包内 aigc_label_check 自查）。许可：MIT-0。"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b6sjwend7cwwmhf7mdg1wzx82dfbf\",\n  \"slug\": \"paper-polisher-pro\",\n  \"version\": \"5.1.0\",\n  \"publishedAt\": 1791565192833\n}"},{"path":"references/ai_patterns_en.json","content":"{\n  \"version\": \"3.1.0\",\n  \"description\": \"English AI writing pattern library\",\n  \"categories\": {\n    \"filler_phrases\": {\n      \"weight\": 3,\n      \"description\": \"High-frequency AI filler phrases (strong signal)\",\n      \"patterns\": [\n        \"it is worth noting that\",\n        \"it is important to note that\",\n        \"it should be noted that\",\n        \"it is worth mentioning that\",\n        \"it is crucial to understand\",\n        \"it is essential to recognize\",\n        \"it is imperative to\",\n        \"it is paramount to\",\n        \"in conclusion\",\n        \"to summarize\",\n        \"in summary\",\n        \"to sum up\",\n        \"all in all\",\n        \"at the end of the day\",\n        \"when all is said and done\",\n        \"needless to say\",\n        \"it goes without saying\",\n        \"as a matter of fact\",\n        \"in today's world\",\n        \"in this day and age\",\n        \"in the modern world\",\n        \"in recent years\",\n        \"with the development of\",\n        \"with the advancement of\",\n        \"in the era of\",\n        \"in the age of\",\n        \"plays a crucial role\",\n        \"plays a vital role\",\n        \"plays an important role\",\n        \"plays a significant role\",\n        \"is of paramount importance\",\n        \"is of great significance\",\n        \"has gained significant attention\",\n        \"has attracted considerable attention\",\n        \"has been widely studied\",\n        \"has been extensively investigated\",\n        \"a growing body of evidence\",\n        \"an increasing number of\",\n        \"a wide range of\",\n        \"a variety of\",\n        \"a plethora of\",\n        \"a myriad of\",\n        \"delve into\",\n        \"shed light on\",\n        \"pave the way for\",\n        \"open new avenues\",\n        \"bridge the gap\",\n        \"furthermore\",\n        \"moreover\",\n        \"additionally\",\n        \"consequently\",\n        \"nevertheless\",\n        \"nonetheless\",\n        \"subsequently\",\n        \"notably\",\n        \"specifically\",\n        \"fundamentally\",\n        \"intriguingly\",\n        \"notably\",\n        \"remarkably\",\n        \"underscores the importance\",\n        \"highlights the significance\",\n        \"serves as a testament\",\n        \"poised to\",\n        \"tailored to\",\n        \"groundbreaking\",\n        \"pivotal\",\n        \"indispensable\",\n        \"multifaceted\",\n        \"nuanced\",\n        \"comprehensive\",\n        \"robust\",\n        \"state-of-the-art\",\n        \"cutting-edge\"\n      ]\n    },\n    \"conclusion_patterns\": {\n      \"weight\": 4,\n      \"description\": \"AI-typical conclusion patterns\",\n      \"patterns\": [\n        \"in conclusion,?.{0,30}(important|significant|crucial)\",\n        \"this study (demonstrates|shows|reveals|suggests) that\",\n        \"these findings (suggest|indicate|demonstrate|highlight)\",\n        \"future research (should|could|may|might) (focus|explore|investigate|examine)\",\n        \"the results of this study (provide|offer|contribute)\",\n        \"this research (contributes|adds|provides) (to|a|new)\",\n        \"taken together,?.{0,20}(suggest|indicate|demonstrate)\",\n        \"over"},{"path":"references/ai_patterns_zh.json","content":"{\n  \"version\": \"4.0.0\",\n  \"description\": \"v3数据驱动重校准版 (基线语料实测FPR过滤)\",\n  \"recalibration\": {\n    \"date\": \"2026-08-15\",\n    \"corpus\": \"eval/corpora_small/eval_zh.jsonl\",\n    \"humans\": 245,\n    \"ais\": 1645,\n    \"dropped\": 568,\n    \"demoted\": 5\n  },\n  \"categories\": {\n    \"weak_signals\": {\n      \"weight\": 1,\n      \"description\": \"v3重校准: 人类语料命中3%+且lift<2的弱信号, 仅组合计分(单独命中不足以判AI)\",\n      \"patterns\": [\n        \"然后\",\n        \"因此\",\n        \"再次\",\n        \"似乎\",\n        \"先.*再.*\"\n      ]\n    },\n    \"filler_phrases\": {\n      \"weight\": 3,\n      \"description\": \"各模型通用AI高频填充短语\",\n      \"patterns\": [\n        \"值得注意的是\",\n        \"综上所述\",\n        \"总而言之\",\n        \"具体而言\",\n        \"深入探讨\",\n        \"深入理解\",\n        \"系统梳理\",\n        \"深入分析\",\n        \"旨在探讨\",\n        \"具有重要意义\",\n        \"提供了新的视角\",\n        \"奠定了基础\",\n        \"本文旨在\",\n        \"本研究旨在\",\n        \"本研究通过\",\n        \"本研究为\",\n        \"与此同时\",\n        \"在此基础上\",\n        \"尤其是\",\n        \"特别是\",\n        \"不容忽视\",\n        \"至关重要\",\n        \"关键作用\",\n        \"重要意义\",\n        \"不可忽视\"\n      ]\n    },\n    \"ai_cliches\": {\n      \"weight\": 4,\n      \"description\": \"AI标志性套话和陈词滥调（最强信号）\",\n      \"patterns\": [\n        \"令人印象深刻\",\n        \"令人满意\",\n        \"新思路\",\n        \"日益受到\",\n        \"不断完善\",\n        \"展现出.*潜力\",\n        \"展现出.*价值\",\n        \"为.*提供了新\",\n        \"值得关注\",\n        \"引起了广泛关注\",\n        \"为.{2,15}(铺平|奠定|打开)了.{2,10}(道路|基础|大门|新方向)\"\n      ]\n    },\n    \"conclusion_patterns\": {\n      \"weight\": 4,\n      \"description\": \"AI标志性结尾模式\",\n      \"patterns\": [\n        \"本研究(为|对).{0,30}(提供了|具有一定的)\",\n        \"研究结果.{0,20}(表明|显示|证实)\"\n      ]\n    },\n    \"transition_words\": {\n      \"weight\": 2,\n      \"description\": \"AI过度使用的过渡词\",\n      \"patterns\": [\n        \"此外\",\n        \"同时\",\n        \"然而\",\n        \"与此同时\",\n        \"另一方面\",\n        \"进一步地\",\n        \"总的来说\",\n        \"首先\",\n        \"其次\"\n      ]\n    },\n    \"sentence_structure\": {\n      \"weight\": 2,\n      \"description\": \"AI典型句式特征\",\n      \"patterns\": [\n        \"不仅.{2,20}而且.{2,30}\",\n        \"不仅.{2,20}还.{2,30}\",\n        \"既.{2,15}又.{2,15}\",\n        \"一方面.{5,40}另一方面.{5,40}\",\n        \"通过.{2,30}实现了.{2,30}\",\n        \"通过.{2,30}为.{2,30}(提供了|奠定了)\",\n        \"随着.{2,20}的(发展|深入|推进|不断|广泛)\",\n        \"对于.*来说\",\n        \"对.*而言\",\n        \"作为.*的.*之一\",\n        \"在.*中(扮演|发挥|起到)\"\n      ]\n    },\n    \"vague_expressions\": {\n      \"weight\": 2,\n      \"description\": \"AI倾向使用的模糊表达\",\n      \"patterns\": [\n        \"在一定程度上\",\n        \"整体而言\",\n        \"可以说\",\n        \"总体而言\",\n        \"总体来说\"\n      ]\n    },\n    \"redundant_expressions\": {\n      \"weight\": 1,\n      \"description\": \"AI冗余表达\",\n      \"patterns\": [\n        \"进行(了)?(深入|详细|全面|系统)的?(分析|研究|探讨|阐述)\",\n        \"对.{2,15}进行了.{2,15}(分析|研究|探讨)\",\n        \"进行了.*分析\",\n        \"进行了.*研究\",\n        \"进行了.*探讨\"\n      ]\n    },\n    \"deepseek_fingerprint\": {\n      \"weight\": 4,\n      \"description\": \"DeepSeek模型指纹：教科书体+数据驱动+编号结构（最强信号）\",\n      \"patterns\": [\n        \"大规模.*研究\"\n      ]\n    },\n    \"glm_fingerprint\": {\n      \"weight\": 3,\n      \"description\": \"GLM/智谱模型指纹：四平八稳综述体\",\n      \"patterns\": [\n        \"关注的焦点\",\n        \"确保.*安全\",\n        \"以确保.*安全\"\n      ]\n    },\n    \""}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell), metaphor audit, quality report, AIGC compliance label check (China 2025-09 labeling rules), paragraph-level attribution, journal precheck, sentence-level rewrite suggestions (locates and advises, never auto-rewrites), plus `--batch DIR` for thesis-scale batch rewriting (per-file AI-rate scores directory-wide). Bilingual CN/EN, 100% local, zero upload, zero credentials; bundled unit-test suite + AST-based zero-network self-verification. v3 delivers a recalibrated multi-layer rule engine (11 core layers + discourse/smoothness heuristics) + token-spectrum layer + length-routed fusion + optional supervised Qwen3-0.6B ONNX layer (AUROC 1.0 on held-out test) + LLM fingerprint attribution (GLM/DeepSeek/Qwen/Kimi/MiniMax/GPT/Claude/Gemini) + freshness pipeline. Base-engine numbers reproduce from the bundled held-out evaluation; supervised columns are author-side measurements (model not bundled). Skill: Paper Polisher Pro — AI Detector & Academic Polishing Owner: docsor1212 Summary: AI-rate self-check for academic writing, polish guidance (style, terminology, translation-smell), metaphor audit, quality report, AIGC compliance label check (China 2025-09 labeling rules), paragraph-level attribution, journal precheck, sentence-level rewrite suggestions (locates and advises, never auto-rewrites), plus --batch DIR","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":2144,"uniquenessScore":46,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T01:31:08.446Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T01:31:08.446Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T07:42:05.088Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}