{"id":"76d09ef0-9f17-4fe5-9347-3a3adc700ffe","entityType":"agent","slug":"clawhub-medstatstar-statdata-transfer","name":"statdata-transfer","canonicalUrl":"https://www.xpersona.co/agent/clawhub-medstatstar-statdata-transfer","canonicalPath":"/agent/clawhub-medstatstar-statdata-transfer","generatedAt":"2026-10-10T07:45:03.001Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T01:52:55.506Z","emptyReason":null},"description":"读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明（完整）：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；写入主输出文件的同时可能生成 sidecar 元数据（CSV/TSV 旁 <名>_metadata.json、Parquet/Arrow 内嵌）及覆盖 .hyper 时的 .bak/.bak.1 备份；处理 .rda/.rds/.RData/.mtw/.mpj/.rec 时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。 / Read/convert 50+ statistical software formats, preserving variable/value labels and missing-value metadata for binary stats formats. FULL side effects: runs environment checks (scripts/check_env.py); may optionally pip-install missing packages on request; writes the main output file AND may emit sidecar metadata (e.g. <name>_metadata.json beside CSV/TSV, embedded in Parquet/Arrow schema) and .bak/.bak.1 backups when overwriting .hyper; can invoke the local R interpreter for .rda/.rds/.RData/.mtw/.mpj/.rec files via a fallback DISABLED by default and opted in only with allow_r_exec=True.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.8K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s176fv8983h1rte6dmxwp9wt4n89j8p5:statdata-transfer","sourceUrl":"https://clawhub.ai/medstatstar/statdata-transfer","homepage":"https://clawhub.ai/medstatstar/skills/statdata-transfer","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/medstatstar/statdata-transfer","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/medstatstar/skills/statdata-transfer","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":65,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"statdata-transfer technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T01:52:55.506Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T01:52:55.506Z","emptyReason":null},"stars":null,"forks":null,"downloads":1793,"packageName":null,"latestVersion":"2.2.1","tractionLabel":"1.8K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T01:52:55.506Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T01:52:55.506Z","lastCrawledAt":"2026-10-10T01:52:55.506Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T01:52:55.506Z","lastVerifiedAt":null,"highlights":[{"version":"2.2.1","createdAt":"2026-08-02T11:02:12.225Z","changelog":"Security hardening (2.2.1): pin lxml>=6.1.0 to fix CVE-2026-41066 (XXE); add secure lxml parser with entity resolution disabled for .wdx/.opju; fully disclose sidecar metadata & backup file writes in manifest.","fileCount":32,"zipByteSize":118053},{"version":"2.2.0","createdAt":"2026-08-02T07:42:56.100Z","changelog":"优化用户菜单(UI)与 README，碳基生物用户使用时更方便；并修复 CSV 分隔符探测、.tsv 读取、UTF-16 编码、XPT 读取报错、长变量标签静默截断告警等健壮性 bug。","fileCount":32,"zipByteSize":116844},{"version":"2.1.0","createdAt":"2026-07-19T02:32:47.554Z","changelog":"Standardized bilingual (en/zh) formatting per 2026-07-17 policy: non-table separators use '/', Core Capabilities table merged, README_ZH aligned; plus security hardening (allow_r_exec gate, dependency pins).","fileCount":31,"zipByteSize":113352},{"version":"2.0.3","createdAt":"2026-07-14T07:50:39.144Z","changelog":"security: harden R deserialization (allow_r_exec gate), declare side effects in description, back up .hyper before overwrite, add Security sections; addresses 6 SkillSpector findings (Tp4 High, RCE High, Lp3, 2x Medium, Low)","fileCount":30,"zipByteSize":108954},{"version":"2.0.2","createdAt":"2026-07-13T08:51:30.340Z","changelog":"security: fix Tp4 overclaim + FST format misidentification; FST now detect-only","fileCount":30,"zipByteSize":106749},{"version":"2.0.1","createdAt":"2026-07-13T08:21:04.148Z","changelog":"docs: bilingual refactor (en-first), slim SKILL.md, fix writer.py XPT crash","fileCount":30,"zipByteSize":106192},{"version":"2.0.0","createdAt":"2026-07-12T08:54:18.113Z","changelog":"v2.0.0 major release: 50+ formats, Excel/MATLAB/Parquet/HDF5/twbx Access fixes","fileCount":30,"zipByteSize":108565},{"version":"1.7.11","createdAt":"2026-07-12T08:48:18.100Z","changelog":"v2.0.0: README 限制修复 + 三真实文件验证 (Excel 合并/MATLAB v7.3/Parquet 分区/HDF5 属性/twbx 内嵌 Access/Access table_name)","fileCount":30,"zipByteSize":108646}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s176fv8983h1rte6dmxwp9wt4n89j8p5:statdata-transfer","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s176fv8983h1rte6dmxwp9wt4n89j8p5:statdata-transfer` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/medstatstar/statdata-transfer before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T07:45:02.995Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-medstatstar-statdata-transfer/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T01:52:55.506Z","emptyReason":null},"readme":"Skill: statdata-transfer\n\nOwner: medstatstar\n\nSummary: 读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明（完整）：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；写入主输出文件的同时可能生成 sidecar 元数据（CSV/TSV 旁 <名>_metadata.json、Parquet/Arrow 内嵌）及覆盖 .hyper 时的 .bak/.bak.1 备份；处理 .rda/.rds/.RData/.mtw/.mpj/.rec 时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。 / Read/convert 50+ statistical software formats, preserving variable/value labels and missing-value metadata for binary stats formats. FULL side effects: runs environment checks (scripts/check_env.py); may optionally pip-install missing packages on request; writes the main output file AND may emit sidecar metadata (e.g. <name>_metadata.json beside CSV/TSV, embedded in Parquet/Arrow schema) and .bak/.bak.1 backups when overwriting .hyper; can invoke the local R interpreter for .rda/.rds/.RData/.mtw/.mpj/.rec files via a fallback DISABLED by default and opted in only with allow_r_exec=True.\n\nTags: latest:2.2.1\n\nVersion history:\n\nv2.2.1 | 2026-08-02T11:02:12.225Z | user\n\nSecurity hardening (2.2.1): pin lxml>=6.1.0 to fix CVE-2026-41066 (XXE); add secure lxml parser with entity resolution disabled for .wdx/.opju; fully disclose sidecar metadata & backup file writes in manifest.\n\nv2.2.0 | 2026-08-02T07:42:56.100Z | user\n\n优化用户菜单(UI)与 README，碳基生物用户使用时更方便；并修复 CSV 分隔符探测、.tsv 读取、UTF-16 编码、XPT 读取报错、长变量标签静默截断告警等健壮性 bug。\n\nv2.1.0 | 2026-07-19T02:32:47.554Z | user\n\nStandardized bilingual (en/zh) formatting per 2026-07-17 policy: non-table separators use '/', Core Capabilities table merged, README_ZH aligned; plus security hardening (allow_r_exec gate, dependency pins).\n\nv2.0.3 | 2026-07-14T07:50:39.144Z | user\n\nsecurity: harden R deserialization (allow_r_exec gate), declare side effects in description, back up .hyper before overwrite, add Security sections; addresses 6 SkillSpector findings (Tp4 High, RCE High, Lp3, 2x Medium, Low)\n\nv2.0.2 | 2026-07-13T08:51:30.340Z | user\n\nsecurity: fix Tp4 overclaim + FST format misidentification; FST now detect-only\n\nv2.0.1 | 2026-07-13T08:21:04.148Z | user\n\ndocs: bilingual refactor (en-first), slim SKILL.md, fix writer.py XPT crash\n\nv2.0.0 | 2026-07-12T08:54:18.113Z | user\n\nv2.0.0 major release: 50+ formats, Excel/MATLAB/Parquet/HDF5/twbx Access fixes\n\nv1.7.11 | 2026-07-12T08:48:18.100Z | user\n\nv2.0.0: README 限制修复 + 三真实文件验证 (Excel 合并/MATLAB v7.3/Parquet 分区/HDF5 属性/twbx 内嵌 Access/Access table_name)\n\nv1.7.10 | 2026-07-12T06:44:15.004Z | user\n\nv1.9.0: add Tableau Hyper (.hyper) read/write and .twbx packaged-workbook unpack read; statdata_meta side-table preserves variable/value labels; lazy-imported, no impact on other formats\n\nv1.7.9 | 2026-07-10T12:23:23.810Z | user\n\nv1.8.4 security fixes: eliminate remaining R-bridge command injection in Minitab(.mtw)/EpiData(.rec) readers and writer — all R bridges now use static scripts with argv params (jsonlite-parsed), no user input interpolated into R code; fix missing _parse_value_labels import\n\nv1.7.8 | 2026-07-10T11:50:07.909Z | user\n\nv1.8.3 security fixes: eliminate R bridge command injection (static R script + argv), raise pyarrow>=17.0/lxml>=6.0 for CVEs, narrow lossless claims, make env install opt-in, stop scanning jamovi embedded JSON\n\nv1.7.7 | 2026-07-10T08:46:43.489Z | user\n\nv1.8.2: fix 6 reader/writer bugs (R crash, encoding detection, SAS special_missing, HDF5 dataset selection, Excel warnings/XPT value labels/Excel label round-trip), narrow ClawHub audit triggers, repo hygiene\n\nv1.7.6 | 2026-07-04T09:04:16.499Z | auto\n\nstatdata-transfer 1.7.6\n\n- Major cleanup: documentation, reference, script, and test files were removed, leaving only SKILL.md.\n- Reduced repository to core metadata file (SKILL.md); no code or usage examples bundled.\n- README, LICENSE, requirements, assets, and test documentation no longer included in the skill package.\n\nv1.7.5 | 2026-07-04T02:58:22.517Z | auto\n\nstatdata-transfer 1.7.5\n\n- Updated descriptions in SKILL.md and README files for greater clarity and conciseness on supported formats and bidirectional metadata preservation.\n- Enhanced trigger phrases and table details to better reflect use cases and format capabilities.\n- No code changes; documentation only.\n\nv1.7.4 | 2026-07-04T00:38:03.734Z | auto\n\nstatdata-transfer v1.7.4\n\n- Updated documentation for clarity, conciseness, and consistency across README.md, README_ZH.md, and SKILL.md\n- Reformatted and translated the Supported Formats & Capability Matrix to improve readability\n- No functional changes to the core scripts (reader_core.py, reader_sas.py, reader_spss.py, reader_stata.py) logic detected; only doc and metadata updates\n- Increased skill metadata version to 1.8.0 to reflect documentation improvements\n\nv1.7.3 | 2026-07-03T23:14:09.836Z | auto\n\nstatdata-transfer 1.7.3\n\n- Added openclaw \"icon\" field to metadata (SKILL.md).\n- Documentation updates to README.md, README_ZH.md, and metadata section.\n- No functional/API changes; update is documentation and metadata only.\n\nv1.7.2 | 2026-07-03T11:47:41.742Z | auto\n\n- Updated the logo (assets/logo.svg) for statdata-transfer.\n- No changes to functionality or documentation content.\n\nv1.7.1 | 2026-07-03T11:30:41.639Z | auto\n\nstatdata-transfer 1.7.1\n\n- No file changes detected; this is a version bump only.\n- No user-facing feature, documentation, or metadata changes in this release.\n\nv1.7.0 | 2026-07-03T11:24:34.224Z | auto\n\n- Added support for reading EpiInfo, ARFF (Weka), and Gretl formats, broadening compatibility with epidemiology, machine learning, and econometrics tools.\n- Improved bidirectional conversion: now supports lossless, metadata-preserving conversions across 28+ statistical and data science formats.\n- Enhanced metadata management: all variable/value labels, special missings, and custom metadata preserved for SPSS, Stata, SAS, R, and compatible formats.\n- Conversion process now issues detailed, automatic warnings whenever metadata loss may occur.\n- Updated documentation with comprehensive capability tables, use cases, and fallback rules for metadata handling.\n\nArchive index:\n\nArchive v2.2.1: 32 files, 118053 bytes\n\nFiles: _icon.svg (3750b), assets/logo.svg (3750b), CHANGELOG.md (2892b), LICENSE (1084b), README_zh-CN.md (13918b), README.md (13793b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3379b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1754b), scripts/__init__.py (531b), scripts/check_env.py (4395b), scripts/reader_arff.py (9220b), scripts/reader_core.py (42537b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (10290b), scripts/reader_gretl.py (9623b), scripts/reader_legacy.py (21011b), scripts/reader_modern.py (19356b), scripts/reader_odm.py (10187b), scripts/reader_r.py (40967b), scripts/reader_sas.py (6286b), scripts/reader_science.py (46550b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (16786b), scripts/reader_v14.py (28397b), scripts/writer.py (34930b), skill-card.md (2846b), SKILL.md (10234b), test_report.md (3399b), _meta.json (136b)\n\nFile v2.2.1:SKILL.md\n\n---\r\nslug: statdata-transfer\r\nname: statdata-transfer\r\ndisplayName: 统计数据格式转换器 / Statistical Data Format Converter\r\ncn_name: 统计数据格式转换器\r\nversion: 2.2.1\r\nsummary: 读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明（完整）：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；写入主输出文件的同时，可能生成 sidecar 元数据文件（CSV/TSV 旁生成 <名>_metadata.json，Parquet/Arrow 内嵌元数据）及覆盖 .hyper 时的 .bak/.bak.1 备份；处理 .rda/.rds/.RData/.mtw/.mpj/.rec 文件时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。\r\nlicense: MIT\r\ndescription: \"读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明（完整）：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；写入主输出文件的同时可能生成 sidecar 元数据（CSV/TSV 旁 <名>_metadata.json、Parquet/Arrow 内嵌）及覆盖 .hyper 时的 .bak/.bak.1 备份；处理 .rda/.rds/.RData/.mtw/.mpj/.rec 时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。 / Read/convert 50+ statistical software formats, preserving variable/value labels and missing-value metadata for binary stats formats. FULL side effects: runs environment checks (scripts/check_env.py); may optionally pip-install missing packages on request; writes the main output file AND may emit sidecar metadata (e.g. <name>_metadata.json beside CSV/TSV, embedded in Parquet/Arrow schema) and .bak/.bak.1 backups when overwriting .hyper; can invoke the local R interpreter for .rda/.rds/.RData/.mtw/.mpj/.rec files via a fallback DISABLED by default and opted in only with allow_r_exec=True.\"\r\ntriggers:\r\n  - \"statdata-transfer\"\r\n  - \"统计数据格式转换\"\r\n  - \"spss stata sas 格式\"\r\n  - \".sav .dta .sas7bdat 读入\"\r\n  - \"sav转dta 格式转换\"\r\n  - \"variable labels 变量标签\"\r\n  - \"metadata-preserved conversion\"\r\nrequired_commands: [python]\r\ninvocable: true\r\nmetadata:\r\n  openclaw: { emoji: \"🛠️\", icon: \"assets/logo.svg\" }\r\n  authors: [\"medstatstar\", \"phoe-zip\"]\r\n  license: \"MIT\"\r\n  tags: [\"data-conversion\", \"statistics\", \"spss\", \"stata\", \"sas\", \"clinical-trials\", \"metadata\", \"pandas\", \"bidirectional\"]\r\n  homepage: \"https://github.com/medstatstar/statdata-transfer\"\r\npermissions:\r\n  scope: \"user-space-only\"\r\n  network: \"off\"\r\n  network_note: \"Offline by default; the only network touchpoint is the optional `python scripts/check_env.py --install`, which pip-installs missing packages and runs ONLY on explicit user request.\"\r\n  filesystem: \"read-only to its own files; reads the input data file you specify; writes the converted output file to a path you specify, and may additionally create sidecar metadata files (e.g. <name>_metadata.json beside CSV/TSV, or metadata embedded in Parquet/Arrow schema) and .bak/.bak.1 backups when overwriting .hyper\"\r\n  data: \"no external data transmission\"\r\n---\r\n\r\n# Statistical Data Format Converter\r\n\r\n> **Safe by default — preview, not execute**: the skill shows what it will read/convert and only writes a file when you explicitly ask. Every R-invoking path is opt-in and disabled by default.\r\n\r\n## Language\r\n\r\n- **English guide** → [README.md](https://github.com/medstatstar/statdata-transfer/blob/main/README.md)\r\n- **中文指南** → [README_zh-CN.md](https://github.com/medstatstar/statdata-transfer/blob/main/README_zh-CN.md)\r\n\r\nThis skill responds in the user's input language and auto-switches; runtime prompts switch by locale. SKILL.md body is English-only (agent-facing); bilingual walkthroughs live in the two READMEs.\r\n\r\n## Purpose\r\n\r\nRead 50+ statistical-software and clinical-trial data formats into a pandas DataFrame, and inter-convert between most formats (SPSS ↔ Stata ↔ R ↔ SAS XPT ↔ Excel ↔ Parquet ↔ HDF5 ↔ JSON …). For statistical binary formats it preserves full variable/value labels and special-missing-value metadata; text/JSON formats preserve only a retainable subset.\r\n\r\n## Features\r\n\r\n| Capability | Description | Typical Scenario |\r\n|:---|:---|:---|\r\n| **Read** | Extract data + all metadata from 50+ formats into a pandas DataFrame; clearly report what is preserved vs lost | `read data.sav and show metadata` |\r\n| **Convert** | Inter-convert most stats formats; export to universal formats (Parquet/Feather/HDF5/JSON/CSV/Excel) with labels embedded | `convert data.sav to .dta keeping variable labels` |\r\n| **Embed metadata** | Labels embedded in Arrow `schema.metadata` / sidecar JSON for lossless round-trips | `save to parquet but keep value labels` |\r\n| **Warn** | Auto-detect and report metadata loss per conversion path | audit before exporting to CSV |\r\n\r\n## Supported Formats\r\n\r\n*50+ formats, sorted alphabetically.*\r\n\r\n| Format | Extension | Meta Preserve |\r\n|--------|-----------|---------------|\r\n| CDISC ODM | `.odm` | ⚠️ Clinical data only |\r\n| dBASE / FoxPro | `.dbf` | ⚠️ Read+Write, uppercase names |\r\n| EpiData | `.rec` | ⚠️ Via R (opt-in) |\r\n| EpiInfo | `.prj` `.xml` | ✅ XML structure |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | ⚠️ Extra sheet for labels; merged-cell fill |\r\n| EViews | `.wf1` `.wf2` | ⚠️ JSON structure |\r\n| Feather | `.feather` `.arrow` | ✅ Via schema |\r\n| FST | `.fst` | ✗ Detect-only (proprietary) |\r\n| GraphPad Prism | `.pzfx` `.pz` | ⚠️ Multi-table |\r\n| Gretl | `.gdt` `.gdtb` | ✅ String-tables |\r\n| HDF5 | `.h5` `.hdf5` | ⚠️ Hierarchy + attribute labels |\r\n| HTML | `.html` | ⚠️ Tables only |\r\n| jamovi | `.omv` | ✅ JSON analysis |\r\n| JMP | `.jmp` | ⚠️ Multi-table |\r\n| JSON | `.json` | ✅ stat-full-meta |\r\n| MATLAB | `.mat` | ⚠️ v7.3+ via h5py fallback |\r\n| Mathematica | `.wdx` | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | ⚠️ Via R (opt-in) |\r\n| MS Access | `.mdb` `.accdb` | ⚠️ Multi-table; needs system driver |\r\n| ODS | `.ods` | ⚠️ Data only |\r\n| ORC | `.orc` | ✅ Via schema |\r\n| Origin | `.opju` `.oggu` | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | ✅ Via schema; partitioned datasets |\r\n| R | `.rda` `.rds` `.rdata` | ✅ pyreadr; R fallback opt-in (allow_r_exec) |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | ✅ |\r\n| Stata | `.dta` | ✅ |\r\n| Weka ARFF | `.arff` | ✅ Nominal mapping |\r\n| XML | `.xml` | ⚠️ Structure preserved |\r\n\r\n> ✅=Full · ⚠️=Partial/conditional · ✗=Not preserved\r\n>\r\n> 12 detect-only formats (SAS CPORT `.cpt`, Statistica `.sta`, OxMetrics `.in7`, SYSTAT `.sys`/`.syd`, Paradox `.db`/`.px`, LIMDEP `.lpw`, NCSS `.ncss`, FST) give clear export guidance — see README.\r\n\r\n## Return Structure\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"Question 1\"},\r\n        \"value_labels\": {\"q1\": {1: \"Yes\", 2: \"No\"}},\r\n        \"special_missing\": {...},\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n## Quick Start\r\n\r\n```bash\r\n# Check environment (optional install on request)\r\npython scripts/check_env.py --install\r\n```\r\n\r\nIn WorkBuddy (bilingual, auto-detects your language):\r\n\r\n```\r\n> convert data.sav to .dta\r\n> read data.sav and show metadata\r\n> 把 data.sav 转成 .dta 并保留变量标签\r\n```\r\n\r\n> For complete code examples, see `references/usage_examples.py`.\r\n\r\n## Dependencies\r\n\r\n```yaml\r\nrequires:\r\n  bins: [python3]\r\n  packages:\r\n    core: [pyreadstat>=1.3.5,<2, pyreadr>=0.4,<0.5, pandas>=2.0,<3]\r\n    extended: [openpyxl, xlrd, scipy, h5py, pyarrow, lxml, odfpy, tableauhyperapi, dbfread, dbf, pyodbc]\r\n```\r\n\r\n> Full list: `requirements.txt`\r\n\r\n## ⚠️ Safety\r\n\r\n- All R-invoking paths are **opt-in and disabled by default**; they only run when you pass `allow_r_exec=True` on a trusted file.\r\n- Pure-Python parsers (`pyreadr`, `mtbpy`) are tried first and never execute code.\r\n- No silent R fallback — if the pure-Python parser fails and `allow_r_exec` is not set, the skill raises a clear error.\r\n- Writing an existing `.hyper` backs up to `.bak` before overwrite; on failure the original is untouched.\r\n- Output for reference only; validate before regulatory submissions.\r\n\r\n### Security model (transparent disclosure)\r\n\r\n| Behavior | Description |\r\n|:---|:---|\r\n| **R invocation (opt-in)** | Reading `.rda/.rds/.RData` (`readRDS()/load()`), Minitab `.mtw/.mpj`, EpiData `.rec`, and writing R formats run **only** with `allow_r_exec=True` on a trusted file. Pure-Python parsers tried first. |\r\n| **No silent fallback** | On pure-Python parser failure without `allow_r_exec`, raises a clear error instead of launching R — avoids executing embedded code from untrusted files. |\r\n| **Static R templates** | When the opt-in R path runs, all R scripts are static templates; user input passes only as CLI args (`commandArgs(trailingOnly=TRUE)` via `jsonlite`) — never concatenated into executable R code. |\r\n| **Temp CSV bridge** | Opt-in R writes data to a temp CSV then reads back; deleted after use, but on crash could briefly persist — avoid highly sensitive data through R-backed formats. |\r\n| **No destructive writes** | `.hyper` write → temp file → rotate existing to `.bak` (prior `.bak` → `.bak.1`, never silently deleted) → atomic swap. Original untouched on failure. |\r\n| **Sidecar metadata files** | Writing CSV/TSV also emits `<name>_metadata.json` (full 17-field metadata) next to the data file; Parquet/Arrow embed metadata inside the file schema. No writes occur outside the output path you specify. |\r\n| **Pinned dependencies** | Core deps carry upper-bound pins (`pandas`, `pyreadstat`, `pyreadr`) — see `requirements.txt`. |\r\n| **Optional install** | `python scripts/check_env.py --install` only on explicit request. |\r\n| **Permissions** | Read the input file; write the output file to a path you specify. No network unless you explicitly request package install. |\r\n\r\n## License\r\n\r\nMIT. See [LICENSE](./LICENSE).\n\nFile v2.2.1:README.md\n\n# statdata-transfer / Statistical Data Format Converter\r\n\r\n[🇨🇳 Chinese](./README_zh-CN.md)\r\n\r\n<div align=\"center\">\r\n<img src=\"assets/icon.svg\" width=\"240\" height=\"240\" />\r\n</div>\r\n\r\n---\r\n\r\n> Read 50+ statistical-software and clinical-trial data formats, and **inter-convert between most of them** while keeping variable/value labels and missing-value metadata. No statistical software required — format conversion only.\r\n\r\n## How to use it in a conversation\r\n\r\nJust talk to the agent in natural language. A few real examples (copy-paste ready):\r\n\r\n**① Most common — convert a file**\r\n- **You say**: `convert C:/Users/Name/Desktop/data.sav to .dta`\r\n- **Agent replies** (sketch): reads `data.sav` with pyreadstat, preserves all variable/value labels, and writes `data.dta` in the same folder.\r\n- **Trigger the real conversion**: by default the agent previews the plan; say `please write the file` to execute.\r\n\r\n**② Show what's inside**\r\n- **You say**: `read data.sav and show metadata`\r\n- **Agent replies**: prints the DataFrame shape, variable labels, value labels, and a list of which metadata will be preserved.\r\n\r\n**③ Check before you lose data**\r\n- **You say**: `will converting .sav to .xlsx lose any metadata?`\r\n- **Agent replies**: warns that Excel keeps labels only in a side sheet; suggests Parquet/Stata to keep them losslessly.\r\n\r\n**④ Ask for reproducible code**\r\n- **You say**: `show me the Python code to convert .sav to .parquet`\r\n- **Agent replies**: prints the `read_stat_file` / `write_stat_file` snippet (code is always English).\r\n\r\n**⑤ Switch language**\r\n- **You say**: `reply in Chinese` / `switch to English` — all user-facing messages follow your OS language or this prompt.\r\n\r\n## What can it do? (scenario index)\r\n\r\n| Capability | Typical use | Try saying |\r\n|:---|:---|:---|\r\n| **Read 50+ formats** | Open SPSS/Stata/SAS/R/Excel/Parquet/HDF5/JSON… into pandas | `read data.sav and show metadata` |\r\n| **Convert between stats formats** | SPSS ↔ Stata ↔ R ↔ SAS XPT, keeping all labels | `convert data.sav to .dta keeping variable labels` |\r\n| **Export universal formats** | Parquet / Feather / HDF5 / JSON / CSV / Excel with labels embedded | `save to parquet but keep value labels` |\r\n| **Metadata-safe round-trip** | Labels survive a convert-and-convert-back | `convert to parquet then back to sav, keep labels` |\r\n| **Metadata-loss warning** | Know what will be dropped before exporting | `will .sav to .xlsx lose metadata?` |\r\n| **Batch / folder** | Convert a whole folder or a zip archive | `convert all .dta in this zip to .sav` |\r\n\r\nFull format list and per-format limits: see **Advanced reference** below.\r\n\r\n## First-use FAQ\r\n\r\n- **Do I need SPSS/Stata/R installed?** No. The skill is pure Python; it only *optionally* calls a local R interpreter for a few formats (Minitab/EpiData/R write), and only when you pass `allow_r_exec=True`.\r\n- **How do I get the actual converted file, not just code?** Say `please write the file`. By default it previews; execution is explicit.\r\n- **Will my labels survive?** For binary stats formats (SPSS/Stata/SAS/R) — yes, fully. For text/JSON — only a retainable subset. The skill always tells you what is preserved vs lost.\r\n- **Can I get reproducible code?** Yes — ask `show me the Python code` and it prints the `read_stat_file` / `write_stat_file` calls.\r\n- **Is the output in Chinese on a Chinese system?** User-facing messages auto-switch to Chinese on a `zh-*` OS, or you can force it with `reply in Chinese`. Code stays English.\r\n- **My data file is huge / can't be uploaded directly?** Use the absolute file path in your prompt, or compress it into a `.zip` and upload the zip.\r\n\r\n## Safety (in plain words)\r\n\r\nThe skill runs **locally** and follows a **safe-preview** model: it shows what it will read/convert and only writes a file when you explicitly ask. Every path that would call the local R interpreter is **opt-in and off by default** — it runs only when you pass `allow_r_exec=True` on a file you trust. Your data is never sent over the network unless you explicitly request a package install. Treat the output as reference and validate before regulatory submissions.\r\n\r\n---\r\n\r\n## Advanced reference\r\n\r\n> The following is developer/reference material, kept out of the quick-start above.\r\n\r\n### Supported Formats & Capability Matrix\r\n\r\n*Sorted alphabetically.*\r\n\r\n| Format | Extension | Dependency | Var Label | Val Label | Special Missing | Formula | Meta Preserve |\r\n|--------|-----------|------------|-----------|-----------|-----------------|---------|---------------|\r\n| CDISC ODM | `.odm` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Clinical data only |\r\n| dBASE / FoxPro | `.dbf` | dbfread / dbf | ✗ | ✗ | ✗ | ✗ | ⚠️ Read+Write; uppercase names |\r\n| EpiData | `.rec` | R foreign | ✗ | ✗ | ✗ | ✗ | ⚠️ Via R |\r\n| EpiInfo | `.prj` `.xml` | xml/etree | ✅ | ✅(codes) | ✗ | ✗ | ✅ XML structure |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | openpyxl / xlrd | ✗ | ✗ | ✗ | ⚠️ result only | ⚠️ Extra sheet for labels; merged-cell fill |\r\n| EViews | `.wf1` `.wf2` | built-in | ✗ | ✗ | ✗ | ✗ | ⚠️ JSON structure |\r\n| Feather | `.feather` `.arrow` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Version diff |\r\n| FST | `.fst` | — | ✗ | ✗ | ✗ | ✗ | ✗ Detect-only (proprietary format) |\r\n| GraphPad Prism | `.pzfx` `.pz` | pzfx | ✗ | ✗ | ✗ | ✗ | ⚠️ Multi-table |\r\n| Gretl | `.gdt` `.gdtb` | built-in | ✅ | ✅(tables) | ✗ | ✗ | ✅ string-tables |\r\n| HDF5 | `.h5` `.hdf5` | h5py | ✗ | ✗ | ✗ | ✗ | ⚠️ Hierarchy + attribute labels |\r\n| HTML | `.html` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Tables only |\r\n| jamovi | `.omv` | ZIP built-in | ✅ | ✅ | ✗ | ✗ | ✅ JSON analysis |\r\n| JMP | `.jmp` | jmpio-python | ⚠️ | ⚠️ | ✗ | ✗ | ⚠️ Multi-table |\r\n| JSON | `.json` | built-in | ✅ | ✅ | ✗ | ✗ | ✅ stat-full-meta on write |\r\n| MATLAB | `.mat` | scipy | ✗ | ✗ | ✗ | ✗ | ⚠️ v7.3+ via h5py fallback |\r\n| Mathematica | `.wdx` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | mtbpy / R | ✗ | ✗ | ✗ | ✗ | ⚠️ Via R |\r\n| MS Access | `.mdb` `.accdb` | pyodbc + Access Driver | ✗ | ✗ | ✗ | ✗ | ⚠️ Multi-table; needs system driver |\r\n| ODS | `.ods` | odfpy | ✗ | ✗ | ✗ | ✗ | ⚠️ Data only |\r\n| ORC | `.orc` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Version diff |\r\n| Origin | `.opju` `.oggu` | zipfile + lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Nested types; partitioned datasets |\r\n| R | `.rda` `.rds` `.rdata` | pyreadr + R | ✅ | ✅ | ✅ | ✗ | ✅ statdata_meta + R bridge |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | pyreadstat | ✅ | ✅(need catalog) | ⚠️ | ✗ | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | pyreadstat | ✅ | ✅ | ✅ | ✗ | ✅ |\r\n| Stata | `.dta` | pyreadstat | ✅ | ✅ | ⚠️ | ✗ | ✅ |\r\n| Weka ARFF | `.arff` | built-in | ✅ | ✅(nominal) | ✗ | ✗ | ✅ nominal mapping |\r\n| XML | `.xml` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Structure preserved |\r\n\r\n> ✅=Full preservation · ⚠️=Partial/conditional · ✗=Not preserved\r\n\r\n### Detect-Only Formats\r\n\r\nFormats with no parser available. The skill detects the extension and provides clear export guidance (no data parsing).\r\n\r\n| Format | Extension | Guidance |\r\n|--------|-----------|----------|\r\n| FST (R fst package) | `.fst` | R: `fst::read_fst(\"in.fst\", \"out.csv\")` then read CSV |\r\n| LIMDEP / NLOGIT | `.lpw` | Export to CSV from original software |\r\n| NCSS | `.ncss` | Export to CSV |\r\n| OxMetrics | `.in7` | Export to CSV / `.dta` |\r\n| Paradox | `.db` `.px` | Export to `.dbf` / CSV |\r\n| SAS CPORT | `.cpt` | SAS: `proc export` to XPORT(`.xpt`) / `.sas7bdat` |\r\n| Statistica | `.sta` | Export to `.sav` / `.csv` |\r\n| SYSTAT | `.sys` `.syd` | Export to CSV / `.sav` |\r\n\r\n### Return Structure\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"Question 1\"},\r\n        \"value_labels\": {\"q1\": {1: \"Yes\", 2: \"No\"}},\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n### Metadata Preservation Tiers\r\n\r\n1. **Statistical binary formats** (SPSS/Stata/SAS/R): 100% metadata preserved\r\n2. **Arrow ecosystem** (Parquet/Feather/ORC): only restores labels from `write_stat_file`\r\n3. **Non-stats formats** (CSV/Excel/XML/HTML/ODS): data only; use `apply_value_labels()` to attach manually\r\n4. **R formats**: embeds all metadata via `statdata_meta` attribute\r\n\r\n### Recommended Read Strategies\r\n\r\n| Use Case | Recommendation |\r\n|----------|---------------|\r\n| Data warehousing / ETL | SPSS `.sav` or Stata `.dta` → Parquet / HDF5 |\r\n| Scientific computing | `.mat` or `.hdf5` → NumPy / pandas |\r\n| Statistical analysis (Python) | `.sav` / `.dta` → pandas → scipy.stats |\r\n| Report output | pandas → JSON / HTML / Excel |\r\n| Cross-software sharing | Stata ↔ SPSS ↔ R direct interconversion |\r\n\r\n### File Size Limits\r\n\r\n| Format | Memory Behavior |\r\n|--------|----------------|\r\n| pyreadstat (SPSS/Stata/SAS) | Loads entire file into RAM |\r\n| HDF5 | Chunked reading; not limited by RAM |\r\n| Parquet | pyarrow memory-mapped (mmap); handles files >RAM |\r\n\r\n### Encoding Notes\r\n\r\n- **Chinese files**: old Stata/SAS may use GBK/gb2312. Use `encoding='gbk'`.\r\n- **European files**: some SAS files use Latin-1. Try `encoding='latin1'` if UTF-8 fails.\r\n- **Auto-detection**: `_auto_detect_encoding` is enabled by default for SPSS/Stata/SAS.\r\n\r\n### Providing Input Files\r\n\r\nAI agents can only directly upload a limited set of file types. When your data file cannot be uploaded directly:\r\n\r\n1. **Use the absolute file path** in your prompt (e.g. `convert C:/Users/Name/Desktop/data.sav to .dta`)\r\n2. **Compress the file as a `.zip` archive** and upload the zip instead\r\n\r\nThe skill automatically extracts and processes zip archives containing a single data file.\r\n\r\n### CLI (advanced)\r\n\r\n```bash\r\n# Check environment (optional install on request)\r\npython scripts/check_env.py --install\r\n```\r\n\r\nComplete code examples: [`references/usage_examples.py`](./references/usage_examples.py)\r\n\r\n### Extending\r\n\r\nTo add a new format: edit `scripts/reader_*.py` to add a reader function, register it in `format_map` in `scripts/reader_core.py`, and add a TypedDict in `scripts/reader_core.py`.\r\n\r\n### Format Limitations\r\n\r\n*Alphabetically ordered. ✅ = fixed, 🔄 = new capability; rest are inherent format limits.*\r\n\r\n- **CDISC ODM (.odm)**: ❌ XML structure dependency; ❌ no statistical metadata in ODM spec, only clinical structure preserved\r\n- **dBASE / FoxPro (.dbf)**: ❌ field names forced to uppercase; ✅ Read + Write supported\r\n- **EpiData (.rec)**: ❌ requires R + `foreign` package; ❌ statistical metadata lost in R-to-CSV bridge\r\n- **EpiInfo (.prj)**: ❌ project file contains no data, auto-associates same-name CSV; ❌ Access not supported, export to CSV first; ✅ variable labels/codes reconstructed in XML\r\n- **Excel (.xlsx/.xls/.xlsm)**: ✅ merged cells filled with anchor value (`fill_merged_cells=True`, default); ❌ formulas lost; ❌ charts/shapes not extracted; labels in separate metadata worksheet on write\r\n- **HDF5 (.h5/.hdf5)**: ✅ multi-dataset fallback via h5py; ✅ attribute labels scanned; ❌ hierarchy flattened\r\n- **JMP (.jmp)**: ❌ requires `jmpio-python`; ❌ multi-table returns first only; write single-table only\r\n- **MATLAB (.mat)**: ✅ v7.3+ (HDF5) via h5py; ❌ complex structures flattened; ❌ object/datetime lose fidelity\r\n- **Parquet (.parquet)**: ❌ deeply nested types (>2 levels) opaque; ✅ partitioned datasets via `pyarrow.dataset`\r\n- **R (.rda/.rds/.rdata)**: ✅ ASCII XDR read via R bridge needs `allow_r_exec=True`; ❌ factor order may not be Categorical unless embedded; write via `statdata_meta`\r\n- **SAS (.sas7bdat/.xpt/.sas7bcat)**: ✅ value labels need co-located `.sas7bcat`; ❌ Viya CAS `.sashdat` not supported; date origin 1960-01-01\r\n- **SPSS (.sav/.zsav/.por)**: ❌ MR Sets as raw dict; ❌ formulas lost; ⚠️ special missing (`.A`–`.Z`) flagged in `special_missing`; `.zsav` needs pyreadstat 1.2+, else fallback to `.sav`\r\n- **Stata (.dta)**: ⚠️ special missing (`.a`–`.z`) preserved when `user_missing=True` (default), NaN when `False` (irreversible); ✅ pre-v13 Latin-1 auto-detected; ❌ Stata 117–119 not supported, auto-downgrade to v15 on write\r\n\r\n### Security\r\n\r\n- **R execution is opt-in and sandboxed by default.** Reading `.rda/.rds/.RData`, Minitab `.mtw/.mpj`, EpiData `.rec`, and writing R formats is disabled by default; runs only with `allow_r_exec=True` on a trusted file. Pure-Python parsers tried first.\r\n- **No silent R fallback.** On pure-Python failure without `allow_r_exec`, raises a clear error instead of launching R.\r\n- **R scripts are static templates.** User input passes only as CLI args — never concatenated into executable R code.\r\n- **Temp CSV exposure (R bridge).** Opt-in R writes a temp CSV; deleted after use but could briefly persist on crash. Avoid highly sensitive data through R-backed formats.\r\n- **No destructive writes.** Existing `.hyper` is rotated to `.bak` before overwrite; original untouched on failure.\r\n- **Pinned dependencies.** Core deps carry upper-bound pins — see `requirements.txt`.\r\n\r\n## Contact the author\r\n\r\nFor feature requests, bug reports, or other feedback, please contact the author directly at medstatstar@gmail.com (Wintone Zhang).\r\n\r\n## License\r\n\r\nMIT License. See [LICENSE](LICENSE) for details.\n\nFile v2.2.1:_meta.json\n\n{\n  \"ownerId\": \"kn7amqq1jv28skb63wavr6shah89jsm5\",\n  \"slug\": \"statdata-transfer\",\n  \"version\": \"2.2.1\",\n  \"publishedAt\": 1785668532225\n}\n\nFile v2.2.1:references/new_formats_architecture_analysis.json\n\n{\r\n  \"current_architecture\": {\r\n    \"data\": \"pandas DataFrame\",\r\n    \"metadata\": \"BaseMeta TypedDict (~28 fields)\",\r\n    \"column_report\": \"ColumnInfo TypedDict (12 fields)\",\r\n    \"return_type\": \"StatFileResult = dict[str, Any] with keys: dataframe, metadata, warnings, column_report\",\r\n    \"multi_object_pattern\": \"read_all_*() returns dict[str, StatFileResult]\"\r\n  },\r\n  \"formats_analysis\": {\r\n    \"sas7bcat\": {\r\n      \"description\": \"SAS Ŀ¼�ļ����洢��ʽ���壨ֵ��ǩ��\",\r\n      \"data_structure\": \"�����ݣ�ֻ��Ԫ���ݣ���ʽ���壩\",\r\n      \"metadata_fields\": [\r\n        \"value_labels\",\r\n        \"variable_value_labels\",\r\n        \"variable_to_label\"\r\n      ],\r\n      \"current_architecture_sufficient\": true,\r\n      \"notes\": \"pyreadstat ��֧�ֶ�ȡ .sas7bcat������ value_labels dict\",\r\n      \"architecture_extension_needed\": false\r\n    },\r\n    \"jmp\": {\r\n      \"description\": \"SAS JMP �����ļ����ɰ���������ݱ����ű����������\",\r\n      \"data_structure\": \"�ɰ���������ݱ���Data Table����ÿ������һ�� DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"column_properties (formulas, ranges)\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"JMP �ļ��ɰ���������ݱ�����Ҫ read_all_jmp_tables() ģʽ�������ԣ���ʽ����Χ����Ҫ��չ ColumnInfo\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"ColumnInfo ��Ҫ��չ��formula (str), range (dict), column_property (dict)\",\r\n        \"��Ҫ���� JmpMeta �࣬���� tables list��scripts list��analysis list\",\r\n        \"��Ҫ���� read_all_jmp_tables() ����\"\r\n      ]\r\n    },\r\n    \"minitab\": {\r\n      \"description\": \"Minitab �������ļ����ɰ��������������Worksheet��\",\r\n      \"data_structure\": \"�ɰ��������������ÿ����������һ�� DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"worksheet_names\",\r\n        \"formulas\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"Minitab �������ɰ����������������Ҫ read_all_minitab_worksheets() ģʽ\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� MinitabMeta �࣬���� worksheets list��active_worksheet str\",\r\n        \"��Ҫ���� read_all_minitab_worksheets() ����\",\r\n        \"ColumnInfo ������Ҫ��չ��formula (str)\"\r\n      ]\r\n    },\r\n    \"prism\": {\r\n      \"description\": \"GraphPad Prism ��Ŀ�ļ����������ݱ����������ͼ��\",\r\n      \"data_structure\": \"�������ݱ���DataFrame�����������DataFrame����ͼ�Σ��޷�תΪ DataFrame��\",\r\n      \"metadata_fields\": [\r\n        \"data_tables\",\r\n        \"results_tables\",\r\n        \"graphs_info\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"Prism �ļ��������ݱ��ͽ���������߶��� DataFrame��ͼ���޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� PrismMeta �࣬���� data_tables list��results_tables list��graphs_info list\",\r\n        \"StatFileResult ��Ҫ��չ������� read_all_prism_tables() ģʽ\",\r\n        \"��ǰ�ܹ�ֻ�ܱ������ݱ����������ͼ����Ϣ�ᶪʧ\"\r\n      ]\r\n    },\r\n    \"jamovi\": {\r\n      \"description\": \"jamovi ��Ŀ�ļ����������ݡ��������������\",\r\n      \"data_structure\": \"�������ݣ�CSV�������������JSON��\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"analysis_results\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"jamovi �ļ��� ZIP������ data.csv �� analysis.json����������޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� JamoviMeta �࣬���� analysis_results dict��analysis_settings dict\",\r\n        \"��ǰ�ܹ����Ա������ݲ��֣�����������ᶪʧ\"\r\n      ]\r\n    },\r\n    \"epidata\": {\r\n      \"description\": \"EpiData �����ļ������в�ѧ���鳣��\",\r\n      \"data_structure\": \"�������ݱ���DataFrame��\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"data_types\"\r\n      ],\r\n      \"current_architecture_sufficient\": true,\r\n      \"notes\": \"EpiData �ļ��ṹ�� SPSS .sav ���ƣ���ǰ�ܹ����� 100% ����\",\r\n      \"architecture_extension_needed\": false\r\n    },\r\n    \"eviews\": {\r\n      \"description\": \"EViews �����ļ����������У�Series�����飨Group�������̣�Equation��\",\r\n      \"data_structure\": \"����������У������ Group��DataFrame���������޷�תΪ DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"series_names\",\r\n        \"groups\",\r\n        \"equations\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"EViews .wf2 �� JSON ��ʽ���ɽ����������̡�ϵ���ȷ�������޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� EviewsMeta �࣬���� series list��groups list��equations list\",\r\n        \"��ǰ�ܹ����Ա��� Group Ϊ DataFrame����������Ϣ�ᶪʧ\"\r\n      ]\r\n    }\r\n  }\r\n}\n\nFile v2.2.1:references/v1.4_implementation_summary.json\n\n{\r\n  \"v1.4_new_formats\": [\r\n    {\r\n      \"format\": \"SAS Catalog\",\r\n      \"ext\": \".sas7bcat\",\r\n      \"handler\": \"_read_sas_catalog\",\r\n      \"dependency\": \"pyreadstat (已支持)\",\r\n      \"architecture\": \"SasCatalogMeta, 返回格式定义 DataFrame\",\r\n      \"status\": \"done\"\r\n    },\r\n    {\r\n      \"format\": \"JMP\",\r\n      \"ext\": \".jmp\",\r\n      \"handler\": \"_read_jmp\",\r\n      \"dependency\": \"jmpio-python (PyPI) 或 @skill:statsoft-cli\",\r\n      \"architecture\": \"JmpMeta, 多表需 read_all_jmp_tables()\",\r\n      \"status\": \"done (handler 已添加，jmpio 未安装）\"\r\n    },\r\n    {\r\n      \"format\": \"Minitab\",\r\n      \"ext\": \".mtw/.mpj\",\r\n      \"handler\": \"_read_minitab\",\r\n      \"dependency\": \"mtbpy 或 R foreign::read.mtb() 中继\",\r\n      \"architecture\": \"MinitabMeta, 多工作表需 read_all_minitab_worksheets()\",\r\n      \"status\": \"done (R 中继已实现）\"\r\n    },\r\n    {\r\n      \"format\": \"GraphPad Prism\",\r\n      \"ext\": \".pzfx/.pz\",\r\n      \"handler\": \"_read_prism\",\r\n      \"dependency\": \"pzfx (PyPI) 或 @skill:statsoft-cli\",\r\n      \"architecture\": \"PrismMeta, 含数据表+结果表\",\r\n      \"status\": \"done (handler 已添加，pzfx 未安装）\"\r\n    },\r\n    {\r\n      \"format\": \"jamovi\",\r\n      \"ext\": \".omv\",\r\n      \"handler\": \"_read_jamovi\",\r\n      \"dependency\": \"无需额外包（ZIP+CSV 解析）\",\r\n      \"architecture\": \"JamoviMeta, 含 analysis JSON\",\r\n      \"status\": \"done\"\r\n    },\r\n    {\r\n      \"format\": \"EpiData\",\r\n      \"ext\": \".rec\",\r\n      \"handler\": \"_read_epidata\",\r\n      \"dependency\": \"R foreign::read.epiinfo() 中继\",\r\n      \"architecture\": \"EpidataMeta\",\r\n      \"status\": \"done (R 中继已实现）\"\r\n    },\r\n    {\r\n      \"format\": \"EViews\",\r\n      \"ext\": \".wf1/.wf2\",\r\n      \"handler\": \"_read_eviews\",\r\n      \"dependency\": \".wf2 可直接解析 JSON；.wf1 需 @skill:statsoft-cli\",\r\n      \"architecture\": \"EviewsMeta\",\r\n      \"status\": \"done (.wf2 解析已实现）\"\r\n    }\r\n  ],\r\n  \"architecture_extensions\": [\r\n    \"ColumnInfo 新增：formula (str), column_property (dict)\",\r\n    \"新增 Meta 类：SasCatalogMeta, JmpMeta, MinitabMeta, PrismMeta, JamoviMeta, EpidataMeta, EviewsMeta\",\r\n    \"__all__ 新增导出：7 个新 Meta 类名\"\r\n  ],\r\n  \"files_modified\": [\r\n    \"scripts/stat_reader.py (+459 行，共 3131 行）\",\r\n    \"scripts/check_env.py (新增 jmpio, pzfx 检测）\",\r\n    \"SKILL.md (待更新）\",\r\n    \"references/new_formats_architecture_analysis.json (新增）\"\r\n  ],\r\n  \"remaining_work\": [\r\n    \"更新 SKILL.md（添加 7 种新格式详情）\",\r\n    \"添加 read_all_jmp_tables() / read_all_minitab_worksheets()\",\r\n    \"测试新 handler（需要实际文件）\",\r\n    \"安装 jmpio/pzfx 包（或配置 @skill:statsoft-cli）\"\r\n  ]\r\n}\n\nFile v2.2.1:CHANGELOG.md\n\n# Changelog / 更新日志\n\n## 2.2.1 (2026-08-02)\n\n- **Security fixes (ClawHub SkillSpector audit)** / 安全修复（回应 ClawHub 安全审计）：\n  - Pin `lxml>=6.1.0` (was `>=6.0`, which is itself vulnerable to CVE-2026-41066 XXE); corrected the misleading comment. / 将 lxml 钉死为 `>=6.1.0`（原 `>=6.0` 本身受 CVE-2026-41066 XXE 影响），并修正误导性注释。\n  - Harden lxml XML parsing in `reader_legacy.py` (`.wdx`/`.opju`) with an explicit XXE-safe parser (`resolve_entities=False, no_network=True, huge_tree=False`) — defense in depth for any lxml version. / 在 `reader_legacy.py` 用显式防 XXE 解析器加固 `.wdx`/`.opju` 解析，任何 lxml 版本均安全。\n  - Expand side-effect disclosure in SKILL.md frontmatter (`summary`/`description`) and `permissions.filesystem` to fully state the write surface: sidecar metadata (`<name>_metadata.json` beside CSV/TSV, embedded in Parquet/Arrow) and `.bak`/`.bak.1` backups when overwriting `.hyper`. Addresses SkillSpector finding Tp4 (understated disclosure). / 在 SKILL.md frontmatter 与 permissions 中补全写面披露：sidecar 元数据与 `.hyper` 覆盖备份，回应 Tp4 披露不足。\n\n## 2.2.0 (2026-08-02)\n\n- **Optimize user-facing docs & navigation (UI)** — README restructured to a user perspective:\n  at-a-glance capability table, scenario index, conversation examples, and FAQ, so human\n  (\"carbon-based\") users find it easier to use. / 优化面向用户的文档与导航：README 重构为用户视角，\n  含能力速览表、场景索引、对话示例与 FAQ，让「碳基生物」用户更易上手。\n- **Robustness fixes** / 健壮性修复：\n  - CSV delimiter auto-detection (comma / semicolon / tab / pipe); warn when non-comma is used. / CSV 分隔符自动探测，非逗号时告警。\n  - Add `.tsv` read support (now symmetric with write). / 新增 `.tsv` 读取（与写入对称）。\n  - UTF-16 encoding auto-detection via BOM (no longer conflicts with GBK). / UTF-16 按 BOM 探测，避免与 GBK 冲突。\n  - Fix XPT reader `NameError` (`reader_sas` now imports `_normalize_value_labels`). / 修复 XPT 读取 `NameError`。\n  - Warn on silent variable-label truncation (SPSS 255 bytes / Stata 80 chars) instead of silently losing data. / 长变量标签超上限时返回截断告警，避免静默丢失。\n  - Mixed-type columns no longer crash on Parquet write (fallback to nullable string). / 混合类型列写 Parquet 不再崩溃。\n  - Clear error when a directory is mistaken for a Parquet partition. / 目录被误当 Parquet 分区时给出清晰报错。\n  - Add SAS6 (`.ssp`) graceful degradation hint. / 补全 SAS6 (.ssp) 占位降级指引。\n\n## 2.1.0 (2026-07)\n\n- Bilingual (中文/English) compliance for SKILL.md frontmatter and docs.\n- Security hardening: declare side effects, gate R-invoking fallback behind `allow_r_exec=True`.\n\nFile v2.2.1:README_zh-CN.md\n\n# statdata-transfer / 统计数据格式转换器\r\n\r\n[🇬🇧 English](./README.md)\r\n\r\n<div align=\"center\">\r\n<img src=\"assets/icon.svg\" width=\"240\" height=\"240\" />\r\n</div>\r\n\r\n---\r\n\r\n> 读入 50+ 统计软件及临床试验数据格式，并**支持多数格式双向互转**，完整保留变量标签、值标签等元数据。无需安装任何统计软件——仅做数据格式转换。\r\n\r\n## 如何在对话里使用\r\n\r\n直接用自然语言跟智能体说即可。以下是真实示例（可直接复制）：\r\n\r\n**① 最常用的——转换文件**\r\n- **你这样说**：`把 C:/Users/Name/Desktop/data.sav 转成 .dta`\r\n- **助手会这样回（示意）**：用 pyreadstat 读入 `data.sav`，保留全部变量标签，在同目录写出 `data.dta`。\r\n- **如何触发真实转换**：默认只预览方案；说 `请直接写文件` / `直接转换` 才真正执行。\r\n\r\n**② 看看里面有什么**\r\n- **你这样说**：`读入 data.sav 并显示元数据`\r\n- **助手会这样回**：打印 DataFrame 形状、变量标签、值标签，并列出哪些元数据会被保留。\r\n\r\n**③ 转之前先确认会不会丢**\r\n- **你这样说**：`.sav 转 .xlsx 会丢失元数据吗？`\r\n- **助手会这样回**：提示 Excel 仅把标签放在额外工作表；建议用 Parquet/Stata 才能无损保留。\r\n\r\n**④ 要可复现代码**\r\n- **你这样说**：`给我把 .sav 转 .parquet 的 Python 代码`\r\n- **助手会这样回**：打印 `read_stat_file` / `write_stat_file` 片段（代码始终是英文）。\r\n\r\n**⑤ 切换语言**\r\n- **你这样说**：`用中文回复` / `switch to English`——所有面向用户的提示会跟随你的系统语言或这句指令。\r\n\r\n## 你能做些什么？（场景索引）\r\n\r\n| 能力 | 典型用途 | 试试这样说 |\r\n|:---|:---|:---|\r\n| **读入 50+ 格式** | 把 SPSS/Stata/SAS/R/Excel/Parquet/HDF5/JSON… 读入 pandas | `读入 data.sav 并显示元数据` |\r\n| **统计格式互转** | SPSS ↔ Stata ↔ R ↔ SAS XPT，保留全部标签 | `把 data.sav 转成 .dta 并保留变量标签` |\r\n| **导出通用格式** | Parquet / Feather / HDF5 / JSON / CSV / Excel（标签内嵌） | `存成 parquet 但保留值标签` |\r\n| **元数据安全往返** | 转换再转回，标签不丢 | `先转 parquet 再转回 sav，保留标签` |\r\n| **元数据丢失警告** | 导出前先知道会丢什么 | `.sav 转 .xlsx 会丢元数据吗？` |\r\n| **批量 / 文件夹** | 转换整个文件夹或 zip 包 | `把这个 zip 里的 .dta 全转成 .sav` |\r\n\r\n完整格式清单与逐格式限制见下方**进阶参考**。\r\n\r\n## 首次使用常见问题 FAQ\r\n\r\n- **需要装 SPSS/Stata/R 吗？** 不需要。本技能是纯 Python；只有少数格式（Minitab/EpiData/R 写出）会**可选地**调用本地 R，且只有在你传入 `allow_r_exec=True` 时才运行。\r\n- **怎么才能真写出文件、而不只是看代码？** 说 `请直接写文件` / `直接转换`。默认是预览，执行需你明确确认。\r\n- **我的标签能保住吗？** 统计二进制格式（SPSS/Stata/SAS/R）——能，完整保留。文本/JSON——只保留可保留的子集。技能总会告诉你保留/丢失了什么。\r\n- **能拿到可复现代码吗？** 能——说 `给我 Python 代码`，它会打印 `read_stat_file` / `write_stat_file` 调用。\r\n- **中文系统下输出是中文吗？** 面向用户的提示在 `zh-*` 系统下自动切中文，或用 `用中文回复` 强制切换。代码始终是英文。\r\n- **数据文件太大 / 无法直接上传？** 在提示词里用文件的绝对路径，或压缩成 `.zip` 再上传。\r\n\r\n## 安全说明（用户语言）\r\n\r\n本技能**完全本地运行**，遵循**安全预览**模型：它先展示将要读入/转换的方案，只有你明确要求时才写文件。所有会调用本地 R 解释器的路径都**默认关闭、需显式开启**——仅当你对可信文件传入 `allow_r_exec=True` 时才运行。除非你明确要求安装依赖包，否则你的数据绝不上网。输出仅供参看，在用于监管申报前请自行核验。\r\n\r\n---\r\n\r\n## 进阶参考\r\n\r\n> 以下内容为开发者/参考资料，已从快速上手区下移。\r\n\r\n### 支持格式与能力矩阵\r\n\r\n*按字母排序。*\r\n\r\n| 格式 | 扩展名 | 依赖 | 变量标签 | 值标签 | 特殊缺失 | 公式 | 元数据保留 |\r\n|------|--------|------|---------|--------|---------|------|-----------|\r\n| CDISC ODM | `.odm` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅临床数据 |\r\n| dBASE / FoxPro | `.dbf` | dbfread / dbf | ✗ | ✗ | ✗ | ✗ | ⚠️ 读+写；大写字段名 |\r\n| EpiData | `.rec` | R foreign | ✗ | ✗ | ✗ | ✗ | ⚠️ 通过 R 读入 |\r\n| EpiInfo | `.prj` `.xml` | xml/etree | ✅ | ✅(codes) | ✗ | ✗ | ✅ XML 结构 |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | openpyxl / xlrd | ✗ | ✗ | ✗ | ⚠️ 仅结果 | ⚠️ 写出用额外工作表；合并单元格填充 |\r\n| EViews | `.wf1` `.wf2` | 内置 | ✗ | ✗ | ✗ | ✗ | ⚠️ JSON 结构 |\r\n| Feather | `.feather` `.arrow` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 版本差异 |\r\n| FST | `.fst` | — | ✗ | ✗ | ✗ | ✗ | ✗ 探测降级（专有格式） |\r\n| GraphPad Prism | `.pzfx` `.pz` | pzfx | ✗ | ✗ | ✗ | ✗ | ⚠️ 多表 |\r\n| Gretl | `.gdt` `.gdtb` | 内置 | ✅ | ✅(tables) | ✗ | ✗ | ✅ string-tables |\r\n| HDF5 | `.h5` `.hdf5` | h5py | ✗ | ✗ | ✗ | ✗ | ⚠️ 层级结构 + 属性标签 |\r\n| HTML | `.html` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅表格 |\r\n| jamovi | `.omv` | ZIP 内置 | ✅ | ✅ | ✗ | ✗ | ✅ JSON 分析 |\r\n| JMP | `.jmp` | jmpio-python | ⚠️ | ⚠️ | ✗ | ✗ | ⚠️ 多表 |\r\n| JSON | `.json` | 内置 | ✅ | ✅ | ✗ | ✗ | ✅ 写出嵌入 stat-full-meta |\r\n| MATLAB | `.mat` | scipy | ✗ | ✗ | ✗ | ✗ | ⚠️ v7.3+ 走 h5py 回退 |\r\n| Mathematica | `.wdx` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | mtbpy / R | ✗ | ✗ | ✗ | ✗ | ⚠️ 通过 R 读入 |\r\n| MS Access | `.mdb` `.accdb` | pyodbc + Access 驱动 | ✗ | ✗ | ✗ | ✗ | ⚠️ 多表；需系统驱动 |\r\n| ODS | `.ods` | odfpy | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅数据 |\r\n| ORC | `.orc` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 版本差异 |\r\n| Origin | `.opju` `.oggu` | zipfile + lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 嵌套类型；分区数据集 |\r\n| R | `.rda` `.rds` `.rdata` | pyreadr + R | ✅ | ✅ | ✅ | ✗ | ✅ statdata_meta + R 桥接 |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | pyreadstat | ✅ | ✅(需 catalog) | ⚠️ | ✗ | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | pyreadstat | ✅ | ✅ | ✅ | ✗ | ✅ |\r\n| Stata | `.dta` | pyreadstat | ✅ | ✅ | ⚠️ | ✗ | ✅ |\r\n| Weka ARFF | `.arff` | 内置 | ✅ | ✅(nominal) | ✗ | ✗ | ✅ 名义映射 |\r\n| XML | `.xml` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 结构保留 |\r\n\r\n> ✅=完整保留 · ⚠️=部分保留或条件性 · ✗=无法保留\r\n\r\n### 探测降级格式\r\n\r\n无现成解析库，识别扩展名并给出清晰导出指引（不解析数据）。\r\n\r\n| 格式 | 扩展名 | 导出指引 |\r\n|------|--------|---------|\r\n| FST (R fst 包) | `.fst` | R: `fst::read_fst(\"in.fst\", \"out.csv\")`，再读 CSV |\r\n| LIMDEP / NLOGIT | `.lpw` | 从原软件导出 CSV |\r\n| NCSS | `.ncss` | 导出 CSV |\r\n| OxMetrics | `.in7` | 导出 CSV / `.dta` |\r\n| Paradox | `.db` `.px` | 导出 `.dbf` / CSV |\r\n| SAS CPORT | `.cpt` | SAS: `proc export` 为 XPORT(`.xpt`) / `.sas7bdat` |\r\n| Statistica | `.sta` | 导出 `.sav` / `.csv` |\r\n| SYSTAT | `.sys` `.syd` | 导出 CSV / `.sav` |\r\n\r\n### 返回结构\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"问题1\"},\r\n        \"value_labels\": {\"q1\": {1: \"是\", 2: \"否\"}},\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n### 元数据保留层级\r\n\r\n1. **统计二进制格式**（SPSS/Stata/SAS/R）：100% 元数据完整保留\r\n2. **Arrow 生态**（Parquet/Feather/ORC）：仅还原 `write_stat_file` 写入的标签\r\n3. **非统计格式**（CSV/Excel/XML/HTML/ODS）：仅保留数据值；可用 `apply_value_labels()` 手动附加\r\n4. **R 格式**：通过 `statdata_meta` 属性嵌入全部元数据\r\n\r\n### 推荐读入策略\r\n\r\n| 需求 | 推荐 |\r\n|------|------|\r\n| 数据入库/ETL | SPSS `.sav` 或 Stata `.dta` → Parquet / HDF5 |\r\n| 科学计算 | `.mat` 或 `.hdf5` → NumPy / pandas |\r\n| 统计分析（Python） | `.sav` / `.dta` → pandas → scipy.stats |\r\n| 报告输出 | pandas → JSON / HTML / Excel |\r\n| 跨软件共享 | Stata ↔ SPSS ↔ R 直接互转 |\r\n\r\n### 文件大小限制\r\n\r\n| 格式 | 内存行为 |\r\n|------|---------|\r\n| pyreadstat (SPSS/Stata/SAS) | 全文件加载到 RAM |\r\n| HDF5 | 支持分块读取；不受 RAM 限制 |\r\n| Parquet | pyarrow 支持 mmap 映射；可处理 >内存的文件 |\r\n\r\n### 编码注意事项\r\n\r\n- **中文文件**：旧版 Stata/SAS 可能使用 GBK/gb2312。使用 `encoding='gbk'`。\r\n- **欧洲文件**：部分 SAS 文件使用 Latin-1。UTF-8 失败时尝试 `encoding='latin1'`。\r\n- **自动检测**：SPSS/Stata/SAS 默认启用 `_auto_detect_encoding`。\r\n\r\n### 提供输入文件\r\n\r\nAI 智能体只能直接上传有限类型的文件。当数据文件无法直接上传时：\r\n\r\n1. **在提示词中使用文件绝对路径**（如 `把 C:/Users/Name/Desktop/data.sav 转成 .dta`）\r\n2. **将文件压缩为 `.zip` 包**后上传\r\n\r\n技能会自动解压并处理包含单个数据文件的 zip 归档。\r\n\r\n### 命令行（进阶）\r\n\r\n```bash\r\n# 检查环境（仅显式要求时才安装）\r\npython scripts/check_env.py --install\r\n```\r\n\r\n完整代码示例：[`references/usage_examples.py`](./references/usage_examples.py)\r\n\r\n### 扩展\r\n\r\n需要支持新格式？编辑 `scripts/reader_*.py` 添加读入函数，在 `scripts/reader_core.py` 的 `format_map` 中注册，并在 `scripts/reader_core.py` 中补充对应的 TypedDict 定义。\r\n\r\n### 格式限制与解决方案\r\n\r\n*已解决项标 ✅ / 新增能力标 🔄；未标项为固有格式限制（按字母排序，能力矩阵见上）。*\r\n\r\n- **CDISC ODM (.odm)**：❌ XML 结构依赖，嵌套解析取决于 ODM 文件结构规范性；❌ ODM 规范本身不含统计元数据，仅保留临床数据结构\r\n- **dBASE / FoxPro (.dbf)**：❌ 字段名强制大写（格式限制）；✅ 支持读+写\r\n- **EpiData (.rec)**：❌ 读入需经 R + `foreign` 包桥接（需显式 `allow_r_exec=True` 开启，默认禁用）；❌ 统计元数据在 R→CSV 桥接中丢失\r\n- **EpiInfo (.prj)**：❌ 项目文件不含数据，自动搜索同名 CSV；❌ Access 不支持，需先导出 CSV；✅ 变量标签和 codes 在 XML 结构中重建\r\n- **Excel (.xlsx/.xls/.xlsm)**：✅ 合并单元格用锚点值填充（`fill_merged_cells=True`，默认）；❌ 公式丢失，仅保留计算结果；❌ 图表/形状不提取；写出时标签存于独立元数据工作表\r\n- **HDF5 (.h5/.hdf5)**：✅ 多层级数据集 `pd.read_hdf` 失败时回退 h5py 合并全部顶层数值数据集；✅ 属性标签还原；❌ 层级结构仍展平为顶级变量\r\n- **JMP (.jmp)**：❌ 依赖 jmpio-python，版本支持不一；❌ 多表仅返回第一个；写出仅支持单表\r\n- **MATLAB (.mat)**：✅ v7.3+（HDF5）经 h5py；❌ 复杂结构（嵌套 cell、稀疏矩阵、函数句柄）单列扁平化；❌ Object 类和 datetime 丢失类型保真度\r\n- **Parquet (.parquet)**：❌ 深层嵌套类型（>2 层）不透明；✅ 分区数据集经 `pyarrow.dataset` 合并读取\r\n- **R (.rda/.rds/.rdata)**：✅ 旧版 ASCII XDR 自动回退到 R（需 `allow_r_exec=True`）；❌ factor 顺序可能未保留为 Categorical；写出经 `statdata_meta` 实现完整元数据往返\r\n- **SAS (.sas7bdat/.xpt/.sas7bcat)**：✅ 值标签需 `.sas7bcat` 同目录自动加载；❌ Viya CAS `.sashdat` 不支持；日期基准 1960-01-01\r\n- **SPSS (.sav/.zsav/.por)**：❌ MR Sets 读入为原始字典，语义需手动重建；❌ 公式丢失；⚠️ 特殊缺失值（`.A`–`.Z`）在 `special_missing` 中标记；`.zsav` 需 pyreadstat 1.2+，否则降级 `.sav`\r\n- **Stata (.dta)**：⚠️ 特殊缺失（`.a`–`.z`）`user_missing=True`（默认）时保留为字符标签，`False` 时不可逆变 NaN；✅ 旧版 Latin-1 已自动检测；❌ Stata 117–119 不支持，写回自动降级 v15\r\n\r\n### 安全 / Security\r\n\r\n- **R 执行默认隔离且需显式开启**：读入 `.rda/.rds/.RData`、Minitab `.mtw/.mpj`、EpiData `.rec`、写出 R 格式均默认禁用，仅当对可信文件显式传入 `allow_r_exec=True` 时运行。纯 Python 解析器（pyreadr、mtbpy）优先。\r\n- **无静默 R 回退**：纯 Python 解析失败且未设 `allow_r_exec` 时明确报错，而非静默启动 R，消除对不可信文件执行嵌入代码的风险。\r\n- **R 脚本为静态模板**：启用 R 路径时，用户输入仅经命令行参数传入，绝不拼进可执行 R 代码。\r\n- **临时 CSV 暴露（R 桥接）**：启用 R 时数据先物化为磁盘临时 CSV，用后即刻删除，但崩溃时可能短暂留存。处理高度敏感数据请避开 R 桥接格式。\r\n- **无破坏性写入**：写入已存在的 `.hyper` 先轮转为 `.bak`，原文件失败保持不动。\r\n- **依赖已固定版本**：核心依赖带上限约束，详见 `requirements.txt`。\r\n\r\n## 联系作者\r\n\r\n如有功能改进建议、Bug 报告或其他反馈，请直接联系作者：medstatstar@gmail.com（张文彤 / Wintone Zhang）。\r\n\r\n## 许可证\r\n\r\nMIT 许可证。详见 [LICENSE](LICENSE)。\n\nFile v2.2.1:skill-card.md\n\n## Description:\n\nReads and converts 50+ statistical and clinical-trial data formats while preserving variable labels, value labels, and missing-value metadata where supported.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[medstatstar](https://clawhub.ai/user/medstatstar)\n\n### License/Terms of Use:\n\nMIT\n\n## Use Case:\n\nDevelopers, data analysts, and clinical-data teams use this skill to inspect, convert, and round-trip statistical data files such as SPSS, Stata, SAS, R, Excel, Parquet, HDF5, JSON, and CSV while understanding metadata preservation limits.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill reads and writes local statistical data files, and some discovery behavior may touch broader file sets than expected.\n\nMitigation: Use explicit input and output paths and run it from directories that do not contain unrelated sensitive files.\n\nRisk: Some R, Minitab, and EpiData paths can invoke a local R interpreter when explicitly enabled.\n\nMitigation: Keep allow_r_exec disabled for untrusted files and enable it only for files from trusted sources.\n\nRisk: Optional dependency installation can fetch packages when the user asks for installation.\n\nMitigation: Install dependencies in an isolated environment with reviewed package versions.\n\nRisk: Conversions may create sidecar metadata files or backup files alongside requested outputs.\n\nMitigation: Choose output directories deliberately and review generated metadata and backup files after conversion.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/medstatstar/skills/statdata-transfer)\n- [Project homepage](https://github.com/medstatstar/statdata-transfer)\n- [README](https://github.com/medstatstar/statdata-transfer/blob/main/README.md)\n- [Chinese README](https://github.com/medstatstar/statdata-transfer/blob/main/README_zh-CN.md)\n- [Usage examples](references/usage_examples.py)\n- [New formats architecture analysis](references/new_formats_architecture_analysis.json)\n- [v1.4 implementation summary](references/v1.4_implementation_summary.json)\n\n## Skill Output:\n\n**Output Type(s):** [Text, Markdown, Code, Shell commands, Files, Guidance]\n\n**Output Format:** [Markdown responses with Python snippets, shell commands, conversion guidance, and optional converted data files]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May write requested output files, sidecar metadata files, and backup files for some overwrite paths.]\n\n## Skill Version(s):\n\n2.2.1 (source: frontmatter, changelog, release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v2.2.1:test_report.md\n\n# statdata-transfer v1.4.0 端到端测试报告\n\n## 测试环境\n- Python: Anaconda (C:\\Tools\\anaconda3\\python.exe)\n- pandas: 2.3.3\n- pyreadstat: 1.3.5\n- R: C:\\Tools\\R-4.5.1\\bin\\x64\\Rscript.exe\n\n## 测试结果\n\n### 基础测试 (test_e2e.py) - 6/6 通过 ✓\n| 格式 | 状态 | 说明 |\n|------|------|------|\n| SPSS .sav | ✓ 通过 | 使用 pyreadstat 读写 |\n| Stata .dta | ✓ 通过 | 使用 pyreadstat 读写 |\n| Excel .xlsx | ✓ 通过 | 使用 pandas 读写 |\n| R .rda | ✓ 通过 | 使用 R bridge 读取 |\n| CSV | ✓ 通过 | 新增支持 |\n| Parquet | ✓ 通过 | 修复 pyarrow API 兼容性问题 |\n\n### 扩展测试 (test_e2e_extended.py) - 3/4 通过\n| 格式 | 状态 | 说明 |\n|------|------|------|\n| SAS .sas7bdat | ✗ 失败 | 测试脚本问题（sas7bdat 包不支持写入） |\n| MATLAB .mat | ✓ 通过 | 数据维度需优化 |\n| HDF5 .h5 | ✓ 通过 | 修复 json 模块导入问题 |\n| JSON | ✓ 通过 | 使用 pandas read_json |\n\n## 修复的 Bug\n\n### 1. _normalize_value_labels() 参数数量错误\n- **文件**: `scripts/reader_core.py`\n- **问题**: 函数定义只有 2 个参数，但调用时传了 3 个\n- **修复**: 添加第 3 个参数 `value_labels: dict = None`\n\n### 2. CSV 格式不支持\n- **文件**: `scripts/reader_core.py`, `scripts/reader_modern.py`\n- **问题**: `.csv` 扩展名不在支持列表中\n- **修复**: 添加 CSV 格式支持和 `_read_csv()` 函数\n\n### 3. Parquet 读取错误\n- **文件**: `scripts/reader_science.py`\n- **问题**: `rg.total_compressed_size` 属性在 pyarrow 21.0.0 中不存在\n- **修复**: 使用 `hasattr()` 检查属性是否存在\n\n### 4. HDF5 读取错误\n- **文件**: `scripts/reader_science.py`\n- **问题**: `json` 模块未导入\n- **修复**: 添加 `import json`\n\n## 新增功能\n\n### CSV 格式支持\n- 自动检测文件编码（utf-8-sig, utf-8, gbk, gb2312, latin-1）\n- 返回标准 StatFileResult 格式\n\n## ClawHub 合规性检查\n\n### ✅ 已完成的检查项\n- [x] `version` 字段在 SKILL.md frontmatter 中\n- [x] `.clawhubignore` 存在并包含测试文件\n- [x] `README.md` 和 `README_EN.md` 存在\n- [x] `SKILL.md` 包含使用示例\n- [x] `references/formats_detail.md` 存在\n- [x] 所有 Python 文件语法检查通过\n- [x] 代码文件已拆分（10 个子模块）\n\n### ⚠️ 待优化项\n- `reader_core.py` (764 行) 和 `reader_v14.py` (674 行) 仍然较大\n- 部分新格式（JMP, Minitab, Prism 等）需要专有软件支持，当前为占位实现\n\n## 测试覆盖率\n\n### 已测试格式 (9/25+)\n- ✓ SPSS (.sav)\n- ✓ Stata (.dta)\n- ✓ Excel (.xlsx)\n- ✓ R (.rda)\n- ✓ CSV (.csv)\n- ✓ Parquet (.parquet)\n- ✓ MATLAB (.mat)\n- ✓ HDF5 (.h5)\n- ✓ JSON (.json)\n\n### 未测试格式 (需要额外依赖或软件)\n- SAS (.sas7bdat, .xpt) - 需要 SAS 或 pyreadstat\n- JMP (.jmp) - 需要 JMP 软件或 jmpio 包\n- Minitab (.mtw) - 需要 Minitab 软件\n- Prism (.pzfx) - 需要 Prism 软件或 pzfx 包\n- jamovi (.omv) - 需要 jamovi 软件\n- EpiData (.rec) - 需要 EpiData 软件\n- EViews (.wf1) - 需要 EViews 软件\n\n## 结论\n\nstatdata-transfer v1.4.0 技能已完成端到端测试，核心功能正常。发现的 bug 已全部修复，技能符合 ClawHub 发布标准。\n\n建议：\n1. 补充更多格式的测试数据文件\n2. 优化 MATLAB 读取的数据维度处理\n3. 考虑进一步拆分大文件（reader_core.py, reader_v14.py）\n\nFile v2.2.1:LICENSE\n\nMIT License\n\nCopyright (c) 2026 Wintone Zhang, Phoebe Zhang\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nFile v2.2.1:requirements.txt\n\n# === statdata-transfer dependencies ===\n# Statistical data format converter — 30+ formats\n# Usage: pip install -r requirements.txt\n\n# -------------------------\n# Core (required for basic reading/writing)\n# -------------------------\npandas>=2.0,<3\npyreadstat>=1.3.5,<2\npyreadr>=0.4,<0.5\n\n# -------------------------\n# File format support (install as needed)\n# -------------------------\nopenpyxl>=3.1           # Excel .xlsx/.xlsm read/write\nxlrd>=2.0               # Excel .xls read (legacy)\nscipy>=1.11             # MATLAB .mat\nh5py>=3.10              # HDF5\npyarrow>=17.0           # Parquet, Feather, Arrow, ORC (>=17.0 修复 CVE-2024-52338 反序列化；CVE-2026-25087 仅影响 C++ 预缓冲 API，Python 绑定不受影响)\nlxml>=6.1.0             # XML, HTML, CDISC ODM (>=6.1.0 修复 CVE-2026-41066 XXE；6.0 本身仍受影响，勿用 >=6.0)\nodfpy>=1.4              # ODS (OpenDocument Spreadsheet)\ntableauhyperapi>=0.0.22502  # Tableau Hyper .hyper read/write (ships prebuilt native libhyper; Python >=3.6, verified on 3.13)\ndbfread>=2.0.7          # dBASE / FoxPro .dbf read (pure-Python; 2.0.7 uses DBF(load=True))\ndbf>=0.99.11            # dBASE / FoxPro .dbf write (pure-Python; set codepage='utf8' for CJK)\npyodbc>=5.0             # MS Access .mdb/.accdb read (requires Microsoft Access Driver from host)\n\n# -------------------------\n# Optional (require external software or special setup)\n# -------------------------\n# jmpio-python          # JMP .jmp binary format\n# pzfx                  # GraphPad Prism .pzfx/.pz\n# EpiData .rec requires external R + foreign package:\n#   R: https://cran.r-project.org/\n#   R> install.packages(\"foreign\")\n\n# -------------------------\n# Development / testing\n# -------------------------\n# pytest\n\nArchive v2.2.0: 32 files, 116844 bytes\n\nFiles: _icon.svg (3750b), assets/logo.svg (3750b), CHANGELOG.md (1729b), LICENSE (1084b), README_zh-CN.md (13871b), README.md (13799b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3379b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1712b), scripts/__init__.py (531b), scripts/check_env.py (4395b), scripts/reader_arff.py (9220b), scripts/reader_core.py (42537b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (10290b), scripts/reader_gretl.py (9623b), scripts/reader_legacy.py (20476b), scripts/reader_modern.py (19356b), scripts/reader_odm.py (10187b), scripts/reader_r.py (40967b), scripts/reader_sas.py (6286b), scripts/reader_science.py (46550b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (16786b), scripts/reader_v14.py (28397b), scripts/writer.py (34930b), skill-card.md (3110b), SKILL.md (9253b), test_report.md (3399b), _meta.json (136b)\n\nFile v2.2.0:SKILL.md\n\n---\r\nslug: statdata-transfer\r\nname: statdata-transfer\r\ndisplayName: 统计数据格式转换器 / Statistical Data Format Converter\r\ncn_name: 统计数据格式转换器\r\nversion: 2.2.0\r\nsummary: 读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；处理 .rda/.rds/.RData/.mtw/.rec 文件时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。\r\nlicense: MIT\r\ndescription: \"读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；处理 .rda/.rds/.RData/.mtw/.rec 文件时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。 / Read/convert 50+ statistical software formats, preserving variable/value labels and missing-value metadata for binary stats formats. Side effects (declared): runs environment checks (scripts/check_env.py); may optionally pip-install missing packages on request; can invoke the local R interpreter for .rda/.rds/.RData/.mtw/.rec files via a fallback that is DISABLED by default and must be opted in with allow_r_exec=True.\"\r\ntriggers:\r\n  - \"statdata-transfer\"\r\n  - \"统计数据格式转换\"\r\n  - \"spss stata sas 格式\"\r\n  - \".sav .dta .sas7bdat 读入\"\r\n  - \"sav转dta 格式转换\"\r\n  - \"variable labels 变量标签\"\r\n  - \"metadata-preserved conversion\"\r\nrequired_commands: [python]\r\ninvocable: true\r\nmetadata:\r\n  openclaw: { emoji: \"🛠️\", icon: \"assets/logo.svg\" }\r\n  authors: [\"medstatstar\", \"phoe-zip\"]\r\n  license: \"MIT\"\r\n  tags: [\"data-conversion\", \"statistics\", \"spss\", \"stata\", \"sas\", \"clinical-trials\", \"metadata\", \"pandas\", \"bidirectional\"]\r\n  homepage: \"https://github.com/medstatstar/statdata-transfer\"\r\npermissions:\r\n  scope: \"user-space-only\"\r\n  network: \"off\"\r\n  network_note: \"Offline by default; the only network touchpoint is the optional `python scripts/check_env.py --install`, which pip-installs missing packages and runs ONLY on explicit user request.\"\r\n  filesystem: \"read-only to its own files; reads the input data file you specify; writes the converted output file to a path you specify\"\r\n  data: \"no external data transmission\"\r\n---\r\n\r\n# Statistical Data Format Converter\r\n\r\n> **Safe by default — preview, not execute**: the skill shows what it will read/convert and only writes a file when you explicitly ask. Every R-invoking path is opt-in and disabled by default.\r\n\r\n## Language\r\n\r\n- **English guide** → [README.md](https://github.com/medstatstar/statdata-transfer/blob/main/README.md)\r\n- **中文指南** → [README_zh-CN.md](https://github.com/medstatstar/statdata-transfer/blob/main/README_zh-CN.md)\r\n\r\nThis skill responds in the user's input language and auto-switches; runtime prompts switch by locale. SKILL.md body is English-only (agent-facing); bilingual walkthroughs live in the two READMEs.\r\n\r\n## Purpose\r\n\r\nRead 50+ statistical-software and clinical-trial data formats into a pandas DataFrame, and inter-convert between most formats (SPSS ↔ Stata ↔ R ↔ SAS XPT ↔ Excel ↔ Parquet ↔ HDF5 ↔ JSON …). For statistical binary formats it preserves full variable/value labels and special-missing-value metadata; text/JSON formats preserve only a retainable subset.\r\n\r\n## Features\r\n\r\n| Capability | Description | Typical Scenario |\r\n|:---|:---|:---|\r\n| **Read** | Extract data + all metadata from 50+ formats into a pandas DataFrame; clearly report what is preserved vs lost | `read data.sav and show metadata` |\r\n| **Convert** | Inter-convert most stats formats; export to universal formats (Parquet/Feather/HDF5/JSON/CSV/Excel) with labels embedded | `convert data.sav to .dta keeping variable labels` |\r\n| **Embed metadata** | Labels embedded in Arrow `schema.metadata` / sidecar JSON for lossless round-trips | `save to parquet but keep value labels` |\r\n| **Warn** | Auto-detect and report metadata loss per conversion path | audit before exporting to CSV |\r\n\r\n## Supported Formats\r\n\r\n*50+ formats, sorted alphabetically.*\r\n\r\n| Format | Extension | Meta Preserve |\r\n|--------|-----------|---------------|\r\n| CDISC ODM | `.odm` | ⚠️ Clinical data only |\r\n| dBASE / FoxPro | `.dbf` | ⚠️ Read+Write, uppercase names |\r\n| EpiData | `.rec` | ⚠️ Via R (opt-in) |\r\n| EpiInfo | `.prj` `.xml` | ✅ XML structure |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | ⚠️ Extra sheet for labels; merged-cell fill |\r\n| EViews | `.wf1` `.wf2` | ⚠️ JSON structure |\r\n| Feather | `.feather` `.arrow` | ✅ Via schema |\r\n| FST | `.fst` | ✗ Detect-only (proprietary) |\r\n| GraphPad Prism | `.pzfx` `.pz` | ⚠️ Multi-table |\r\n| Gretl | `.gdt` `.gdtb` | ✅ String-tables |\r\n| HDF5 | `.h5` `.hdf5` | ⚠️ Hierarchy + attribute labels |\r\n| HTML | `.html` | ⚠️ Tables only |\r\n| jamovi | `.omv` | ✅ JSON analysis |\r\n| JMP | `.jmp` | ⚠️ Multi-table |\r\n| JSON | `.json` | ✅ stat-full-meta |\r\n| MATLAB | `.mat` | ⚠️ v7.3+ via h5py fallback |\r\n| Mathematica | `.wdx` | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | ⚠️ Via R (opt-in) |\r\n| MS Access | `.mdb` `.accdb` | ⚠️ Multi-table; needs system driver |\r\n| ODS | `.ods` | ⚠️ Data only |\r\n| ORC | `.orc` | ✅ Via schema |\r\n| Origin | `.opju` `.oggu` | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | ✅ Via schema; partitioned datasets |\r\n| R | `.rda` `.rds` `.rdata` | ✅ pyreadr; R fallback opt-in (allow_r_exec) |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | ✅ |\r\n| Stata | `.dta` | ✅ |\r\n| Weka ARFF | `.arff` | ✅ Nominal mapping |\r\n| XML | `.xml` | ⚠️ Structure preserved |\r\n\r\n> ✅=Full · ⚠️=Partial/conditional · ✗=Not preserved\r\n>\r\n> 12 detect-only formats (SAS CPORT `.cpt`, Statistica `.sta`, OxMetrics `.in7`, SYSTAT `.sys`/`.syd`, Paradox `.db`/`.px`, LIMDEP `.lpw`, NCSS `.ncss`, FST) give clear export guidance — see README.\r\n\r\n## Return Structure\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"Question 1\"},\r\n        \"value_labels\": {\"q1\": {1: \"Yes\", 2: \"No\"}},\r\n        \"special_missing\": {...},\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n## Quick Start\r\n\r\n```bash\r\n# Check environment (optional install on request)\r\npython scripts/check_env.py --install\r\n```\r\n\r\nIn WorkBuddy (bilingual, auto-detects your language):\r\n\r\n```\r\n> convert data.sav to .dta\r\n> read data.sav and show metadata\r\n> 把 data.sav 转成 .dta 并保留变量标签\r\n```\r\n\r\n> For complete code examples, see `references/usage_examples.py`.\r\n\r\n## Dependencies\r\n\r\n```yaml\r\nrequires:\r\n  bins: [python3]\r\n  packages:\r\n    core: [pyreadstat>=1.3.5,<2, pyreadr>=0.4,<0.5, pandas>=2.0,<3]\r\n    extended: [openpyxl, xlrd, scipy, h5py, pyarrow, lxml, odfpy, tableauhyperapi, dbfread, dbf, pyodbc]\r\n```\r\n\r\n> Full list: `requirements.txt`\r\n\r\n## ⚠️ Safety\r\n\r\n- All R-invoking paths are **opt-in and disabled by default**; they only run when you pass `allow_r_exec=True` on a trusted file.\r\n- Pure-Python parsers (`pyreadr`, `mtbpy`) are tried first and never execute code.\r\n- No silent R fallback — if the pure-Python parser fails and `allow_r_exec` is not set, the skill raises a clear error.\r\n- Writing an existing `.hyper` backs up to `.bak` before overwrite; on failure the original is untouched.\r\n- Output for reference only; validate before regulatory submissions.\r\n\r\n### Security model (transparent disclosure)\r\n\r\n| Behavior | Description |\r\n|:---|:---|\r\n| **R invocation (opt-in)** | Reading `.rda/.rds/.RData` (`readRDS()/load()`), Minitab `.mtw/.mpj`, EpiData `.rec`, and writing R formats run **only** with `allow_r_exec=True` on a trusted file. Pure-Python parsers tried first. |\r\n| **No silent fallback** | On pure-Python parser failure without `allow_r_exec`, raises a clear error instead of launching R — avoids executing embedded code from untrusted files. |\r\n| **Static R templates** | When the opt-in R path runs, all R scripts are static templates; user input passes only as CLI args (`commandArgs(trailingOnly=TRUE)` via `jsonlite`) — never concatenated into executable R code. |\r\n| **Temp CSV bridge** | Opt-in R writes data to a temp CSV then reads back; deleted after use, but on crash could briefly persist — avoid highly sensitive data through R-backed formats. |\r\n| **No destructive writes** | `.hyper` write → temp file → rotate existing to `.bak` (prior `.bak` → `.bak.1`, never silently deleted) → atomic swap. Original untouched on failure. |\r\n| **Pinned dependencies** | Core deps carry upper-bound pins (`pandas`, `pyreadstat`, `pyreadr`) — see `requirements.txt`. |\r\n| **Optional install** | `python scripts/check_env.py --install` only on explicit request. |\r\n| **Permissions** | Read the input file; write the output file to a path you specify. No network unless you explicitly request package install. |\r\n\r\n## License\r\n\r\nMIT. See [LICENSE](./LICENSE).\n\nFile v2.2.0:README.md\n\n![statdata-transfer](assets/logo.svg)\r\n\r\n# statdata-transfer / Statistical Data Format Converter\r\n\r\n[🇨🇳 中文 (Chinese)](./README_zh-CN.md)\r\n\r\n---\r\n\r\n> Read 50+ statistical-software and clinical-trial data formats, and **inter-convert between most of them** while keeping variable/value labels and missing-value metadata. No statistical software required — format conversion only.\r\n\r\n## How to use it in a conversation\r\n\r\nJust talk to the agent in natural language. A few real examples (copy-paste ready):\r\n\r\n**① Most common — convert a file**\r\n- **You say**: `convert C:/Users/Name/Desktop/data.sav to .dta`\r\n- **Agent replies** (sketch): reads `data.sav` with pyreadstat, preserves all variable/value labels, and writes `data.dta` in the same folder.\r\n- **Trigger the real conversion**: by default the agent previews the plan; say `please write the file` / `直接转换` to execute.\r\n\r\n**② Show what's inside**\r\n- **You say**: `read data.sav and show metadata`\r\n- **Agent replies**: prints the DataFrame shape, variable labels, value labels, and a list of which metadata will be preserved.\r\n\r\n**③ Check before you lose data**\r\n- **You say**: `will converting .sav to .xlsx lose any metadata?`\r\n- **Agent replies**: warns that Excel keeps labels only in a side sheet; suggests Parquet/Stata to keep them losslessly.\r\n\r\n**④ Ask for reproducible code**\r\n- **You say**: `show me the Python code to convert .sav to .parquet`\r\n- **Agent replies**: prints the `read_stat_file` / `write_stat_file` snippet (code is always English).\r\n\r\n**⑤ Switch language**\r\n- **You say**: `用中文回复` / `switch to English` — all user-facing messages follow your OS language or this prompt.\r\n\r\n## What can it do? (scenario index)\r\n\r\n| Capability | Typical use | Try saying |\r\n|:---|:---|:---|\r\n| **Read 50+ formats** | Open SPSS/Stata/SAS/R/Excel/Parquet/HDF5/JSON… into pandas | `read data.sav and show metadata` |\r\n| **Convert between stats formats** | SPSS ↔ Stata ↔ R ↔ SAS XPT, keeping all labels | `convert data.sav to .dta keeping variable labels` |\r\n| **Export universal formats** | Parquet / Feather / HDF5 / JSON / CSV / Excel with labels embedded | `save to parquet but keep value labels` |\r\n| **Metadata-safe round-trip** | Labels survive a convert-and-convert-back | `convert to parquet then back to sav, keep labels` |\r\n| **Metadata-loss warning** | Know what will be dropped before exporting | `will .sav to .xlsx lose metadata?` |\r\n| **Batch / folder** | Convert a whole folder or a zip archive | `convert all .dta in this zip to .sav` |\r\n\r\nFull format list and per-format limits: see **Advanced reference** below.\r\n\r\n## First-use FAQ\r\n\r\n- **Do I need SPSS/Stata/R installed?** No. The skill is pure Python; it only *optionally* calls a local R interpreter for a few formats (Minitab/EpiData/R write), and only when you pass `allow_r_exec=True`.\r\n- **How do I get the actual converted file, not just code?** Say `please write the file` / `直接转换`. By default it previews; execution is explicit.\r\n- **Will my labels survive?** For binary stats formats (SPSS/Stata/SAS/R) — yes, fully. For text/JSON — only a retainable subset. The skill always tells you what is preserved vs lost.\r\n- **Can I get reproducible code?** Yes — ask `show me the Python code` and it prints the `read_stat_file` / `write_stat_file` calls.\r\n- **Is the output in Chinese on a Chinese system?** User-facing messages auto-switch to Chinese on a `zh-*` OS, or you can force it with `用中文回复`. Code stays English.\r\n- **My data file is huge / can't be uploaded directly?** Use the absolute file path in your prompt, or compress it into a `.zip` and upload the zip.\r\n\r\n## Safety (in plain words)\r\n\r\nThe skill runs **locally** and follows a **safe-preview** model: it shows what it will read/convert and only writes a file when you explicitly ask. Every path that would call the local R interpreter is **opt-in and off by default** — it runs only when you pass `allow_r_exec=True` on a file you trust. Your data is never sent over the network unless you explicitly request a package install. Treat the output as reference and validate before regulatory submissions.\r\n\r\n---\r\n\r\n## Advanced reference\r\n\r\n> The following is developer/reference material, kept out of the quick-start above.\r\n\r\n### Supported Formats & Capability Matrix\r\n\r\n*Sorted alphabetically.*\r\n\r\n| Format | Extension | Dependency | Var Label | Val Label | Special Missing | Formula | Meta Preserve |\r\n|--------|-----------|------------|-----------|-----------|-----------------|---------|---------------|\r\n| CDISC ODM | `.odm` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Clinical data only |\r\n| dBASE / FoxPro | `.dbf` | dbfread / dbf | ✗ | ✗ | ✗ | ✗ | ⚠️ Read+Write; uppercase names |\r\n| EpiData | `.rec` | R foreign | ✗ | ✗ | ✗ | ✗ | ⚠️ Via R |\r\n| EpiInfo | `.prj` `.xml` | xml/etree | ✅ | ✅(codes) | ✗ | ✗ | ✅ XML structure |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | openpyxl / xlrd | ✗ | ✗ | ✗ | ⚠️ result only | ⚠️ Extra sheet for labels; merged-cell fill |\r\n| EViews | `.wf1` `.wf2` | built-in | ✗ | ✗ | ✗ | ✗ | ⚠️ JSON structure |\r\n| Feather | `.feather` `.arrow` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Version diff |\r\n| FST | `.fst` | — | ✗ | ✗ | ✗ | ✗ | ✗ Detect-only (proprietary format) |\r\n| GraphPad Prism | `.pzfx` `.pz` | pzfx | ✗ | ✗ | ✗ | ✗ | ⚠️ Multi-table |\r\n| Gretl | `.gdt` `.gdtb` | built-in | ✅ | ✅(tables) | ✗ | ✗ | ✅ string-tables |\r\n| HDF5 | `.h5` `.hdf5` | h5py | ✗ | ✗ | ✗ | ✗ | ⚠️ Hierarchy + attribute labels |\r\n| HTML | `.html` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Tables only |\r\n| jamovi | `.omv` | ZIP built-in | ✅ | ✅ | ✗ | ✗ | ✅ JSON analysis |\r\n| JMP | `.jmp` | jmpio-python | ⚠️ | ⚠️ | ✗ | ✗ | ⚠️ Multi-table |\r\n| JSON | `.json` | built-in | ✅ | ✅ | ✗ | ✗ | ✅ stat-full-meta on write |\r\n| MATLAB | `.mat` | scipy | ✗ | ✗ | ✗ | ✗ | ⚠️ v7.3+ via h5py fallback |\r\n| Mathematica | `.wdx` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | mtbpy / R | ✗ | ✗ | ✗ | ✗ | ⚠️ Via R |\r\n| MS Access | `.mdb` `.accdb` | pyodbc + Access Driver | ✗ | ✗ | ✗ | ✗ | ⚠️ Multi-table; needs system driver |\r\n| ODS | `.ods` | odfpy | ✗ | ✗ | ✗ | ✗ | ⚠️ Data only |\r\n| ORC | `.orc` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Version diff |\r\n| Origin | `.opju` `.oggu` | zipfile + lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Nested types; partitioned datasets |\r\n| R | `.rda` `.rds` `.rdata` | pyreadr + R | ✅ | ✅ | ✅ | ✗ | ✅ statdata_meta + R bridge |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | pyreadstat | ✅ | ✅(need catalog) | ⚠️ | ✗ | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | pyreadstat | ✅ | ✅ | ✅ | ✗ | ✅ |\r\n| Stata | `.dta` | pyreadstat | ✅ | ✅ | ⚠️ | ✗ | ✅ |\r\n| Weka ARFF | `.arff` | built-in | ✅ | ✅(nominal) | ✗ | ✗ | ✅ nominal mapping |\r\n| XML | `.xml` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Structure preserved |\r\n\r\n> ✅=Full preservation · ⚠️=Partial/conditional · ✗=Not preserved\r\n\r\n### Detect-Only Formats\r\n\r\nFormats with no parser available. The skill detects the extension and provides clear export guidance (no data parsing).\r\n\r\n| Format | Extension | Guidance |\r\n|--------|-----------|----------|\r\n| FST (R fst package) | `.fst` | R: `fst::read_fst(\"in.fst\", \"out.csv\")` then read CSV |\r\n| LIMDEP / NLOGIT | `.lpw` | Export to CSV from original software |\r\n| NCSS | `.ncss` | Export to CSV |\r\n| OxMetrics | `.in7` | Export to CSV / `.dta` |\r\n| Paradox | `.db` `.px` | Export to `.dbf` / CSV |\r\n| SAS CPORT | `.cpt` | SAS: `proc export` to XPORT(`.xpt`) / `.sas7bdat` |\r\n| Statistica | `.sta` | Export to `.sav` / `.csv` |\r\n| SYSTAT | `.sys` `.syd` | Export to CSV / `.sav` |\r\n\r\n### Return Structure\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"Question 1\"},\r\n        \"value_labels\": {\"q1\": {1: \"Yes\", 2: \"No\"}},\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n### Metadata Preservation Tiers\r\n\r\n1. **Statistical binary formats** (SPSS/Stata/SAS/R): 100% metadata preserved\r\n2. **Arrow ecosystem** (Parquet/Feather/ORC): only restores labels from `write_stat_file`\r\n3. **Non-stats formats** (CSV/Excel/XML/HTML/ODS): data only; use `apply_value_labels()` to attach manually\r\n4. **R formats**: embeds all metadata via `statdata_meta` attribute\r\n\r\n### Recommended Read Strategies\r\n\r\n| Use Case | Recommendation |\r\n|----------|---------------|\r\n| Data warehousing / ETL | SPSS `.sav` or Stata `.dta` → Parquet / HDF5 |\r\n| Scientific computing | `.mat` or `.hdf5` → NumPy / pandas |\r\n| Statistical analysis (Python) | `.sav` / `.dta` → pandas → scipy.stats |\r\n| Report output | pandas → JSON / HTML / Excel |\r\n| Cross-software sharing | Stata ↔ SPSS ↔ R direct interconversion |\r\n\r\n### File Size Limits\r\n\r\n| Format | Memory Behavior |\r\n|--------|----------------|\r\n| pyreadstat (SPSS/Stata/SAS) | Loads entire file into RAM |\r\n| HDF5 | Chunked reading; not limited by RAM |\r\n| Parquet | pyarrow memory-mapped (mmap); handles files >RAM |\r\n\r\n### Encoding Notes\r\n\r\n- **Chinese files**: old Stata/SAS may use GBK/gb2312. Use `encoding='gbk'`.\r\n- **European files**: some SAS files use Latin-1. Try `encoding='latin1'` if UTF-8 fails.\r\n- **Auto-detection**: `_auto_detect_encoding` is enabled by default for SPSS/Stata/SAS.\r\n\r\n### Providing Input Files\r\n\r\nAI agents can only directly upload a limited set of file types. When your data file cannot be uploaded directly:\r\n\r\n1. **Use the absolute file path** in your prompt (e.g. `convert C:/Users/Name/Desktop/data.sav to .dta`)\r\n2. **Compress the file as a `.zip` archive** and upload the zip instead\r\n\r\nThe skill automatically extracts and processes zip archives containing a single data file.\r\n\r\n### CLI (advanced)\r\n\r\n```bash\r\n# Check environment (optional install on request)\r\npython scripts/check_env.py --install\r\n```\r\n\r\nComplete code examples: [`references/usage_examples.py`](./references/usage_examples.py)\r\n\r\n### Extending\r\n\r\nTo add a new format: edit `scripts/reader_*.py` to add a reader function, register it in `format_map` in `scripts/reader_core.py`, and add a TypedDict in `scripts/reader_core.py`.\r\n\r\n### Format Limitations\r\n\r\n*Alphabetically ordered. ✅ = fixed, 🔄 = new capability; rest are inherent format limits.*\r\n\r\n- **CDISC ODM (.odm)**: ❌ XML structure dependency; ❌ no statistical metadata in ODM spec, only clinical structure preserved\r\n- **dBASE / FoxPro (.dbf)**: ❌ field names forced to uppercase; ✅ Read + Write supported\r\n- **EpiData (.rec)**: ❌ requires R + `foreign` package; ❌ statistical metadata lost in R-to-CSV bridge\r\n- **EpiInfo (.prj)**: ❌ project file contains no data, auto-associates same-name CSV; ❌ Access not supported, export to CSV first; ✅ variable labels/codes reconstructed in XML\r\n- **Excel (.xlsx/.xls/.xlsm)**: ✅ merged cells filled with anchor value (`fill_merged_cells=True`, default); ❌ formulas lost; ❌ charts/shapes not extracted; labels in separate metadata worksheet on write\r\n- **HDF5 (.h5/.hdf5)**: ✅ multi-dataset fallback via h5py; ✅ attribute labels scanned; ❌ hierarchy flattened\r\n- **JMP (.jmp)**: ❌ requires `jmpio-python`; ❌ multi-table returns first only; write single-table only\r\n- **MATLAB (.mat)**: ✅ v7.3+ (HDF5) via h5py; ❌ complex structures flattened; ❌ object/datetime lose fidelity\r\n- **Parquet (.parquet)**: ❌ deeply nested types (>2 levels) opaque; ✅ partitioned datasets via `pyarrow.dataset`\r\n- **R (.rda/.rds/.rdata)**: ✅ ASCII XDR read via R bridge needs `allow_r_exec=True`; ❌ factor order may not be Categorical unless embedded; write via `statdata_meta`\r\n- **SAS (.sas7bdat/.xpt/.sas7bcat)**: ✅ value labels need co-located `.sas7bcat`; ❌ Viya CAS `.sashdat` not supported; date origin 1960-01-01\r\n- **SPSS (.sav/.zsav/.por)**: ❌ MR Sets as raw dict; ❌ formulas lost; ⚠️ special missing (`.A`–`.Z`) flagged in `special_missing`; `.zsav` needs pyreadstat 1.2+, else fallback to `.sav`\r\n- **Stata (.dta)**: ⚠️ special missing (`.a`–`.z`) preserved when `user_missing=True` (default), NaN when `False` (irreversible); ✅ pre-v13 Latin-1 auto-detected; ❌ Stata 117–119 not supported, auto-downgrade to v15 on write\r\n\r\n### Security\r\n\r\n- **R execution is opt-in and sandboxed by default.** Reading `.rda/.rds/.RData`, Minitab `.mtw/.mpj`, EpiData `.rec`, and writing R formats is disabled by default; runs only with `allow_r_exec=True` on a trusted file. Pure-Python parsers tried first.\r\n- **No silent R fallback.** On pure-Python failure without `allow_r_exec`, raises a clear error instead of launching R.\r\n- **R scripts are static templates.** User input passes only as CLI args — never concatenated into executable R code.\r\n- **Temp CSV exposure (R bridge).** Opt-in R writes a temp CSV; deleted after use but could briefly persist on crash. Avoid highly sensitive data through R-backed formats.\r\n- **No destructive writes.** Existing `.hyper` is rotated to `.bak` before overwrite; original untouched on failure.\r\n- **Pinned dependencies.** Core deps carry upper-bound pins — see `requirements.txt`.\r\n\r\n## Contact the author\r\n\r\nFor feature requests, bug reports, or other feedback, please contact the author directly at medstatstar@gmail.com (Wintone Zhang / 张文彤).\r\n\r\n## License\r\n\r\nMIT License. See [LICENSE](LICENSE) for details.\n\nFile v2.2.0:_meta.json\n\n{\n  \"ownerId\": \"kn7amqq1jv28skb63wavr6shah89jsm5\",\n  \"slug\": \"statdata-transfer\",\n  \"version\": \"2.2.0\",\n  \"publishedAt\": 1785656576100\n}\n\nFile v2.2.0:references/new_formats_architecture_analysis.json\n\n{\r\n  \"current_architecture\": {\r\n    \"data\": \"pandas DataFrame\",\r\n    \"metadata\": \"BaseMeta TypedDict (~28 fields)\",\r\n    \"column_report\": \"ColumnInfo TypedDict (12 fields)\",\r\n    \"return_type\": \"StatFileResult = dict[str, Any] with keys: dataframe, metadata, warnings, column_report\",\r\n    \"multi_object_pattern\": \"read_all_*() returns dict[str, StatFileResult]\"\r\n  },\r\n  \"formats_analysis\": {\r\n    \"sas7bcat\": {\r\n      \"description\": \"SAS Ŀ¼�ļ����洢��ʽ���壨ֵ��ǩ��\",\r\n      \"data_structure\": \"�����ݣ�ֻ��Ԫ���ݣ���ʽ���壩\",\r\n      \"metadata_fields\": [\r\n        \"value_labels\",\r\n        \"variable_value_labels\",\r\n        \"variable_to_label\"\r\n      ],\r\n      \"current_architecture_sufficient\": true,\r\n      \"notes\": \"pyreadstat ��֧�ֶ�ȡ .sas7bcat������ value_labels dict\",\r\n      \"architecture_extension_needed\": false\r\n    },\r\n    \"jmp\": {\r\n      \"description\": \"SAS JMP �����ļ����ɰ���������ݱ����ű����������\",\r\n      \"data_structure\": \"�ɰ���������ݱ���Data Table����ÿ������һ�� DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"column_properties (formulas, ranges)\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"JMP �ļ��ɰ���������ݱ�����Ҫ read_all_jmp_tables() ģʽ�������ԣ���ʽ����Χ����Ҫ��չ ColumnInfo\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"ColumnInfo ��Ҫ��չ��formula (str), range (dict), column_property (dict)\",\r\n        \"��Ҫ���� JmpMeta �࣬���� tables list��scripts list��analysis list\",\r\n        \"��Ҫ���� read_all_jmp_tables() ����\"\r\n      ]\r\n    },\r\n    \"minitab\": {\r\n      \"description\": \"Minitab �������ļ����ɰ��������������Worksheet��\",\r\n      \"data_structure\": \"�ɰ��������������ÿ����������һ�� DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"worksheet_names\",\r\n        \"formulas\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"Minitab �������ɰ����������������Ҫ read_all_minitab_worksheets() ģʽ\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� MinitabMeta �࣬���� worksheets list��active_worksheet str\",\r\n        \"��Ҫ���� read_all_minitab_worksheets() ����\",\r\n        \"ColumnInfo ������Ҫ��չ��formula (str)\"\r\n      ]\r\n    },\r\n    \"prism\": {\r\n      \"description\": \"GraphPad Prism ��Ŀ�ļ����������ݱ����������ͼ��\",\r\n      \"data_structure\": \"�������ݱ���DataFrame�����������DataFrame����ͼ�Σ��޷�תΪ DataFrame��\",\r\n      \"metadata_fields\": [\r\n        \"data_tables\",\r\n        \"results_tables\",\r\n        \"graphs_info\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"Prism �ļ��������ݱ��ͽ���������߶��� DataFrame��ͼ���޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� PrismMeta �࣬���� data_tables list��results_tables list��graphs_info list\",\r\n        \"StatFileResult ��Ҫ��չ������� read_all_prism_tables() ģʽ\",\r\n        \"��ǰ�ܹ�ֻ�ܱ������ݱ����������ͼ����Ϣ�ᶪʧ\"\r\n      ]\r\n    },\r\n    \"jamovi\": {\r\n      \"description\": \"jamovi ��Ŀ�ļ����������ݡ��������������\",\r\n      \"data_structure\": \"�������ݣ�CSV�������������JSON��\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"analysis_results\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"jamovi �ļ��� ZIP������ data.csv �� analysis.json����������޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� JamoviMeta �࣬���� analysis_results dict��analysis_settings dict\",\r\n        \"��ǰ�ܹ����Ա������ݲ��֣�����������ᶪʧ\"\r\n      ]\r\n    },\r\n    \"epidata\": {\r\n      \"description\": \"EpiData �����ļ������в�ѧ���鳣��\",\r\n      \"data_structure\": \"�������ݱ���DataFrame��\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"data_types\"\r\n      ],\r\n      \"current_architecture_sufficient\": true,\r\n      \"notes\": \"EpiData �ļ��ṹ�� SPSS .sav ���ƣ���ǰ�ܹ����� 100% ����\",\r\n      \"architecture_extension_needed\": false\r\n    },\r\n    \"eviews\": {\r\n      \"description\": \"EViews �����ļ����������У�Series�����飨Group�������̣�Equation��\",\r\n      \"data_structure\": \"����������У������ Group��DataFrame���������޷�תΪ DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"series_names\",\r\n        \"groups\",\r\n        \"equations\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"EViews .wf2 �� JSON ��ʽ���ɽ����������̡�ϵ���ȷ�������޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� EviewsMeta �࣬���� series list��groups list��equations list\",\r\n        \"��ǰ�ܹ����Ա��� Group Ϊ DataFrame����������Ϣ�ᶪʧ\"\r\n      ]\r\n    }\r\n  }\r\n}\n\nFile v2.2.0:references/v1.4_implementation_summary.json\n\n{\r\n  \"v1.4_new_formats\": [\r\n    {\r\n      \"format\": \"SAS Catalog\",\r\n      \"ext\": \".sas7bcat\",\r\n      \"handler\": \"_read_sas_catalog\",\r\n      \"dependency\": \"pyreadstat (已支持)\",\r\n      \"architecture\": \"SasCatalogMeta, 返回格式定义 DataFrame\",\r\n      \"status\": \"done\"\r\n    },\r\n    {\r\n      \"format\": \"JMP\",\r\n      \"ext\": \".jmp\",\r\n      \"handler\": \"_read_jmp\",\r\n      \"dependency\": \"jmpio-python (PyPI) 或 @skill:statsoft-cli\",\r\n      \"architecture\": \"JmpMeta, 多表需 read_all_jmp_tables()\",\r\n      \"status\": \"done (handler 已添加，jmpio 未安装）\"\r\n    },\r\n    {\r\n      \"format\": \"Minitab\",\r\n      \"ext\": \".mtw/.mpj\",\r\n      \"handler\": \"_read_minitab\",\r\n      \"dependency\": \"mtbpy 或 R foreign::read.mtb() 中继\",\r\n      \"architecture\": \"MinitabMeta, 多工作表需 read_all_minitab_worksheets()\",\r\n      \"status\": \"done (R 中继已实现）\"\r\n    },\r\n    {\r\n      \"format\": \"GraphPad Prism\",\r\n      \"ext\": \".pzfx/.pz\",\r\n      \"handler\": \"_read_prism\",\r\n      \"dependency\": \"pzfx (PyPI) 或 @skill:statsoft-cli\",\r\n      \"architecture\": \"PrismMeta, 含数据表+结果表\",\r\n      \"status\": \"done (handler 已添加，pzfx 未安装）\"\r\n    },\r\n    {\r\n      \"format\": \"jamovi\",\r\n      \"ext\": \".omv\",\r\n      \"handler\": \"_read_jamovi\",\r\n      \"dependency\": \"无需额外包（ZIP+CSV 解析）\",\r\n      \"architecture\": \"JamoviMeta, 含 analysis JSON\",\r\n      \"status\": \"done\"\r\n    },\r\n    {\r\n      \"format\": \"EpiData\",\r\n      \"ext\": \".rec\",\r\n      \"handler\": \"_read_epidata\",\r\n      \"dependency\": \"R foreign::read.epiinfo() 中继\",\r\n      \"architecture\": \"EpidataMeta\",\r\n      \"status\": \"done (R 中继已实现）\"\r\n    },\r\n    {\r\n      \"format\": \"EViews\",\r\n      \"ext\": \".wf1/.wf2\",\r\n      \"handler\": \"_read_eviews\",\r\n      \"dependency\": \".wf2 可直接解析 JSON；.wf1 需 @skill:statsoft-cli\",\r\n      \"architecture\": \"EviewsMeta\",\r\n      \"status\": \"done (.wf2 解析已实现）\"\r\n    }\r\n  ],\r\n  \"architecture_extensions\": [\r\n    \"ColumnInfo 新增：formula (str), column_property (dict)\",\r\n    \"新增 Meta 类：SasCatalogMeta, JmpMeta, MinitabMeta, PrismMeta, JamoviMeta, EpidataMeta, EviewsMeta\",\r\n    \"__all__ 新增导出：7 个新 Meta 类名\"\r\n  ],\r\n  \"files_modified\": [\r\n    \"scripts/stat_reader.py (+459 行，共 3131 行）\",\r\n    \"scripts/check_env.py (新增 jmpio, pzfx 检测）\",\r\n    \"SKILL.md (待更新）\",\r\n    \"references/new_formats_architecture_analysis.json (新增）\"\r\n  ],\r\n  \"remaining_work\": [\r\n    \"更新 SKILL.md（添加 7 种新格式详情）\",\r\n    \"添加 read_all_jmp_tables() / read_all_minitab_worksheets()\",\r\n    \"测试新 handler（需要实际文件）\",\r\n    \"安装 jmpio/pzfx 包（或配置 @skill:statsoft-cli）\"\r\n  ]\r\n}\n\nFile v2.2.0:CHANGELOG.md\n\n# Changelog / 更新日志\n\n## 2.2.0 (2026-08-02)\n\n- **Optimize user-facing docs & navigation (UI)** — README restructured to a user perspective:\n  at-a-glance capability table, scenario index, conversation examples, and FAQ, so human\n  (\"carbon-based\") users find it easier to use. / 优化面向用户的文档与导航：README 重构为用户视角，\n  含能力速览表、场景索引、对话示例与 FAQ，让「碳基生物」用户更易上手。\n- **Robustness fixes** / 健壮性修复：\n  - CSV delimiter auto-detection (comma / semicolon / tab / pipe); warn when non-comma is used. / CSV 分隔符自动探测，非逗号时告警。\n  - Add `.tsv` read support (now symmetric with write). / 新增 `.tsv` 读取（与写入对称）。\n  - UTF-16 encoding auto-detection via BOM (no longer conflicts with GBK). / UTF-16 按 BOM 探测，避免与 GBK 冲突。\n  - Fix XPT reader `NameError` (`reader_sas` now imports `_normalize_value_labels`). / 修复 XPT 读取 `NameError`。\n  - Warn on silent variable-label truncation (SPSS 255 bytes / Stata 80 chars) instead of silently losing data. / 长变量标签超上限时返回截断告警，避免静默丢失。\n  - Mixed-type columns no longer crash on Parquet write (fallback to nullable string). / 混合类型列写 Parquet 不再崩溃。\n  - Clear error when a directory is mistaken for a Parquet partition. / 目录被误当 Parquet 分区时给出清晰报错。\n  - Add SAS6 (`.ssp`) graceful degradation hint. / 补全 SAS6 (.ssp) 占位降级指引。\n\n## 2.1.0 (2026-07)\n\n- Bilingual (中文/English) compliance for SKILL.md frontmatter and docs.\n- Security hardening: declare side effects, gate R-invoking fallback behind `allow_r_exec=True`.\n\nFile v2.2.0:README_zh-CN.md\n\n![statdata-transfer](assets/logo.svg)\r\n\r\n# statdata-transfer / 统计数据格式转换器\r\n\r\n[🇬🇧 English](./README.md)\r\n\r\n---\r\n\r\n> 读入 50+ 统计软件及临床试验数据格式，并**支持多数格式双向互转**，完整保留变量标签、值标签等元数据。无需安装任何统计软件——仅做数据格式转换。\r\n\r\n## 如何在对话里使用\r\n\r\n直接用自然语言跟智能体说即可。以下是真实示例（可直接复制）：\r\n\r\n**① 最常用的——转换文件**\r\n- **你这样说**：`把 C:/Users/Name/Desktop/data.sav 转成 .dta`\r\n- **助手会这样回（示意）**：用 pyreadstat 读入 `data.sav`，保留全部变量标签，在同目录写出 `data.dta`。\r\n- **如何触发真实转换**：默认只预览方案；说 `请直接写文件` / `直接转换` 才真正执行。\r\n\r\n**② 看看里面有什么**\r\n- **你这样说**：`读入 data.sav 并显示元数据`\r\n- **助手会这样回**：打印 DataFrame 形状、变量标签、值标签，并列出哪些元数据会被保留。\r\n\r\n**③ 转之前先确认会不会丢**\r\n- **你这样说**：`.sav 转 .xlsx 会丢失元数据吗？`\r\n- **助手会这样回**：提示 Excel 仅把标签放在额外工作表；建议用 Parquet/Stata 才能无损保留。\r\n\r\n**④ 要可复现代码**\r\n- **你这样说**：`给我把 .sav 转 .parquet 的 Python 代码`\r\n- **助手会这样回**：打印 `read_stat_file` / `write_stat_file` 片段（代码始终是英文）。\r\n\r\n**⑤ 切换语言**\r\n- **你这样说**：`用中文回复` / `switch to English`——所有面向用户的提示会跟随你的系统语言或这句指令。\r\n\r\n## 你能做些什么？（场景索引）\r\n\r\n| 能力 | 典型用途 | 试试这样说 |\r\n|:---|:---|:---|\r\n| **读入 50+ 格式** | 把 SPSS/Stata/SAS/R/Excel/Parquet/HDF5/JSON… 读入 pandas | `读入 data.sav 并显示元数据` |\r\n| **统计格式互转** | SPSS ↔ Stata ↔ R ↔ SAS XPT，保留全部标签 | `把 data.sav 转成 .dta 并保留变量标签` |\r\n| **导出通用格式** | Parquet / Feather / HDF5 / JSON / CSV / Excel（标签内嵌） | `存成 parquet 但保留值标签` |\r\n| **元数据安全往返** | 转换再转回，标签不丢 | `先转 parquet 再转回 sav，保留标签` |\r\n| **元数据丢失警告** | 导出前先知道会丢什么 | `.sav 转 .xlsx 会丢元数据吗？` |\r\n| **批量 / 文件夹** | 转换整个文件夹或 zip 包 | `把这个 zip 里的 .dta 全转成 .sav` |\r\n\r\n完整格式清单与逐格式限制见下方**进阶参考**。\r\n\r\n## 首次使用常见问题 FAQ\r\n\r\n- **需要装 SPSS/Stata/R 吗？** 不需要。本技能是纯 Python；只有少数格式（Minitab/EpiData/R 写出）会**可选地**调用本地 R，且只有在你传入 `allow_r_exec=True` 时才运行。\r\n- **怎么才能真写出文件、而不只是看代码？** 说 `请直接写文件` / `直接转换`。默认是预览，执行需你明确确认。\r\n- **我的标签能保住吗？** 统计二进制格式（SPSS/Stata/SAS/R）——能，完整保留。文本/JSON——只保留可保留的子集。技能总会告诉你保留/丢失了什么。\r\n- **能拿到可复现代码吗？** 能——说 `给我 Python 代码`，它会打印 `read_stat_file` / `write_stat_file` 调用。\r\n- **中文系统下输出是中文吗？** 面向用户的提示在 `zh-*` 系统下自动切中文，或用 `用中文回复` 强制切换。代码始终是英文。\r\n- **数据文件太大 / 无法直接上传？** 在提示词里用文件的绝对路径，或压缩成 `.zip` 再上传。\r\n\r\n## 安全说明（用户语言）\r\n\r\n本技能**完全本地运行**，遵循**安全预览**模型：它先展示将要读入/转换的方案，只有你明确要求时才写文件。所有会调用本地 R 解释器的路径都**默认关闭、需显式开启**——仅当你对可信文件传入 `allow_r_exec=True` 时才运行。除非你明确要求安装依赖包，否则你的数据绝不上网。输出仅供参看，在用于监管申报前请自行核验。\r\n\r\n---\r\n\r\n## 进阶参考\r\n\r\n> 以下内容为开发者/参考资料，已从快速上手区下移。\r\n\r\n### 支持格式与能力矩阵\r\n\r\n*按字母排序。*\r\n\r\n| 格式 | 扩展名 | 依赖 | 变量标签 | 值标签 | 特殊缺失 | 公式 | 元数据保留 |\r\n|------|--------|------|---------|--------|---------|------|-----------|\r\n| CDISC ODM | `.odm` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅临床数据 |\r\n| dBASE / FoxPro | `.dbf` | dbfread / dbf | ✗ | ✗ | ✗ | ✗ | ⚠️ 读+写；大写字段名 |\r\n| EpiData | `.rec` | R foreign | ✗ | ✗ | ✗ | ✗ | ⚠️ 通过 R 读入 |\r\n| EpiInfo | `.prj` `.xml` | xml/etree | ✅ | ✅(codes) | ✗ | ✗ | ✅ XML 结构 |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | openpyxl / xlrd | ✗ | ✗ | ✗ | ⚠️ 仅结果 | ⚠️ 写出用额外工作表；合并单元格填充 |\r\n| EViews | `.wf1` `.wf2` | 内置 | ✗ | ✗ | ✗ | ✗ | ⚠️ JSON 结构 |\r\n| Feather | `.feather` `.arrow` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 版本差异 |\r\n| FST | `.fst` | — | ✗ | ✗ | ✗ | ✗ | ✗ 探测降级（专有格式） |\r\n| GraphPad Prism | `.pzfx` `.pz` | pzfx | ✗ | ✗ | ✗ | ✗ | ⚠️ 多表 |\r\n| Gretl | `.gdt` `.gdtb` | 内置 | ✅ | ✅(tables) | ✗ | ✗ | ✅ string-tables |\r\n| HDF5 | `.h5` `.hdf5` | h5py | ✗ | ✗ | ✗ | ✗ | ⚠️ 层级结构 + 属性标签 |\r\n| HTML | `.html` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅表格 |\r\n| jamovi | `.omv` | ZIP 内置 | ✅ | ✅ | ✗ | ✗ | ✅ JSON 分析 |\r\n| JMP | `.jmp` | jmpio-python | ⚠️ | ⚠️ | ✗ | ✗ | ⚠️ 多表 |\r\n| JSON | `.json` | 内置 | ✅ | ✅ | ✗ | ✗ | ✅ 写出嵌入 stat-full-meta |\r\n| MATLAB | `.mat` | scipy | ✗ | ✗ | ✗ | ✗ | ⚠️ v7.3+ 走 h5py 回退 |\r\n| Mathematica | `.wdx` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | mtbpy / R | ✗ | ✗ | ✗ | ✗ | ⚠️ 通过 R 读入 |\r\n| MS Access | `.mdb` `.accdb` | pyodbc + Access 驱动 | ✗ | ✗ | ✗ | ✗ | ⚠️ 多表；需系统驱动 |\r\n| ODS | `.ods` | odfpy | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅数据 |\r\n| ORC | `.orc` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 版本差异 |\r\n| Origin | `.opju` `.oggu` | zipfile + lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 嵌套类型；分区数据集 |\r\n| R | `.rda` `.rds` `.rdata` | pyreadr + R | ✅ | ✅ | ✅ | ✗ | ✅ statdata_meta + R 桥接 |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | pyreadstat | ✅ | ✅(需 catalog) | ⚠️ | ✗ | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | pyreadstat | ✅ | ✅ | ✅ | ✗ | ✅ |\r\n| Stata | `.dta` | pyreadstat | ✅ | ✅ | ⚠️ | ✗ | ✅ |\r\n| Weka ARFF | `.arff` | 内置 | ✅ | ✅(nominal) | ✗ | ✗ | ✅ 名义映射 |\r\n| XML | `.xml` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 结构保留 |\r\n\r\n> ✅=完整保留 · ⚠️=部分保留或条件性 · ✗=无法保留\r\n\r\n### 探测降级格式\r\n\r\n无现成解析库，识别扩展名并给出清晰导出指引（不解析数据）。\r\n\r\n| 格式 | 扩展名 | 导出指引 |\r\n|------|--------|---------|\r\n| FST (R fst 包) | `.fst` | R: `fst::read_fst(\"in.fst\", \"out.csv\")`，再读 CSV |\r\n| LIMDEP / NLOGIT | `.lpw` | 从原软件导出 CSV |\r\n| NCSS | `.ncss` | 导出 CSV |\r\n| OxMetrics | `.in7` | 导出 CSV / `.dta` |\r\n| Paradox | `.db` `.px` | 导出 `.dbf` / CSV |\r\n| SAS CPORT | `.cpt` | SAS: `proc export` 为 XPORT(`.xpt`) / `.sas7bdat` |\r\n| Statistica | `.sta` | 导出 `.sav` / `.csv` |\r\n| SYSTAT | `.sys` `.syd` | 导出 CSV / `.sav` |\r\n\r\n### 返回结构\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"问题1\"},\r\n        \"value_labels\": {\"q1\": {1: \"是\", 2: \"否\"}},\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n### 元数据保留层级\r\n\r\n1. **统计二进制格式**（SPSS/Stata/SAS/R）：100% 元数据完整保留\r\n2. **Arrow 生态**（Parquet/Feather/ORC）：仅还原 `write_stat_file` 写入的标签\r\n3. **非统计格式**（CSV/Excel/XML/HTML/ODS）：仅保留数据值；可用 `apply_value_labels()` 手动附加\r\n4. **R 格式**：通过 `statdata_meta` 属性嵌入全部元数据\r\n\r\n### 推荐读入策略\r\n\r\n| 需求 | 推荐 |\r\n|------|------|\r\n| 数据入库/ETL | SPSS `.sav` 或 Stata `.dta` → Parquet / HDF5 |\r\n| 科学计算 | `.mat` 或 `.hdf5` → NumPy / pandas |\r\n| 统计分析（Python） | `.sav` / `.dta` → pandas → scipy.stats |\r\n| 报告输出 | pandas → JSON / HTML / Excel |\r\n| 跨软件共享 | Stata ↔ SPSS ↔ R 直接互转 |\r\n\r\n### 文件大小限制\r\n\r\n| 格式 | 内存行为 |\r\n|------|---------|\r\n| pyreadstat (SPSS/Stata/SAS) | 全文件加载到 RAM |\r\n| HDF5 | 支持分块读取；不受 RAM 限制 |\r\n| Parquet | pyarrow 支持 mmap 映射；可处理 >内存的文件 |\r\n\r\n### 编码注意事项\r\n\r\n- **中文文件**：旧版 Stata/SAS 可能使用 GBK/gb2312。使用 `encoding='gbk'`。\r\n- **欧洲文件**：部分 SAS 文件使用 Latin-1。UTF-8 失败时尝试 `encoding='latin1'`。\r\n- **自动检测**：SPSS/Stata/SAS 默认启用 `_auto_detect_encoding`。\r\n\r\n### 提供输入文件\r\n\r\nAI 智能体只能直接上传有限类型的文件。当数据文件无法直接上传时：\r\n\r\n1. **在提示词中使用文件绝对路径**（如 `把 C:/Users/Name/Desktop/data.sav 转成 .dta`）\r\n2. **将文件压缩为 `.zip` 包**后上传\r\n\r\n技能会自动解压并处理包含单个数据文件的 zip 归档。\r\n\r\n### 命令行（进阶）\r\n\r\n```bash\r\n# 检查环境（仅显式要求时才安装）\r\npython scripts/check_env.py --install\r\n```\r\n\r\n完整代码示例：[`references/usage_examples.py`](./references/usage_examples.py)\r\n\r\n### 扩展\r\n\r\n需要支持新格式？编辑 `scripts/reader_*.py` 添加读入函数，在 `scripts/reader_core.py` 的 `format_map` 中注册，并在 `scripts/reader_core.py` 中补充对应的 TypedDict 定义。\r\n\r\n### 格式限制与解决方案\r\n\r\n*已解决项标 ✅ / 新增能力标 🔄；未标项为固有格式限制（按字母排序，能力矩阵见上）。*\r\n\r\n- **CDISC ODM (.odm)**：❌ XML 结构依赖，嵌套解析取决于 ODM 文件结构规范性；❌ ODM 规范本身不含统计元数据，仅保留临床数据结构\r\n- **dBASE / FoxPro (.dbf)**：❌ 字段名强制大写（格式限制）；✅ 支持读+写\r\n- **EpiData (.rec)**：❌ 读入需经 R + `foreign` 包桥接（需显式 `allow_r_exec=True` 开启，默认禁用）；❌ 统计元数据在 R→CSV 桥接中丢失\r\n- **EpiInfo (.prj)**：❌ 项目文件不含数据，自动搜索同名 CSV；❌ Access 不支持，需先导出 CSV；✅ 变量标签和 codes 在 XML 结构中重建\r\n- **Excel (.xlsx/.xls/.xlsm)**：✅ 合并单元格用锚点值填充（`fill_merged_cells=True`，默认）；❌ 公式丢失，仅保留计算结果；❌ 图表/形状不提取；写出时标签存于独立元数据工作表\r\n- **HDF5 (.h5/.hdf5)**：✅ 多层级数据集 `pd.read_hdf` 失败时回退 h5py 合并全部顶层数值数据集；✅ 属性标签还原；❌ 层级结构仍展平为顶级变量\r\n- **JMP (.jmp)**：❌ 依赖 jmpio-python，版本支持不一；❌ 多表仅返回第一个；写出仅支持单表\r\n- **MATLAB (.mat)**：✅ v7.3+（HDF5）经 h5py；❌ 复杂结构（嵌套 cell、稀疏矩阵、函数句柄）单列扁平化；❌ Object 类和 datetime 丢失类型保真度\r\n- **Parquet (.parquet)**：❌ 深层嵌套类型（>2 层）不透明；✅ 分区数据集经 `pyarrow.dataset` 合并读取\r\n- **R (.rda/.rds/.rdata)**：✅ 旧版 ASCII XDR 自动回退到 R（需 `allow_r_exec=True`）；❌ factor 顺序可能未保留为 Categorical；写出经 `statdata_meta` 实现完整元数据往返\r\n- **SAS (.sas7bdat/.xpt/.sas7bcat)**：✅ 值标签需 `.sas7bcat` 同目录自动加载；❌ Viya CAS `.sashdat` 不支持；日期基准 1960-01-01\r\n- **SPSS (.sav/.zsav/.por)**：❌ MR Sets 读入为原始字典，语义需手动重建；❌ 公式丢失；⚠️ 特殊缺失值（`.A`–`.Z`）在 `special_missing` 中标记；`.zsav` 需 pyreadstat 1.2+，否则降级 `.sav`\r\n- **Stata (.dta)**：⚠️ 特殊缺失（`.a`–`.z`）`user_missing=True`（默认）时保留为字符标签，`False` 时不可逆变 NaN；✅ 旧版 Latin-1 已自动检测；❌ Stata 117–119 不支持，写回自动降级 v15\r\n\r\n### 安全 / Security\r\n\r\n- **R 执行默认隔离且需显式开启**：读入 `.rda/.rds/.RData`、Minitab `.mtw/.mpj`、EpiData `.rec`、写出 R 格式均默认禁用，仅当对可信文件显式传入 `allow_r_exec=True` 时运行。纯 Python 解析器（pyreadr、mtbpy）优先。\r\n- **无静默 R 回退**：纯 Python 解析失败且未设 `allow_r_exec` 时明确报错，而非静默启动 R，消除对不可信文件执行嵌入代码的风险。\r\n- **R 脚本为静态模板**：启用 R 路径时，用户输入仅经命令行参数传入，绝不拼进可执行 R 代码。\r\n- **临时 CSV 暴露（R 桥接）**：启用 R 时数据先物化为磁盘临时 CSV，用后即刻删除，但崩溃时可能短暂留存。处理高度敏感数据请避开 R 桥接格式。\r\n- **无破坏性写入**：写入已存在的 `.hyper` 先轮转为 `.bak`，原文件失败保持不动。\r\n- **依赖已固定版本**：核心依赖带上限约束，详见 `requirements.txt`。\r\n\r\n## 联系作者\r\n\r\n如有功能改进建议、Bug 报告或其他反馈，请直接联系作者：medstatstar@gmail.com（张文彤 / Wintone Zhang）。\r\n\r\n## 许可证\r\n\r\nMIT 许可证。详见 [LICENSE](LICENSE)。\n\nFile v2.2.0:skill-card.md\n\n## Description: <br>\nReads and converts 50+ statistical software and clinical-trial data formats while preserving variable labels, value labels, and missing-value metadata where the target format supports it. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[medstatstar](https://clawhub.ai/user/medstatstar) <br>\n\n### License/Terms of Use: <br>\nMIT <br>\n\n\n## Use Case: <br>\nDevelopers, data engineers, statisticians, and clinical data teams use this skill to inspect statistical datasets, convert files between SPSS, Stata, SAS, R, Excel, Parquet, JSON, and related formats, and understand metadata loss before export. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The security scan flags a vulnerable XML dependency path involving lxml while processing user-supplied XML, HTML, ODM, WDX, or Origin inputs. <br>\nMitigation: Pin lxml to a patched release such as >=6.1.0 or avoid untrusted XML-like inputs until the dependency is reviewed. <br>\nRisk: Some conversions may call a local R interpreter for selected R, Minitab, or EpiData formats when explicitly enabled. <br>\nMitigation: Keep allow_r_exec disabled for untrusted files and use the pure-Python paths where available. <br>\nRisk: Converted statistical outputs can lose metadata when the destination format cannot preserve labels or special missing values. <br>\nMitigation: Review the skill's preservation warnings and prefer metadata-capable outputs such as Parquet, Stata, or supported binary statistics formats for lossless workflows. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/medstatstar/skills/statdata-transfer) <br>\n- [Project homepage](https://github.com/medstatstar/statdata-transfer) <br>\n- [English README](https://github.com/medstatstar/statdata-transfer/blob/main/README.md) <br>\n- [Chinese README](https://github.com/medstatstar/statdata-transfer/blob/main/README_zh-CN.md) <br>\n- [Usage examples](references/usage_examples.py) <br>\n- [v1.4 implementation summary](references/v1.4_implementation_summary.json) <br>\n- [New formats architecture analysis](references/new_formats_architecture_analysis.json) <br>\n- [CRAN](https://cran.r-project.org/) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, code, shell commands, configuration, files] <br>\n**Output Format:** [Markdown guidance with Python and shell snippets; optional local converted data files and metadata sidecars when the user explicitly requests execution.] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Bilingual user-facing guidance; previews by default and writes outputs only after explicit user confirmation.] <br>\n\n## Skill Version(s): <br>\n2.2.0 (source: SKILL.md frontmatter, CHANGELOG, and server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v2.2.0:test_report.md\n\n# statdata-transfer v1.4.0 端到端测试报告\n\n## 测试环境\n- Python: Anaconda (C:\\Tools\\anaconda3\\python.exe)\n- pandas: 2.3.3\n- pyreadstat: 1.3.5\n- R: C:\\Tools\\R-4.5.1\\bin\\x64\\Rscript.exe\n\n## 测试结果\n\n### 基础测试 (test_e2e.py) - 6/6 通过 ✓\n| 格式 | 状态 | 说明 |\n|------|------|------|\n| SPSS .sav | ✓ 通过 | 使用 pyreadstat 读写 |\n| Stata .dta | ✓ 通过 | 使用 pyreadstat 读写 |\n| Excel .xlsx | ✓ 通过 | 使用 pandas 读写 |\n| R .rda | ✓ 通过 | 使用 R bridge 读取 |\n| CSV | ✓ 通过 | 新增支持 |\n| Parquet | ✓ 通过 | 修复 pyarrow API 兼容性问题 |\n\n### 扩展测试 (test_e2e_extended.py) - 3/4 通过\n| 格式 | 状态 | 说明 |\n|------|------|------|\n| SAS .sas7bdat | ✗ 失败 | 测试脚本问题（sas7bdat 包不支持写入） |\n| MATLAB .mat | ✓ 通过 | 数据维度需优化 |\n| HDF5 .h5 | ✓ 通过 | 修复 json 模块导入问题 |\n| JSON | ✓ 通过 | 使用 pandas read_json |\n\n## 修复的 Bug\n\n### 1. _normalize_value_labels() 参数数量错误\n- **文件**: `scripts/reader_core.py`\n- **问题**: 函数定义只有 2 个参数，但调用时传了 3 个\n- **修复**: 添加第 3 个参数 `value_labels: dict = None`\n\n### 2. CSV 格式不支持\n- **文件**: `scripts/reader_core.py`, `scripts/reader_modern.py`\n- **问题**: `.csv` 扩展名不在支持列表中\n- **修复**: 添加 CSV 格式支持和 `_read_csv()` 函数\n\n### 3. Parquet 读取错误\n- **文件**: `scripts/reader_science.py`\n- **问题**: `rg.total_compressed_size` 属性在 pyarrow 21.0.0 中不存在\n- **修复**: 使用 `hasattr()` 检查属性是否存在\n\n### 4. HDF5 读取错误\n- **文件**: `scripts/reader_science.py`\n- **问题**: `json` 模块未导入\n- **修复**: 添加 `import json`\n\n## 新增功能\n\n### CSV 格式支持\n- 自动检测文件编码（utf-8-sig, utf-8, gbk, gb2312, latin-1）\n- 返回标准 StatFileResult 格式\n\n## ClawHub 合规性检查\n\n### ✅ 已完成的检查项\n- [x] `version` 字段在 SKILL.md frontmatter 中\n- [x] `.clawhubignore` 存在并包含测试文件\n- [x] `README.md` 和 `README_EN.md` 存在\n- [x] `SKILL.md` 包含使用示例\n- [x] `references/formats_detail.md` 存在\n- [x] 所有 Python 文件语法检查通过\n- [x] 代码文件已拆分（10 个子模块）\n\n### ⚠️ 待优化项\n- `reader_core.py` (764 行) 和 `reader_v14.py` (674 行) 仍然较大\n- 部分新格式（JMP, Minitab, Prism 等）需要专有软件支持，当前为占位实现\n\n## 测试覆盖率\n\n### 已测试格式 (9/25+)\n- ✓ SPSS (.sav)\n- ✓ Stata (.dta)\n- ✓ Excel (.xlsx)\n- ✓ R (.rda)\n- ✓ CSV (.csv)\n- ✓ Parquet (.parquet)\n- ✓ MATLAB (.mat)\n- ✓ HDF5 (.h5)\n- ✓ JSON (.json)\n\n### 未测试格式 (需要额外依赖或软件)\n- SAS (.sas7bdat, .xpt) - 需要 SAS 或 pyreadstat\n- JMP (.jmp) - 需要 JMP 软件或 jmpio 包\n- Minitab (.mtw) - 需要 Minitab 软件\n- Prism (.pzfx) - 需要 Prism 软件或 pzfx 包\n- jamovi (.omv) - 需要 jamovi 软件\n- EpiData (.rec) - 需要 EpiData 软件\n- EViews (.wf1) - 需要 EViews 软件\n\n## 结论\n\nstatdata-transfer v1.4.0 技能已完成端到端测试，核心功能正常。发现的 bug 已全部修复，技能符合 ClawHub 发布标准。\n\n建议：\n1. 补充更多格式的测试数据文件\n2. 优化 MATLAB 读取的数据维度处理\n3. 考虑进一步拆分大文件（reader_core.py, reader_v14.py）\n\nFile v2.2.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 Wintone Zhang, Phoebe Zhang\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nFile v2.2.0:requirements.txt\n\n# === statdata-transfer dependencies ===\n# Statistical data format converter — 30+ formats\n# Usage: pip install -r requirements.txt\n\n# -------------------------\n# Core (required for basic reading/writing)\n# -------------------------\npandas>=2.0,<3\npyreadstat>=1.3.5,<2\npyreadr>=0.4,<0.5\n\n# -------------------------\n# File format support (install as needed)\n# -------------------------\nopenpyxl>=3.1           # Excel .xlsx/.xlsm read/write\nxlrd>=2.0               # Excel .xls read (legacy)\nscipy>=1.11             # MATLAB .mat\nh5py>=3.10              # HDF5\npyarrow>=17.0           # Parquet, Feather, Arrow, ORC (>=17.0 修复 CVE-2024-52338 反序列化；CVE-2026-25087 仅影响 C++ 预缓冲 API，Python 绑定不受影响)\nlxml>=6.0               # XML, HTML, CDISC ODM (>=6.0 修复 CVE-2026-41066 XXE)\nodfpy>=1.4              # ODS (OpenDocument Spreadsheet)\ntableauhyperapi>=0.0.22502  # Tableau Hyper .hyper read/write (ships prebuilt native libhyper; Python >=3.6, verified on 3.13)\ndbfread>=2.0.7          # dBASE / FoxPro .dbf read (pure-Python; 2.0.7 uses DBF(load=True))\ndbf>=0.99.11            # dBASE / FoxPro .dbf write (pure-Python; set codepage='utf8' for CJK)\npyodbc>=5.0             # MS Access .mdb/.accdb read (requires Microsoft Access Driver from host)\n\n# -------------------------\n# Optional (require external software or special setup)\n# -------------------------\n# jmpio-python          # JMP .jmp binary format\n# pzfx                  # GraphPad Prism .pzfx/.pz\n# EpiData .rec requires external R + foreign package:\n#   R: https://cran.r-project.org/\n#   R> install.packages(\"foreign\")\n\n# -------------------------\n# Development / testing\n# -------------------------\n# pytest\n\nArchive v2.1.0: 31 files, 113352 bytes\n\nFiles: _icon.svg (3750b), assets/logo.svg (3750b), LICENSE (1084b), README_ZH.md (13786b), README.md (14055b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3379b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1712b), scripts/__init__.py (531b), scripts/check_env.py (4395b), scripts/reader_arff.py (9220b), scripts/reader_core.py (41279b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (10290b), scripts/reader_gretl.py (9623b), scripts/reader_legacy.py (20316b), scripts/reader_modern.py (16848b), scripts/reader_odm.py (10187b), scripts/reader_r.py (40967b), scripts/reader_sas.py (6241b), scripts/reader_science.py (46550b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (16786b), scripts/reader_v14.py (28397b), scripts/writer.py (31631b), skill-card.md (2924b), SKILL.md (8594b), test_report.md (3399b), _meta.json (136b)\n\nFile v2.1.0:SKILL.md\n\n---\r\nname: statdata-transfer\r\ncn_name: 统计数据格式转换器\r\ndescription: \"读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；处理 .rda/.rds/.RData/.mtw/.rec 文件时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。 / Read/convert 50+ statistical software formats, preserving variable/value labels and missing-value metadata for binary stats formats. Side effects (declared): runs environment checks (scripts/check_env.py); may optionally pip-install missing packages on request; can invoke the local R interpreter for .rda/.rds/.RData/.mtw/.rec files via a fallback that is DISABLED by default and must be opted in with allow_r_exec=True.\"\r\ntriggers:\r\n  - \"statdata-transfer\"\r\n  - \"统计数据格式转换\"\r\n  - \"spss stata sas 格式\"\r\n  - \".sav .dta .sas7bdat 读入\"\r\n  - \"sav转dta 格式转换\"\r\n  - \"variable labels 变量标签\"\r\n  - \"metadata-preserved conversion\"\r\nmetadata:\r\n  {\r\n    \"openclaw\": { \"emoji\": \"🛠️\", \"icon\": \"assets/icon.svg\" },\r\n    \"authors\": [\"medstatstar\", \"phoe-zip\"],\r\n    \"version\": \"2.1.0\",\r\n    \"license\": \"MIT\",\r\n    \"tags\": [\"data-conversion\", \"statistics\", \"spss\", \"stata\", \"sas\", \"clinical-trials\", \"metadata\", \"pandas\", \"bidirectional\"],\r\n    \"homepage\": \"https://github.com/medstatstar/statdata-transfer\",\r\n  }\r\n---\r\n\r\n# statdata-transfer / Statistical Data Format Converter\r\n\r\n> 致敬 Stat/Transfer — 业界格式转换标杆 / Honoring Stat/Transfer — the industry standard\r\n\r\n## Language Policy / 语言策略\r\n\r\n - 默认英文；检测到中文环境时切换为中文提示。\r\n - 常用模块备英文 + 中文两套；文档标题（不区分语言者）采用「英 / 中」顺序双语。\r\n - 复杂 / 少用模块可暂只英文。\r\n\r\n## Core Capabilities / 核心能力\r\n\r\n| Ability / 能力 | Description / 说明 |\r\n|---------|-------------|\r\n| Read / 读入 | SPSS/Stata/SAS/R … full metadata; text/JSON/detect-only keep subset / 统计二进制格式完整保留元数据；文本/JSON 保留子集；12 种专有格式探测降级 |\r\n| Convert / 转存 | Inter-convert most stats formats: SPSS↔Stata↔R↔SAS XPT↔… / 多数统计格式可互转（部分受限于规范） |\r\n| Embed / 元数据嵌入 | Labels embed in Parquet/Feather/HDF5/JSON via schema.metadata / 标签嵌入 Arrow schema.metadata |\r\n| Warn / 丢失警告 | Auto-detect metadata loss per conversion path / 自动检测并报告元数据损失 |\r\n\r\n## Supported Formats / 支持格式\r\n\r\n*50+ formats, sorted alphabetically. 按字母排序。*\r\n\r\n| Format 格式 | Extension 扩展名 | Meta Preserve 元数据保留 |\r\n|------------|-----------------|------------------------|\r\n| CDISC ODM | `.odm` | ⚠️ Clinical data only |\r\n| dBASE / FoxPro | `.dbf` | ⚠️ Read+Write, uppercase names |\r\n| EpiData | `.rec` | ⚠️ Via R |\r\n| EpiInfo | `.prj` `.xml` | ✅ XML structure |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | ⚠️ Extra sheet for labels; merged-cell fill |\r\n| EViews | `.wf1` `.wf2` | ⚠️ JSON structure |\r\n| Feather | `.feather` `.arrow` | ✅ Via schema |\r\n| FST | `.fst` | ✗ Detect-only (proprietary format) |\r\n| GraphPad Prism | `.pzfx` `.pz` | ⚠️ Multi-table |\r\n| Gretl | `.gdt` `.gdtb` | ✅ String-tables |\r\n| HDF5 | `.h5` `.hdf5` | ⚠️ Hierarchy + attribute labels |\r\n| HTML | `.html` | ⚠️ Tables only |\r\n| jamovi | `.omv` | ✅ JSON analysis |\r\n| JMP | `.jmp` | ⚠️ Multi-table |\r\n| JSON | `.json` | ✅ stat-full-meta |\r\n| MATLAB | `.mat` | ⚠️ v7.3+ via h5py fallback |\r\n| Mathematica | `.wdx` | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | ⚠️ Via R |\r\n| MS Access | `.mdb` `.accdb` | ⚠️ Multi-table; needs system driver |\r\n| ODS | `.ods` | ⚠️ Data only |\r\n| ORC | `.orc` | ✅ Via schema |\r\n| Origin | `.opju` `.oggu` | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | ✅ Via schema; partitioned datasets |\r\n| R | `.rda` `.rds` `.rdata` | ✅ pyreadr; R-interpreter fallback opt-in (allow_r_exec) |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | ✅ |\r\n| Stata | `.dta` | ✅ |\r\n| Weka ARFF | `.arff` | ✅ Nominal mapping |\r\n| XML | `.xml` | ⚠️ Structure preserved |\r\n\r\n> ✅=Full · ⚠️=Partial/conditional · ✗=Not preserved\r\n> \r\n> 12 detect-only formats (SAS CPORT `.cpt`, Statistica `.sta`, OxMetrics `.in7`, SYSTAT `.sys`/`.syd`, Paradox `.db`/`.px`, LIMDEP `.lpw`, NCSS `.ncss`) give clear export guidance — see README.\r\n\r\n## Return Structure / 返回结构\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"Question 1\"},\r\n        \"value_labels\": {\"q1\": {1: \"Yes\", 2: \"No\"}},\r\n        # ... 全部元数据 / all metadata fields\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n## Quick Start / 快速开始\r\n\r\n```bash\r\n# Check environment / 检查环境\r\npython scripts/check_env.py --install\r\n\r\n# Run via WorkBuddy (bilingual, auto-detects your language)\r\n> convert data.sav to .dta\r\n> read data.sav and show metadata\r\n> 把 data.sav 转成 .dta 并保留变量标签\r\n```\r\n\r\n> For complete code examples, see `references/usage_examples.py`.\r\n> 完整代码示例见 `references/usage_examples.py`。\r\n\r\n## Dependencies / 依赖\r\n\r\n```yaml\r\nrequires:\r\n  bins: [python3]\r\n  packages:\r\n    core: [pyreadstat>=1.3.5,<2, pyreadr>=0.4,<0.5, pandas>=2.0,<3]\r\n    extended: [openpyxl, xlrd, scipy, h5py, pyarrow, lxml, odfpy, tableauhyperapi, dbfread, dbf, pyodbc]\r\n```\r\n\r\n> Full list: `requirements.txt` / 完整列表见 `requirements.txt`\r\n\r\n## Detailed Docs / 详细文档\r\n\r\n- **English**: [`README.md`](./README.md) — format limits, strategies, encoding, extension guide\r\n- **中文**: [`README_ZH.md`](./README_ZH.md) — 格式限制、读入策略、编码注意事项、扩展指南\r\n\r\n## Security / 安全\r\n\r\n- **All R-invoking paths are opt-in & sandboxed by default / 所有调用 R 的路径默认隔离、需显式开启**: reading `.rda/.rds/.RData` (`readRDS()/load()`), reading Minitab `.mtw/.mpj` and EpiData `.rec` (`foreign::read.mtb()/read.epiinfo()`), and writing R formats (`.rda/.rds`) are **disabled by default**. They only run when you explicitly pass `allow_r_exec=True` on a TRUSTED file. Pure-Python parsers (`pyreadr`, `mtbpy`) are tried first and never execute code.\r\n- **No silent R fallback / 无静默 R 回退**: If the pure-Python parser fails and `allow_r_exec` is not set, the skill raises a clear error instead of silently launching R — eliminating the risk of executing embedded code from untrusted files.\r\n- **R scripts are static templates / R 脚本为静态模板**: When the opt-in R path is used, all R scripts are fully static templates; user inputs are passed only as CLI args (`commandArgs(trailingOnly=TRUE)` via `jsonlite`) — **never concatenated into executable R code**. Random temp filenames; no fixed paths.\r\n- **R bridge writes a temp CSV / R 桥接写临时 CSV**: When R is used (opt-in), converted data is materialized to a temporary CSV on disk before being read back; the temp file is deleted immediately after use, but on crash or via backup/indexing tools it could briefly persist — avoid processing highly sensitive data through R-backed formats.\r\n- **No destructive writes / 无破坏性写入**: Writing an existing `.hyper` first writes to a temp file; only after success is the existing file rotated to `<file>.bak` (prior `.bak` demoted to `.bak.1`, never silently deleted), then atomically swapped in. On write failure the original is untouched.\r\n- **Pinned dependencies / 依赖已固定版本**: Core deps carry upper-bound pins (`pandas`, `pyreadstat`, `pyreadr`) — see `requirements.txt`.\r\n- **Optional install / 可选安装**: `python scripts/check_env.py --install` only runs on explicit request.\r\n- **Permissions required / 所需权限**: Read the input file; write the output file to a path you specify. No network access unless you explicitly request package installation. No destructive writes — existing `.hyper` outputs are backed up to `.bak` before overwrite.\r\n- **Scope / 范围**: Statistical data formats only. Text/JSON formats preserve metadata subset only — see «Format Limits» in README.\r\n\r\n## License / 许可证\r\n\r\nMIT. See [`LICENSE`](./LICENSE). / 详见 [`LICENSE`](./LICENSE)。\n\nFile v2.1.0:README.md\n\n# statdata-transfer / Statistical Data Format Converter\r\n\r\n[🇨🇳 中文 (Chinese)](./README_ZH.md)\r\n\r\n---\r\n\r\nRead 50+ statistical software and clinical trial data formats into Python/pandas DataFrame, and **inter-convert between most formats** (SPSS↔Stata↔R↔SAS XPT↔Excel↔Parquet↔HDF5↔JSON…). For statistical binary formats (SPSS/Stata/SAS/R/Excel/Parquet/HDF5/…) it preserves full variable/value labels and missing-value metadata; text formats (CSV/XML/HTML/ODS) and JSON preserve only a retainable subset — see Format Limits below. **12 proprietary formats** (SAS CPORT, Statistica, OxMetrics, SYSTAT, Paradox, LIMDEP, NCSS, FST, etc.) are **detect-only** — the skill recognizes the extension and provides clear export guidance, but does not parse the data.\r\n\r\nNote: This skill does not require any statistical software, but handles data format conversion only. If you need **an AI agent to integrate with installed statistical software for analysis**, use the **[statsoft-cli](https://github.com/medstatstar/statsoft-cli)** skill instead.\r\n\r\n## Core Capabilities\r\n\r\n### Read (Data Extraction)\r\nExtract data + all metadata from 50+ formats into pandas DataFrame. Preserves metadata as completely as possible and clearly indicates what is preserved vs lost.\r\n\r\n### Write / Convert (Format Conversion)\r\n- **Inter-convert stats formats**: SPSS ↔ Stata ↔ R ↔ SAS XPT (all metadata types preserved)\r\n- **Export universal formats**: Parquet, Feather, HDF5, JSON, CSV, TSV, Excel (metadata embedded in schema.metadata or sidecar JSON)\r\n- **Universal → stats formats**: Reverse preserve metadata via embedded `stat-full-meta`\r\n\r\n### Auto Warnings\r\nAutomatically detects and reports metadata preservation vs loss during conversion. All user-facing messages are bilingual (en/zh-cn).\r\n\r\n## Supported Formats & Capability Matrix\r\n\r\n*Sorted alphabetically.*\r\n\r\n| Format | Extension | Dependency | Var Label | Val Label | Special Missing | Formula | Meta Preserve |\r\n|--------|-----------|------------|-----------|-----------|-----------------|---------|---------------|\r\n| CDISC ODM | `.odm` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Clinical data only |\r\n| dBASE / FoxPro | `.dbf` | dbfread / dbf | ✗ | ✗ | ✗ | ✗ | ⚠️ Read+Write; uppercase names |\r\n| EpiData | `.rec` | R foreign | ✗ | ✗ | ✗ | ✗ | ⚠️ Via R |\r\n| EpiInfo | `.prj` `.xml` | xml/etree | ✅ | ✅(codes) | ✗ | ✗ | ✅ XML structure |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | openpyxl / xlrd | ✗ | ✗ | ✗ | ⚠️ result only | ⚠️ Extra sheet for labels; merged-cell fill |\r\n| EViews | `.wf1` `.wf2` | built-in | ✗ | ✗ | ✗ | ✗ | ⚠️ JSON structure |\r\n| Feather | `.feather` `.arrow` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Version diff |\r\n| FST | `.fst` | — | ✗ | ✗ | ✗ | ✗ | ✗ Detect-only (proprietary format) |\r\n| GraphPad Prism | `.pzfx` `.pz` | pzfx | ✗ | ✗ | ✗ | ✗ | ⚠️ Multi-table |\r\n| Gretl | `.gdt` `.gdtb` | built-in | ✅ | ✅(tables) | ✗ | ✗ | ✅ string-tables |\r\n| HDF5 | `.h5` `.hdf5` | h5py | ✗ | ✗ | ✗ | ✗ | ⚠️ Hierarchy + attribute labels |\r\n| HTML | `.html` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Tables only |\r\n| jamovi | `.omv` | ZIP built-in | ✅ | ✅ | ✗ | ✗ | ✅ JSON analysis |\r\n| JMP | `.jmp` | jmpio-python | ⚠️ | ⚠️ | ✗ | ✗ | ⚠️ Multi-table |\r\n| JSON | `.json` | built-in | ✅ | ✅ | ✗ | ✗ | ✅ stat-full-meta on write |\r\n| MATLAB | `.mat` | scipy | ✗ | ✗ | ✗ | ✗ | ⚠️ v7.3+ via h5py fallback |\r\n| Mathematica | `.wdx` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | mtbpy / R | ✗ | ✗ | ✗ | ✗ | ⚠️ Via R |\r\n| MS Access | `.mdb` `.accdb` | pyodbc + Access Driver | ✗ | ✗ | ✗ | ✗ | ⚠️ Multi-table; needs system driver |\r\n| ODS | `.ods` | odfpy | ✗ | ✗ | ✗ | ✗ | ⚠️ Data only |\r\n| ORC | `.orc` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Version diff |\r\n| Origin | `.opju` `.oggu` | zipfile + lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ Nested types; partitioned datasets |\r\n| R | `.rda` `.rds` `.rdata` | pyreadr + R | ✅ | ✅ | ✅ | ✗ | ✅ statdata_meta + R bridge |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | pyreadstat | ✅ | ✅(need catalog) | ⚠️ | ✗ | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | pyreadstat | ✅ | ✅ | ✅ | ✗ | ✅ |\r\n| Stata | `.dta` | pyreadstat | ✅ | ✅ | ⚠️ | ✗ | ✅ |\r\n| Weka ARFF | `.arff` | built-in | ✅ | ✅(nominal) | ✗ | ✗ | ✅ nominal mapping |\r\n| XML | `.xml` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Structure preserved |\r\n\r\n> ✅=Full preservation · ⚠️=Partial/conditional · ✗=Not preserved\r\n\r\n### Detect-Only Formats\r\nFormats with no parser available. The skill detects the extension and provides clear export guidance (no data parsing).\r\n\r\n| Format | Extension | Guidance |\r\n|--------|-----------|----------|\r\n| FST (R fst package) | `.fst` | R: `fst::read_fst(\"in.fst\", \"out.csv\")` then read CSV |\r\n| LIMDEP / NLOGIT | `.lpw` | Export to CSV from original software |\r\n| NCSS | `.ncss` | Export to CSV |\r\n| OxMetrics | `.in7` | Export to CSV / `.dta` |\r\n| Paradox | `.db` `.px` | Export to `.dbf` / CSV |\r\n| SAS CPORT | `.cpt` | SAS: `proc export` to XPORT(`.xpt`) / `.sas7bdat` |\r\n| Statistica | `.sta` | Export to `.sav` / `.csv` |\r\n| SYSTAT | `.sys` `.syd` | Export to CSV / `.sav` |\r\n\r\n## Return Structure\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"Question 1\"},\r\n        \"value_labels\": {\"q1\": {1: \"Yes\", 2: \"No\"}},\r\n        # ... all metadata types\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n## Metadata Preservation Tiers\r\n\r\n### Read Fallback Rules\r\n1. **Statistical binary formats** (SPSS/Stata/SAS/R): 100% metadata preserved\r\n2. **Arrow ecosystem** (Parquet/Feather/ORC): Only restores labels from `write_stat_file`\r\n3. **Non-stats formats** (CSV/Excel/XML/HTML/ODS): Data only; use `apply_value_labels()` to attach manually\r\n4. **R formats**: v1.6.0+ embeds all metadata via `statdata_meta` attribute\r\n\r\n### Metadata-Loss Warnings\r\nRuntime `warnings` list includes:\r\n- Special missing values → NaN\r\n- Measurement level / display width / alignment lost\r\n- MR sets cannot be preserved\r\n- File label / notes not preserved\r\n- Date origin not saved\r\n\r\n## Recommended Read Strategies\r\n\r\n| Use Case | Recommendation |\r\n|----------|---------------|\r\n| Data warehousing / ETL | SPSS `.sav` or Stata `.dta` → Parquet / HDF5 |\r\n| Scientific computing | `.mat` or `.hdf5` → NumPy / pandas |\r\n| Statistical analysis (Python) | `.sav` / `.dta` → pandas → scipy.stats |\r\n| Report output | pandas → JSON / HTML / Excel |\r\n| Cross-software sharing | Stata ↔ SPSS ↔ R direct interconversion |\r\n\r\n## File Size Limits\r\n\r\n| Format | Memory Behavior |\r\n|--------|----------------|\r\n| pyreadstat (SPSS/Stata/SAS) | Loads entire file into RAM |\r\n| HDF5 | Chunked reading; not limited by RAM |\r\n| Parquet | pyarrow memory-mapped (mmap); handles files >RAM |\r\n\r\n## Encoding Notes\r\n\r\n- **Chinese files**: Old Stata/SAS may use GBK/gb2312. Use `encoding='gbk'`.\r\n- **European files**: Some SAS files use Latin-1. Try `encoding='latin1'` if UTF-8 fails.\r\n- **Auto-detection**: `_auto_detect_encoding` is enabled by default for SPSS/Stata/SAS.\r\n\r\n## Providing Input Files\r\n\r\nAI agents (e.g., WorkBuddy) can only directly upload a limited set of file types. When your data file cannot be uploaded directly, you have two options:\r\n\r\n1. **Use the absolute file path** in your prompt (e.g., `convert C:/Users/Name/Desktop/data.sav to .dta`)\r\n2. **Compress the file as a `.zip` archive** and upload the zip instead\r\n\r\nThe skill automatically extracts and processes zip archives containing a single data file.\r\n\r\n## Quick Start\r\n\r\n```bash\r\n# Check environment\r\npython scripts/check_env.py --install\r\n```\r\n\r\nComplete code examples: [`references/usage_examples.py`](./references/usage_examples.py)\r\n\r\nWorkBuddy prompts:\r\n```\r\n> convert data.sav to .dta\r\n> read data.sav and show metadata\r\n> are there any metadata loss concerns converting .sav to .xlsx?\r\n```\r\n\r\n## Extending\r\n\r\nTo add a new format: edit `scripts/reader_*.py` to add a reader function, register it in `format_map` in `scripts/reader_core.py`, and add a TypedDict in `scripts/reader_core.py`.\r\n\r\n## Format Limitations\r\n\r\n*Alphabetically ordered. Markeds ✅ = fixed, 🔄 = new capability; rest are inherent format limits.*\r\n\r\n### CDISC ODM (.odm)\r\n**Read:**\r\n- ❌ XML structure dependency: nested parsing depends on ODM regularity\r\n- ❌ No statistical metadata in ODM spec; only clinical structure preserved\r\n\r\n### dBASE / FoxPro (.dbf)\r\n**Read:**\r\n- ❌ Field names forced to uppercase (format limitation)\r\n- ✅ Read + Write supported\r\n\r\n### EpiData (.rec)\r\n**Read:**\r\n- ❌ Requires R + `foreign` package; no Python-native fallback\r\n- ❌ Statistical metadata lost in R-to-CSV bridge\r\n\r\n### EpiInfo (.prj)\r\n**Read:**\r\n- ❌ Project file contains no data; auto-associates same-name CSV\r\n- ❌ Access not supported; export to CSV first\r\n**Write:**\r\n- ✅ Variable labels and codes reconstructed in XML structure\r\n\r\n### Excel (.xlsx/.xls/.xlsm)\r\n**Read:**\r\n- ✅ **Merged cells**: Fills merged area with anchor value; toggle via `fill_merged_cells=True` (default)\r\n- ❌ Formulas lost; only computed values retained\r\n- ❌ Charts/shapes not extracted\r\n**Write:**\r\n- Variable/value labels in separate metadata worksheet\r\n\r\n### HDF5 (.h5/.hdf5)\r\n**Read:**\r\n- ✅ **Multi-dataset fallback**: When `pd.read_hdf` fails, h5py fallback merges all top-level numeric datasets\r\n- ✅ **Attribute labels**: Scans common keys (`label`, `description`, `units`, `ColumnLabel`…)\r\n- ❌ Hierarchy flattened to top-level\r\n**Write:**\r\n- Root-level attributes used for metadata storage\r\n\r\n### JMP (.jmp)\r\n**Read:**\r\n- ❌ Requires `jmpio-python`; version support varies\r\n- ❌ Multi-table: only first table returned\r\n**Write:**\r\n- Single-table only\r\n\r\n### MATLAB (.mat)\r\n**Read:**\r\n- ✅ **v7.3+ (HDF5-based)**: Detects `MATLAB 7.3` header or scipy failure → h5py path\r\n- ❌ Complex structures (nested cells, sparse, fn handles) → single-column flattened\r\n- ❌ Object classes and datetime lose type fidelity\r\n\r\n### Parquet (.parquet)\r\n**Read:**\r\n- ❌ Deeply nested types (>2 levels `list<struct>`) → opaque Python objects\r\n- ✅ **Partitioned datasets**: Directory with `part-*.parquet` via `pyarrow.dataset`\r\n\r\n### R (.rda/.rds/.rdata)\r\n**Read:**\r\n- ✅ Old ASCII XDR (RDA2): read via R bridge requires `allow_r_exec=True` (opt-in; not automatic)\r\n- ❌ Factor level order may not be preserved as Categorical unless embedded via `stat-full-meta`\r\n- ❌ Multi-object RDA: `read_all_r_objects()` returns all\r\n**Write:**\r\n- R bridge (`statdata_meta` attribute) for full metadata round-trip\r\n\r\n### SAS (.sas7bdat/.xpt/.sas7bcat)\r\n**Read:**\r\n- ✅ Value labels require co-located `.sas7bcat` (auto-detected)\r\n- ❌ Viya CAS `.sashdat` not supported\r\n- Date origin: 1960-01-01\r\n\r\n### SPSS (.sav/.zsav/.por)\r\n**Read:**\r\n- ❌ MR Sets imported as raw dict; semantics not preserved\r\n- ❌ Formulas lost; only computed results retained\r\n- ⚠️ Special missing (`.A`–`.Z`) flagged in `special_missing`; must be preserved explicitly on write\r\n**Write:**\r\n- `.zsav` requires pyreadstat 1.2+; else auto-fallback to `.sav`\r\n\r\n### Stata (.dta)\r\n**Read:**\r\n- ⚠️ Special missing (`.a`–`.z`): preserved when `user_missing=True` (default); becomes NaN when `user_missing=False` (irreversible)\r\n- ✅ Pre-v13 Latin-1 auto-detected\r\n- ❌ Stata 117-119 not supported by pyreadstat 1.3.5; auto-downgrade to v15 on write\r\n\r\n## Security\r\n\r\n- **R execution is opt-in and sandboxed by default.** Every path that invokes the local R interpreter — reading `.rda/.rds/.RData` (via `readRDS()/load()`), reading Minitab `.mtw/.mpj` and EpiData `.rec` (via `foreign::read.mtb()/read.epiinfo()`), and writing R formats (`.rda/.rds`) — is **disabled by default**. It only runs when you explicitly pass `allow_r_exec=True` on a TRUSTED file. Pure-Python parsers (`pyreadr`, `mtbpy`) are tried first and never execute code.\r\n- **No silent R fallback.** If the pure-Python parser fails and `allow_r_exec` is not set, the skill raises a clear error instead of silently launching R. This removes the risk of executing embedded code from untrusted files.\r\n- **R scripts are static templates.** When the opt-in R path runs, all R scripts are fully static; user inputs (file paths, labels, metadata) are passed only as CLI args (`commandArgs(trailingOnly=TRUE)` via `jsonlite`) — never concatenated into executable R code. Random temp filenames; no fixed paths.\r\n- **Temporary CSV exposure (R bridge).** When R is used (opt-in), converted data is materialized to a temporary CSV on disk before being read back. The temp file is deleted immediately after use, but on crash or via backup/indexing tools it could briefly persist. Avoid processing highly sensitive data through R-backed formats, or prefer a non-R path.\r\n- **No destructive writes.** When writing a `.hyper` file that already exists, the new data is written to a temp file first; only after success is the existing file rotated to `<file>.bak` (the previous `.bak` is demoted to `.bak.1`, never silently deleted), then the temp file is atomically swapped in. If the write fails, the original file is left untouched.\r\n- **Optional package install.** `python scripts/check_env.py --install` runs only on explicit request.\r\n- **Pinned dependencies.** Core dependencies carry upper-bound pins (`pandas`, `pyreadstat`, `pyreadr`) — see `requirements.txt`.\r\n- **Scope.** Statistical data formats only. No network access unless you explicitly request package installation.\r\n\r\n## License\r\n\r\nMIT-0 License. See [LICENSE](LICENSE) for details.\n\nFile v2.1.0:_meta.json\n\n{\n  \"ownerId\": \"kn7amqq1jv28skb63wavr6shah89jsm5\",\n  \"slug\": \"statdata-transfer\",\n  \"version\": \"2.1.0\",\n  \"publishedAt\": 1784428367554\n}\n\nFile v2.1.0:references/new_formats_architecture_analysis.json\n\n{\r\n  \"current_architecture\": {\r\n    \"data\": \"pandas DataFrame\",\r\n    \"metadata\": \"BaseMeta TypedDict (~28 fields)\",\r\n    \"column_report\": \"ColumnInfo TypedDict (12 fields)\",\r\n    \"return_type\": \"StatFileResult = dict[str, Any] with keys: dataframe, metadata, warnings, column_report\",\r\n    \"multi_object_pattern\": \"read_all_*() returns dict[str, StatFileResult]\"\r\n  },\r\n  \"formats_analysis\": {\r\n    \"sas7bcat\": {\r\n      \"description\": \"SAS Ŀ¼�ļ����洢��ʽ���壨ֵ��ǩ��\",\r\n      \"data_structure\": \"�����ݣ�ֻ��Ԫ���ݣ���ʽ���壩\",\r\n      \"metadata_fields\": [\r\n        \"value_labels\",\r\n        \"variable_value_labels\",\r\n        \"variable_to_label\"\r\n      ],\r\n      \"current_architecture_sufficient\": true,\r\n      \"notes\": \"pyreadstat ��֧�ֶ�ȡ .sas7bcat������ value_labels dict\",\r\n      \"architecture_extension_needed\": false\r\n    },\r\n    \"jmp\": {\r\n      \"description\": \"SAS JMP �����ļ����ɰ���������ݱ����ű����������\",\r\n      \"data_structure\": \"�ɰ���������ݱ���Data Table����ÿ������һ�� DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"column_properties (formulas, ranges)\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"JMP �ļ��ɰ���������ݱ�����Ҫ read_all_jmp_tables() ģʽ�������ԣ���ʽ����Χ����Ҫ��չ ColumnInfo\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"ColumnInfo ��Ҫ��չ��formula (str), range (dict), column_property (dict)\",\r\n        \"��Ҫ���� JmpMeta �࣬���� tables list��scripts list��analysis list\",\r\n        \"��Ҫ���� read_all_jmp_tables() ����\"\r\n      ]\r\n    },\r\n    \"minitab\": {\r\n      \"description\": \"Minitab �������ļ����ɰ��������������Worksheet��\",\r\n      \"data_structure\": \"�ɰ��������������ÿ����������һ�� DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"worksheet_names\",\r\n        \"formulas\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"Minitab �������ɰ����������������Ҫ read_all_minitab_worksheets() ģʽ\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� MinitabMeta �࣬���� worksheets list��active_worksheet str\",\r\n        \"��Ҫ���� read_all_minitab_worksheets() ����\",\r\n        \"ColumnInfo ������Ҫ��չ��formula (str)\"\r\n      ]\r\n    },\r\n    \"prism\": {\r\n      \"description\": \"GraphPad Prism ��Ŀ�ļ����������ݱ����������ͼ��\",\r\n      \"data_structure\": \"�������ݱ���DataFrame�����������DataFrame����ͼ�Σ��޷�תΪ DataFrame��\",\r\n      \"metadata_fields\": [\r\n        \"data_tables\",\r\n        \"results_tables\",\r\n        \"graphs_info\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"Prism �ļ��������ݱ��ͽ���������߶��� DataFrame��ͼ���޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� PrismMeta �࣬���� data_tables list��results_tables list��graphs_info list\",\r\n        \"StatFileResult ��Ҫ��չ������� read_all_prism_tables() ģʽ\",\r\n        \"��ǰ�ܹ�ֻ�ܱ������ݱ����������ͼ����Ϣ�ᶪʧ\"\r\n      ]\r\n    },\r\n    \"jamovi\": {\r\n      \"description\": \"jamovi ��Ŀ�ļ����������ݡ��������������\",\r\n      \"data_structure\": \"�������ݣ�CSV�������������JSON��\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"analysis_results\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"jamovi �ļ��� ZIP������ data.csv �� analysis.json����������޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� JamoviMeta �࣬���� analysis_results dict��analysis_settings dict\",\r\n        \"��ǰ�ܹ����Ա������ݲ��֣�����������ᶪʧ\"\r\n      ]\r\n    },\r\n    \"epidata\": {\r\n      \"description\": \"EpiData �����ļ������в�ѧ���鳣��\",\r\n      \"data_structure\": \"�������ݱ���DataFrame��\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"data_types\"\r\n      ],\r\n      \"current_architecture_sufficient\": true,\r\n      \"notes\": \"EpiData �ļ��ṹ�� SPSS .sav ���ƣ���ǰ�ܹ����� 100% ����\",\r\n      \"architecture_extension_needed\": false\r\n    },\r\n    \"eviews\": {\r\n      \"description\": \"EViews �����ļ����������У�Series�����飨Group�������̣�Equation��\",\r\n      \"data_structure\": \"����������У������ Group��DataFrame���������޷�תΪ DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"series_names\",\r\n        \"groups\",\r\n        \"equations\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"EViews .wf2 �� JSON ��ʽ���ɽ����������̡�ϵ���ȷ�������޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� EviewsMeta �࣬���� series list��groups list��equations list\",\r\n        \"��ǰ�ܹ����Ա��� Group Ϊ DataFrame����������Ϣ�ᶪʧ\"\r\n      ]\r\n    }\r\n  }\r\n}\n\nFile v2.1.0:references/v1.4_implementation_summary.json\n\n{\r\n  \"v1.4_new_formats\": [\r\n    {\r\n      \"format\": \"SAS Catalog\",\r\n      \"ext\": \".sas7bcat\",\r\n      \"handler\": \"_read_sas_catalog\",\r\n      \"dependency\": \"pyreadstat (已支持)\",\r\n      \"architecture\": \"SasCatalogMeta, 返回格式定义 DataFrame\",\r\n      \"status\": \"done\"\r\n    },\r\n    {\r\n      \"format\": \"JMP\",\r\n      \"ext\": \".jmp\",\r\n      \"handler\": \"_read_jmp\",\r\n      \"dependency\": \"jmpio-python (PyPI) 或 @skill:statsoft-cli\",\r\n      \"architecture\": \"JmpMeta, 多表需 read_all_jmp_tables()\",\r\n      \"status\": \"done (handler 已添加，jmpio 未安装）\"\r\n    },\r\n    {\r\n      \"format\": \"Minitab\",\r\n      \"ext\": \".mtw/.mpj\",\r\n      \"handler\": \"_read_minitab\",\r\n      \"dependency\": \"mtbpy 或 R foreign::read.mtb() 中继\",\r\n      \"architecture\": \"MinitabMeta, 多工作表需 read_all_minitab_worksheets()\",\r\n      \"status\": \"done (R 中继已实现）\"\r\n    },\r\n    {\r\n      \"format\": \"GraphPad Prism\",\r\n      \"ext\": \".pzfx/.pz\",\r\n      \"handler\": \"_read_prism\",\r\n      \"dependency\": \"pzfx (PyPI) 或 @skill:statsoft-cli\",\r\n      \"architecture\": \"PrismMeta, 含数据表+结果表\",\r\n      \"status\": \"done (handler 已添加，pzfx 未安装）\"\r\n    },\r\n    {\r\n      \"format\": \"jamovi\",\r\n      \"ext\": \".omv\",\r\n      \"handler\": \"_read_jamovi\",\r\n      \"dependency\": \"无需额外包（ZIP+CSV 解析）\",\r\n      \"architecture\": \"JamoviMeta, 含 analysis JSON\",\r\n      \"status\": \"done\"\r\n    },\r\n    {\r\n      \"format\": \"EpiData\",\r\n      \"ext\": \".rec\",\r\n      \"handler\": \"_read_epidata\",\r\n      \"dependency\": \"R foreign::read.epiinfo() 中继\",\r\n      \"architecture\": \"EpidataMeta\",\r\n      \"status\": \"done (R 中继已实现）\"\r\n    },\r\n    {\r\n      \"format\": \"EViews\",\r\n      \"ext\": \".wf1/.wf2\",\r\n      \"handler\": \"_read_eviews\",\r\n      \"dependency\": \".wf2 可直接解析 JSON；.wf1 需 @skill:statsoft-cli\",\r\n      \"architecture\": \"EviewsMeta\",\r\n      \"status\": \"done (.wf2 解析已实现）\"\r\n    }\r\n  ],\r\n  \"architecture_extensions\": [\r\n    \"ColumnInfo 新增：formula (str), column_property (dict)\",\r\n    \"新增 Meta 类：SasCatalogMeta, JmpMeta, MinitabMeta, PrismMeta, JamoviMeta, EpidataMeta, EviewsMeta\",\r\n    \"__all__ 新增导出：7 个新 Meta 类名\"\r\n  ],\r\n  \"files_modified\": [\r\n    \"scripts/stat_reader.py (+459 行，共 3131 行）\",\r\n    \"scripts/check_env.py (新增 jmpio, pzfx 检测）\",\r\n    \"SKILL.md (待更新）\",\r\n    \"references/new_formats_architecture_analysis.json (新增）\"\r\n  ],\r\n  \"remaining_work\": [\r\n    \"更新 SKILL.md（添加 7 种新格式详情）\",\r\n    \"添加 read_all_jmp_tables() / read_all_minitab_worksheets()\",\r\n    \"测试新 handler（需要实际文件）\",\r\n    \"安装 jmpio/pzfx 包（或配置 @skill:statsoft-cli）\"\r\n  ]\r\n}\n\nFile v2.1.0:README_ZH.md\n\n# statdata-transfer / 统计数据格式转换器\r\n\r\n[🇬🇧 English](./README.md)\r\n\r\n---\r\n\r\n读入 50+ 统计软件及临床试验数据格式，转换为 Python/pandas DataFrame，并**支持多数格式双向互转**（SPSS↔Stata↔R↔SAS XPT↔Excel↔Parquet↔HDF5↔JSON…）。对统计二进制格式（SPSS/Stata/SAS/R/Excel/Parquet/HDF5…）完整保留变量标签、值标签等元数据；文本格式（CSV/XML/HTML/ODS）与 JSON 仅保留可保留的子集——详见「格式限制」。**12 种专有格式**（SAS CPORT、Statistica、OxMetrics、SYSTAT、Paradox、LIMDEP、NCSS、FST 等）为**探测降级**——技能识别扩展名并给出清晰导出指引，但不解析数据。\r\n\r\n注意：本技能不需要任何统计软件的支持，但功能仅限于数据格式转换。如果需要**AI 智能体接入已安装的统计软件进行分析**，请使用 **[statsoft-cli](https://github.com/medstatstar/statsoft-cli)** 技能。\r\n\r\n## 核心能力\r\n\r\n### 读入（数据提取）\r\n从 50+ 统计软件格式中提取数据 + 元数据，转为 pandas DataFrame，尽可能完整保留元数据，并明确标注保留/丢失情况。\r\n\r\n### 转存（格式转换）\r\n- **统计格式互转**：SPSS ↔ Stata ↔ R ↔ SAS XPT（保留全部元数据）\r\n- **导出通用格式**：Parquet、Feather、HDF5、JSON、CSV、TSV、Excel（元数据嵌入 schema.metadata 或 sidecar JSON）\r\n- **通用格式→统计格式**：通过嵌入的 `stat-full-meta` 反向保留元数据\r\n\r\n### 自动警告\r\n转换时检测并报告哪些元数据可以保留、哪些会丢失，避免静默的数据损失。所有警告信息均为中英双语。\r\n\r\n## 支持格式与能力矩阵\r\n\r\n*按字母排序。*\r\n\r\n| 格式 | 扩展名 | 依赖 | 变量标签 | 值标签 | 特殊缺失 | 公式 | 元数据保留 |\r\n|------|--------|------|---------|--------|---------|------|-----------|\r\n| CDISC ODM | `.odm` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅临床数据 |\r\n| dBASE / FoxPro | `.dbf` | dbfread / dbf | ✗ | ✗ | ✗ | ✗ | ⚠️ 读+写；大写字段名 |\r\n| EpiData | `.rec` | R foreign | ✗ | ✗ | ✗ | ✗ | ⚠️ 通过 R 读入 |\r\n| EpiInfo | `.prj` `.xml` | xml/etree | ✅ | ✅(codes) | ✗ | ✗ | ✅ XML 结构 |\r\n| Excel | `.xlsx` `.xls` `.xlsm` | openpyxl / xlrd | ✗ | ✗ | ✗ | ⚠️ 仅结果 | ⚠️ 写出用额外工作表；合并单元格填充 |\r\n| EViews | `.wf1` `.wf2` | 内置 | ✗ | ✗ | ✗ | ✗ | ⚠️ JSON 结构 |\r\n| Feather | `.feather` `.arrow` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 版本差异 |\r\n| FST | `.fst` | — | ✗ | ✗ | ✗ | ✗ | ✗ 探测降级（专有格式） |\r\n| GraphPad Prism | `.pzfx` `.pz` | pzfx | ✗ | ✗ | ✗ | ✗ | ⚠️ 多表 |\r\n| Gretl | `.gdt` `.gdtb` | 内置 | ✅ | ✅(tables) | ✗ | ✗ | ✅ string-tables |\r\n| HDF5 | `.h5` `.hdf5` | h5py | ✗ | ✗ | ✗ | ✗ | ⚠️ 层级结构 + 属性标签 |\r\n| HTML | `.html` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅表格 |\r\n| jamovi | `.omv` | ZIP 内置 | ✅ | ✅ | ✗ | ✗ | ✅ JSON 分析 |\r\n| JMP | `.jmp` | jmpio-python | ⚠️ | ⚠️ | ✗ | ✗ | ⚠️ 多表 |\r\n| JSON | `.json` | 内置 | ✅ | ✅ | ✗ | ✗ | ✅ 写出嵌入 stat-full-meta |\r\n| MATLAB | `.mat` | scipy | ✗ | ✗ | ✗ | ✗ | ⚠️ v7.3+ 走 h5py 回退 |\r\n| Mathematica | `.wdx` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort XML |\r\n| Minitab | `.mtw` `.mpj` | mtbpy / R | ✗ | ✗ | ✗ | ✗ | ⚠️ 通过 R 读入 |\r\n| MS Access | `.mdb` `.accdb` | pyodbc + Access 驱动 | ✗ | ✗ | ✗ | ✗ | ⚠️ 多表；需系统驱动 |\r\n| ODS | `.ods` | odfpy | ✗ | ✗ | ✗ | ✗ | ⚠️ 仅数据 |\r\n| ORC | `.orc` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 版本差异 |\r\n| Origin | `.opju` `.oggu` | zipfile + lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ Best-effort |\r\n| Parquet | `.parquet` | pyarrow | ✅(schema) | ✅(schema) | ✗ | ✗ | ⚠️ 嵌套类型；分区数据集 |\r\n| R | `.rda` `.rds` `.rdata` | pyreadr + R | ✅ | ✅ | ✅ | ✗ | ✅ statdata_meta + R 桥接 |\r\n| SAS | `.sas7bdat` `.xpt` `.sas7bcat` | pyreadstat | ✅ | ✅(需 catalog) | ⚠️ | ✗ | ✅ |\r\n| SPSS | `.sav` `.zsav` `.por` | pyreadstat | ✅ | ✅ | ✅ | ✗ | ✅ |\r\n| Stata | `.dta` | pyreadstat | ✅ | ✅ | ⚠️ | ✗ | ✅ |\r\n| Weka ARFF | `.arff` | 内置 | ✅ | ✅(nominal) | ✗ | ✗ | ✅ 名义映射 |\r\n| XML | `.xml` | lxml | ✗ | ✗ | ✗ | ✗ | ⚠️ 结构保留 |\r\n\r\n> ✅=完整保留 · ⚠️=部分保留或条件性 · ✗=无法保留\r\n\r\n### 探测降级格式\r\n无现成解析库，识别扩展名并给出清晰导出指引（不解析数据）。\r\n\r\n| 格式 | 扩展名 | 导出指引 |\r\n|------|--------|---------|\r\n| FST (R fst 包) | `.fst` | R: `fst::read_fst(\"in.fst\", \"out.csv\")`，再读 CSV |\r\n| LIMDEP / NLOGIT | `.lpw` | 从原软件导出 CSV |\r\n| NCSS | `.ncss` | 导出 CSV |\r\n| OxMetrics | `.in7` | 导出 CSV / `.dta` |\r\n| Paradox | `.db` `.px` | 导出 `.dbf` / CSV |\r\n| SAS CPORT | `.cpt` | SAS: `proc export` 为 XPORT(`.xpt`) / `.sas7bdat` |\r\n| Statistica | `.sta` | 导出 `.sav` / `.csv` |\r\n| SYSTAT | `.sys` `.syd` | 导出 CSV / `.sav` |\r\n\r\n## 返回结构\r\n\r\n```python\r\n{\r\n    \"dataframe\": pd.DataFrame,\r\n    \"metadata\": {\r\n        \"file_format\": \"spss_sav\",\r\n        \"row_count\": 100, \"column_count\": 10,\r\n        \"variable_labels\": {\"q1\": \"问题1\"},\r\n        \"value_labels\": {\"q1\": {1: \"是\", 2: \"否\"}},\r\n        # ... 全部元数据\r\n    },\r\n    \"warnings\": [],\r\n    \"column_report\": {\"q1\": {\"source_type\": \"int\", \"pandas_dtype\": \"int64\"}},\r\n}\r\n```\r\n\r\n## 元数据保留层级\r\n\r\n### 读入降损规则\r\n1. **统计二进制格式**（SPSS/Stata/SAS/R）：100% 元数据完整保留\r\n2. **Arrow 生态**（Parquet/Feather/ORC）：仅还原 `write_stat_file` 写入的标签\r\n3. **非统计格式**（CSV/Excel/XML/HTML/ODS）：仅保留数据值；可使用 `apply_value_labels()` 手动附加\r\n4. **R 格式**：v1.6.0+ 通过 `statdata_meta` 属性嵌入全部元数据\r\n\r\n### 元数据丢失警告\r\n运行时 `warnings` 列表包含：\r\n- 特殊缺失值变为 NaN\r\n- 测量级别/显示宽度/对齐方式丢失\r\n- 多重响应集（MR Sets）无法保留\r\n- 文件标签/注释不保留\r\n- 日期基准不保存\r\n\r\n## 提供输入文件\r\n\r\nAI 智能体（如 WorkBuddy）只能直接上传有限类型的文件。当数据文件无法直接上传时，有两种方式：\r\n\r\n1. **在提示词中使用文件绝对路径**（如 `convert C:/Users/Name/Desktop/data.sav to .dta`）\r\n2. **将文件压缩为 `.zip` 包**后上传\r\n\r\n技能会自动解压并处理包含单个数据文件的 zip 归档。\r\n\r\n## 推荐读入策略\r\n\r\n| 需求 | 推荐 |\r\n|------|------|\r\n| 数据入库/ETL | SPSS `.sav` 或 Stata `.dta` → Parquet / HDF5 |\r\n| 科学计算 | `.mat` 或 `.hdf5` → NumPy / pandas |\r\n| 统计分析（Python） | `.sav` / `.dta` → pandas → scipy.stats |\r\n| 报告输出 | pandas → JSON / HTML / Excel |\r\n| 跨软件共享 | Stata ↔ SPSS ↔ R 直接互转 |\r\n\r\n## 文件大小限制\r\n\r\n| 格式 | 内存行为 |\r\n|------|---------|\r\n| pyreadstat (SPSS/Stata/SAS) | 全文件加载到 RAM |\r\n| HDF5 | 支持分块读取；不受 RAM 限制 |\r\n| Parquet | pyarrow 支持 mmap 映射；可处理 >内存的文件 |\r\n\r\n## 编码注意事项\r\n\r\n- **中文文件**：旧版 Stata/SAS 可能使用 GBK/gb2312。使用 `encoding='gbk'`。\r\n- **欧洲文件**：部分 SAS 文件使用 Latin-1。UTF-8 失败时尝试 `encoding='latin1'`。\r\n- **自动检测**：SPSS/Stata/SAS 默认启用 `_auto_detect_encoding`。\r\n\r\n## 快速开始\r\n\r\n```bash\r\n# 检查环境\r\npython scripts/check_env.py --install\r\n```\r\n\r\n完整代码示例：[`references/usage_examples.py`](./references/usage_examples.py)\r\n\r\nWorkBuddy 对话示例：\r\n```\r\n> convert data.sav to .dta\r\n> 读入 data.sav 并显示元数据\r\n> .sav 转 .xlsx 会丢失元数据吗？\r\n```\r\n\r\n## 扩展\r\n\r\n需要支持新格式？编辑 `scripts/reader_*.py` 添加读入函数，在 `scripts/reader_core.py` 的 `format_map` 中注册，并在 `scripts/reader_core.py` 中补充对应的 TypedDict 定义。\r\n\r\n## 格式限制与解决方案\r\n\r\n*已解决项标 ✅ / 新增能力标 🔄；未标项为固有格式限制（按字母排序，能力矩阵见上）。*\r\n\r\n### CDISC ODM (.odm)\r\n**读入：**\r\n- ❌ XML 结构依赖，嵌套解析取决于 ODM 文件结构规范性\r\n- ❌ ODM 规范本身不含统计元数据，仅保留临床数据结构\r\n\r\n### dBASE / FoxPro (.dbf)\r\n**读入：**\r\n- ❌ 字段名强制大写（格式限制）\r\n- ✅ 支持读+写\r\n\r\n### EpiData (.rec)\r\n**读入：**\r\n- ❌ 读入需经 R + `foreign` 包桥接（需显式 `allow_r_exec=True` 开启，默认禁用）\r\n- ❌ 统计元数据在 R → CSV 桥接过程中丢失\r\n\r\n### EpiInfo (.prj)\r\n**读入：**\r\n- ❌ 项目文件不含数据；自动搜索同名 CSV\r\n- ❌ Access 不支持，需先导出 CSV\r\n**写出：**\r\n- ✅ 变量标签和 codes 在 XML 结构中重建\r\n\r\n### Excel (.xlsx/.xls/.xlsm)\r\n**读入：**\r\n- ✅ **合并单元格**：用锚点值填充合并区域；通过 `fill_merged_cells=True`（默认）启用，可传 `False` 关闭\r\n- ❌ 公式丢失，仅保留计算结果\r\n- ❌ 图表/形状不提取\r\n**写出：**\r\n- 变量/值标签存储在独立元数据工作表中\r\n\r\n### HDF5 (.h5/.hdf5)\r\n**读入：**\r\n- ✅ **多层级数据集**：`pd.read_hdf` 失败时自动回退到 h5py，合并全部顶层数值数据集\r\n- ✅ **属性标签还原**：扫描常用属性键（`label`、`description`、`units`、`ColumnLabel`…）\r\n- ❌ 层级结构仍展平为顶级变量\r\n**写出：**\r\n- 写入时使用文件级属性存储元数据\r\n\r\n### JMP (.jmp)\r\n**读入：**\r\n- ❌ 依赖 jmpio-python，版本支持不一\r\n- ❌ 多表 JMP 文件仅返回第一个数据表\r\n**写出：**\r\n- 仅支持单表写出\r\n\r\n### MATLAB (.mat)\r\n**读入：**\r\n- ✅ **v7.3+（HDF5 格式）**：检测 `MATLAB 7.3` 文件头或 scipy 失败 → h5py 路径\r\n- ❌ 复杂结构（嵌套 cell、稀疏矩阵、函数句柄）→ 单列扁平化输出\r\n- ❌ Object 类和 datetime 列丢失类型保真度\r\n\r\n### Parquet (.parquet)\r\n**读入：**\r\n- ❌ 深层嵌套类型（>2 层 list<struct>）→ 不透明的 Python 对象列\r\n- ✅ **分区数据集**：含 `part-*.parquet` 或 Hive 分区子目录 → `pyarrow.dataset` 合并读取\r\n\r\n### R (.rda/.rds/.rdata)\r\n**读入：**\r\n- ✅ 旧版 ASCII XDR (RDA2)：**自动回退到 R**（推荐安装 R）\r\n- ❌ factor 顺序可能未保留为 Categorical 顺序\r\n- ❌ 多对象 RDA 文件：`read_all_r_objects()` 返回全部对象列表\r\n**写出：**\r\n- 通过 R 桥接（statdata_meta 属性）实现完整元数据往返\r\n\r\n### SAS (.sas7bdat/.xpt/.sas7bcat)\r\n**读入：**\r\n- ✅ 值标签定义在 `.sas7bcat`，需与数据文件同目录自动加载\r\n- ❌ Viya CAS `.sashdat` 不支持\r\n- 日期基准：1960-01-01\r\n\r\n### SPSS (.sav/.zsav/.por)\r\n**读入：**\r\n- ❌ MR Sets 读入为原始字典，语义需手动重建\r\n- ❌ 公式丢失，仅保留计算结果\r\n- ⚠️ 特殊缺失值（`.A`–`.Z`）在 `special_missing` 中标记；写出时需显式保留\r\n**写出：**\r\n- `.zsav` 写出需 pyreadstat 1.2+ 支持，否则自动降级为 `.sav`\r\n\r\n### Stata (.dta)\r\n**读入：**\r\n- ⚠️ 特殊缺失值（`.a`–`.z`）：`user_missing=True`（默认）时保留为字符标签；`user_missing=False` 时不可逆地变为 NaN\r\n- ✅ 旧版 (pre-v13) Latin-1 编码已自动检测\r\n- ❌ Stata 117-119：pyreadstat 1.3.5 不支持，写回时自动降级为 version 15\r\n\r\n## 安全 / Security\r\n\r\n- **R 执行默认隔离且需显式开启 / R execution opt-in & sandboxed**: 所有调用本地 R 解释器的路径——读入 `.rda/.rds/.RData`（`readRDS()/load()`）、读入 Minitab `.mtw/.mpj` 与 EpiData `.rec`（`foreign::read.mtb()/read.epiinfo()`）、写出 R 格式（`.rda/.rds`）——**默认禁用**。仅当对可信文件显式传入 `allow_r_exec=True` 时才会运行。纯 Python 解析器（`pyreadr`、`mtbpy`）优先使用，绝不执行代码。\r\n- **无静默 R 回退 / No silent R fallback**: 若纯 Python 解析失败且未设 `allow_r_exec`，技能会明确报错而非静默启动 R，从而消除对不可信文件执行嵌入代码的风险。\r\n- **R 脚本为静态模板 / R scripts are static templates**: 启用 R 路径时，所有 R 脚本均为完全静态模板；用户输入（文件路径、标签、元数据）仅经命令行参数（`commandArgs(trailingOnly=TRUE)` + `jsonlite`）传入，绝不拼进可执行 R 代码。临时文件名随机，无固定路径。\r\n- **临时 CSV 暴露（R 桥接）| Temporary CSV exposure (R bridge)**: 启用 R 时（opt-in），转换数据会先物化为磁盘临时 CSV 再读回。临时文件在用后即刻删除，但崩溃时或经备份/索引工具可能短暂留存。处理高度敏感数据请避开 R 桥接格式，或优先选用非 R 路径。\r\n- **无破坏性写入 / No destructive writes**: 写入已存在的 `.hyper` 文件时，新数据先写入临时文件；仅当写入成功后，原文件才轮转为 `<file>.bak`（上一份 `.bak` 降级为 `.bak.1`，绝不静默删除），随后临时文件原子替换入位。若写入失败，原文件保持不动。\r\n- **可选安装 / Optional install**: `python scripts/check_env.py --install` 仅在显式要求时运行。\r\n- **依赖已固定版本 / Pinned dependencies**: 核心依赖均带上限约束（`pandas`、`pyreadstat`、`pyreadr`），详见 `requirements.txt`。\r\n- **范围 / Scope**: 仅统计数据格式。除非你显式要求安装包，否则不访问网络。\r\n\r\n## 许可证\r\n\r\nMIT 许可证。详见 [LICENSE](LICENSE)。\n\nFile v2.1.0:skill-card.md\n\n## Description: <br>\nRead and convert 50+ statistical software and clinical trial data formats into Python/pandas, preserving variable labels, value labels, and missing-value metadata where the source format supports them. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[medstatstar](https://clawhub.ai/user/medstatstar) <br>\n\n### License/Terms of Use: <br>\nMIT <br>\n\n\n## Use Case: <br>\nDevelopers, data engineers, statisticians, and clinical data teams use this skill to inspect, read, and convert statistical datasets across formats such as SPSS, Stata, SAS, R, Excel, Parquet, HDF5, JSON, and CSV while surfacing metadata-loss warnings. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Optional environment setup can modify the local Python environment when package installation is explicitly requested. <br>\nMitigation: Run environment checks without installation by default and use `python scripts/check_env.py --install` only in an environment where dependency changes are acceptable. <br>\nRisk: Some R-backed formats can invoke a local R interpreter when `allow_r_exec=True` is enabled for trusted files. <br>\nMitigation: Keep R execution disabled for untrusted files and prefer pure-Python parsing paths for sensitive or unknown datasets. <br>\nRisk: The opt-in R bridge may briefly materialize converted data in a temporary CSV file. <br>\nMitigation: Avoid R-backed conversion paths for highly sensitive data unless the working environment and temporary-file handling are acceptable. <br>\n\n\n## Reference(s): <br>\n- [ClawHub Skill Page](https://clawhub.ai/medstatstar/skills/statdata-transfer) <br>\n- [Project Homepage](https://github.com/medstatstar/statdata-transfer) <br>\n- [README](artifact/README.md) <br>\n- [Chinese README](artifact/README_ZH.md) <br>\n- [Usage Examples](artifact/references/usage_examples.py) <br>\n- [New Formats Architecture Analysis](artifact/references/new_formats_architecture_analysis.json) <br>\n- [Version 1.4 Implementation Summary](artifact/references/v1.4_implementation_summary.json) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Text, Markdown, Code, Shell commands, Configuration, Guidance] <br>\n**Output Format:** [Markdown guidance with Python examples, shell commands, and structured conversion results] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [May produce pandas DataFrames, metadata dictionaries, warning lists, and converted data files when used with local files.] <br>\n\n## Skill Version(s): <br>\n2.1.0 (source: server release metadata and SKILL.md frontmatter) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v2.1.0:test_report.md\n\n# statdata-transfer v1.4.0 端到端测试报告\n\n## 测试环境\n- Python: Anaconda (C:\\Tools\\anaconda3\\python.exe)\n- pandas: 2.3.3\n- pyreadstat: 1.3.5\n- R: C:\\Tools\\R-4.5.1\\bin\\x64\\Rscript.exe\n\n## 测试结果\n\n### 基础测试 (test_e2e.py) - 6/6 通过 ✓\n| 格式 | 状态 | 说明 |\n|------|------|------|\n| SPSS .sav | ✓ 通过 | 使用 pyreadstat 读写 |\n| Stata .dta | ✓ 通过 | 使用 pyreadstat 读写 |\n| Excel .xlsx | ✓ 通过 | 使用 pandas 读写 |\n| R .rda | ✓ 通过 | 使用 R bridge 读取 |\n| CSV | ✓ 通过 | 新增支持 |\n| Parquet | ✓ 通过 | 修复 pyarrow API 兼容性问题 |\n\n### 扩展测试 (test_e2e_extended.py) - 3/4 通过\n| 格式 | 状态 | 说明 |\n|------|------|------|\n| SAS .sas7bdat | ✗ 失败 | 测试脚本问题（sas7bdat 包不支持写入） |\n| MATLAB .mat | ✓ 通过 | 数据维度需优化 |\n| HDF5 .h5 | ✓ 通过 | 修复 json 模块导入问题 |\n| JSON | ✓ 通过 | 使用 pandas read_json |\n\n## 修复的 Bug\n\n### 1. _normalize_value_labels() 参数数量错误\n- **文件**: `scripts/reader_core.py`\n- **问题**: 函数定义只有 2 个参数，但调用时传了 3 个\n- **修复**: 添加第 3 个参数 `value_labels: dict = None`\n\n### 2. CSV 格式不支持\n- **文件**: `scripts/reader_core.py`, `scripts/reader_modern.py`\n- **问题**: `.csv` 扩展名不在支持列表中\n- **修复**: 添加 CSV 格式支持和 `_read_csv()` 函数\n\n### 3. Parquet 读取错误\n- **文件**: `scripts/reader_science.py`\n- **问题**: `rg.total_compressed_size` 属性在 pyarrow 21.0.0 中不存在\n- **修复**: 使用 `hasattr()` 检查属性是否存在\n\n### 4. HDF5 读取错误\n- **文件**: `scripts/reader_science.py`\n- **问题**: `json` 模块未导入\n- **修复**: 添加 `import json`\n\n## 新增功能\n\n### CSV 格式支持\n- 自动检测文件编码（utf-8-sig, utf-8, gbk, gb2312, latin-1）\n- 返回标准 StatFileResult 格式\n\n## ClawHub 合规性检查\n\n### ✅ 已完成的检查项\n- [x] `version` 字段在 SKILL.md frontmatter 中\n- [x] `.clawhubignore` 存在并包含测试文件\n- [x] `README.md` 和 `README_EN.md` 存在\n- [x] `SKILL.md` 包含使用示例\n- [x] `references/formats_detail.md` 存在\n- [x] 所有 Python 文件语法检查通过\n- [x] 代码文件已拆分（10 个子模块）\n\n### ⚠️ 待优化项\n- `reader_core.py` (764 行) 和 `reader_v14.py` (674 行) 仍然较大\n- 部分新格式（JMP, Minitab, Prism 等）需要专有软件支持，当前为占位实现\n\n## 测试覆盖率\n\n### 已测试格式 (9/25+)\n- ✓ SPSS (.sav)\n- ✓ Stata (.dta)\n- ✓ Excel (.xlsx)\n- ✓ R (.rda)\n- ✓ CSV (.csv)\n- ✓ Parquet (.parquet)\n- ✓ MATLAB (.mat)\n- ✓ HDF5 (.h5)\n- ✓ JSON (.json)\n\n### 未测试格式 (需要额外依赖或软件)\n- SAS (.sas7bdat, .xpt) - 需要 SAS 或 pyreadstat\n- JMP (.jmp) - 需要 JMP 软件或 jmpio 包\n- Minitab (.mtw) - 需要 Minitab 软件\n- Prism (.pzfx) - 需要 Prism 软件或 pzfx 包\n- jamovi (.omv) - 需要 jamovi 软件\n- EpiData (.rec) - 需要 EpiData 软件\n- EViews (.wf1) - 需要 EViews 软件\n\n## 结论\n\nstatdata-transfer v1.4.0 技能已完成端到端测试，核心功能正常。发现的 bug 已全部修复，技能符合 ClawHub 发布标准。\n\n建议：\n1. 补充更多格式的测试数据文件\n2. 优化 MATLAB 读取的数据维度处理\n3. 考虑进一步拆分大文件（reader_core.py, reader_v14.py）\n\nFile v2.1.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 Wintone Zhang, Phoebe Zhang\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALIN\n\nArchive v2.0.3: 30 files, 108954 bytes\n\nFiles: assets/logo.svg (3750b), LICENSE (1084b), README_ZH.md (12432b), README.md (12770b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3379b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1701b), scripts/__init__.py (531b), scripts/check_env.py (4395b), scripts/reader_arff.py (9220b), scripts/reader_core.py (40977b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (10290b), scripts/reader_gretl.py (9623b), scripts/reader_legacy.py (20316b), scripts/reader_modern.py (16848b), scripts/reader_odm.py (10187b), scripts/reader_r.py (40967b), scripts/reader_sas.py (6241b), scripts/reader_science.py (46550b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (15814b), scripts/reader_v14.py (26984b), scripts/writer.py (30651b), skill-card.md (3155b), SKILL.md (7298b), test_report.md (3399b), _meta.json (136b)\n\nArchive v2.0.2: 30 files, 106749 bytes\n\nFiles: assets/logo.svg (3750b), LICENSE (1084b), README_ZH.md (11650b), README.md (11925b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3379b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1701b), scripts/__init__.py (531b), scripts/check_env.py (4395b), scripts/reader_arff.py (9220b), scripts/reader_core.py (40890b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (10290b), scripts/reader_gretl.py (9623b), scripts/reader_legacy.py (20316b), scripts/reader_modern.py (16848b), scripts/reader_odm.py (10187b), scripts/reader_r.py (38673b), scripts/reader_sas.py (6241b), scripts/reader_science.py (46560b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (15573b), scripts/reader_v14.py (26984b), scripts/writer.py (30651b), skill-card.md (2800b), SKILL.md (6081b), test_report.md (3399b), _meta.json (136b)\n\nArchive v2.0.1: 30 files, 106192 bytes\n\nFiles: assets/logo.svg (3750b), LICENSE (1084b), README_ZH.md (11554b), README.md (11781b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3379b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1701b), scripts/__init__.py (531b), scripts/check_env.py (4395b), scripts/reader_arff.py (9220b), scripts/reader_core.py (40890b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (10290b), scripts/reader_gretl.py (9623b), scripts/reader_legacy.py (20316b), scripts/reader_modern.py (16848b), scripts/reader_odm.py (10187b), scripts/reader_r.py (38673b), scripts/reader_sas.py (6241b), scripts/reader_science.py (46041b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (15573b), scripts/reader_v14.py (26984b), scripts/writer.py (30651b), skill-card.md (2604b), SKILL.md (6083b), test_report.md (3399b), _meta.json (136b)\n\nArchive v2.0.0: 30 files, 108565 bytes\n\nFiles: assets/logo.svg (3750b), LICENSE (1084b), README_ZH.md (11500b), README.md (12041b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3378b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1701b), scripts/__init__.py (531b), scripts/check_env.py (4395b), scripts/reader_arff.py (9220b), scripts/reader_core.py (40859b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (10290b), scripts/reader_gretl.py (9623b), scripts/reader_legacy.py (20316b), scripts/reader_modern.py (16848b), scripts/reader_odm.py (10187b), scripts/reader_r.py (38673b), scripts/reader_sas.py (6241b), scripts/reader_science.py (46041b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (15573b), scripts/reader_v14.py (26984b), scripts/writer.py (30647b), skill-card.md (2735b), SKILL.md (12676b), test_report.md (3399b), _meta.json (136b)\n\nArchive v1.7.11: 30 files, 108646 bytes\n\nFiles: assets/logo.svg (3750b), LICENSE (1084b), README_ZH.md (11500b), README.md (12041b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3378b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1701b), scripts/__init__.py (531b), scripts/check_env.py (4395b), scripts/reader_arff.py (9220b), scripts/reader_core.py (40859b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (10290b), scripts/reader_gretl.py (9623b), scripts/reader_legacy.py (20316b), scripts/reader_modern.py (16848b), scripts/reader_odm.py (10187b), scripts/reader_r.py (38673b), scripts/reader_sas.py (6241b), scripts/reader_science.py (46041b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (15573b), scripts/reader_v14.py (26984b), scripts/writer.py (30647b), skill-card.md (2903b), SKILL.md (12676b), test_report.md (3399b), _meta.json (137b)\n\nArchive v1.7.10: 29 files, 96531 bytes\n\nFiles: assets/logo.svg (3750b), LICENSE (1084b), README_ZH.md (10429b), README.md (12041b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3378b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1416b), scripts/__init__.py (531b), scripts/check_env.py (4152b), scripts/reader_arff.py (9220b), scripts/reader_core.py (38154b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (8184b), scripts/reader_gretl.py (9623b), scripts/reader_modern.py (16848b), scripts/reader_odm.py (10187b), scripts/reader_r.py (38673b), scripts/reader_sas.py (6241b), scripts/reader_science.py (35355b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_tableau.py (14339b), scripts/reader_v14.py (26984b), scripts/writer.py (30359b), skill-card.md (3226b), SKILL.md (10617b), test_report.md (3399b), _meta.json (137b)\n\nArchive v1.7.9: 28 files, 90796 bytes\n\nFiles: assets/logo.svg (3750b), LICENSE (1084b), README_ZH.md (10429b), README.md (12041b), references/new_formats_architecture_analysis.json (4852b), references/usage_examples.py (3378b), references/v1.4_implementation_summary.json (2703b), requirements.txt (1289b), scripts/__init__.py (531b), scripts/check_env.py (4035b), scripts/reader_arff.py (9220b), scripts/reader_core.py (37596b), scripts/reader_epinfo.py (13497b), scripts/reader_excel.py (8184b), scripts/reader_gretl.py (9623b), scripts/reader_modern.py (16848b), scripts/reader_odm.py (10187b), scripts/reader_r.py (38673b), scripts/reader_sas.py (6241b), scripts/reader_science.py (35355b), scripts/reader_spss.py (4591b), scripts/reader_stata.py (4588b), scripts/reader_v14.py (26984b), scripts/writer.py (30044b), skill-card.md (2894b), SKILL.md (10086b), test_report.md (3399b), _meta.json (136b)","readmeExcerpt":"Skill: statdata-transfer Owner: medstatstar Summary: 读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明（完整）：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；写入主输出文件的同时可能生成 sidecar 元数据（CSV/TSV 旁 <名>_metadata.json、Parquet/Arrow 内嵌）及覆盖 .hyper 时的 .bak/.bak.1 备份；处理 .rda/.rds/.RData/.mtw/.mpj/.rec 时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。 / Read/convert 50+ statistical software formats, preserving variable/value labels","codeSnippets":[],"executableExamples":[],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\r\nslug: statdata-transfer\r\nname: statdata-transfer\r\ndisplayName: 统计数据格式转换器 / Statistical Data Format Converter\r\ncn_name: 统计数据格式转换器\r\nversion: 2.2.1\r\nsummary: 读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明（完整）：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；写入主输出文件的同时，可能生成 sidecar 元数据文件（CSV/TSV 旁生成 <名>_metadata.json，Parquet/Arrow 内嵌元数据）及覆盖 .hyper 时的 .bak/.bak.1 备份；处理 .rda/.rds/.RData/.mtw/.mpj/.rec 文件时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。\r\nlicense: MIT\r\ndescription: \"读入/转存 50+ 统计软件格式，对统计二进制格式完整保留变量标签/值标签/特殊缺失值等元数据。副作用声明（完整）：运行环境检查（scripts/check_env.py）；可应要求 pip 安装缺失包；写入主输出文件的同时可能生成 sidecar 元数据（CSV/TSV 旁 <名>_metadata.json、Parquet/Arrow 内嵌）及覆盖 .hyper 时的 .bak/.bak.1 备份；处理 .rda/.rds/.RData/.mtw/.mpj/.rec 时可调用本地 R 解释器，但该回退默认禁用，需 allow_r_exec=True 显式开启。 / Read/convert 50+ statistical software formats, preserving variable/value labels and missing-value metadata for binary stats formats. FULL side effects: runs environment checks (scripts/check_env.py); may optionally pip-install missing packages on request; writes the main output file AND may emit sidecar metadata (e.g. <name>_metadata.json beside CSV/TSV, embedded in Parquet/Arrow schema) and .bak/.bak.1 backups when overwriting .hyper; can invoke the local R interpreter for .rda/.rds/.RData/.mtw/.mpj/.rec files via a fallback DISABLED by default and opted in only with allow_r_exec=True.\"\r\ntriggers:\r\n  - \"statdata-transfer\"\r\n  - \"统计数据格式转换\"\r\n  - \"spss stata sas 格式\"\r\n  - \".sav .dta .sas7bdat 读入\"\r\n  - \"sav转dta 格式转换\"\r\n  - \"variable labels 变量标签\"\r\n  - \"metadata-preserved conversion\"\r\nrequired_commands: [python]\r\ninvocable: true\r\nmetadata:\r\n  openclaw: { emoji: \"🛠️\", icon: \"assets/logo.svg\" }\r\n  authors: [\"medstatstar\", \"phoe-zip\"]\r\n  license: \"MIT\"\r\n  tags: [\"data-conversion\", \"statistics\", \"spss\", \"stata\", \"sas\", \"clinical-trials\", \"metadata\", \"pandas\", \"bidirectional\"]\r\n  homepage: \"https://github.com/medstatstar/statdata-transfer\"\r\npermissions:\r\n  scope: \"user-space-only\"\r\n  network: \"off\"\r\n  network_note: \"Offline by default; the only network touchpoint is the optional `python scripts/check_env.py --install`, which pip-installs missing packages and runs ONLY on explicit user request.\"\r\n  filesystem: \"read-only to its own files; reads the input data file you specify; writes the converted output file to a path you specify, and may additionally create sidecar metadata files (e.g. <name>_metadata.json beside CSV/TSV, or metadata embedded in Parquet/Arrow schema) and .bak/.bak.1 backups when overwriting .hyper\"\r\n  data: \"no external data transmission\"\r\n---\r\n\r\n# Statistical Data Format Converter\r\n\r\n> **Safe by default — preview, not execute**: the skill shows what it will read/convert and only writes a file when you explicitly ask. Every R-invoking path is opt-in and disabled by default.\r\n\r\n## Language\r\n\r\n- **English guide** → [README.md](https://github.com/medstatstar/statdata-transfer/blob/main/README.md)\r\n- **中文指南** → [README_zh-CN.md](https://github.com/medstatstar/statdata-transfer/blob/main"},{"path":"README.md","content":"# statdata-transfer / Statistical Data Format Converter\r\n\r\n[🇨🇳 Chinese](./README_zh-CN.md)\r\n\r\n<div align=\"center\">\r\n<img src=\"assets/icon.svg\" width=\"240\" height=\"240\" />\r\n</div>\r\n\r\n---\r\n\r\n> Read 50+ statistical-software and clinical-trial data formats, and **inter-convert between most of them** while keeping variable/value labels and missing-value metadata. No statistical software required — format conversion only.\r\n\r\n## How to use it in a conversation\r\n\r\nJust talk to the agent in natural language. A few real examples (copy-paste ready):\r\n\r\n**① Most common — convert a file**\r\n- **You say**: `convert C:/Users/Name/Desktop/data.sav to .dta`\r\n- **Agent replies** (sketch): reads `data.sav` with pyreadstat, preserves all variable/value labels, and writes `data.dta` in the same folder.\r\n- **Trigger the real conversion**: by default the agent previews the plan; say `please write the file` to execute.\r\n\r\n**② Show what's inside**\r\n- **You say**: `read data.sav and show metadata`\r\n- **Agent replies**: prints the DataFrame shape, variable labels, value labels, and a list of which metadata will be preserved.\r\n\r\n**③ Check before you lose data**\r\n- **You say**: `will converting .sav to .xlsx lose any metadata?`\r\n- **Agent replies**: warns that Excel keeps labels only in a side sheet; suggests Parquet/Stata to keep them losslessly.\r\n\r\n**④ Ask for reproducible code**\r\n- **You say**: `show me the Python code to convert .sav to .parquet`\r\n- **Agent replies**: prints the `read_stat_file` / `write_stat_file` snippet (code is always English).\r\n\r\n**⑤ Switch language**\r\n- **You say**: `reply in Chinese` / `switch to English` — all user-facing messages follow your OS language or this prompt.\r\n\r\n## What can it do? (scenario index)\r\n\r\n| Capability | Typical use | Try saying |\r\n|:---|:---|:---|\r\n| **Read 50+ formats** | Open SPSS/Stata/SAS/R/Excel/Parquet/HDF5/JSON… into pandas | `read data.sav and show metadata` |\r\n| **Convert between stats formats** | SPSS ↔ Stata ↔ R ↔ SAS XPT, keeping all labels | `convert data.sav to .dta keeping variable labels` |\r\n| **Export universal formats** | Parquet / Feather / HDF5 / JSON / CSV / Excel with labels embedded | `save to parquet but keep value labels` |\r\n| **Metadata-safe round-trip** | Labels survive a convert-and-convert-back | `convert to parquet then back to sav, keep labels` |\r\n| **Metadata-loss warning** | Know what will be dropped before exporting | `will .sav to .xlsx lose metadata?` |\r\n| **Batch / folder** | Convert a whole folder or a zip archive | `convert all .dta in this zip to .sav` |\r\n\r\nFull format list and per-format limits: see **Advanced reference** below.\r\n\r\n## First-use FAQ\r\n\r\n- **Do I need SPSS/Stata/R installed?** No. The skill is pure Python; it only *optionally* calls a local R interpreter for a few formats (Minitab/EpiData/R write), and only when you pass `allow_r_exec=True`.\r\n- **How do I get the actual converted file, not just code?** Say `please write the file`. By default it previews; execution is e"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7amqq1jv28skb63wavr6shah89jsm5\",\n  \"slug\": \"statdata-transfer\",\n  \"version\": \"2.2.1\",\n  \"publishedAt\": 1785668532225\n}"},{"path":"references/new_formats_architecture_analysis.json","content":"{\r\n  \"current_architecture\": {\r\n    \"data\": \"pandas DataFrame\",\r\n    \"metadata\": \"BaseMeta TypedDict (~28 fields)\",\r\n    \"column_report\": \"ColumnInfo TypedDict (12 fields)\",\r\n    \"return_type\": \"StatFileResult = dict[str, Any] with keys: dataframe, metadata, warnings, column_report\",\r\n    \"multi_object_pattern\": \"read_all_*() returns dict[str, StatFileResult]\"\r\n  },\r\n  \"formats_analysis\": {\r\n    \"sas7bcat\": {\r\n      \"description\": \"SAS Ŀ¼�ļ����洢��ʽ���壨ֵ��ǩ��\",\r\n      \"data_structure\": \"�����ݣ�ֻ��Ԫ���ݣ���ʽ���壩\",\r\n      \"metadata_fields\": [\r\n        \"value_labels\",\r\n        \"variable_value_labels\",\r\n        \"variable_to_label\"\r\n      ],\r\n      \"current_architecture_sufficient\": true,\r\n      \"notes\": \"pyreadstat ��֧�ֶ�ȡ .sas7bcat������ value_labels dict\",\r\n      \"architecture_extension_needed\": false\r\n    },\r\n    \"jmp\": {\r\n      \"description\": \"SAS JMP �����ļ����ɰ���������ݱ����ű����������\",\r\n      \"data_structure\": \"�ɰ���������ݱ���Data Table����ÿ������һ�� DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"value_labels\",\r\n        \"column_properties (formulas, ranges)\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"JMP �ļ��ɰ���������ݱ�����Ҫ read_all_jmp_tables() ģʽ�������ԣ���ʽ����Χ����Ҫ��չ ColumnInfo\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"ColumnInfo ��Ҫ��չ��formula (str), range (dict), column_property (dict)\",\r\n        \"��Ҫ���� JmpMeta �࣬���� tables list��scripts list��analysis list\",\r\n        \"��Ҫ���� read_all_jmp_tables() ����\"\r\n      ]\r\n    },\r\n    \"minitab\": {\r\n      \"description\": \"Minitab �������ļ����ɰ��������������Worksheet��\",\r\n      \"data_structure\": \"�ɰ��������������ÿ����������һ�� DataFrame\",\r\n      \"metadata_fields\": [\r\n        \"variable_labels\",\r\n        \"worksheet_names\",\r\n        \"formulas\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"Minitab �������ɰ����������������Ҫ read_all_minitab_worksheets() ģʽ\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� MinitabMeta �࣬���� worksheets list��active_worksheet str\",\r\n        \"��Ҫ���� read_all_minitab_worksheets() ����\",\r\n        \"ColumnInfo ������Ҫ��չ��formula (str)\"\r\n      ]\r\n    },\r\n    \"prism\": {\r\n      \"description\": \"GraphPad Prism ��Ŀ�ļ����������ݱ����������ͼ��\",\r\n      \"data_structure\": \"�������ݱ���DataFrame�����������DataFrame����ͼ�Σ��޷�תΪ DataFrame��\",\r\n      \"metadata_fields\": [\r\n        \"data_tables\",\r\n        \"results_tables\",\r\n        \"graphs_info\"\r\n      ],\r\n      \"current_architecture_sufficient\": false,\r\n      \"notes\": \"Prism �ļ��������ݱ��ͽ���������߶��� DataFrame��ͼ���޷�����Ϊ DataFrame\",\r\n      \"architecture_extension_needed\": true,\r\n      \"extension_details\": [\r\n        \"��Ҫ���� PrismMeta �࣬���� data_tables list��results_tables list��graphs_info list\",\r\n        \"StatFileResult ��Ҫ��չ������� read_all_prism_tables() ģʽ\",\r\n        \"��ǰ�ܹ�ֻ�ܱ������ݱ����������ͼ����Ϣ�ᶪʧ\"\r\n      ]\r\n    },\r\n    \"jamovi\": {\r\n  "},{"path":"references/v1.4_implementation_summary.json","content":"{\r\n  \"v1.4_new_formats\": [\r\n    {\r\n      \"format\": \"SAS Catalog\",\r\n      \"ext\": \".sas7bcat\",\r\n      \"handler\": \"_read_sas_catalog\",\r\n      \"dependency\": \"pyreadstat (已支持)\",\r\n      \"architecture\": \"SasCatalogMeta, 返回格式定义 DataFrame\",\r\n      \"status\": \"done\"\r\n    },\r\n    {\r\n      \"format\": \"JMP\",\r\n      \"ext\": \".jmp\",\r\n      \"handler\": \"_read_jmp\",\r\n      \"dependency\": \"jmpio-python (PyPI) 或 @skill:statsoft-cli\",\r\n      \"architecture\": \"JmpMeta, 多表需 read_all_jmp_tables()\",\r\n      \"status\": \"done (handler 已添加，jmpio 未安装）\"\r\n    },\r\n    {\r\n      \"format\": \"Minitab\",\r\n      \"ext\": \".mtw/.mpj\",\r\n      \"handler\": \"_read_minitab\",\r\n      \"dependency\": \"mtbpy 或 R foreign::read.mtb() 中继\",\r\n      \"architecture\": \"MinitabMeta, 多工作表需 read_all_minitab_worksheets()\",\r\n      \"status\": \"done (R 中继已实现）\"\r\n    },\r\n    {\r\n      \"format\": \"GraphPad Prism\",\r\n      \"ext\": \".pzfx/.pz\",\r\n      \"handler\": \"_read_prism\",\r\n      \"dependency\": \"pzfx (PyPI) 或 @skill:statsoft-cli\",\r\n      \"architecture\": \"PrismMeta, 含数据表+结果表\",\r\n      \"status\": \"done (handler 已添加，pzfx 未安装）\"\r\n    },\r\n    {\r\n      \"format\": \"jamovi\",\r\n      \"ext\": \".omv\",\r\n      \"handler\": \"_read_jamovi\",\r\n      \"dependency\": \"无需额外包（ZIP+CSV 解析）\",\r\n      \"architecture\": \"JamoviMeta, 含 analysis JSON\",\r\n      \"status\": \"done\"\r\n    },\r\n    {\r\n      \"format\": \"EpiData\",\r\n      \"ext\": \".rec\",\r\n      \"handler\": \"_read_epidata\",\r\n      \"dependency\": \"R foreign::read.epiinfo() 中继\",\r\n      \"architecture\": \"EpidataMeta\",\r\n      \"status\": \"done (R 中继已实现）\"\r\n    },\r\n    {\r\n      \"format\": \"EViews\",\r\n      \"ext\": \".wf1/.wf2\",\r\n      \"handler\": \"_read_eviews\",\r\n      \"dependency\": \".wf2 可直接解析 JSON；.wf1 需 @skill:statsoft-cli\",\r\n      \"architecture\": \"EviewsMeta\",\r\n      \"status\": \"done (.wf2 解析已实现）\"\r\n    }\r\n  ],\r\n  \"architecture_extensions\": [\r\n    \"ColumnInfo 新增：formula (str), column_property (dict)\",\r\n    \"新增 Meta 类：SasCatalogMeta, JmpMeta, MinitabMeta, PrismMeta, JamoviMeta, EpidataMeta, EviewsMeta\",\r\n    \"__all__ 新增导出：7 个新 Meta 类名\"\r\n  ],\r\n  \"files_modified\": [\r\n    \"scripts/stat_reader.py (+459 行，共 3131 行）\",\r\n    \"scripts/check_env.py (新增 jmpio, pzfx 检测）\",\r\n    \"SKILL.md (待更新）\",\r\n    \"references/new_formats_architecture_analysis.json (新增）\"\r\n  ],\r\n  \"remaining_work\": [\r\n    \"更新 SKILL.md（添加 7 种新格式详情）\",\r\n    \"添加 read_all_jmp_tables() / read_all_minitab_worksheets()\",\r\n    \"测试新 handler（需要实际文件）\",\r\n    \"安装 jmpio/pzfx 包（或配置 @skill:statsoft-cli）\"\r\n  ]\r\n}"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1795,"uniquenessScore":40,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T01:52:55.506Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T01:52:55.506Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T07:45:03.001Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}