{"id":"01fad05b-2880-4598-98eb-65d1702ca7e9","entityType":"agent","slug":"clawhub-gopendrasharma89-tech-clean-csv-toolkit","name":"Clean CSV Toolkit","canonicalUrl":"https://www.xpersona.co/agent/clawhub-gopendrasharma89-tech-clean-csv-toolkit","canonicalPath":"/agent/clawhub-gopendrasharma89-tech-clean-csv-toolkit","generatedAt":"2026-10-11T07:39:55.845Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-11T05:03:04.697Z","emptyReason":null},"description":"Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate... Skill: Clean CSV Toolkit Owner: gopendrasharma89-tech Summary: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate... Tags: aggregation:0.3.0, column:0.5.0, concat:0.4.0, convert:0.2.0, csv:0.5.0, data:0.3.0, dedupe:0.2.0, derive:0.5.0, diff:0.2.0, expression:0.5.0, filter:0.5.0, groupby:0.4.0, head:0.2.0, join:0","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17cp87fy279ggwne1mqcb6675843tdy:clean-csv-toolkit","sourceUrl":"https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit","homepage":"https://clawhub.ai/gopendrasharma89-tech/skills/clean-csv-toolkit","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/gopendrasharma89-tech/skills/clean-csv-toolkit","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":61,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate..."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T05:03:04.697Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T05:03:04.697Z","emptyReason":null},"stars":null,"forks":null,"downloads":1152,"packageName":null,"latestVersion":"0.5.0","tractionLabel":"1.2K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T05:03:04.617Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T05:03:04.697Z","lastCrawledAt":"2026-10-11T05:03:04.617Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T05:03:04.617Z","lastVerifiedAt":null,"highlights":[{"version":"0.5.0","createdAt":"2026-05-21T07:16:07.997Z","changelog":"v0.5.0: Add scripts/transform.py - derived columns + schema operations. Hand-rolled tokenizer + recursive-descent parser (no eval, no subprocess) supports arithmetic, string concat, parens, comparisons, function calls (upper/lower/strip/len/abs/round/int/float/str/replace/split/join/coalesce/year/month/day). Six ops: --add, --set, --drop, --rename, --cast (int/float/bool/string), --keep. Schema computed symbolically before streaming so empty cells in arithmetic don't crash the pipeline; per-row errors leave the derived value empty but continue processing. Bug fixed: schema-detection pass was evaluating expressions against empty rows. All 14 scripts now use the same safe-path policy and 0/1/2 exit codes.","fileCount":21,"zipByteSize":55570},{"version":"0.4.0","createdAt":"2026-05-18T11:48:46.067Z","changelog":"v0.4.0: Add scripts/filter.py (safe-predicate row filter, no eval), scripts/sort.py (type-aware stable multi-column sort, numeric auto-detect), scripts/concat.py (vertical UNION ALL with header union or strict-match, --add-source, --dedupe). Bug fix: filter.py 'in' operator now slurps comma-separated RHS tokens correctly. All scripts stream where possible; 100k-row sort in ~0.3s. Safe-path policy and 0/1/2 exit-code contract preserved across all 13 scripts.","fileCount":19,"zipByteSize":47298},{"version":"0.3.0","createdAt":"2026-05-17T03:59:33.629Z","changelog":"v0.3.0: Add scripts/merge.py (CSV/TSV/JSONL joins, inner/left/right/outer, --on / --left-on / --right-on, suffix disambiguation) and scripts/pivot.py (group-by aggregations: count/sum/avg/min/max/first/last/nunique; wide pivots via --pivot-on; numeric-aware --sort-by; CSV/TSV/JSONL/MD output). Both scripts stream rows where possible: 100k-row pivot < 1s, 50k x 200k merge ~1.3s. Safe-path policy, exit-code contract, and zero third-party dependencies preserved.","fileCount":16,"zipByteSize":37031},{"version":"0.2.0","createdAt":"2026-05-15T15:12:35.117Z","changelog":"v0.2.0 adds three new preview helpers: head.py prints first N rows in csv/tsv/jsonl/md/aligned format with optional --columns subset; tail.py prints last N rows using a bounded ring buffer that works on multi-GB files without loading them; sample.py picks a uniformly random sample of N rows via reservoir sampling (algorithm R, single pass, O(N) memory) with optional --seed for reproducibility and --preserve-order to keep original row order. All three share the same --as csv|tsv|jsonl|md|aligned, --output, and --columns flags, mirroring the convention already used by convert.py. Default output format is aligned, a fixed-width text table that an agent can paste straight into a reply. Performance: 100k-row / 1.6 MB CSV processed by head in 50ms, tail in 180ms, sample in 260ms. 13 end-to-end tests cover all three scripts across 5 output formats, column subset, seed reproducibility, preserve-order, --output to file, and 5 error paths (missing input, unsafe path, negative n, n=0, invalid format). No breaking changes - every v0.1.0 CLI flag, output format, and 0/1/2 exit-code contract preserved.","fileCount":14,"zipByteSize":28963},{"version":"0.1.0","createdAt":"2026-05-12T12:13:30.830Z","changelog":"v0.1.0 initial release. Local CSV/TSV/JSONL toolkit. Five scripts: inspect.py profiles a tabular file with auto-detected column types (int/float/bool/date/datetime/string), null counts, distincts, samples, encoding, and dialect. validate.py checks a file against a small JSON schema (required columns, per-column type, min/max, enum, regex, unique). dedupe.py removes duplicates by full-row or key columns with optional --keep first/last, --case-insensitive, --trim, and JSONL removed-rows report. diff.py compares two files by key column(s) and classifies rows as added/removed/changed/unchanged with per-column before/after for changed rows. convert.py converts between csv/tsv/jsonl/json/md. Pure Python 3 standard library, no pandas, no numpy, no subprocess, no remote calls. Consistent 0/1/2 exit codes across all scripts. 26 end-to-end tests covering CSV/TSV/JSONL inputs, schema validation, full-row and keyed dedup, file diffing, round-trip format conversion, 10k-row performance, error paths.","fileCount":10,"zipByteSize":21140}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17cp87fy279ggwne1mqcb6675843tdy:clean-csv-toolkit","setupComplexity":"low","setupSteps":["Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T07:39:55.842Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-csv-toolkit/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-11T05:03:04.697Z","emptyReason":null},"readme":"Skill: Clean CSV Toolkit\n\nOwner: gopendrasharma89-tech\n\nSummary: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate...\n\nTags: aggregation:0.3.0, column:0.5.0, concat:0.4.0, convert:0.2.0, csv:0.5.0, data:0.3.0, dedupe:0.2.0, derive:0.5.0, diff:0.2.0, expression:0.5.0, filter:0.5.0, groupby:0.4.0, head:0.2.0, join:0.4.0, jsonl:0.5.0, latest:0.5.0, markdown:0.2.0, merge:0.5.0, pivot:0.5.0, sample:0.2.0, sort:0.5.0, stdlib:0.5.0, tail:0.2.0, transform:0.5.0, tsv:0.5.0, union:0.4.0, validation:0.4.0\n\nVersion history:\n\nv0.5.0 | 2026-05-21T07:16:07.997Z | user\n\nv0.5.0: Add scripts/transform.py - derived columns + schema operations. Hand-rolled tokenizer + recursive-descent parser (no eval, no subprocess) supports arithmetic, string concat, parens, comparisons, function calls (upper/lower/strip/len/abs/round/int/float/str/replace/split/join/coalesce/year/month/day). Six ops: --add, --set, --drop, --rename, --cast (int/float/bool/string), --keep. Schema computed symbolically before streaming so empty cells in arithmetic don't crash the pipeline; per-row errors leave the derived value empty but continue processing. Bug fixed: schema-detection pass was evaluating expressions against empty rows. All 14 scripts now use the same safe-path policy and 0/1/2 exit codes.\n\nv0.4.0 | 2026-05-18T11:48:46.067Z | user\n\nv0.4.0: Add scripts/filter.py (safe-predicate row filter, no eval), scripts/sort.py (type-aware stable multi-column sort, numeric auto-detect), scripts/concat.py (vertical UNION ALL with header union or strict-match, --add-source, --dedupe). Bug fix: filter.py 'in' operator now slurps comma-separated RHS tokens correctly. All scripts stream where possible; 100k-row sort in ~0.3s. Safe-path policy and 0/1/2 exit-code contract preserved across all 13 scripts.\n\nv0.3.0 | 2026-05-17T03:59:33.629Z | user\n\nv0.3.0: Add scripts/merge.py (CSV/TSV/JSONL joins, inner/left/right/outer, --on / --left-on / --right-on, suffix disambiguation) and scripts/pivot.py (group-by aggregations: count/sum/avg/min/max/first/last/nunique; wide pivots via --pivot-on; numeric-aware --sort-by; CSV/TSV/JSONL/MD output). Both scripts stream rows where possible: 100k-row pivot < 1s, 50k x 200k merge ~1.3s. Safe-path policy, exit-code contract, and zero third-party dependencies preserved.\n\nv0.2.0 | 2026-05-15T15:12:35.117Z | user\n\nv0.2.0 adds three new preview helpers: head.py prints first N rows in csv/tsv/jsonl/md/aligned format with optional --columns subset; tail.py prints last N rows using a bounded ring buffer that works on multi-GB files without loading them; sample.py picks a uniformly random sample of N rows via reservoir sampling (algorithm R, single pass, O(N) memory) with optional --seed for reproducibility and --preserve-order to keep original row order. All three share the same --as csv|tsv|jsonl|md|aligned, --output, and --columns flags, mirroring the convention already used by convert.py. Default output format is aligned, a fixed-width text table that an agent can paste straight into a reply. Performance: 100k-row / 1.6 MB CSV processed by head in 50ms, tail in 180ms, sample in 260ms. 13 end-to-end tests cover all three scripts across 5 output formats, column subset, seed reproducibility, preserve-order, --output to file, and 5 error paths (missing input, unsafe path, negative n, n=0, invalid format). No breaking changes - every v0.1.0 CLI flag, output format, and 0/1/2 exit-code contract preserved.\n\nv0.1.0 | 2026-05-12T12:13:30.830Z | user\n\nv0.1.0 initial release. Local CSV/TSV/JSONL toolkit. Five scripts: inspect.py profiles a tabular file with auto-detected column types (int/float/bool/date/datetime/string), null counts, distincts, samples, encoding, and dialect. validate.py checks a file against a small JSON schema (required columns, per-column type, min/max, enum, regex, unique). dedupe.py removes duplicates by full-row or key columns with optional --keep first/last, --case-insensitive, --trim, and JSONL removed-rows report. diff.py compares two files by key column(s) and classifies rows as added/removed/changed/unchanged with per-column before/after for changed rows. convert.py converts between csv/tsv/jsonl/json/md. Pure Python 3 standard library, no pandas, no numpy, no subprocess, no remote calls. Consistent 0/1/2 exit codes across all scripts. 26 end-to-end tests covering CSV/TSV/JSONL inputs, schema validation, full-row and keyed dedup, file diffing, round-trip format conversion, 10k-row performance, error paths.\n\nArchive index:\n\nArchive v0.5.0: 21 files, 55570 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (8641b), scripts/_preview.py (4531b), scripts/check_deps.sh (844b), scripts/concat.py (5933b), scripts/convert.py (4988b), scripts/dedupe.py (8022b), scripts/diff.py (7311b), scripts/filter.py (11595b), scripts/head.py (3735b), scripts/inspect.py (5051b), scripts/merge.py (10003b), scripts/pivot.py (12249b), scripts/sample.py (4364b), scripts/sort.py (6843b), scripts/tail.py (3756b), scripts/transform.py (19920b), scripts/validate.py (9961b), skill-card.md (2224b), SKILL.md (22811b), _meta.json (136b)\n\nFile v0.5.0:SKILL.md\n\n---\nname: clean-csv-toolkit\ndescription: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate language, sort, concat, merge (inner/left/right/outer joins), pivot (group-by aggregations + wide cross-tabs), transform (derived columns with safe expression evaluator: profit = revenue - cost, name = upper(first) + last, year(date), coalesce()), and convert between csv/tsv/jsonl/json/markdown. Pure Python 3 standard library, no pandas, no remote calls. Handles 100k+ rows in under 1 second.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit\"}}\n---\n\n# clean-csv-toolkit\n\nv0.5.0\n\nA small honest toolkit for the work agents end up doing constantly: read a CSV someone sent you, work out what's in it, clean it up, and forward only the safe rows downstream. Built on Python 3 standard library only. No `pandas`, no `numpy`, no pip installs, no remote calls.\n\n## What this skill does\n\n- `scripts/inspect.py` — profile a `.csv` / `.tsv` / `.jsonl` file: row count, auto-detected column types (`int`, `float`, `bool`, `date`, `datetime`, `string`, `empty`), null counts per column, distinct value counts (capped), three sample values per column, file size, and detected encoding.\n- `scripts/validate.py` — check the file against a small JSON schema (required columns, per-column type, min/max, enum, regex, unique). Exits 0/1 so it slots into CI.\n- `scripts/dedupe.py` — remove duplicate rows by full-row match or by key columns. Optional `--keep first|last`, `--case-insensitive`, `--trim`, and a JSONL report of every removed row.\n- `scripts/diff.py` — compare two files by key column(s) and classify every row as added / removed / changed / unchanged, with a per-column before/after diff for changed rows.\n- `scripts/convert.py` — convert between CSV, TSV, JSON Lines, JSON array, and GitHub-flavored Markdown table.\n- `scripts/head.py` (NEW in v0.2.0) — print the first N rows in csv / tsv / jsonl / md / aligned format, with optional column subset.\n- `scripts/tail.py` (NEW in v0.2.0) — print the last N rows using a streaming ring buffer (works on multi-gigabyte files without loading them).\n- `scripts/sample.py` (NEW in v0.2.0) — pick a uniformly random sample of N rows via reservoir sampling. Single-pass, O(N) memory, optional `--seed` for reproducibility, optional `--preserve-order` to keep original row order.\n- `scripts/merge.py` (NEW in v0.3.0) — join two files on one or more key columns. Supports `inner` / `left` / `right` / `outer` joins, separate key names per side via `--left-on` and `--right-on`, and duplicate-column disambiguation via `--suffix-left` / `--suffix-right`. Streams the LEFT side, indexes the RIGHT side (peak memory ≈ size of right file).\n- `scripts/pivot.py` (NEW in v0.3.0) — group-by aggregations and wide pivot tables. Aggregations: `count`, `sum`, `avg`/`mean`, `min`, `max`, `first`, `last`, `nunique`. Set `--pivot-on COL` to produce a wide cross-tab (e.g. region × product, sum of revenue). Numeric-aware `--sort-by` so `--sort-by revenue_sum --desc` orders correctly.\n- `scripts/filter.py` (NEW in v0.4.0) — keep rows that match a safe predicate (`amount > 100`, `status in pending,approved`, `email =~ @example\\.com$`, `name is_not_empty`). Supports `==`, `!=`, `<`, `<=`, `>`, `>=`, `=~` (regex), `in`, `contains`, `is_empty` / `is_not_empty` / `is_number` / `is_not_number`. `and` / `or` / parentheses / `not`. NO Python eval — a hand-rolled tokenizer + recursive-descent parser. Optional `--invert`, `--limit`, `--columns`.\n- `scripts/sort.py` (NEW in v0.4.0) — stable, type-aware sort. Auto-detects which columns are numeric and sorts them numerically (so `1200 > 899 > 100 > 50`, not `\"50\" > \"1200\"`). Per-column direction with `--by amount:desc,region:asc`. Optional `--case-insensitive`, `--limit`, `--numeric` (force numeric on all sort cols).\n- `scripts/concat.py` (NEW in v0.4.0) — stack files vertically (UNION ALL). Default mode unions the headers of all inputs; `--strict` requires identical headers; `--add-source COL` tags each row with its source filename; `--dedupe` drops exact-duplicate rows across inputs. Streams one file at a time.\n- `scripts/transform.py` (NEW in v0.5.0) — add, modify, rename, drop, cast, or keep columns. Derived columns via a safe expression language (no `eval`): `--add 'profit = revenue - cost'`, `--add 'full_name = upper(first) + \" \" + upper(last)'`, `--add 'year = year(signup_date)'`, `--add 'safe = coalesce(value, default)'`. Built-in functions: upper, lower, strip, len, abs, round, int, float, str, replace, split, join, coalesce, year, month, day. Chainable with `--cast COL:int|float|bool|string`, `--rename OLD=NEW`, `--drop COL[,...]`, `--keep COL[,...]`.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load a full dataframe into memory just to do simple structural work; the helpers stream rows where possible.\n- It does not write outside the input/output paths the caller provides.\n- It does not do statistical analysis (mean, percentile, correlation). For that, use a dataframe library.\n- It does not parse Excel files (.xls / .xlsx). Export to CSV first.\n\n## Required dependencies\n\n```bash\nbash scripts/check_deps.sh\n```\n\nOnly `python3` is required. The skill uses `csv`, `json`, `re`, `pathlib`, `argparse`, `datetime`, `collections` — all stdlib.\n\n## Workflows\n\n### 0. Quickly preview an unknown CSV (NEW in v0.2.0)\n\n```bash\n# First 10 rows in a clean aligned table\npython3 scripts/head.py mystery.csv\n\n# Last 5 rows of a multi-GB log\npython3 scripts/tail.py huge.csv -n 5\n\n# A reproducible random sample for spot-checking\npython3 scripts/sample.py customers.csv -n 20 --seed 42\n\n# Preview only specific columns\npython3 scripts/head.py customers.csv --columns id,email,status\n\n# Emit a previewable Markdown table for an agent's reply\npython3 scripts/head.py customers.csv -n 5 --as md\n```\n\nAll three scripts accept `-n N`, `--as csv|tsv|jsonl|md|aligned`, `--output file`, and `--columns col1,col2,...`. `sample.py` additionally accepts `--seed INT` and `--preserve-order`. Default output format is `aligned` — a fixed-width text table sized to the actual data, which is what an agent usually wants to show inline. Default `N` is 10.\n\nStreaming guarantees:\n- `head.py` reads at most N+1 rows from the file.\n- `tail.py` keeps a bounded `deque(maxlen=N)` and emits only the last N rows.\n- `sample.py` uses reservoir sampling (algorithm R): single pass, O(N) memory regardless of file size.\n\nOn a 100,000-row / 1.6 MB CSV: `head -n 3` runs in ~50 ms, `tail -n 3` in ~180 ms, `sample -n 5` in ~260 ms.\n\n### 1. Profile an unknown CSV\n\n```bash\npython3 scripts/inspect.py customers.csv\n```\n\nOutput:\n\n```\nfile:      /path/customers.csv\nsize:      284 B (284 bytes)\nencoding:  utf-8\nkind:      csv\nrows:      5\ncolumns:   6\n\n  #  name                          type           nulls   null%    distinct  sample\n----------------------------------------------------------------------------------------------------\n  1  id                            int                0    0.00           5  '1', '2', '3'\n  2  email                         string             0    0.00           5  'alice@example.com', ...\n  3  name                          string             0    0.00           5  'Alice', 'Bob', 'Carol'\n  4  amount                        float              1   20.00           4  '42.50', '100.00', '7.25'\n  5  status                        string             0    0.00           3  'approved', 'pending', ...\n  6  signup_date                   date               0    0.00           5  '2025-01-15', ...\n```\n\nPass `--json` for machine-readable output that pipes into other tools.\n\nThe script auto-detects the dialect (CSV vs TSV vs JSON Lines) and a sensible encoding (`utf-8`, `utf-8-sig`, `cp1252`, `latin-1`). Type inference takes up to 1000 non-empty values per column and picks the most specific type that fits all of them.\n\n### 2. Validate against a schema\n\nWrite a `schema.json`:\n\n```json\n{\n  \"required_columns\": [\"id\", \"email\", \"amount\", \"status\"],\n  \"columns\": {\n    \"id\":     {\"type\": \"int\", \"required\": true, \"unique\": true, \"min\": 1},\n    \"email\":  {\"type\": \"string\", \"required\": true, \"regex\": \".+@.+\\\\..+\"},\n    \"amount\": {\"type\": \"float\", \"min\": 0, \"max\": 100000},\n    \"status\": {\"type\": \"string\", \"enum\": [\"pending\", \"approved\", \"rejected\"]},\n    \"signup_date\": {\"type\": \"date\"}\n  }\n}\n```\n\nThen:\n\n```bash\npython3 scripts/validate.py customers.csv --schema schema.json\n```\n\nA clean file exits 0 with `verdict: pass`. A bad file exits 1 with a detailed error table:\n\n```\n   row  column                  kind                    detail\n------------------------------------------------------------------------------------------------\n     2  email                   regex_mismatch          value did not match regex | value='not-an-email'\n     2  amount                  bad_type                value does not match type 'float' | value='abc'\n     3  amount                  below_min               value -50.0 < min 0 | value='-50.00'\n     3  status                  not_in_enum             value not in allowed set | value='unknown_status'\n     4  id                      duplicate_unique        value already seen earlier in this column | value='1'\n```\n\nPass `--json` for a structured report and `--max-errors N` to cap collection on huge files.\n\n### 3. Remove duplicates\n\nBy full-row match (any two rows identical in every column):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv\n```\n\nBy a key column (only one canonical row per `id`):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv --key id \\\n  --removed-report removed.jsonl\n```\n\n`--keep first` (default) keeps the earlier-occurring row; `--keep last` keeps the later one — useful when later rows are corrections. `--case-insensitive` and `--trim` normalise key values before comparison so `\" alice@example.com\"` and `\"ALICE@example.com\"` collapse to one row.\n\nThe `--removed-report` writes one JSON object per removed row, with the original 1-based row index, the key tuple that was duplicated, and the full row, so the dedup decision is auditable.\n\n### 4. Diff two files\n\n```bash\npython3 scripts/diff.py customers_old.csv customers_new.csv --key id\n```\n\nOutput:\n\n```\nadded:      1\nremoved:    1\nchanged:    1\n\n--- ADDED (1) ---\n  + 6\n--- REMOVED (1) ---\n  - 4\n--- CHANGED (1) ---\n  ~ 2\n      amount: '100.00' -> '150.00'\n      status: 'pending' -> 'approved'\n```\n\nMulti-column keys are supported: `--key customer_id,date`. Exit codes are 0 if the files are identical on the key columns, 1 if they differ — so this also works as a CI guard (\"fail the build if the snapshot file changed\").\n\n### 5. Convert between formats\n\n```bash\npython3 scripts/convert.py data.csv data.jsonl       # row -> JSON Lines\npython3 scripts/convert.py data.jsonl data.csv       # back\npython3 scripts/convert.py data.csv data.json --pretty\npython3 scripts/convert.py data.csv data.md          # GitHub-flavored table\npython3 scripts/convert.py data.tsv data.csv         # delimiter change\n```\n\nOutput format is picked from the extension. Allowed extensions: `.csv`, `.tsv`, `.jsonl`, `.json`, `.md`. The Markdown writer escapes `|` and `\\n` in cell values so the table stays well-formed.\n\n### 6. Join two files (NEW in v0.3.0)\n\n```bash\n# Inner join: only users with at least one order\npython3 scripts/merge.py users.csv orders.csv joined.csv \\\n    --left-on id --right-on user_id\n\n# Left join: keep every user, fill unmatched with empty strings\npython3 scripts/merge.py users.csv orders.csv left.csv \\\n    --left-on id --right-on user_id --how left\n\n# Same key name on both sides: --on shorthand\npython3 scripts/merge.py users.csv orders.csv out.csv --on user_id\n\n# Outer join into JSON Lines, machine-readable summary on stdout\npython3 scripts/merge.py a.csv b.csv full.jsonl --on key --how outer --json\n```\n\nDuplicate non-key columns are auto-renamed with `--suffix-left` / `--suffix-right` (defaults `_x` / `_y`).\n\n### 7. Group-by aggregations and wide pivots (NEW in v0.3.0)\n\n```bash\n# Sum revenue per region\npython3 scripts/pivot.py sales.csv by_region.csv \\\n    --group-by region --agg revenue:sum --sort-by revenue_sum --desc\n\n# Multiple aggregations per group\npython3 scripts/pivot.py sales.csv detail.csv \\\n    --group-by region,product \\\n    --agg \"units:sum,revenue:sum,revenue:avg,product:nunique\"\n\n# Wide cross-tab: region × product matrix of revenue\npython3 scripts/pivot.py sales.csv crosstab.csv \\\n    --group-by region --pivot-on product --agg revenue:sum --fill 0\n\n# Same wide pivot rendered as Markdown for a report\npython3 scripts/pivot.py sales.csv crosstab.md \\\n    --group-by region --pivot-on product --agg revenue:sum --fill \"-\"\n```\n\nAggregation functions: `count`, `sum`, `avg`/`mean`, `min`, `max`, `first`, `last`, `nunique`. Output column names follow `<col>_<func>` (e.g. `revenue_sum`). `--sort-by` is numeric-aware: numeric columns are ordered numerically, string columns lexicographically.\n\n### 8. Filter rows (NEW in v0.4.0)\n\n```bash\n# Numeric comparison\npython3 scripts/filter.py orders.csv big.csv --where \"amount > 100\"\n\n# Combine boolean conditions\npython3 scripts/filter.py orders.csv top.csv \\\n    --where \"status == approved and amount >= 50\"\n\n# Set membership (commas are part of the value, not separators)\npython3 scripts/filter.py users.csv targeted.csv \\\n    --where \"country in IN,US,UK and signup_year >= 2024\"\n\n# Regex match\npython3 scripts/filter.py users.csv company.csv \\\n    --where 'email =~ @example\\.com$'\n\n# Null / type checks (no right-hand side)\npython3 scripts/filter.py users.csv missing.csv --where \"phone is_empty\"\n\n# Invert the predicate, write only specific columns, cap at N matches\npython3 scripts/filter.py log.csv non_errors.csv \\\n    --where \"level == ERROR\" --invert --columns ts,msg --limit 1000\n```\n\nThe expression language is deliberately small and is parsed by a hand-rolled tokenizer + recursive-descent parser. There is no `eval`, no shell, no subprocess.\n\n### 9. Sort by one or more columns (NEW in v0.4.0)\n\n```bash\n# Auto-numeric sort, descending\npython3 scripts/sort.py sales.csv s.csv --by amount:desc\n\n# Multi-key: country ascending, signup_date descending (stable)\npython3 scripts/sort.py users.csv s.csv --by country:asc,signup_date:desc\n\n# Top 10 by revenue\npython3 scripts/sort.py sales.csv top10.csv --by revenue:desc --limit 10\n\n# Case-insensitive string sort\npython3 scripts/sort.py contacts.csv s.csv --by name --case-insensitive\n```\n\nEach `--by` column is treated numerically when every value parses as a number, otherwise string. `--numeric` forces numeric on all sort columns (non-numeric rows sort last).\n\n### 10. Concatenate multiple files (NEW in v0.4.0)\n\n```bash\n# Stack monthly shards into one CSV (header union)\npython3 scripts/concat.py all_quarter.csv jan.csv feb.csv mar.csv\n\n# Strict mode: require every input to have an identical header\npython3 scripts/concat.py all.csv jan.csv feb.csv mar.csv --strict\n\n# Tag each row with the source filename (without extension)\npython3 scripts/concat.py tagged.csv shard_*.csv --add-source origin --source-stem\n\n# Stack + drop duplicate rows across files\npython3 scripts/concat.py all.csv jan.csv feb.csv apr.csv --dedupe\n```\n\n### 11. Transform columns (NEW in v0.5.0)\n\n```bash\n# Add a derived column\npython3 scripts/transform.py orders.csv with_profit.csv \\\n    --add 'profit = revenue - cost'\n\n# Multiple --add operations + cast + final column selection\npython3 scripts/transform.py sales.csv clean.csv \\\n    --add 'profit = revenue - cost' \\\n    --add 'margin_pct = round(profit / revenue * 100, 1)' \\\n    --add 'name = lower(strip(first_name)) + \"_\" + lower(strip(last_name))' \\\n    --add 'signup_year = year(signup)' \\\n    --cast revenue:float --cast cost:float \\\n    --keep id,name,country,revenue,profit,margin_pct,signup_year\n\n# Rename and drop\npython3 scripts/transform.py users.csv clean.csv \\\n    --rename 'signup=joined_date' --drop password_hash\n\n# Boolean comparisons produce 0/1 columns\npython3 scripts/transform.py orders.csv flagged.csv \\\n    --add 'is_high_value = amount > 1000'\n\n# Fallback for missing values\npython3 scripts/transform.py users.csv filled.csv \\\n    --add 'safe_email = coalesce(email, \"unknown@example.com\")'\n```\n\nThe expression language is intentionally small: arithmetic (`+ - * / %`), string concat (`+`), comparisons (`== != < <= > >=` → yield 0/1), parentheses, identifiers (column references), string and number literals, and function calls. **No `eval`, no `subprocess`, no shell.** Empty cells that propagate into arithmetic leave the derived value empty for that row instead of crashing the pipeline.\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / validation pass / files identical |\n| 1 | validation fail / files differ / no rows in input |\n| 2 | bad arguments / unsafe path / missing input / unsupported extension / schema malformed |\n\nThis 0/1/2 split is consistent across all five scripts, so they slot into shell pipelines cleanly:\n\n```bash\npython3 scripts/validate.py incoming.csv --schema schema.json \\\n  && python3 scripts/dedupe.py incoming.csv clean.csv --key id \\\n  && python3 scripts/inspect.py clean.csv\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, backslash-newline, etc.).\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides. No temp files outside the system's tempdir.\n- All inputs and outputs use UTF-8 by default; CSV reads auto-fall-back through utf-8-sig, cp1252, and latin-1 when the file's encoding is non-UTF-8.\n- Deterministic: the same input produces the same output every time.\n\n## Performance\n\n- `inspect.py` profiles 10,000 rows in well under one second on a single core (single-pass streaming read).\n- All scripts stream rows; they do not load the entire file into memory for processing. The exception is `dedupe.py` and `diff.py`, which build an in-memory dict keyed by row identity — fine for hundreds of thousands of rows on a typical laptop.\n- No background threads, no process pool, no caching.\n\n## Known limitations\n\n- Type inference uses regex-shape matching, not locale-aware parsing. `\"1,234.56\"` is detected as `string`, not `float`. Re-export with a different number format if you need different inference.\n- The Markdown writer flattens multi-line cells to single lines (newlines become spaces).\n- JSON Lines input must have one JSON object per line. Multi-line JSON arrays are not supported; use the regular CSV/JSONL pipeline.\n\n## v0.5.0 changes\n\n- Added `scripts/transform.py`: derived columns + schema operations. Hand-rolled tokenizer + recursive-descent parser (no `eval`, no subprocess) supports arithmetic (+/-/* / / %), string concat (+), parentheses, function calls, and boolean comparisons that yield 0/1. Built-in functions: upper, lower, strip, len, abs, round, int, float, str, replace, split, join, coalesce, year, month, day. Six op kinds: `--add`, `--set`, `--drop`, `--rename`, `--cast`, `--keep`. Schema is computed symbolically before the streaming pass, so empty cells in arithmetic columns don't crash the whole pipeline — they just leave the derived value empty for that single row.\n- Bug fixed during testing: the schema-detection pass was running expressions against an empty-string row, which broke arithmetic. Replaced with a purely-structural schema walk.\n\n## v0.4.0 changes\n\n- Added `scripts/filter.py`: safe-predicate row filter. Hand-rolled tokenizer + recursive-descent parser, NO `eval` and no `subprocess`. Numeric and string compare, regex (`=~`), `in COMMA,LIST`, `contains`, `is_empty` / `is_not_empty` / `is_number` / `is_not_number`. Boolean `and` / `or` / `not` with parentheses. `--invert`, `--limit`, `--columns`.\n- Added `scripts/sort.py`: type-aware stable sort with per-column direction (`--by amount:desc,region:asc`). Auto-detects numeric columns. Optional `--case-insensitive`, `--limit`, `--numeric`. 100k rows sorted in ~0.3 s.\n- Added `scripts/concat.py`: vertical UNION ALL of multiple CSV / TSV / JSONL files. Header union by default, `--strict` for exact-match check, `--add-source` to tag rows, `--dedupe` to drop exact-duplicate rows. Streams one input at a time, memory does not grow with the number of inputs (unless `--dedupe` is set).\n- All three scripts honor the existing safe-path policy and the 0 / 1 / 2 exit-code contract.\n\n## v0.3.0 changes\n\n- Added `scripts/merge.py`: join two CSV / TSV / JSONL files on one or more key columns. Supports inner, left, right, outer joins; separate key names per side; duplicate-column suffixing; CSV / TSV / JSONL output. Single-pass over LEFT, peak memory ≈ size of RIGHT. Merged 50k × 200k rows in ~1.3 s.\n- Added `scripts/pivot.py`: group-by aggregations and wide pivot tables. Functions: count, sum, avg, min, max, first, last, nunique. Wide mode produces region × product cross-tabs. Numeric-aware sort. Streamed 100k rows in ~0.7 s.\n- All new scripts honor the existing safe-path policy (no shell metacharacters), use exit code 2 for bad arguments / missing files / missing columns, exit 1 for empty results, exit 0 for success. Output extension is validated against an explicit allow-list per script.\n\n## v0.2.0 changes\n\n**Three new preview helpers** (`head.py`, `tail.py`, `sample.py`):\n\n- `head.py` and `tail.py` give shell-style preview that is format-aware and never mangles quoting the way a naive `head` / `tail` would. They auto-detect dialect (csv/tsv/jsonl), let you pick the output format with `--as`, and can re-emit any subset of columns with `--columns`.\n- `sample.py` runs reservoir sampling (algorithm R): a single streaming pass, O(N) memory regardless of file size. `--seed INT` makes the sample reproducible so it slots into test suites and CI; `--preserve-order` re-sorts the reservoir back into original row order.\n- All three share the same `--as csv|tsv|jsonl|md|aligned`, `--output`, and `--columns` flags, mirroring the convention already used by `convert.py`.\n- Default output format is `aligned`, a fixed-width text table that an agent can paste straight into a reply.\n\n**Performance**: on a 100,000-row / 1.6 MB CSV, `head` runs in ~50 ms, `tail` in ~180 ms, `sample` in ~260 ms.\n\n**No breaking changes**: every v0.1.0 CLI flag, output format, and exit-code contract is preserved.\n\n## License\n\nMIT. See `LICENSE`.\n\nFile v0.5.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-csv-toolkit\",\n  \"version\": \"0.5.0\",\n  \"publishedAt\": 1779347767997\n}\n\nFile v0.5.0:skill-card.md\n\n## Description:\n\nLocal CSV, TSV, and JSONL inspection and cleanup toolkit for profiling, validation, deduplication, diffing, previewing, filtering, sorting, concatenation, joins, pivots, transformations, and format conversion using Python 3 standard library tools.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[gopendrasharma89-tech](https://clawhub.ai/user/gopendrasharma89-tech)\n\n### License/Terms of Use:\n\nMIT\n\n## Use Case:\n\nDevelopers, data operators, and agents use this skill to inspect, clean, validate, reshape, compare, and convert local tabular datasets before sending safe rows, reports, or derived files downstream.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Data-writing commands can erase source files when an output path points to an input file.\n\nMitigation: Use separate approved output paths, keep backups, and avoid reusing any input path as an output path until the aliasing issue is fixed.\n\nRisk: CSV previews or reports may expose PII or secrets from local datasets.\n\nMitigation: Use the skill only on approved local datasets and avoid printing or saving previews and reports that contain sensitive data.\n\n## Reference(s):\n\n- [Clean CSV Toolkit ClawHub Skill Page](https://clawhub.ai/gopendrasharma89-tech/skills/clean-csv-toolkit)\n- [Clean CSV Toolkit Homepage](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance, files]\n\n**Output Format:** [Markdown guidance with shell command examples; script outputs include CSV, TSV, JSON, JSONL, Markdown tables, and aligned text tables.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Local file-processing outputs depend on the selected command and caller-provided input and output paths; no remote calls are used.]\n\n## Skill Version(s):\n\n0.5.0 (source: server release metadata and SKILL.md)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v0.5.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nArchive v0.4.0: 19 files, 47298 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (8641b), scripts/_preview.py (4531b), scripts/check_deps.sh (844b), scripts/concat.py (5933b), scripts/convert.py (4988b), scripts/dedupe.py (8022b), scripts/diff.py (7311b), scripts/filter.py (11595b), scripts/head.py (3735b), scripts/inspect.py (5051b), scripts/merge.py (10003b), scripts/pivot.py (12249b), scripts/sample.py (4364b), scripts/sort.py (6843b), scripts/tail.py (3756b), scripts/validate.py (9961b), SKILL.md (19834b), _meta.json (136b)\n\nFile v0.4.0:SKILL.md\n\n---\nname: clean-csv-toolkit\ndescription: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate language, sort, concat, merge (inner/left/right/outer joins), pivot (group-by aggregations + wide cross-tabs), and convert between csv/tsv/jsonl/json/markdown. Pure Python 3 standard library, no pandas, no remote calls. Handles 100k+ rows in under 1 second.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit\"}}\n---\n\n# clean-csv-toolkit\n\nv0.4.0\n\nA small honest toolkit for the work agents end up doing constantly: read a CSV someone sent you, work out what's in it, clean it up, and forward only the safe rows downstream. Built on Python 3 standard library only. No `pandas`, no `numpy`, no pip installs, no remote calls.\n\n## What this skill does\n\n- `scripts/inspect.py` — profile a `.csv` / `.tsv` / `.jsonl` file: row count, auto-detected column types (`int`, `float`, `bool`, `date`, `datetime`, `string`, `empty`), null counts per column, distinct value counts (capped), three sample values per column, file size, and detected encoding.\n- `scripts/validate.py` — check the file against a small JSON schema (required columns, per-column type, min/max, enum, regex, unique). Exits 0/1 so it slots into CI.\n- `scripts/dedupe.py` — remove duplicate rows by full-row match or by key columns. Optional `--keep first|last`, `--case-insensitive`, `--trim`, and a JSONL report of every removed row.\n- `scripts/diff.py` — compare two files by key column(s) and classify every row as added / removed / changed / unchanged, with a per-column before/after diff for changed rows.\n- `scripts/convert.py` — convert between CSV, TSV, JSON Lines, JSON array, and GitHub-flavored Markdown table.\n- `scripts/head.py` (NEW in v0.2.0) — print the first N rows in csv / tsv / jsonl / md / aligned format, with optional column subset.\n- `scripts/tail.py` (NEW in v0.2.0) — print the last N rows using a streaming ring buffer (works on multi-gigabyte files without loading them).\n- `scripts/sample.py` (NEW in v0.2.0) — pick a uniformly random sample of N rows via reservoir sampling. Single-pass, O(N) memory, optional `--seed` for reproducibility, optional `--preserve-order` to keep original row order.\n- `scripts/merge.py` (NEW in v0.3.0) — join two files on one or more key columns. Supports `inner` / `left` / `right` / `outer` joins, separate key names per side via `--left-on` and `--right-on`, and duplicate-column disambiguation via `--suffix-left` / `--suffix-right`. Streams the LEFT side, indexes the RIGHT side (peak memory ≈ size of right file).\n- `scripts/pivot.py` (NEW in v0.3.0) — group-by aggregations and wide pivot tables. Aggregations: `count`, `sum`, `avg`/`mean`, `min`, `max`, `first`, `last`, `nunique`. Set `--pivot-on COL` to produce a wide cross-tab (e.g. region × product, sum of revenue). Numeric-aware `--sort-by` so `--sort-by revenue_sum --desc` orders correctly.\n- `scripts/filter.py` (NEW in v0.4.0) — keep rows that match a safe predicate (`amount > 100`, `status in pending,approved`, `email =~ @example\\.com$`, `name is_not_empty`). Supports `==`, `!=`, `<`, `<=`, `>`, `>=`, `=~` (regex), `in`, `contains`, `is_empty` / `is_not_empty` / `is_number` / `is_not_number`. `and` / `or` / parentheses / `not`. NO Python eval — a hand-rolled tokenizer + recursive-descent parser. Optional `--invert`, `--limit`, `--columns`.\n- `scripts/sort.py` (NEW in v0.4.0) — stable, type-aware sort. Auto-detects which columns are numeric and sorts them numerically (so `1200 > 899 > 100 > 50`, not `\"50\" > \"1200\"`). Per-column direction with `--by amount:desc,region:asc`. Optional `--case-insensitive`, `--limit`, `--numeric` (force numeric on all sort cols).\n- `scripts/concat.py` (NEW in v0.4.0) — stack files vertically (UNION ALL). Default mode unions the headers of all inputs; `--strict` requires identical headers; `--add-source COL` tags each row with its source filename; `--dedupe` drops exact-duplicate rows across inputs. Streams one file at a time.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load a full dataframe into memory just to do simple structural work; the helpers stream rows where possible.\n- It does not write outside the input/output paths the caller provides.\n- It does not do statistical analysis (mean, percentile, correlation). For that, use a dataframe library.\n- It does not parse Excel files (.xls / .xlsx). Export to CSV first.\n\n## Required dependencies\n\n```bash\nbash scripts/check_deps.sh\n```\n\nOnly `python3` is required. The skill uses `csv`, `json`, `re`, `pathlib`, `argparse`, `datetime`, `collections` — all stdlib.\n\n## Workflows\n\n### 0. Quickly preview an unknown CSV (NEW in v0.2.0)\n\n```bash\n# First 10 rows in a clean aligned table\npython3 scripts/head.py mystery.csv\n\n# Last 5 rows of a multi-GB log\npython3 scripts/tail.py huge.csv -n 5\n\n# A reproducible random sample for spot-checking\npython3 scripts/sample.py customers.csv -n 20 --seed 42\n\n# Preview only specific columns\npython3 scripts/head.py customers.csv --columns id,email,status\n\n# Emit a previewable Markdown table for an agent's reply\npython3 scripts/head.py customers.csv -n 5 --as md\n```\n\nAll three scripts accept `-n N`, `--as csv|tsv|jsonl|md|aligned`, `--output file`, and `--columns col1,col2,...`. `sample.py` additionally accepts `--seed INT` and `--preserve-order`. Default output format is `aligned` — a fixed-width text table sized to the actual data, which is what an agent usually wants to show inline. Default `N` is 10.\n\nStreaming guarantees:\n- `head.py` reads at most N+1 rows from the file.\n- `tail.py` keeps a bounded `deque(maxlen=N)` and emits only the last N rows.\n- `sample.py` uses reservoir sampling (algorithm R): single pass, O(N) memory regardless of file size.\n\nOn a 100,000-row / 1.6 MB CSV: `head -n 3` runs in ~50 ms, `tail -n 3` in ~180 ms, `sample -n 5` in ~260 ms.\n\n### 1. Profile an unknown CSV\n\n```bash\npython3 scripts/inspect.py customers.csv\n```\n\nOutput:\n\n```\nfile:      /path/customers.csv\nsize:      284 B (284 bytes)\nencoding:  utf-8\nkind:      csv\nrows:      5\ncolumns:   6\n\n  #  name                          type           nulls   null%    distinct  sample\n----------------------------------------------------------------------------------------------------\n  1  id                            int                0    0.00           5  '1', '2', '3'\n  2  email                         string             0    0.00           5  'alice@example.com', ...\n  3  name                          string             0    0.00           5  'Alice', 'Bob', 'Carol'\n  4  amount                        float              1   20.00           4  '42.50', '100.00', '7.25'\n  5  status                        string             0    0.00           3  'approved', 'pending', ...\n  6  signup_date                   date               0    0.00           5  '2025-01-15', ...\n```\n\nPass `--json` for machine-readable output that pipes into other tools.\n\nThe script auto-detects the dialect (CSV vs TSV vs JSON Lines) and a sensible encoding (`utf-8`, `utf-8-sig`, `cp1252`, `latin-1`). Type inference takes up to 1000 non-empty values per column and picks the most specific type that fits all of them.\n\n### 2. Validate against a schema\n\nWrite a `schema.json`:\n\n```json\n{\n  \"required_columns\": [\"id\", \"email\", \"amount\", \"status\"],\n  \"columns\": {\n    \"id\":     {\"type\": \"int\", \"required\": true, \"unique\": true, \"min\": 1},\n    \"email\":  {\"type\": \"string\", \"required\": true, \"regex\": \".+@.+\\\\..+\"},\n    \"amount\": {\"type\": \"float\", \"min\": 0, \"max\": 100000},\n    \"status\": {\"type\": \"string\", \"enum\": [\"pending\", \"approved\", \"rejected\"]},\n    \"signup_date\": {\"type\": \"date\"}\n  }\n}\n```\n\nThen:\n\n```bash\npython3 scripts/validate.py customers.csv --schema schema.json\n```\n\nA clean file exits 0 with `verdict: pass`. A bad file exits 1 with a detailed error table:\n\n```\n   row  column                  kind                    detail\n------------------------------------------------------------------------------------------------\n     2  email                   regex_mismatch          value did not match regex | value='not-an-email'\n     2  amount                  bad_type                value does not match type 'float' | value='abc'\n     3  amount                  below_min               value -50.0 < min 0 | value='-50.00'\n     3  status                  not_in_enum             value not in allowed set | value='unknown_status'\n     4  id                      duplicate_unique        value already seen earlier in this column | value='1'\n```\n\nPass `--json` for a structured report and `--max-errors N` to cap collection on huge files.\n\n### 3. Remove duplicates\n\nBy full-row match (any two rows identical in every column):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv\n```\n\nBy a key column (only one canonical row per `id`):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv --key id \\\n  --removed-report removed.jsonl\n```\n\n`--keep first` (default) keeps the earlier-occurring row; `--keep last` keeps the later one — useful when later rows are corrections. `--case-insensitive` and `--trim` normalise key values before comparison so `\" alice@example.com\"` and `\"ALICE@example.com\"` collapse to one row.\n\nThe `--removed-report` writes one JSON object per removed row, with the original 1-based row index, the key tuple that was duplicated, and the full row, so the dedup decision is auditable.\n\n### 4. Diff two files\n\n```bash\npython3 scripts/diff.py customers_old.csv customers_new.csv --key id\n```\n\nOutput:\n\n```\nadded:      1\nremoved:    1\nchanged:    1\n\n--- ADDED (1) ---\n  + 6\n--- REMOVED (1) ---\n  - 4\n--- CHANGED (1) ---\n  ~ 2\n      amount: '100.00' -> '150.00'\n      status: 'pending' -> 'approved'\n```\n\nMulti-column keys are supported: `--key customer_id,date`. Exit codes are 0 if the files are identical on the key columns, 1 if they differ — so this also works as a CI guard (\"fail the build if the snapshot file changed\").\n\n### 5. Convert between formats\n\n```bash\npython3 scripts/convert.py data.csv data.jsonl       # row -> JSON Lines\npython3 scripts/convert.py data.jsonl data.csv       # back\npython3 scripts/convert.py data.csv data.json --pretty\npython3 scripts/convert.py data.csv data.md          # GitHub-flavored table\npython3 scripts/convert.py data.tsv data.csv         # delimiter change\n```\n\nOutput format is picked from the extension. Allowed extensions: `.csv`, `.tsv`, `.jsonl`, `.json`, `.md`. The Markdown writer escapes `|` and `\\n` in cell values so the table stays well-formed.\n\n### 6. Join two files (NEW in v0.3.0)\n\n```bash\n# Inner join: only users with at least one order\npython3 scripts/merge.py users.csv orders.csv joined.csv \\\n    --left-on id --right-on user_id\n\n# Left join: keep every user, fill unmatched with empty strings\npython3 scripts/merge.py users.csv orders.csv left.csv \\\n    --left-on id --right-on user_id --how left\n\n# Same key name on both sides: --on shorthand\npython3 scripts/merge.py users.csv orders.csv out.csv --on user_id\n\n# Outer join into JSON Lines, machine-readable summary on stdout\npython3 scripts/merge.py a.csv b.csv full.jsonl --on key --how outer --json\n```\n\nDuplicate non-key columns are auto-renamed with `--suffix-left` / `--suffix-right` (defaults `_x` / `_y`).\n\n### 7. Group-by aggregations and wide pivots (NEW in v0.3.0)\n\n```bash\n# Sum revenue per region\npython3 scripts/pivot.py sales.csv by_region.csv \\\n    --group-by region --agg revenue:sum --sort-by revenue_sum --desc\n\n# Multiple aggregations per group\npython3 scripts/pivot.py sales.csv detail.csv \\\n    --group-by region,product \\\n    --agg \"units:sum,revenue:sum,revenue:avg,product:nunique\"\n\n# Wide cross-tab: region × product matrix of revenue\npython3 scripts/pivot.py sales.csv crosstab.csv \\\n    --group-by region --pivot-on product --agg revenue:sum --fill 0\n\n# Same wide pivot rendered as Markdown for a report\npython3 scripts/pivot.py sales.csv crosstab.md \\\n    --group-by region --pivot-on product --agg revenue:sum --fill \"-\"\n```\n\nAggregation functions: `count`, `sum`, `avg`/`mean`, `min`, `max`, `first`, `last`, `nunique`. Output column names follow `<col>_<func>` (e.g. `revenue_sum`). `--sort-by` is numeric-aware: numeric columns are ordered numerically, string columns lexicographically.\n\n### 8. Filter rows (NEW in v0.4.0)\n\n```bash\n# Numeric comparison\npython3 scripts/filter.py orders.csv big.csv --where \"amount > 100\"\n\n# Combine boolean conditions\npython3 scripts/filter.py orders.csv top.csv \\\n    --where \"status == approved and amount >= 50\"\n\n# Set membership (commas are part of the value, not separators)\npython3 scripts/filter.py users.csv targeted.csv \\\n    --where \"country in IN,US,UK and signup_year >= 2024\"\n\n# Regex match\npython3 scripts/filter.py users.csv company.csv \\\n    --where 'email =~ @example\\.com$'\n\n# Null / type checks (no right-hand side)\npython3 scripts/filter.py users.csv missing.csv --where \"phone is_empty\"\n\n# Invert the predicate, write only specific columns, cap at N matches\npython3 scripts/filter.py log.csv non_errors.csv \\\n    --where \"level == ERROR\" --invert --columns ts,msg --limit 1000\n```\n\nThe expression language is deliberately small and is parsed by a hand-rolled tokenizer + recursive-descent parser. There is no `eval`, no shell, no subprocess.\n\n### 9. Sort by one or more columns (NEW in v0.4.0)\n\n```bash\n# Auto-numeric sort, descending\npython3 scripts/sort.py sales.csv s.csv --by amount:desc\n\n# Multi-key: country ascending, signup_date descending (stable)\npython3 scripts/sort.py users.csv s.csv --by country:asc,signup_date:desc\n\n# Top 10 by revenue\npython3 scripts/sort.py sales.csv top10.csv --by revenue:desc --limit 10\n\n# Case-insensitive string sort\npython3 scripts/sort.py contacts.csv s.csv --by name --case-insensitive\n```\n\nEach `--by` column is treated numerically when every value parses as a number, otherwise string. `--numeric` forces numeric on all sort columns (non-numeric rows sort last).\n\n### 10. Concatenate multiple files (NEW in v0.4.0)\n\n```bash\n# Stack monthly shards into one CSV (header union)\npython3 scripts/concat.py all_quarter.csv jan.csv feb.csv mar.csv\n\n# Strict mode: require every input to have an identical header\npython3 scripts/concat.py all.csv jan.csv feb.csv mar.csv --strict\n\n# Tag each row with the source filename (without extension)\npython3 scripts/concat.py tagged.csv shard_*.csv --add-source origin --source-stem\n\n# Stack + drop duplicate rows across files\npython3 scripts/concat.py all.csv jan.csv feb.csv apr.csv --dedupe\n```\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / validation pass / files identical |\n| 1 | validation fail / files differ / no rows in input |\n| 2 | bad arguments / unsafe path / missing input / unsupported extension / schema malformed |\n\nThis 0/1/2 split is consistent across all five scripts, so they slot into shell pipelines cleanly:\n\n```bash\npython3 scripts/validate.py incoming.csv --schema schema.json \\\n  && python3 scripts/dedupe.py incoming.csv clean.csv --key id \\\n  && python3 scripts/inspect.py clean.csv\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, backslash-newline, etc.).\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides. No temp files outside the system's tempdir.\n- All inputs and outputs use UTF-8 by default; CSV reads auto-fall-back through utf-8-sig, cp1252, and latin-1 when the file's encoding is non-UTF-8.\n- Deterministic: the same input produces the same output every time.\n\n## Performance\n\n- `inspect.py` profiles 10,000 rows in well under one second on a single core (single-pass streaming read).\n- All scripts stream rows; they do not load the entire file into memory for processing. The exception is `dedupe.py` and `diff.py`, which build an in-memory dict keyed by row identity — fine for hundreds of thousands of rows on a typical laptop.\n- No background threads, no process pool, no caching.\n\n## Known limitations\n\n- Type inference uses regex-shape matching, not locale-aware parsing. `\"1,234.56\"` is detected as `string`, not `float`. Re-export with a different number format if you need different inference.\n- The Markdown writer flattens multi-line cells to single lines (newlines become spaces).\n- JSON Lines input must have one JSON object per line. Multi-line JSON arrays are not supported; use the regular CSV/JSONL pipeline.\n\n## v0.4.0 changes\n\n- Added `scripts/filter.py`: safe-predicate row filter. Hand-rolled tokenizer + recursive-descent parser, NO `eval` and no `subprocess`. Numeric and string compare, regex (`=~`), `in COMMA,LIST`, `contains`, `is_empty` / `is_not_empty` / `is_number` / `is_not_number`. Boolean `and` / `or` / `not` with parentheses. `--invert`, `--limit`, `--columns`.\n- Added `scripts/sort.py`: type-aware stable sort with per-column direction (`--by amount:desc,region:asc`). Auto-detects numeric columns. Optional `--case-insensitive`, `--limit`, `--numeric`. 100k rows sorted in ~0.3 s.\n- Added `scripts/concat.py`: vertical UNION ALL of multiple CSV / TSV / JSONL files. Header union by default, `--strict` for exact-match check, `--add-source` to tag rows, `--dedupe` to drop exact-duplicate rows. Streams one input at a time, memory does not grow with the number of inputs (unless `--dedupe` is set).\n- All three scripts honor the existing safe-path policy and the 0 / 1 / 2 exit-code contract.\n\n## v0.3.0 changes\n\n- Added `scripts/merge.py`: join two CSV / TSV / JSONL files on one or more key columns. Supports inner, left, right, outer joins; separate key names per side; duplicate-column suffixing; CSV / TSV / JSONL output. Single-pass over LEFT, peak memory ≈ size of RIGHT. Merged 50k × 200k rows in ~1.3 s.\n- Added `scripts/pivot.py`: group-by aggregations and wide pivot tables. Functions: count, sum, avg, min, max, first, last, nunique. Wide mode produces region × product cross-tabs. Numeric-aware sort. Streamed 100k rows in ~0.7 s.\n- All new scripts honor the existing safe-path policy (no shell metacharacters), use exit code 2 for bad arguments / missing files / missing columns, exit 1 for empty results, exit 0 for success. Output extension is validated against an explicit allow-list per script.\n\n## v0.2.0 changes\n\n**Three new preview helpers** (`head.py`, `tail.py`, `sample.py`):\n\n- `head.py` and `tail.py` give shell-style preview that is format-aware and never mangles quoting the way a naive `head` / `tail` would. They auto-detect dialect (csv/tsv/jsonl), let you pick the output format with `--as`, and can re-emit any subset of columns with `--columns`.\n- `sample.py` runs reservoir sampling (algorithm R): a single streaming pass, O(N) memory regardless of file size. `--seed INT` makes the sample reproducible so it slots into test suites and CI; `--preserve-order` re-sorts the reservoir back into original row order.\n- All three share the same `--as csv|tsv|jsonl|md|aligned`, `--output`, and `--columns` flags, mirroring the convention already used by `convert.py`.\n- Default output format is `aligned`, a fixed-width text table that an agent can paste straight into a reply.\n\n**Performance**: on a 100,000-row / 1.6 MB CSV, `head` runs in ~50 ms, `tail` in ~180 ms, `sample` in ~260 ms.\n\n**No breaking changes**: every v0.1.0 CLI flag, output format, and exit-code contract is preserved.\n\n## License\n\nMIT. See `LICENSE`.\n\nFile v0.4.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-csv-toolkit\",\n  \"version\": \"0.4.0\",\n  \"publishedAt\": 1779104926067\n}\n\nFile v0.4.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nArchive v0.3.0: 16 files, 37031 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (8641b), scripts/_preview.py (4531b), scripts/check_deps.sh (844b), scripts/convert.py (4988b), scripts/dedupe.py (8022b), scripts/diff.py (7311b), scripts/head.py (3735b), scripts/inspect.py (5051b), scripts/merge.py (10003b), scripts/pivot.py (12249b), scripts/sample.py (4364b), scripts/tail.py (3756b), scripts/validate.py (9961b), SKILL.md (15606b), _meta.json (136b)\n\nFile v0.3.0:SKILL.md\n\n---\nname: clean-csv-toolkit\ndescription: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile a tabular file (row count, auto-detected column types, nulls, distincts, samples), validate it against a small JSON schema, deduplicate, diff, preview (head/tail/random sample), merge two files on key columns (inner/left/right/outer), pivot with group-by aggregations (sum/avg/min/max/count/nunique) or wide cross-tabs, and convert between csv/tsv/jsonl/json/markdown. Pure Python 3 standard library, no pandas, no remote calls. Handles 100k+ rows in under 1 second.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit\"}}\n---\n\n# clean-csv-toolkit\n\nv0.3.0\n\nA small honest toolkit for the work agents end up doing constantly: read a CSV someone sent you, work out what's in it, clean it up, and forward only the safe rows downstream. Built on Python 3 standard library only. No `pandas`, no `numpy`, no pip installs, no remote calls.\n\n## What this skill does\n\n- `scripts/inspect.py` — profile a `.csv` / `.tsv` / `.jsonl` file: row count, auto-detected column types (`int`, `float`, `bool`, `date`, `datetime`, `string`, `empty`), null counts per column, distinct value counts (capped), three sample values per column, file size, and detected encoding.\n- `scripts/validate.py` — check the file against a small JSON schema (required columns, per-column type, min/max, enum, regex, unique). Exits 0/1 so it slots into CI.\n- `scripts/dedupe.py` — remove duplicate rows by full-row match or by key columns. Optional `--keep first|last`, `--case-insensitive`, `--trim`, and a JSONL report of every removed row.\n- `scripts/diff.py` — compare two files by key column(s) and classify every row as added / removed / changed / unchanged, with a per-column before/after diff for changed rows.\n- `scripts/convert.py` — convert between CSV, TSV, JSON Lines, JSON array, and GitHub-flavored Markdown table.\n- `scripts/head.py` (NEW in v0.2.0) — print the first N rows in csv / tsv / jsonl / md / aligned format, with optional column subset.\n- `scripts/tail.py` (NEW in v0.2.0) — print the last N rows using a streaming ring buffer (works on multi-gigabyte files without loading them).\n- `scripts/sample.py` (NEW in v0.2.0) — pick a uniformly random sample of N rows via reservoir sampling. Single-pass, O(N) memory, optional `--seed` for reproducibility, optional `--preserve-order` to keep original row order.\n- `scripts/merge.py` (NEW in v0.3.0) — join two files on one or more key columns. Supports `inner` / `left` / `right` / `outer` joins, separate key names per side via `--left-on` and `--right-on`, and duplicate-column disambiguation via `--suffix-left` / `--suffix-right`. Streams the LEFT side, indexes the RIGHT side (peak memory ≈ size of right file).\n- `scripts/pivot.py` (NEW in v0.3.0) — group-by aggregations and wide pivot tables. Aggregations: `count`, `sum`, `avg`/`mean`, `min`, `max`, `first`, `last`, `nunique`. Set `--pivot-on COL` to produce a wide cross-tab (e.g. region × product, sum of revenue). Numeric-aware `--sort-by` so `--sort-by revenue_sum --desc` orders correctly.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load a full dataframe into memory just to do simple structural work; the helpers stream rows where possible.\n- It does not write outside the input/output paths the caller provides.\n- It does not do statistical analysis (mean, percentile, correlation). For that, use a dataframe library.\n- It does not parse Excel files (.xls / .xlsx). Export to CSV first.\n\n## Required dependencies\n\n```bash\nbash scripts/check_deps.sh\n```\n\nOnly `python3` is required. The skill uses `csv`, `json`, `re`, `pathlib`, `argparse`, `datetime`, `collections` — all stdlib.\n\n## Workflows\n\n### 0. Quickly preview an unknown CSV (NEW in v0.2.0)\n\n```bash\n# First 10 rows in a clean aligned table\npython3 scripts/head.py mystery.csv\n\n# Last 5 rows of a multi-GB log\npython3 scripts/tail.py huge.csv -n 5\n\n# A reproducible random sample for spot-checking\npython3 scripts/sample.py customers.csv -n 20 --seed 42\n\n# Preview only specific columns\npython3 scripts/head.py customers.csv --columns id,email,status\n\n# Emit a previewable Markdown table for an agent's reply\npython3 scripts/head.py customers.csv -n 5 --as md\n```\n\nAll three scripts accept `-n N`, `--as csv|tsv|jsonl|md|aligned`, `--output file`, and `--columns col1,col2,...`. `sample.py` additionally accepts `--seed INT` and `--preserve-order`. Default output format is `aligned` — a fixed-width text table sized to the actual data, which is what an agent usually wants to show inline. Default `N` is 10.\n\nStreaming guarantees:\n- `head.py` reads at most N+1 rows from the file.\n- `tail.py` keeps a bounded `deque(maxlen=N)` and emits only the last N rows.\n- `sample.py` uses reservoir sampling (algorithm R): single pass, O(N) memory regardless of file size.\n\nOn a 100,000-row / 1.6 MB CSV: `head -n 3` runs in ~50 ms, `tail -n 3` in ~180 ms, `sample -n 5` in ~260 ms.\n\n### 1. Profile an unknown CSV\n\n```bash\npython3 scripts/inspect.py customers.csv\n```\n\nOutput:\n\n```\nfile:      /path/customers.csv\nsize:      284 B (284 bytes)\nencoding:  utf-8\nkind:      csv\nrows:      5\ncolumns:   6\n\n  #  name                          type           nulls   null%    distinct  sample\n----------------------------------------------------------------------------------------------------\n  1  id                            int                0    0.00           5  '1', '2', '3'\n  2  email                         string             0    0.00           5  'alice@example.com', ...\n  3  name                          string             0    0.00           5  'Alice', 'Bob', 'Carol'\n  4  amount                        float              1   20.00           4  '42.50', '100.00', '7.25'\n  5  status                        string             0    0.00           3  'approved', 'pending', ...\n  6  signup_date                   date               0    0.00           5  '2025-01-15', ...\n```\n\nPass `--json` for machine-readable output that pipes into other tools.\n\nThe script auto-detects the dialect (CSV vs TSV vs JSON Lines) and a sensible encoding (`utf-8`, `utf-8-sig`, `cp1252`, `latin-1`). Type inference takes up to 1000 non-empty values per column and picks the most specific type that fits all of them.\n\n### 2. Validate against a schema\n\nWrite a `schema.json`:\n\n```json\n{\n  \"required_columns\": [\"id\", \"email\", \"amount\", \"status\"],\n  \"columns\": {\n    \"id\":     {\"type\": \"int\", \"required\": true, \"unique\": true, \"min\": 1},\n    \"email\":  {\"type\": \"string\", \"required\": true, \"regex\": \".+@.+\\\\..+\"},\n    \"amount\": {\"type\": \"float\", \"min\": 0, \"max\": 100000},\n    \"status\": {\"type\": \"string\", \"enum\": [\"pending\", \"approved\", \"rejected\"]},\n    \"signup_date\": {\"type\": \"date\"}\n  }\n}\n```\n\nThen:\n\n```bash\npython3 scripts/validate.py customers.csv --schema schema.json\n```\n\nA clean file exits 0 with `verdict: pass`. A bad file exits 1 with a detailed error table:\n\n```\n   row  column                  kind                    detail\n------------------------------------------------------------------------------------------------\n     2  email                   regex_mismatch          value did not match regex | value='not-an-email'\n     2  amount                  bad_type                value does not match type 'float' | value='abc'\n     3  amount                  below_min               value -50.0 < min 0 | value='-50.00'\n     3  status                  not_in_enum             value not in allowed set | value='unknown_status'\n     4  id                      duplicate_unique        value already seen earlier in this column | value='1'\n```\n\nPass `--json` for a structured report and `--max-errors N` to cap collection on huge files.\n\n### 3. Remove duplicates\n\nBy full-row match (any two rows identical in every column):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv\n```\n\nBy a key column (only one canonical row per `id`):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv --key id \\\n  --removed-report removed.jsonl\n```\n\n`--keep first` (default) keeps the earlier-occurring row; `--keep last` keeps the later one — useful when later rows are corrections. `--case-insensitive` and `--trim` normalise key values before comparison so `\" alice@example.com\"` and `\"ALICE@example.com\"` collapse to one row.\n\nThe `--removed-report` writes one JSON object per removed row, with the original 1-based row index, the key tuple that was duplicated, and the full row, so the dedup decision is auditable.\n\n### 4. Diff two files\n\n```bash\npython3 scripts/diff.py customers_old.csv customers_new.csv --key id\n```\n\nOutput:\n\n```\nadded:      1\nremoved:    1\nchanged:    1\n\n--- ADDED (1) ---\n  + 6\n--- REMOVED (1) ---\n  - 4\n--- CHANGED (1) ---\n  ~ 2\n      amount: '100.00' -> '150.00'\n      status: 'pending' -> 'approved'\n```\n\nMulti-column keys are supported: `--key customer_id,date`. Exit codes are 0 if the files are identical on the key columns, 1 if they differ — so this also works as a CI guard (\"fail the build if the snapshot file changed\").\n\n### 5. Convert between formats\n\n```bash\npython3 scripts/convert.py data.csv data.jsonl       # row -> JSON Lines\npython3 scripts/convert.py data.jsonl data.csv       # back\npython3 scripts/convert.py data.csv data.json --pretty\npython3 scripts/convert.py data.csv data.md          # GitHub-flavored table\npython3 scripts/convert.py data.tsv data.csv         # delimiter change\n```\n\nOutput format is picked from the extension. Allowed extensions: `.csv`, `.tsv`, `.jsonl`, `.json`, `.md`. The Markdown writer escapes `|` and `\\n` in cell values so the table stays well-formed.\n\n### 6. Join two files (NEW in v0.3.0)\n\n```bash\n# Inner join: only users with at least one order\npython3 scripts/merge.py users.csv orders.csv joined.csv \\\n    --left-on id --right-on user_id\n\n# Left join: keep every user, fill unmatched with empty strings\npython3 scripts/merge.py users.csv orders.csv left.csv \\\n    --left-on id --right-on user_id --how left\n\n# Same key name on both sides: --on shorthand\npython3 scripts/merge.py users.csv orders.csv out.csv --on user_id\n\n# Outer join into JSON Lines, machine-readable summary on stdout\npython3 scripts/merge.py a.csv b.csv full.jsonl --on key --how outer --json\n```\n\nDuplicate non-key columns are auto-renamed with `--suffix-left` / `--suffix-right` (defaults `_x` / `_y`).\n\n### 7. Group-by aggregations and wide pivots (NEW in v0.3.0)\n\n```bash\n# Sum revenue per region\npython3 scripts/pivot.py sales.csv by_region.csv \\\n    --group-by region --agg revenue:sum --sort-by revenue_sum --desc\n\n# Multiple aggregations per group\npython3 scripts/pivot.py sales.csv detail.csv \\\n    --group-by region,product \\\n    --agg \"units:sum,revenue:sum,revenue:avg,product:nunique\"\n\n# Wide cross-tab: region × product matrix of revenue\npython3 scripts/pivot.py sales.csv crosstab.csv \\\n    --group-by region --pivot-on product --agg revenue:sum --fill 0\n\n# Same wide pivot rendered as Markdown for a report\npython3 scripts/pivot.py sales.csv crosstab.md \\\n    --group-by region --pivot-on product --agg revenue:sum --fill \"-\"\n```\n\nAggregation functions: `count`, `sum`, `avg`/`mean`, `min`, `max`, `first`, `last`, `nunique`. Output column names follow `<col>_<func>` (e.g. `revenue_sum`). `--sort-by` is numeric-aware: numeric columns are ordered numerically, string columns lexicographically.\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / validation pass / files identical |\n| 1 | validation fail / files differ / no rows in input |\n| 2 | bad arguments / unsafe path / missing input / unsupported extension / schema malformed |\n\nThis 0/1/2 split is consistent across all five scripts, so they slot into shell pipelines cleanly:\n\n```bash\npython3 scripts/validate.py incoming.csv --schema schema.json \\\n  && python3 scripts/dedupe.py incoming.csv clean.csv --key id \\\n  && python3 scripts/inspect.py clean.csv\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, backslash-newline, etc.).\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides. No temp files outside the system's tempdir.\n- All inputs and outputs use UTF-8 by default; CSV reads auto-fall-back through utf-8-sig, cp1252, and latin-1 when the file's encoding is non-UTF-8.\n- Deterministic: the same input produces the same output every time.\n\n## Performance\n\n- `inspect.py` profiles 10,000 rows in well under one second on a single core (single-pass streaming read).\n- All scripts stream rows; they do not load the entire file into memory for processing. The exception is `dedupe.py` and `diff.py`, which build an in-memory dict keyed by row identity — fine for hundreds of thousands of rows on a typical laptop.\n- No background threads, no process pool, no caching.\n\n## Known limitations\n\n- Type inference uses regex-shape matching, not locale-aware parsing. `\"1,234.56\"` is detected as `string`, not `float`. Re-export with a different number format if you need different inference.\n- The Markdown writer flattens multi-line cells to single lines (newlines become spaces).\n- JSON Lines input must have one JSON object per line. Multi-line JSON arrays are not supported; use the regular CSV/JSONL pipeline.\n\n## v0.3.0 changes\n\n- Added `scripts/merge.py`: join two CSV / TSV / JSONL files on one or more key columns. Supports inner, left, right, outer joins; separate key names per side; duplicate-column suffixing; CSV / TSV / JSONL output. Single-pass over LEFT, peak memory ≈ size of RIGHT. Merged 50k × 200k rows in ~1.3 s.\n- Added `scripts/pivot.py`: group-by aggregations and wide pivot tables. Functions: count, sum, avg, min, max, first, last, nunique. Wide mode produces region × product cross-tabs. Numeric-aware sort. Streamed 100k rows in ~0.7 s.\n- All new scripts honor the existing safe-path policy (no shell metacharacters), use exit code 2 for bad arguments / missing files / missing columns, exit 1 for empty results, exit 0 for success. Output extension is validated against an explicit allow-list per script.\n\n## v0.2.0 changes\n\n**Three new preview helpers** (`head.py`, `tail.py`, `sample.py`):\n\n- `head.py` and `tail.py` give shell-style preview that is format-aware and never mangles quoting the way a naive `head` / `tail` would. They auto-detect dialect (csv/tsv/jsonl), let you pick the output format with `--as`, and can re-emit any subset of columns with `--columns`.\n- `sample.py` runs reservoir sampling (algorithm R): a single streaming pass, O(N) memory regardless of file size. `--seed INT` makes the sample reproducible so it slots into test suites and CI; `--preserve-order` re-sorts the reservoir back into original row order.\n- All three share the same `--as csv|tsv|jsonl|md|aligned`, `--output`, and `--columns` flags, mirroring the convention already used by `convert.py`.\n- Default output format is `aligned`, a fixed-width text table that an agent can paste straight into a reply.\n\n**Performance**: on a 100,000-row / 1.6 MB CSV, `head` runs in ~50 ms, `tail` in ~180 ms, `sample` in ~260 ms.\n\n**No breaking changes**: every v0.1.0 CLI flag, output format, and exit-code contract is preserved.\n\n## License\n\nMIT. See `LICENSE`.\n\nFile v0.3.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-csv-toolkit\",\n  \"version\": \"0.3.0\",\n  \"publishedAt\": 1778990373629\n}\n\nFile v0.3.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nArchive v0.2.0: 14 files, 28963 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (8641b), scripts/_preview.py (4531b), scripts/check_deps.sh (844b), scripts/convert.py (4988b), scripts/dedupe.py (8022b), scripts/diff.py (7311b), scripts/head.py (3735b), scripts/inspect.py (5051b), scripts/sample.py (4364b), scripts/tail.py (3756b), scripts/validate.py (9961b), SKILL.md (12211b), _meta.json (136b)\n\nFile v0.2.0:SKILL.md\n\n---\nname: clean-csv-toolkit\ndescription: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile a tabular file (row count, auto-detected column types, nulls, distincts, samples), validate it against a small JSON schema, deduplicate by full row or key columns, diff two files by key, preview rows with head/tail/sample, and convert between csv/tsv/jsonl/json/markdown. Pure Python 3 standard library, no pandas, no remote calls.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit\"}}\n---\n\n# clean-csv-toolkit\n\nv0.2.0\n\nA small honest toolkit for the work agents end up doing constantly: read a CSV someone sent you, work out what's in it, clean it up, and forward only the safe rows downstream. Built on Python 3 standard library only. No `pandas`, no `numpy`, no pip installs, no remote calls.\n\n## What this skill does\n\n- `scripts/inspect.py` — profile a `.csv` / `.tsv` / `.jsonl` file: row count, auto-detected column types (`int`, `float`, `bool`, `date`, `datetime`, `string`, `empty`), null counts per column, distinct value counts (capped), three sample values per column, file size, and detected encoding.\n- `scripts/validate.py` — check the file against a small JSON schema (required columns, per-column type, min/max, enum, regex, unique). Exits 0/1 so it slots into CI.\n- `scripts/dedupe.py` — remove duplicate rows by full-row match or by key columns. Optional `--keep first|last`, `--case-insensitive`, `--trim`, and a JSONL report of every removed row.\n- `scripts/diff.py` — compare two files by key column(s) and classify every row as added / removed / changed / unchanged, with a per-column before/after diff for changed rows.\n- `scripts/convert.py` — convert between CSV, TSV, JSON Lines, JSON array, and GitHub-flavored Markdown table.\n- `scripts/head.py` (NEW in v0.2.0) — print the first N rows in csv / tsv / jsonl / md / aligned format, with optional column subset.\n- `scripts/tail.py` (NEW in v0.2.0) — print the last N rows using a streaming ring buffer (works on multi-gigabyte files without loading them).\n- `scripts/sample.py` (NEW in v0.2.0) — pick a uniformly random sample of N rows via reservoir sampling. Single-pass, O(N) memory, optional `--seed` for reproducibility, optional `--preserve-order` to keep original row order.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load a full dataframe into memory just to do simple structural work; the helpers stream rows where possible.\n- It does not write outside the input/output paths the caller provides.\n- It does not do statistical analysis (mean, percentile, correlation). For that, use a dataframe library.\n- It does not parse Excel files (.xls / .xlsx). Export to CSV first.\n\n## Required dependencies\n\n```bash\nbash scripts/check_deps.sh\n```\n\nOnly `python3` is required. The skill uses `csv`, `json`, `re`, `pathlib`, `argparse`, `datetime`, `collections` — all stdlib.\n\n## Workflows\n\n### 0. Quickly preview an unknown CSV (NEW in v0.2.0)\n\n```bash\n# First 10 rows in a clean aligned table\npython3 scripts/head.py mystery.csv\n\n# Last 5 rows of a multi-GB log\npython3 scripts/tail.py huge.csv -n 5\n\n# A reproducible random sample for spot-checking\npython3 scripts/sample.py customers.csv -n 20 --seed 42\n\n# Preview only specific columns\npython3 scripts/head.py customers.csv --columns id,email,status\n\n# Emit a previewable Markdown table for an agent's reply\npython3 scripts/head.py customers.csv -n 5 --as md\n```\n\nAll three scripts accept `-n N`, `--as csv|tsv|jsonl|md|aligned`, `--output file`, and `--columns col1,col2,...`. `sample.py` additionally accepts `--seed INT` and `--preserve-order`. Default output format is `aligned` — a fixed-width text table sized to the actual data, which is what an agent usually wants to show inline. Default `N` is 10.\n\nStreaming guarantees:\n- `head.py` reads at most N+1 rows from the file.\n- `tail.py` keeps a bounded `deque(maxlen=N)` and emits only the last N rows.\n- `sample.py` uses reservoir sampling (algorithm R): single pass, O(N) memory regardless of file size.\n\nOn a 100,000-row / 1.6 MB CSV: `head -n 3` runs in ~50 ms, `tail -n 3` in ~180 ms, `sample -n 5` in ~260 ms.\n\n### 1. Profile an unknown CSV\n\n```bash\npython3 scripts/inspect.py customers.csv\n```\n\nOutput:\n\n```\nfile:      /path/customers.csv\nsize:      284 B (284 bytes)\nencoding:  utf-8\nkind:      csv\nrows:      5\ncolumns:   6\n\n  #  name                          type           nulls   null%    distinct  sample\n----------------------------------------------------------------------------------------------------\n  1  id                            int                0    0.00           5  '1', '2', '3'\n  2  email                         string             0    0.00           5  'alice@example.com', ...\n  3  name                          string             0    0.00           5  'Alice', 'Bob', 'Carol'\n  4  amount                        float              1   20.00           4  '42.50', '100.00', '7.25'\n  5  status                        string             0    0.00           3  'approved', 'pending', ...\n  6  signup_date                   date               0    0.00           5  '2025-01-15', ...\n```\n\nPass `--json` for machine-readable output that pipes into other tools.\n\nThe script auto-detects the dialect (CSV vs TSV vs JSON Lines) and a sensible encoding (`utf-8`, `utf-8-sig`, `cp1252`, `latin-1`). Type inference takes up to 1000 non-empty values per column and picks the most specific type that fits all of them.\n\n### 2. Validate against a schema\n\nWrite a `schema.json`:\n\n```json\n{\n  \"required_columns\": [\"id\", \"email\", \"amount\", \"status\"],\n  \"columns\": {\n    \"id\":     {\"type\": \"int\", \"required\": true, \"unique\": true, \"min\": 1},\n    \"email\":  {\"type\": \"string\", \"required\": true, \"regex\": \".+@.+\\\\..+\"},\n    \"amount\": {\"type\": \"float\", \"min\": 0, \"max\": 100000},\n    \"status\": {\"type\": \"string\", \"enum\": [\"pending\", \"approved\", \"rejected\"]},\n    \"signup_date\": {\"type\": \"date\"}\n  }\n}\n```\n\nThen:\n\n```bash\npython3 scripts/validate.py customers.csv --schema schema.json\n```\n\nA clean file exits 0 with `verdict: pass`. A bad file exits 1 with a detailed error table:\n\n```\n   row  column                  kind                    detail\n------------------------------------------------------------------------------------------------\n     2  email                   regex_mismatch          value did not match regex | value='not-an-email'\n     2  amount                  bad_type                value does not match type 'float' | value='abc'\n     3  amount                  below_min               value -50.0 < min 0 | value='-50.00'\n     3  status                  not_in_enum             value not in allowed set | value='unknown_status'\n     4  id                      duplicate_unique        value already seen earlier in this column | value='1'\n```\n\nPass `--json` for a structured report and `--max-errors N` to cap collection on huge files.\n\n### 3. Remove duplicates\n\nBy full-row match (any two rows identical in every column):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv\n```\n\nBy a key column (only one canonical row per `id`):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv --key id \\\n  --removed-report removed.jsonl\n```\n\n`--keep first` (default) keeps the earlier-occurring row; `--keep last` keeps the later one — useful when later rows are corrections. `--case-insensitive` and `--trim` normalise key values before comparison so `\" alice@example.com\"` and `\"ALICE@example.com\"` collapse to one row.\n\nThe `--removed-report` writes one JSON object per removed row, with the original 1-based row index, the key tuple that was duplicated, and the full row, so the dedup decision is auditable.\n\n### 4. Diff two files\n\n```bash\npython3 scripts/diff.py customers_old.csv customers_new.csv --key id\n```\n\nOutput:\n\n```\nadded:      1\nremoved:    1\nchanged:    1\n\n--- ADDED (1) ---\n  + 6\n--- REMOVED (1) ---\n  - 4\n--- CHANGED (1) ---\n  ~ 2\n      amount: '100.00' -> '150.00'\n      status: 'pending' -> 'approved'\n```\n\nMulti-column keys are supported: `--key customer_id,date`. Exit codes are 0 if the files are identical on the key columns, 1 if they differ — so this also works as a CI guard (\"fail the build if the snapshot file changed\").\n\n### 5. Convert between formats\n\n```bash\npython3 scripts/convert.py data.csv data.jsonl       # row -> JSON Lines\npython3 scripts/convert.py data.jsonl data.csv       # back\npython3 scripts/convert.py data.csv data.json --pretty\npython3 scripts/convert.py data.csv data.md          # GitHub-flavored table\npython3 scripts/convert.py data.tsv data.csv         # delimiter change\n```\n\nOutput format is picked from the extension. Allowed extensions: `.csv`, `.tsv`, `.jsonl`, `.json`, `.md`. The Markdown writer escapes `|` and `\\n` in cell values so the table stays well-formed.\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / validation pass / files identical |\n| 1 | validation fail / files differ / no rows in input |\n| 2 | bad arguments / unsafe path / missing input / unsupported extension / schema malformed |\n\nThis 0/1/2 split is consistent across all five scripts, so they slot into shell pipelines cleanly:\n\n```bash\npython3 scripts/validate.py incoming.csv --schema schema.json \\\n  && python3 scripts/dedupe.py incoming.csv clean.csv --key id \\\n  && python3 scripts/inspect.py clean.csv\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, backslash-newline, etc.).\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides. No temp files outside the system's tempdir.\n- All inputs and outputs use UTF-8 by default; CSV reads auto-fall-back through utf-8-sig, cp1252, and latin-1 when the file's encoding is non-UTF-8.\n- Deterministic: the same input produces the same output every time.\n\n## Performance\n\n- `inspect.py` profiles 10,000 rows in well under one second on a single core (single-pass streaming read).\n- All scripts stream rows; they do not load the entire file into memory for processing. The exception is `dedupe.py` and `diff.py`, which build an in-memory dict keyed by row identity — fine for hundreds of thousands of rows on a typical laptop.\n- No background threads, no process pool, no caching.\n\n## Known limitations\n\n- Type inference uses regex-shape matching, not locale-aware parsing. `\"1,234.56\"` is detected as `string`, not `float`. Re-export with a different number format if you need different inference.\n- The Markdown writer flattens multi-line cells to single lines (newlines become spaces).\n- JSON Lines input must have one JSON object per line. Multi-line JSON arrays are not supported; use the regular CSV/JSONL pipeline.\n\n## v0.2.0 changes\n\n**Three new preview helpers** (`head.py`, `tail.py`, `sample.py`):\n\n- `head.py` and `tail.py` give shell-style preview that is format-aware and never mangles quoting the way a naive `head` / `tail` would. They auto-detect dialect (csv/tsv/jsonl), let you pick the output format with `--as`, and can re-emit any subset of columns with `--columns`.\n- `sample.py` runs reservoir sampling (algorithm R): a single streaming pass, O(N) memory regardless of file size. `--seed INT` makes the sample reproducible so it slots into test suites and CI; `--preserve-order` re-sorts the reservoir back into original row order.\n- All three share the same `--as csv|tsv|jsonl|md|aligned`, `--output`, and `--columns` flags, mirroring the convention already used by `convert.py`.\n- Default output format is `aligned`, a fixed-width text table that an agent can paste straight into a reply.\n\n**Performance**: on a 100,000-row / 1.6 MB CSV, `head` runs in ~50 ms, `tail` in ~180 ms, `sample` in ~260 ms.\n\n**No breaking changes**: every v0.1.0 CLI flag, output format, and exit-code contract is preserved.\n\n## License\n\nMIT. See `LICENSE`.\n\nFile v0.2.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-csv-toolkit\",\n  \"version\": \"0.2.0\",\n  \"publishedAt\": 1778857955117\n}\n\nFile v0.2.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nArchive v0.1.0: 10 files, 21140 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (8641b), scripts/check_deps.sh (844b), scripts/convert.py (4988b), scripts/dedupe.py (8022b), scripts/diff.py (7311b), scripts/inspect.py (5051b), scripts/validate.py (9961b), SKILL.md (9319b), _meta.json (136b)\n\nFile v0.1.0:SKILL.md\n\n---\nname: clean-csv-toolkit\ndescription: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile a tabular file (row count, auto-detected column types, nulls, distincts, samples), validate it against a small JSON schema, deduplicate by full row or key columns, diff two files by key, and convert between csv/tsv/jsonl/json/markdown. Pure Python 3 standard library, no pandas, no remote calls.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit\"}}\n---\n\n# clean-csv-toolkit\n\nv0.1.0\n\nA small honest toolkit for the work agents end up doing constantly: read a CSV someone sent you, work out what's in it, clean it up, and forward only the safe rows downstream. Built on Python 3 standard library only. No `pandas`, no `numpy`, no pip installs, no remote calls.\n\n## What this skill does\n\n- `scripts/inspect.py` — profile a `.csv` / `.tsv` / `.jsonl` file: row count, auto-detected column types (`int`, `float`, `bool`, `date`, `datetime`, `string`, `empty`), null counts per column, distinct value counts (capped), three sample values per column, file size, and detected encoding.\n- `scripts/validate.py` — check the file against a small JSON schema (required columns, per-column type, min/max, enum, regex, unique). Exits 0/1 so it slots into CI.\n- `scripts/dedupe.py` — remove duplicate rows by full-row match or by key columns. Optional `--keep first|last`, `--case-insensitive`, `--trim`, and a JSONL report of every removed row.\n- `scripts/diff.py` — compare two files by key column(s) and classify every row as added / removed / changed / unchanged, with a per-column before/after diff for changed rows.\n- `scripts/convert.py` — convert between CSV, TSV, JSON Lines, JSON array, and GitHub-flavored Markdown table.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load a full dataframe into memory just to do simple structural work; the helpers stream rows where possible.\n- It does not write outside the input/output paths the caller provides.\n- It does not do statistical analysis (mean, percentile, correlation). For that, use a dataframe library.\n- It does not parse Excel files (.xls / .xlsx). Export to CSV first.\n\n## Required dependencies\n\n```bash\nbash scripts/check_deps.sh\n```\n\nOnly `python3` is required. The skill uses `csv`, `json`, `re`, `pathlib`, `argparse`, `datetime`, `collections` — all stdlib.\n\n## Workflows\n\n### 1. Profile an unknown CSV\n\n```bash\npython3 scripts/inspect.py customers.csv\n```\n\nOutput:\n\n```\nfile:      /path/customers.csv\nsize:      284 B (284 bytes)\nencoding:  utf-8\nkind:      csv\nrows:      5\ncolumns:   6\n\n  #  name                          type           nulls   null%    distinct  sample\n----------------------------------------------------------------------------------------------------\n  1  id                            int                0    0.00           5  '1', '2', '3'\n  2  email                         string             0    0.00           5  'alice@example.com', ...\n  3  name                          string             0    0.00           5  'Alice', 'Bob', 'Carol'\n  4  amount                        float              1   20.00           4  '42.50', '100.00', '7.25'\n  5  status                        string             0    0.00           3  'approved', 'pending', ...\n  6  signup_date                   date               0    0.00           5  '2025-01-15', ...\n```\n\nPass `--json` for machine-readable output that pipes into other tools.\n\nThe script auto-detects the dialect (CSV vs TSV vs JSON Lines) and a sensible encoding (`utf-8`, `utf-8-sig`, `cp1252`, `latin-1`). Type inference takes up to 1000 non-empty values per column and picks the most specific type that fits all of them.\n\n### 2. Validate against a schema\n\nWrite a `schema.json`:\n\n```json\n{\n  \"required_columns\": [\"id\", \"email\", \"amount\", \"status\"],\n  \"columns\": {\n    \"id\":     {\"type\": \"int\", \"required\": true, \"unique\": true, \"min\": 1},\n    \"email\":  {\"type\": \"string\", \"required\": true, \"regex\": \".+@.+\\\\..+\"},\n    \"amount\": {\"type\": \"float\", \"min\": 0, \"max\": 100000},\n    \"status\": {\"type\": \"string\", \"enum\": [\"pending\", \"approved\", \"rejected\"]},\n    \"signup_date\": {\"type\": \"date\"}\n  }\n}\n```\n\nThen:\n\n```bash\npython3 scripts/validate.py customers.csv --schema schema.json\n```\n\nA clean file exits 0 with `verdict: pass`. A bad file exits 1 with a detailed error table:\n\n```\n   row  column                  kind                    detail\n------------------------------------------------------------------------------------------------\n     2  email                   regex_mismatch          value did not match regex | value='not-an-email'\n     2  amount                  bad_type                value does not match type 'float' | value='abc'\n     3  amount                  below_min               value -50.0 < min 0 | value='-50.00'\n     3  status                  not_in_enum             value not in allowed set | value='unknown_status'\n     4  id                      duplicate_unique        value already seen earlier in this column | value='1'\n```\n\nPass `--json` for a structured report and `--max-errors N` to cap collection on huge files.\n\n### 3. Remove duplicates\n\nBy full-row match (any two rows identical in every column):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv\n```\n\nBy a key column (only one canonical row per `id`):\n\n```bash\npython3 scripts/dedupe.py messy.csv clean.csv --key id \\\n  --removed-report removed.jsonl\n```\n\n`--keep first` (default) keeps the earlier-occurring row; `--keep last` keeps the later one — useful when later rows are corrections. `--case-insensitive` and `--trim` normalise key values before comparison so `\" alice@example.com\"` and `\"ALICE@example.com\"` collapse to one row.\n\nThe `--removed-report` writes one JSON object per removed row, with the original 1-based row index, the key tuple that was duplicated, and the full row, so the dedup decision is auditable.\n\n### 4. Diff two files\n\n```bash\npython3 scripts/diff.py customers_old.csv customers_new.csv --key id\n```\n\nOutput:\n\n```\nadded:      1\nremoved:    1\nchanged:    1\n\n--- ADDED (1) ---\n  + 6\n--- REMOVED (1) ---\n  - 4\n--- CHANGED (1) ---\n  ~ 2\n      amount: '100.00' -> '150.00'\n      status: 'pending' -> 'approved'\n```\n\nMulti-column keys are supported: `--key customer_id,date`. Exit codes are 0 if the files are identical on the key columns, 1 if they differ — so this also works as a CI guard (\"fail the build if the snapshot file changed\").\n\n### 5. Convert between formats\n\n```bash\npython3 scripts/convert.py data.csv data.jsonl       # row -> JSON Lines\npython3 scripts/convert.py data.jsonl data.csv       # back\npython3 scripts/convert.py data.csv data.json --pretty\npython3 scripts/convert.py data.csv data.md          # GitHub-flavored table\npython3 scripts/convert.py data.tsv data.csv         # delimiter change\n```\n\nOutput format is picked from the extension. Allowed extensions: `.csv`, `.tsv`, `.jsonl`, `.json`, `.md`. The Markdown writer escapes `|` and `\\n` in cell values so the table stays well-formed.\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / validation pass / files identical |\n| 1 | validation fail / files differ / no rows in input |\n| 2 | bad arguments / unsafe path / missing input / unsupported extension / schema malformed |\n\nThis 0/1/2 split is consistent across all five scripts, so they slot into shell pipelines cleanly:\n\n```bash\npython3 scripts/validate.py incoming.csv --schema schema.json \\\n  && python3 scripts/dedupe.py incoming.csv clean.csv --key id \\\n  && python3 scripts/inspect.py clean.csv\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, backslash-newline, etc.).\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides. No temp files outside the system's tempdir.\n- All inputs and outputs use UTF-8 by default; CSV reads auto-fall-back through utf-8-sig, cp1252, and latin-1 when the file's encoding is non-UTF-8.\n- Deterministic: the same input produces the same output every time.\n\n## Performance\n\n- `inspect.py` profiles 10,000 rows in well under one second on a single core (single-pass streaming read).\n- All scripts stream rows; they do not load the entire file into memory for processing. The exception is `dedupe.py` and `diff.py`, which build an in-memory dict keyed by row identity — fine for hundreds of thousands of rows on a typical laptop.\n- No background threads, no process pool, no caching.\n\n## Known limitations\n\n- Type inference uses regex-shape matching, not locale-aware parsing. `\"1,234.56\"` is detected as `string`, not `float`. Re-export with a different number format if you need different inference.\n- The Markdown writer flattens multi-line cells to single lines (newlines become spaces).\n- JSON Lines input must have one JSON object per line. Multi-line JSON arrays are not supported; use the regular CSV/JSONL pipeline.\n\n## License\n\nMIT. See `LICENSE`.\n\nFile v0.1.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-csv-toolkit\",\n  \"version\": \"0.1.0\",\n  \"publishedAt\": 1778588010830\n}\n\nFile v0.1.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.","readmeExcerpt":"Skill: Clean CSV Toolkit Owner: gopendrasharma89-tech Summary: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate... Tags: aggregation:0.3.0, column:0.5.0, concat:0.4.0, convert:0.2.0, csv:0.5.0, data:0.3.0, dedupe:0.2.0, derive:0.5.0, diff:0.2.0, expression:0.5.0, filter:0.5.0, groupby:0.4.0, head:0.2.0, join:0","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"bash scripts/check_deps.sh"},{"language":"bash","snippet":"# First 10 rows in a clean aligned table\npython3 scripts/head.py mystery.csv\n\n# Last 5 rows of a multi-GB log\npython3 scripts/tail.py huge.csv -n 5\n\n# A reproducible random sample for spot-checking\npython3 scripts/sample.py customers.csv -n 20 --seed 42\n\n# Preview only specific columns\npython3 scripts/head.py customers.csv --columns id,email,status\n\n# Emit a previewable Markdown table for an agent's reply\npython3 scripts/head.py customers.csv -n 5 --as md"},{"language":"bash","snippet":"python3 scripts/inspect.py customers.csv"},{"language":"text","snippet":"file:      /path/customers.csv\nsize:      284 B (284 bytes)\nencoding:  utf-8\nkind:      csv\nrows:      5\ncolumns:   6\n\n  #  name                          type           nulls   null%    distinct  sample\n----------------------------------------------------------------------------------------------------\n  1  id                            int                0    0.00           5  '1', '2', '3'\n  2  email                         string             0    0.00           5  'alice@example.com', ...\n  3  name                          string             0    0.00           5  'Alice', 'Bob', 'Carol'\n  4  amount                        float              1   20.00           4  '42.50', '100.00', '7.25'\n  5  status                        string             0    0.00           3  'approved', 'pending', ...\n  6  signup_date                   date               0    0.00           5  '2025-01-15', ..."},{"language":"json","snippet":"{\n  \"required_columns\": [\"id\", \"email\", \"amount\", \"status\"],\n  \"columns\": {\n    \"id\":     {\"type\": \"int\", \"required\": true, \"unique\": true, \"min\": 1},\n    \"email\":  {\"type\": \"string\", \"required\": true, \"regex\": \".+@.+\\\\..+\"},\n    \"amount\": {\"type\": \"float\", \"min\": 0, \"max\": 100000},\n    \"status\": {\"type\": \"string\", \"enum\": [\"pending\", \"approved\", \"rejected\"]},\n    \"signup_date\": {\"type\": \"date\"}\n  }\n}"},{"language":"bash","snippet":"python3 scripts/validate.py customers.csv --schema schema.json"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: clean-csv-toolkit\ndescription: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate language, sort, concat, merge (inner/left/right/outer joins), pivot (group-by aggregations + wide cross-tabs), transform (derived columns with safe expression evaluator: profit = revenue - cost, name = upper(first) + last, year(date), coalesce()), and convert between csv/tsv/jsonl/json/markdown. Pure Python 3 standard library, no pandas, no remote calls. Handles 100k+ rows in under 1 second.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit\"}}\n---\n\n# clean-csv-toolkit\n\nv0.5.0\n\nA small honest toolkit for the work agents end up doing constantly: read a CSV someone sent you, work out what's in it, clean it up, and forward only the safe rows downstream. Built on Python 3 standard library only. No `pandas`, no `numpy`, no pip installs, no remote calls.\n\n## What this skill does\n\n- `scripts/inspect.py` — profile a `.csv` / `.tsv` / `.jsonl` file: row count, auto-detected column types (`int`, `float`, `bool`, `date`, `datetime`, `string`, `empty`), null counts per column, distinct value counts (capped), three sample values per column, file size, and detected encoding.\n- `scripts/validate.py` — check the file against a small JSON schema (required columns, per-column type, min/max, enum, regex, unique). Exits 0/1 so it slots into CI.\n- `scripts/dedupe.py` — remove duplicate rows by full-row match or by key columns. Optional `--keep first|last`, `--case-insensitive`, `--trim`, and a JSONL report of every removed row.\n- `scripts/diff.py` — compare two files by key column(s) and classify every row as added / removed / changed / unchanged, with a per-column before/after diff for changed rows.\n- `scripts/convert.py` — convert between CSV, TSV, JSON Lines, JSON array, and GitHub-flavored Markdown table.\n- `scripts/head.py` (NEW in v0.2.0) — print the first N rows in csv / tsv / jsonl / md / aligned format, with optional column subset.\n- `scripts/tail.py` (NEW in v0.2.0) — print the last N rows using a streaming ring buffer (works on multi-gigabyte files without loading them).\n- `scripts/sample.py` (NEW in v0.2.0) — pick a uniformly random sample of N rows via reservoir sampling. Single-pass, O(N) memory, optional `--seed` for reproducibility, optional `--preserve-order` to keep original row order.\n- `scripts/merge.py` (NEW in v0.3.0) — join two files on one or more key columns. Supports `inner` / `left` / `right` / `outer` joins, separate key names per side via `--left-on` and `--right-on`, and duplicate-column disambiguation via `--suffix-left` / `--suffix-right`. Streams the LEFT side, indexes the RIGHT side (peak memory ≈ size of right file).\n- `scripts/pivot.py` (NEW in v0.3.0) — group-by aggregations and wide pivot tables. Aggregations: `count`, `sum`, `avg"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-csv-toolkit\",\n  \"version\": \"0.5.0\",\n  \"publishedAt\": 1779347767997\n}"},{"path":"skill-card.md","content":"## Description:\n\nLocal CSV, TSV, and JSONL inspection and cleanup toolkit for profiling, validation, deduplication, diffing, previewing, filtering, sorting, concatenation, joins, pivots, transformations, and format conversion using Python 3 standard library tools.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[gopendrasharma89-tech](https://clawhub.ai/user/gopendrasharma89-tech)\n\n### License/Terms of Use:\n\nMIT\n\n## Use Case:\n\nDevelopers, data operators, and agents use this skill to inspect, clean, validate, reshape, compare, and convert local tabular datasets before sending safe rows, reports, or derived files downstream.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Data-writing commands can erase source files when an output path points to an input file.\n\nMitigation: Use separate approved output paths, keep backups, and avoid reusing any input path as an output path until the aliasing issue is fixed.\n\nRisk: CSV previews or reports may expose PII or secrets from local datasets.\n\nMitigation: Use the skill only on approved local datasets and avoid printing or saving previews and reports that contain sensitive data.\n\n## Reference(s):\n\n- [Clean CSV Toolkit ClawHub Skill Page](https://clawhub.ai/gopendrasharma89-tech/skills/clean-csv-toolkit)\n- [Clean CSV Toolkit Homepage](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance, files]\n\n**Output Format:** [Markdown guidance with shell command examples; script outputs include CSV, TSV, JSON, JSONL, Markdown tables, and aligned text tables.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Local file-processing outputs depend on the selected command and caller-provided input and output paths; no remote calls are used.]\n\n## Skill Version(s):\n\n0.5.0 (source: server release metadata and SKILL.md)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."},{"path":"LICENSE","content":"MIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate... Skill: Clean CSV Toolkit Owner: gopendrasharma89-tech Summary: Local CSV / TSV / JSONL inspection and cleanup toolkit. Profile, validate, deduplicate, diff, preview (head/tail/random sample), filter with a safe predicate... Tags: aggregation:0.3.0, column:0.5.0, concat:0.4.0, convert:0.2.0, csv:0.5.0, data:0.3.0, dedupe:0.2.0, derive:0.5.0, diff:0.2.0, expression:0.5.0, filter:0.5.0, groupby:0.4.0, head:0.2.0, join:0","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":1844,"uniquenessScore":45,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T05:03:04.697Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T05:03:04.697Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T07:39:55.845Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}