{"id":"bae337c7-a5ee-44f6-b4ad-4828588308a9","entityType":"agent","slug":"clawhub-gopendrasharma89-tech-clean-text-toolkit","name":"Clean Text Toolkit","canonicalUrl":"https://www.xpersona.co/agent/clawhub-gopendrasharma89-tech-clean-text-toolkit","canonicalPath":"/agent/clawhub-gopendrasharma89-tech-clean-text-toolkit","generatedAt":"2026-10-11T10:50:21.594Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-11T06:50:57.039Z","emptyReason":null},"description":"Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-... Skill: Clean Text Toolkit Owner: gopendrasharma89-tech Summary: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-... Tags: agent:0.4.0, dedupe:0.1.0, diff:0.1.0, extract:0.4.0, html:0.4.0, jinja-free:0.3.0, latest:0.4.0, links:0.4.0, markdown:0.3.0, normalize:0.1.0, pii:0.4.0, redact:0.4.0, regex:0.4.0, replace","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17cp87fy279ggwne1mqcb6675843tdy:clean-text-toolkit","sourceUrl":"https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit","homepage":"https://clawhub.ai/gopendrasharma89-tech/skills/clean-text-toolkit","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/gopendrasharma89-tech/skills/clean-text-toolkit","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":61,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-..."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:50:57.039Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:50:57.039Z","emptyReason":null},"stars":null,"forks":null,"downloads":1130,"packageName":null,"latestVersion":"0.4.0","tractionLabel":"1.1K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:50:56.974Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T06:50:57.039Z","lastCrawledAt":"2026-10-11T06:50:56.974Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T06:50:56.974Z","lastVerifiedAt":null,"highlights":[{"version":"0.4.0","createdAt":"2026-05-29T15:25:49.206Z","changelog":"v0.4.0: Add scripts/htmlstrip.py for AI agents that scrape web pages. Three modes - text (HTML to plain readable text, drops script/style content, preserves line breaks at block tags, --keep-links renders anchor as 'text (href)'), html (sanitize - removes script/style/iframe/object/embed/form/input tags + all on* event-handler attributes + inline style attrs, keeps rest of structure intact), extract (pull links/images/headings/tables as JSON/JSONL/TSV with full table cell data). Built on stdlib html.parser, no BeautifulSoup, no lxml. Designed for the single most common agent need: turn raw scraped HTML into something useful in one command. Same safe-path policy and 0/1/2 exit-code contract.","fileCount":17,"zipByteSize":42089},{"version":"0.3.0","createdAt":"2026-05-25T04:23:30.035Z","changelog":"v0.3.0: Add scripts/replace.py - sed-like find-and-replace with regex / literal / --word modes, capture-group back-references (\\1, \\2), multiple --find/--replace pairs in one pass, JSON --rules file with per-rule overrides, --dry-run preview with line:col context, --max N cap. Bug fixes: extract.py --kind url no longer grabs trailing sentence-punctuation (so 'Visit https://example.com.' yields the URL without the period); slug.py --text mode now exits 1 when input slugifies to an empty string (e.g. '!!!'), matching batch mode.","fileCount":16,"zipByteSize":37271},{"version":"0.2.0","createdAt":"2026-05-21T07:16:37.358Z","changelog":"v0.2.0: Three new scripts. template.py: no-Jinja2 placeholder substitution, three syntaxes (mustache/dollar/percent), pipe filters (upper/lower/title/strip/capitalize/reverse/len/escape-html/escape-json/urlencode), default values via ?fallback, optional --strict mode. slug.py: URL-safe slug generator with --text single-string mode + batch mode, Unicode-aware with optional --ascii transliteration via NFKD, --keep-dots for filenames, --dedupe for batch. markdown.py: three modes - text (strip), html (minimal renderer), extract (headings/links/images/code/lists to JSON/JSONL/TSV). All three share safe-path policy and 0/1/2 exit-code contract.","fileCount":14,"zipByteSize":31798},{"version":"0.1.0","createdAt":"2026-05-18T11:58:10.544Z","changelog":"v0.1.0 initial release. Six scripts in pure stdlib: extract.py (URL/email/phone/IP/hashtag/money/iso-date), normalize.py (chainable transforms: BOM/CRLF/smart-quotes/whitespace/tabs/case/Unicode NFC/dehyphenate), redact.py (PII anonymization with Luhn-validated credit-card detection and custom placeholder templates), lines.py (count/dedupe/sort/shuffle/head/tail), wordcount.py (stats + top-N with stopwords), diff_text.py (unified/side/html diffs). Shares the safe-path policy and 0/1/2 exit-code contract with clean-csv-toolkit. Streams where possible; 100k lines deduped in 60ms.","fileCount":11,"zipByteSize":21884}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17cp87fy279ggwne1mqcb6675843tdy:clean-text-toolkit","setupComplexity":"low","setupSteps":["Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T10:50:21.593Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-gopendrasharma89-tech-clean-text-toolkit/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-11T06:50:57.039Z","emptyReason":null},"readme":"Skill: Clean Text Toolkit\n\nOwner: gopendrasharma89-tech\n\nSummary: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-...\n\nTags: agent:0.4.0, dedupe:0.1.0, diff:0.1.0, extract:0.4.0, html:0.4.0, jinja-free:0.3.0, latest:0.4.0, links:0.4.0, markdown:0.3.0, normalize:0.1.0, pii:0.4.0, redact:0.4.0, regex:0.4.0, replace:0.3.0, sanitize:0.4.0, scrape:0.4.0, sed:0.3.0, slug:0.3.0, sort:0.1.0, stdlib:0.4.0, strip:0.4.0, tables:0.4.0, template:0.3.0, text:0.4.0, wordcount:0.1.0\n\nVersion history:\n\nv0.4.0 | 2026-05-29T15:25:49.206Z | user\n\nv0.4.0: Add scripts/htmlstrip.py for AI agents that scrape web pages. Three modes - text (HTML to plain readable text, drops script/style content, preserves line breaks at block tags, --keep-links renders anchor as 'text (href)'), html (sanitize - removes script/style/iframe/object/embed/form/input tags + all on* event-handler attributes + inline style attrs, keeps rest of structure intact), extract (pull links/images/headings/tables as JSON/JSONL/TSV with full table cell data). Built on stdlib html.parser, no BeautifulSoup, no lxml. Designed for the single most common agent need: turn raw scraped HTML into something useful in one command. Same safe-path policy and 0/1/2 exit-code contract.\n\nv0.3.0 | 2026-05-25T04:23:30.035Z | user\n\nv0.3.0: Add scripts/replace.py - sed-like find-and-replace with regex / literal / --word modes, capture-group back-references (\\1, \\2), multiple --find/--replace pairs in one pass, JSON --rules file with per-rule overrides, --dry-run preview with line:col context, --max N cap. Bug fixes: extract.py --kind url no longer grabs trailing sentence-punctuation (so 'Visit https://example.com.' yields the URL without the period); slug.py --text mode now exits 1 when input slugifies to an empty string (e.g. '!!!'), matching batch mode.\n\nv0.2.0 | 2026-05-21T07:16:37.358Z | user\n\nv0.2.0: Three new scripts. template.py: no-Jinja2 placeholder substitution, three syntaxes (mustache/dollar/percent), pipe filters (upper/lower/title/strip/capitalize/reverse/len/escape-html/escape-json/urlencode), default values via ?fallback, optional --strict mode. slug.py: URL-safe slug generator with --text single-string mode + batch mode, Unicode-aware with optional --ascii transliteration via NFKD, --keep-dots for filenames, --dedupe for batch. markdown.py: three modes - text (strip), html (minimal renderer), extract (headings/links/images/code/lists to JSON/JSONL/TSV). All three share safe-path policy and 0/1/2 exit-code contract.\n\nv0.1.0 | 2026-05-18T11:58:10.544Z | user\n\nv0.1.0 initial release. Six scripts in pure stdlib: extract.py (URL/email/phone/IP/hashtag/money/iso-date), normalize.py (chainable transforms: BOM/CRLF/smart-quotes/whitespace/tabs/case/Unicode NFC/dehyphenate), redact.py (PII anonymization with Luhn-validated credit-card detection and custom placeholder templates), lines.py (count/dedupe/sort/shuffle/head/tail), wordcount.py (stats + top-N with stopwords), diff_text.py (unified/side/html diffs). Shares the safe-path policy and 0/1/2 exit-code contract with clean-csv-toolkit. Streams where possible; 100k lines deduped in 60ms.\n\nArchive index:\n\nArchive v0.4.0: 17 files, 42089 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (2147b), scripts/check_deps.sh (305b), scripts/diff_text.py (6697b), scripts/extract.py (6960b), scripts/htmlstrip.py (15103b), scripts/lines.py (7186b), scripts/markdown.py (13603b), scripts/normalize.py (9013b), scripts/redact.py (7347b), scripts/replace.py (10178b), scripts/slug.py (5464b), scripts/template.py (7414b), scripts/wordcount.py (4704b), skill-card.md (2483b), SKILL.md (14709b), _meta.json (137b)\n\nFile v0.4.0:SKILL.md\n\n---\nname: clean-text-toolkit\ndescription: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-with-Luhn/SSN/JWT/AWS keys/UUIDs), normalize (BOM/CRLF/smart-quotes/whitespace/tabs/case/Unicode NFC), line utilities (count/dedupe/sort/shuffle/head/tail), word-frequency stats with stopwords, three-mode text diffs (unified/side/HTML), no-Jinja2 template renderer with filters and defaults, URL-safe slug generator, and Markdown converter (strip-to-text / minimal HTML / extract headings/links/images/code/lists). Pure Python 3 standard library, no third-party dependencies, no remote calls.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit\"}}\n---\n\n# clean-text-toolkit\n\nv0.4.0\n\nA small, honest local toolkit for the work agents end up doing constantly: read some text someone sent you, find the structured bits, clean it up, redact the secrets, and forward it downstream. Built on Python 3 standard library only. No `pandas`, no `nltk`, no pip installs, no remote calls.\n\nThis skill is the companion to [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit): that one handles structured tabular data, this one handles unstructured text.\n\n## What this skill does\n\n- `scripts/extract.py` — pull structured items out of any text file. Kinds: `url`, `email`, `phone`, `ipv4`, `ipv6`, `hashtag`, `mention`, `hex-color`, `money`, `iso-date`. Output to stdout (one-per-line or JSON), or to a `.txt` / `.json` / `.jsonl` file. Optional `--unique`, `--sort`, `--with-line` (prefix with the source line number).\n- `scripts/normalize.py` — clean up messy text. Chainable transforms applied in command-line order: `--trim`, `--collapse-spaces`, `--strip-blank`, `--to-unix`, `--to-crlf`, `--dehyphenate` (rejoin OCR/PDF hyphenated line-breaks), `--unsmart` (smart quotes / em-dashes → ASCII), `--strip-bom`, `--strip-zwsp` (zero-width spaces and joiners), `--tabs-to-spaces N`, `--spaces-to-tabs N`, `--lower` / `--upper` / `--title`, `--normalize-unicode NFC|NFD|NFKC|NFKD`.\n- `scripts/redact.py` — anonymize text by replacing PII-like patterns with placeholder tokens. Kinds: `email`, `phone`, `ipv4`, `ipv6`, `url`, `credit-card` (with Luhn validation to suppress false positives), `ssn-us`, `uuid`, `hex-token` (32+ hex chars, typical for tokens / hashes), `aws-access-key` (AKIA…), `jwt` (three base64url segments with the `eyJ` header). `--keep-counts` makes the same value always get the same placeholder; `--preserve-length` pads/truncates the placeholder to the original length.\n- `scripts/lines.py` — line-oriented utilities. `--op count | dedupe | sort | shuffle | head | tail`. Streams `count`, `head`, `tail`. `dedupe` and `sort` are O(N) memory in the number of lines, but each line is small so 1 M lines is fine on a laptop. `--case-insensitive`, `--keep first|last`, `--numeric`, `--reverse`, `--seed` for deterministic shuffles.\n- `scripts/wordcount.py` — word / character / line / sentence statistics. Optional `--top N` for most-frequent words, `--stopwords PATH`, `--min-length N`, `--ignore-case`, `--regex PATTERN` (default `[A-Za-z']+`).\n- `scripts/diff_text.py` — three-mode text diff using stdlib `difflib`. `--mode unified` (default), `--mode side` (custom two-column layout), `--mode html` (writes a full HTML file with red/green coloring). `--ignore-case`, `--ignore-whitespace`, `--context N`.\n- `scripts/template.py` (NEW in v0.2.0) — substitute placeholders in a text file with values from a JSON object or inline `--set key=value` overrides. Mustache (`{{name}}`), dollar (`${name}`), or percent (`%(name)s`) syntax. Filters: upper, lower, title, strip, capitalize, reverse, len, escape-html, escape-json, urlencode. Default values: `{{name ?Unknown}}`. Strict mode (`--strict`) exits 1 if any placeholder is unresolved. **No Jinja2, no `eval`.**\n- `scripts/slug.py` (NEW in v0.2.0) — turn strings into URL-safe slugs. Single string mode (`--text \"Hello World\"`) or batch mode (line-in-file -> line-out-file). Options: `--separator`, `--max-length`, `--no-lower`, `--ascii` (Unicode -> ASCII transliteration via NFKD), `--keep-dots` (useful for filenames), `--dedupe`.\n- `scripts/markdown.py` (NEW in v0.2.0) — strip Markdown to plain text, render a minimal HTML approximation, or extract structured items (headings, links, images, code blocks, list items) as JSON / JSONL / TSV. For text mode, `--link-style anchor|url|both` controls how `[text](url)` is rendered.\n- `scripts/replace.py` (NEW in v0.3.0) — find-and-replace with regex / literal / word-boundary modes, capture-group back-references (`\\1`, `\\2`), multiple `--find/--replace` pairs in a single pass, or a JSON `--rules` file with per-rule settings. `--dry-run` previews matches with line:col and context; `--max N` caps replacements per rule. Returns exit 1 when zero replacements happen so it slots into CI.\n- `scripts/htmlstrip.py` (NEW in v0.4.0) — strip HTML tags from scraped pages. Three modes: `text` (collapse to plain readable text, drop `<script>`/`<style>` content, preserve line breaks at block tags), `html` (sanitize — remove `script,style,iframe,object,embed,form,input` tags + all `on*` event-handler attributes + inline `style=`, keep the rest intact), `extract` (pull links/images/headings/tables as JSON/JSONL/TSV). Built on Python stdlib `html.parser`. The single most-asked-for agent capability: turn scraped HTML into something useful in one command.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load entire files into memory unless an operation truly needs the whole file (full-content normalization, sort-and-write, diff). Streaming-friendly operations (`extract`, `lines --op count|head|tail`, `wordcount` for chars/lines counters) read one line at a time.\n- It does not write outside the input/output paths the caller provides.\n\n## Quick start\n\n### 1. Pull every email out of a log file\n\n```bash\npython3 scripts/extract.py app.log --kind email --unique --sort\npython3 scripts/extract.py app.log --kind email --output emails.txt --unique\n```\n\n### 2. Find every URL and tag it with the source line\n\n```bash\npython3 scripts/extract.py article.md --kind url --with-line\n```\n\n### 3. Clean up a messy OCR dump\n\n```bash\npython3 scripts/normalize.py scanned.txt clean.txt \\\n    --strip-bom --to-unix --dehyphenate --collapse-spaces \\\n    --unsmart --strip-blank --normalize-unicode NFC\n```\n\nThe transforms run in the order you list them on the command line.\n\n### 4. Redact PII before sharing a transcript\n\n```bash\npython3 scripts/redact.py transcript.txt safe.txt\n# default kinds = all\n# default placeholder = [REDACTED_{kind}_{i}]\n```\n\n```bash\n# Only redact emails and phones, give the same email the same placeholder\npython3 scripts/redact.py transcript.txt safe.txt \\\n    --kinds email,phone --keep-counts\n```\n\n```bash\n# Custom template\npython3 scripts/redact.py log.txt safe.txt \\\n    --token-template \"<<{kind}#{i}>>\"\n```\n\n```bash\n# Pad placeholder to match original length (for fixed-width layouts)\npython3 scripts/redact.py log.txt safe.txt --preserve-length\n```\n\nCredit-card matches are validated against the Luhn checksum so 16 random digits in a row don't trigger a false positive.\n\n### 5. Line utilities\n\n```bash\n# Quick file stats\npython3 scripts/lines.py haystack.txt --op count\n\n# Drop duplicates, case-insensitive\npython3 scripts/lines.py users.txt --op dedupe --case-insensitive --output unique.txt\n\n# Numeric sort (so \"100\" > \"23\" > \"7\")\npython3 scripts/lines.py scores.txt --op sort --numeric --reverse\n\n# Deterministic shuffle\npython3 scripts/lines.py prompts.txt --op shuffle --seed 42\n\n# Look at the head and tail of a multi-gig log\npython3 scripts/lines.py huge.log --op head -n 20\npython3 scripts/lines.py huge.log --op tail -n 20\n```\n\n### 6. Word counts\n\n```bash\n# Basic stats\npython3 scripts/wordcount.py essay.txt\n\n# Top words with stopwords filter\npython3 scripts/wordcount.py essay.txt --top 20 --ignore-case --stopwords stop.txt\n\n# Machine-readable output\npython3 scripts/wordcount.py essay.txt --top 10 --json > stats.json\n```\n\n### 7. Text diff\n\n```bash\n# Standard unified diff\npython3 scripts/diff_text.py before.txt after.txt\n\n# Side-by-side\npython3 scripts/diff_text.py before.txt after.txt --mode side\n\n# HTML report (colorized) for sharing\npython3 scripts/diff_text.py before.txt after.txt --mode html --output diff.html\n\n# Whitespace-insensitive compare\npython3 scripts/diff_text.py before.txt after.txt --ignore-whitespace\n```\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / one or more matches / files identical |\n| 1 | zero matches / zero redactions / files differ / empty input |\n| 2 | bad arguments / unsafe path / missing input / unknown kind / bad regex / unsupported output extension |\n\nThis 0 / 1 / 2 split is consistent across all six scripts so they slot into shell pipelines cleanly:\n\n```bash\n# Normalize, then redact, then count words in one shot\npython3 scripts/normalize.py raw.txt clean.txt --to-unix --dehyphenate \\\n  && python3 scripts/redact.py clean.txt safe.txt \\\n  && python3 scripts/wordcount.py safe.txt --top 10\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies, no `pip install`.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, etc.). The same `safe_path()` helper that powers `clean-csv-toolkit`.\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides.\n- All inputs and outputs default to UTF-8; reads fall back through `utf-8-sig`, `cp1252`, `latin-1` if needed. Writes are always UTF-8.\n- Deterministic where it matters: `shuffle --seed N` is reproducible; `extract` and `wordcount` always emit results in the same order for a given input.\n\n## Performance\n\n- `lines.py --op dedupe` processes 100,000 short lines (500 distinct) in ~0.06 s.\n- `lines.py --op sort` processes 100,000 lines in ~0.10 s.\n- `extract.py` scans the file in a single streaming pass — memory does not grow with file size.\n\n## Known limitations\n\n- The PII patterns are pragmatic heuristics, not strict RFC validators. The `email` regex accepts `user@host.tld` shapes but does not validate that `host.tld` resolves. `phone` accepts three telltale formats (`+<digits>`, `(XXX) XXX-XXXX`, `XXX-XXX-XXXX` / `XXX XXX XXXX`) so it doesn't grab IPs, dates, or credit-card numbers — but it will miss exotic local formats.\n- `credit-card` uses the Luhn checksum, but `hex-token` (and similar high-recall patterns) intentionally over-match; review the count before sharing redacted output publicly.\n- `diff_text.py --mode html` produces the standard `difflib.HtmlDiff` markup, which embeds inline styles. The file is portable but the styling is not customizable.\n\n## v0.4.0 changes\n\n- Added `scripts/htmlstrip.py`: HTML → plain text / sanitized HTML / structured extract. Built on stdlib `html.parser`. Three modes (text / html / extract), keeps links optionally, drops `<script>/<style>/<noscript>` content entirely in text mode, removes `on*` event-handler attributes in sanitize mode. Extract mode pulls links, images, headings, and full table data as JSON.\n- Specifically designed for agents that scrape web pages: one command turns a raw HTML dump into plain text or a structured links/images/tables JSON.\n- Same safe-path policy and 0/1/2 exit-code contract as the rest of the toolkit.\n\n## v0.3.0 changes\n\n- Added `scripts/replace.py`: sed-like find-and-replace with optional regex, capture-group back-references, multiple find/replace pairs in one pass, JSON `--rules` file, `--dry-run` preview with line:col context, `--max N` cap per rule, `--word` boundaries for literal mode.\n- Fixed `extract.py`: `--kind url` was grabbing trailing sentence-punctuation (`.`, `)`, `,`, etc.) as part of the URL. Now strips a single trailing punctuation char so `Visit https://example.com.` correctly extracts `https://example.com` instead of `https://example.com.`.\n- Fixed `slug.py`: `--text` mode with input that slugifies to an empty string (e.g. `\"!!! @@@\"`) now exits 1, matching the existing batch-mode behaviour. Previously it returned 0 silently.\n\n## v0.2.0 changes\n\n- Added `scripts/template.py`: no-Jinja2 template renderer. Three placeholder syntaxes (mustache `{{x}}`, dollar `${x}`, percent `%(x)s`), pipe filters, fallback defaults, and an optional `--strict` mode for CI. **Hand-rolled regex tokenizer, no `eval`, no `subprocess`.**\n- Added `scripts/slug.py`: URL-safe slug generator. Single-string mode (prints to stdout) or batch mode (one slug per input line). Unicode-aware with optional ASCII transliteration via NFKD; `--keep-dots` for filename use; `--dedupe` for batch outputs.\n- Added `scripts/markdown.py`: three-mode Markdown processor. `text` strips all markup; `html` renders a minimal HTML approximation (headings, paragraphs, lists, blockquotes, fenced code, links, images, bold/italic/code); `extract` pulls structured items (headings, links, images, code blocks, list items) as JSON / JSONL / TSV.\n- All three new scripts share the same safe-path policy and 0 / 1 / 2 exit-code contract as the rest of the toolkit.\n\n## v0.1.0 changes\n\n- First public release of clean-text-toolkit.\n- Six scripts: `extract.py`, `normalize.py`, `redact.py`, `lines.py`, `wordcount.py`, `diff_text.py`.\n- Shared `_common.py` with `safe_path`, `read_text`, `iter_lines`, and `write_text` helpers (mirrors the design of `clean-csv-toolkit/scripts/_common.py`).\n- Bug fixed during development: initial `phone` regex was too greedy and matched IPs / ISO dates / credit-card-with-spaces; tightened to three explicit shapes (international, parenthesized, 3-3-4 dashed) that don't collide with those other patterns. Tested against a mixed-content fixture with 5 valid phones and 3 confusable non-phones.\n- Zero third-party dependencies; works on any system that ships Python 3.\n\n## Pairs well with\n\n- [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit) — same author, same design philosophy (pure stdlib, exit-code contract, safe-path policy), for structured tabular data.\n- [`openclaw-prompt-shield`](https://clawhub.ai/gopendrasharma89-tech/openclaw-prompt-shield) — pair `extract.py --kind email,url` with prompt-shield's redaction pipeline to scrub user-supplied text before passing it to an LLM.\n\n## License\n\nMIT\n\nFile v0.4.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-text-toolkit\",\n  \"version\": \"0.4.0\",\n  \"publishedAt\": 1780068349206\n}\n\nFile v0.4.0:skill-card.md\n\n## Description:\n\nLocal text cleanup and inspection toolkit that helps agents extract structured items, redact PII-like patterns, normalize text, inspect lines and word counts, compare text, render simple templates, generate slugs, and process Markdown or HTML using Python 3 standard library scripts with no remote calls.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[gopendrasharma89-tech](https://clawhub.ai/user/gopendrasharma89-tech)\n\n### License/Terms of Use:\n\nMIT\n\n## Use Case:\n\nDevelopers and agents use this skill to clean, inspect, transform, redact, and compare local text files before passing the results into downstream workflows.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: HTML sanitization and Markdown-to-HTML output may not make untrusted content safe for browser or web application use.\n\nMitigation: Use the toolkit for local text processing, and pass untrusted HTML through a maintained allowlist sanitizer before rendering it in a browser or web app.\n\nRisk: The scripts can write or overwrite caller-specified output paths.\n\nMitigation: Review output paths before execution and run commands in a working directory where overwrites are acceptable.\n\n## Reference(s):\n\n- [Clean Text Toolkit ClawHub skill page](https://clawhub.ai/gopendrasharma89-tech/skills/clean-text-toolkit)\n- [Clean Text Toolkit OpenClaw homepage](https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit)\n- [Clean CSV Toolkit companion skill](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit)\n- [OpenClaw Prompt Shield companion skill](https://clawhub.ai/gopendrasharma89-tech/openclaw-prompt-shield)\n\n## Skill Output:\n\n**Output Type(s):** [Text, Markdown, Code, Shell commands, Configuration]\n\n**Output Format:** [Markdown with inline shell commands and local script guidance; scripts can produce plain text, JSON, JSONL, TSV, HTML, and diff output.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Outputs are local file or stdout results from Python 3 standard library utilities; some operations write or overwrite caller-specified paths.]\n\n## Skill Version(s):\n\n0.4.0 (source: server release metadata and SKILL.md body)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v0.4.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nArchive v0.3.0: 16 files, 37271 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (2147b), scripts/check_deps.sh (305b), scripts/diff_text.py (6697b), scripts/extract.py (6960b), scripts/lines.py (7186b), scripts/markdown.py (13603b), scripts/normalize.py (9013b), scripts/redact.py (7347b), scripts/replace.py (10178b), scripts/slug.py (5464b), scripts/template.py (7414b), scripts/wordcount.py (4704b), skill-card.md (2358b), SKILL.md (13510b), _meta.json (137b)\n\nFile v0.3.0:SKILL.md\n\n---\nname: clean-text-toolkit\ndescription: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-with-Luhn/SSN/JWT/AWS keys/UUIDs), normalize (BOM/CRLF/smart-quotes/whitespace/tabs/case/Unicode NFC), line utilities (count/dedupe/sort/shuffle/head/tail), word-frequency stats with stopwords, three-mode text diffs (unified/side/HTML), no-Jinja2 template renderer with filters and defaults, URL-safe slug generator, and Markdown converter (strip-to-text / minimal HTML / extract headings/links/images/code/lists). Pure Python 3 standard library, no third-party dependencies, no remote calls.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit\"}}\n---\n\n# clean-text-toolkit\n\nv0.3.0\n\nA small, honest local toolkit for the work agents end up doing constantly: read some text someone sent you, find the structured bits, clean it up, redact the secrets, and forward it downstream. Built on Python 3 standard library only. No `pandas`, no `nltk`, no pip installs, no remote calls.\n\nThis skill is the companion to [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit): that one handles structured tabular data, this one handles unstructured text.\n\n## What this skill does\n\n- `scripts/extract.py` — pull structured items out of any text file. Kinds: `url`, `email`, `phone`, `ipv4`, `ipv6`, `hashtag`, `mention`, `hex-color`, `money`, `iso-date`. Output to stdout (one-per-line or JSON), or to a `.txt` / `.json` / `.jsonl` file. Optional `--unique`, `--sort`, `--with-line` (prefix with the source line number).\n- `scripts/normalize.py` — clean up messy text. Chainable transforms applied in command-line order: `--trim`, `--collapse-spaces`, `--strip-blank`, `--to-unix`, `--to-crlf`, `--dehyphenate` (rejoin OCR/PDF hyphenated line-breaks), `--unsmart` (smart quotes / em-dashes → ASCII), `--strip-bom`, `--strip-zwsp` (zero-width spaces and joiners), `--tabs-to-spaces N`, `--spaces-to-tabs N`, `--lower` / `--upper` / `--title`, `--normalize-unicode NFC|NFD|NFKC|NFKD`.\n- `scripts/redact.py` — anonymize text by replacing PII-like patterns with placeholder tokens. Kinds: `email`, `phone`, `ipv4`, `ipv6`, `url`, `credit-card` (with Luhn validation to suppress false positives), `ssn-us`, `uuid`, `hex-token` (32+ hex chars, typical for tokens / hashes), `aws-access-key` (AKIA…), `jwt` (three base64url segments with the `eyJ` header). `--keep-counts` makes the same value always get the same placeholder; `--preserve-length` pads/truncates the placeholder to the original length.\n- `scripts/lines.py` — line-oriented utilities. `--op count | dedupe | sort | shuffle | head | tail`. Streams `count`, `head`, `tail`. `dedupe` and `sort` are O(N) memory in the number of lines, but each line is small so 1 M lines is fine on a laptop. `--case-insensitive`, `--keep first|last`, `--numeric`, `--reverse`, `--seed` for deterministic shuffles.\n- `scripts/wordcount.py` — word / character / line / sentence statistics. Optional `--top N` for most-frequent words, `--stopwords PATH`, `--min-length N`, `--ignore-case`, `--regex PATTERN` (default `[A-Za-z']+`).\n- `scripts/diff_text.py` — three-mode text diff using stdlib `difflib`. `--mode unified` (default), `--mode side` (custom two-column layout), `--mode html` (writes a full HTML file with red/green coloring). `--ignore-case`, `--ignore-whitespace`, `--context N`.\n- `scripts/template.py` (NEW in v0.2.0) — substitute placeholders in a text file with values from a JSON object or inline `--set key=value` overrides. Mustache (`{{name}}`), dollar (`${name}`), or percent (`%(name)s`) syntax. Filters: upper, lower, title, strip, capitalize, reverse, len, escape-html, escape-json, urlencode. Default values: `{{name ?Unknown}}`. Strict mode (`--strict`) exits 1 if any placeholder is unresolved. **No Jinja2, no `eval`.**\n- `scripts/slug.py` (NEW in v0.2.0) — turn strings into URL-safe slugs. Single string mode (`--text \"Hello World\"`) or batch mode (line-in-file -> line-out-file). Options: `--separator`, `--max-length`, `--no-lower`, `--ascii` (Unicode -> ASCII transliteration via NFKD), `--keep-dots` (useful for filenames), `--dedupe`.\n- `scripts/markdown.py` (NEW in v0.2.0) — strip Markdown to plain text, render a minimal HTML approximation, or extract structured items (headings, links, images, code blocks, list items) as JSON / JSONL / TSV. For text mode, `--link-style anchor|url|both` controls how `[text](url)` is rendered.\n- `scripts/replace.py` (NEW in v0.3.0) — find-and-replace with regex / literal / word-boundary modes, capture-group back-references (`\\1`, `\\2`), multiple `--find/--replace` pairs in a single pass, or a JSON `--rules` file with per-rule settings. `--dry-run` previews matches with line:col and context; `--max N` caps replacements per rule. Returns exit 1 when zero replacements happen so it slots into CI.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load entire files into memory unless an operation truly needs the whole file (full-content normalization, sort-and-write, diff). Streaming-friendly operations (`extract`, `lines --op count|head|tail`, `wordcount` for chars/lines counters) read one line at a time.\n- It does not write outside the input/output paths the caller provides.\n\n## Quick start\n\n### 1. Pull every email out of a log file\n\n```bash\npython3 scripts/extract.py app.log --kind email --unique --sort\npython3 scripts/extract.py app.log --kind email --output emails.txt --unique\n```\n\n### 2. Find every URL and tag it with the source line\n\n```bash\npython3 scripts/extract.py article.md --kind url --with-line\n```\n\n### 3. Clean up a messy OCR dump\n\n```bash\npython3 scripts/normalize.py scanned.txt clean.txt \\\n    --strip-bom --to-unix --dehyphenate --collapse-spaces \\\n    --unsmart --strip-blank --normalize-unicode NFC\n```\n\nThe transforms run in the order you list them on the command line.\n\n### 4. Redact PII before sharing a transcript\n\n```bash\npython3 scripts/redact.py transcript.txt safe.txt\n# default kinds = all\n# default placeholder = [REDACTED_{kind}_{i}]\n```\n\n```bash\n# Only redact emails and phones, give the same email the same placeholder\npython3 scripts/redact.py transcript.txt safe.txt \\\n    --kinds email,phone --keep-counts\n```\n\n```bash\n# Custom template\npython3 scripts/redact.py log.txt safe.txt \\\n    --token-template \"<<{kind}#{i}>>\"\n```\n\n```bash\n# Pad placeholder to match original length (for fixed-width layouts)\npython3 scripts/redact.py log.txt safe.txt --preserve-length\n```\n\nCredit-card matches are validated against the Luhn checksum so 16 random digits in a row don't trigger a false positive.\n\n### 5. Line utilities\n\n```bash\n# Quick file stats\npython3 scripts/lines.py haystack.txt --op count\n\n# Drop duplicates, case-insensitive\npython3 scripts/lines.py users.txt --op dedupe --case-insensitive --output unique.txt\n\n# Numeric sort (so \"100\" > \"23\" > \"7\")\npython3 scripts/lines.py scores.txt --op sort --numeric --reverse\n\n# Deterministic shuffle\npython3 scripts/lines.py prompts.txt --op shuffle --seed 42\n\n# Look at the head and tail of a multi-gig log\npython3 scripts/lines.py huge.log --op head -n 20\npython3 scripts/lines.py huge.log --op tail -n 20\n```\n\n### 6. Word counts\n\n```bash\n# Basic stats\npython3 scripts/wordcount.py essay.txt\n\n# Top words with stopwords filter\npython3 scripts/wordcount.py essay.txt --top 20 --ignore-case --stopwords stop.txt\n\n# Machine-readable output\npython3 scripts/wordcount.py essay.txt --top 10 --json > stats.json\n```\n\n### 7. Text diff\n\n```bash\n# Standard unified diff\npython3 scripts/diff_text.py before.txt after.txt\n\n# Side-by-side\npython3 scripts/diff_text.py before.txt after.txt --mode side\n\n# HTML report (colorized) for sharing\npython3 scripts/diff_text.py before.txt after.txt --mode html --output diff.html\n\n# Whitespace-insensitive compare\npython3 scripts/diff_text.py before.txt after.txt --ignore-whitespace\n```\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / one or more matches / files identical |\n| 1 | zero matches / zero redactions / files differ / empty input |\n| 2 | bad arguments / unsafe path / missing input / unknown kind / bad regex / unsupported output extension |\n\nThis 0 / 1 / 2 split is consistent across all six scripts so they slot into shell pipelines cleanly:\n\n```bash\n# Normalize, then redact, then count words in one shot\npython3 scripts/normalize.py raw.txt clean.txt --to-unix --dehyphenate \\\n  && python3 scripts/redact.py clean.txt safe.txt \\\n  && python3 scripts/wordcount.py safe.txt --top 10\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies, no `pip install`.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, etc.). The same `safe_path()` helper that powers `clean-csv-toolkit`.\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides.\n- All inputs and outputs default to UTF-8; reads fall back through `utf-8-sig`, `cp1252`, `latin-1` if needed. Writes are always UTF-8.\n- Deterministic where it matters: `shuffle --seed N` is reproducible; `extract` and `wordcount` always emit results in the same order for a given input.\n\n## Performance\n\n- `lines.py --op dedupe` processes 100,000 short lines (500 distinct) in ~0.06 s.\n- `lines.py --op sort` processes 100,000 lines in ~0.10 s.\n- `extract.py` scans the file in a single streaming pass — memory does not grow with file size.\n\n## Known limitations\n\n- The PII patterns are pragmatic heuristics, not strict RFC validators. The `email` regex accepts `user@host.tld` shapes but does not validate that `host.tld` resolves. `phone` accepts three telltale formats (`+<digits>`, `(XXX) XXX-XXXX`, `XXX-XXX-XXXX` / `XXX XXX XXXX`) so it doesn't grab IPs, dates, or credit-card numbers — but it will miss exotic local formats.\n- `credit-card` uses the Luhn checksum, but `hex-token` (and similar high-recall patterns) intentionally over-match; review the count before sharing redacted output publicly.\n- `diff_text.py --mode html` produces the standard `difflib.HtmlDiff` markup, which embeds inline styles. The file is portable but the styling is not customizable.\n\n## v0.3.0 changes\n\n- Added `scripts/replace.py`: sed-like find-and-replace with optional regex, capture-group back-references, multiple find/replace pairs in one pass, JSON `--rules` file, `--dry-run` preview with line:col context, `--max N` cap per rule, `--word` boundaries for literal mode.\n- Fixed `extract.py`: `--kind url` was grabbing trailing sentence-punctuation (`.`, `)`, `,`, etc.) as part of the URL. Now strips a single trailing punctuation char so `Visit https://example.com.` correctly extracts `https://example.com` instead of `https://example.com.`.\n- Fixed `slug.py`: `--text` mode with input that slugifies to an empty string (e.g. `\"!!! @@@\"`) now exits 1, matching the existing batch-mode behaviour. Previously it returned 0 silently.\n\n## v0.2.0 changes\n\n- Added `scripts/template.py`: no-Jinja2 template renderer. Three placeholder syntaxes (mustache `{{x}}`, dollar `${x}`, percent `%(x)s`), pipe filters, fallback defaults, and an optional `--strict` mode for CI. **Hand-rolled regex tokenizer, no `eval`, no `subprocess`.**\n- Added `scripts/slug.py`: URL-safe slug generator. Single-string mode (prints to stdout) or batch mode (one slug per input line). Unicode-aware with optional ASCII transliteration via NFKD; `--keep-dots` for filename use; `--dedupe` for batch outputs.\n- Added `scripts/markdown.py`: three-mode Markdown processor. `text` strips all markup; `html` renders a minimal HTML approximation (headings, paragraphs, lists, blockquotes, fenced code, links, images, bold/italic/code); `extract` pulls structured items (headings, links, images, code blocks, list items) as JSON / JSONL / TSV.\n- All three new scripts share the same safe-path policy and 0 / 1 / 2 exit-code contract as the rest of the toolkit.\n\n## v0.1.0 changes\n\n- First public release of clean-text-toolkit.\n- Six scripts: `extract.py`, `normalize.py`, `redact.py`, `lines.py`, `wordcount.py`, `diff_text.py`.\n- Shared `_common.py` with `safe_path`, `read_text`, `iter_lines`, and `write_text` helpers (mirrors the design of `clean-csv-toolkit/scripts/_common.py`).\n- Bug fixed during development: initial `phone` regex was too greedy and matched IPs / ISO dates / credit-card-with-spaces; tightened to three explicit shapes (international, parenthesized, 3-3-4 dashed) that don't collide with those other patterns. Tested against a mixed-content fixture with 5 valid phones and 3 confusable non-phones.\n- Zero third-party dependencies; works on any system that ships Python 3.\n\n## Pairs well with\n\n- [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit) — same author, same design philosophy (pure stdlib, exit-code contract, safe-path policy), for structured tabular data.\n- [`openclaw-prompt-shield`](https://clawhub.ai/gopendrasharma89-tech/openclaw-prompt-shield) — pair `extract.py --kind email,url` with prompt-shield's redaction pipeline to scrub user-supplied text before passing it to an LLM.\n\n## License\n\nMIT\n\nFile v0.3.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-text-toolkit\",\n  \"version\": \"0.3.0\",\n  \"publishedAt\": 1779683010035\n}\n\nFile v0.3.0:skill-card.md\n\n## Description: <br>\nClean Text Toolkit provides local Python utilities for extracting, normalizing, redacting, diffing, templating, slugifying, and converting unstructured text without third-party dependencies or remote calls. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[gopendrasharma89-tech](https://clawhub.ai/user/gopendrasharma89-tech) <br>\n\n### License/Terms of Use: <br>\nMIT <br>\n\n\n## Use Case: <br>\nDevelopers, engineers, and agent operators use this skill to clean and inspect local text files, redact common sensitive patterns, transform Markdown, and produce deterministic command-line outputs for downstream workflows. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The scripts can read local input files and write or overwrite output files selected by the caller. <br>\nMitigation: Review input and output paths before execution and avoid pointing outputs at files that should be preserved. <br>\nRisk: Pattern-based redaction can miss unusual sensitive-data formats or over-match token-like strings. <br>\nMitigation: Review redaction results before sharing sensitive text and tune the selected redaction kinds for the data being processed. <br>\n\n\n## Reference(s): <br>\n- [Clean Text Toolkit ClawHub Page](https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit) <br>\n- [Clean CSV Toolkit](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit) <br>\n- [OpenClaw Prompt Shield](https://clawhub.ai/gopendrasharma89-tech/openclaw-prompt-shield) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Text, Markdown, Code, Shell commands, Configuration, Guidance] <br>\n**Output Format:** [Markdown guidance with bash commands; scripts emit plain text, JSON, JSONL, TSV, or HTML depending on command options.] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Local Python 3 standard-library command-line workflow; no remote calls are described.] <br>\n\n## Skill Version(s): <br>\n0.3.0 (source: server release metadata and SKILL.md heading) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v0.3.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nArchive v0.2.0: 14 files, 31798 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (2147b), scripts/check_deps.sh (305b), scripts/diff_text.py (6697b), scripts/extract.py (6437b), scripts/lines.py (7186b), scripts/markdown.py (13603b), scripts/normalize.py (9013b), scripts/redact.py (7347b), scripts/slug.py (5311b), scripts/template.py (7414b), scripts/wordcount.py (4704b), SKILL.md (12343b), _meta.json (137b)\n\nFile v0.2.0:SKILL.md\n\n---\nname: clean-text-toolkit\ndescription: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-with-Luhn/SSN/JWT/AWS keys/UUIDs), normalize (BOM/CRLF/smart-quotes/whitespace/tabs/case/Unicode NFC), line utilities (count/dedupe/sort/shuffle/head/tail), word-frequency stats with stopwords, three-mode text diffs (unified/side/HTML), no-Jinja2 template renderer with filters and defaults, URL-safe slug generator, and Markdown converter (strip-to-text / minimal HTML / extract headings/links/images/code/lists). Pure Python 3 standard library, no third-party dependencies, no remote calls.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit\"}}\n---\n\n# clean-text-toolkit\n\nv0.2.0\n\nA small, honest local toolkit for the work agents end up doing constantly: read some text someone sent you, find the structured bits, clean it up, redact the secrets, and forward it downstream. Built on Python 3 standard library only. No `pandas`, no `nltk`, no pip installs, no remote calls.\n\nThis skill is the companion to [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit): that one handles structured tabular data, this one handles unstructured text.\n\n## What this skill does\n\n- `scripts/extract.py` — pull structured items out of any text file. Kinds: `url`, `email`, `phone`, `ipv4`, `ipv6`, `hashtag`, `mention`, `hex-color`, `money`, `iso-date`. Output to stdout (one-per-line or JSON), or to a `.txt` / `.json` / `.jsonl` file. Optional `--unique`, `--sort`, `--with-line` (prefix with the source line number).\n- `scripts/normalize.py` — clean up messy text. Chainable transforms applied in command-line order: `--trim`, `--collapse-spaces`, `--strip-blank`, `--to-unix`, `--to-crlf`, `--dehyphenate` (rejoin OCR/PDF hyphenated line-breaks), `--unsmart` (smart quotes / em-dashes → ASCII), `--strip-bom`, `--strip-zwsp` (zero-width spaces and joiners), `--tabs-to-spaces N`, `--spaces-to-tabs N`, `--lower` / `--upper` / `--title`, `--normalize-unicode NFC|NFD|NFKC|NFKD`.\n- `scripts/redact.py` — anonymize text by replacing PII-like patterns with placeholder tokens. Kinds: `email`, `phone`, `ipv4`, `ipv6`, `url`, `credit-card` (with Luhn validation to suppress false positives), `ssn-us`, `uuid`, `hex-token` (32+ hex chars, typical for tokens / hashes), `aws-access-key` (AKIA…), `jwt` (three base64url segments with the `eyJ` header). `--keep-counts` makes the same value always get the same placeholder; `--preserve-length` pads/truncates the placeholder to the original length.\n- `scripts/lines.py` — line-oriented utilities. `--op count | dedupe | sort | shuffle | head | tail`. Streams `count`, `head`, `tail`. `dedupe` and `sort` are O(N) memory in the number of lines, but each line is small so 1 M lines is fine on a laptop. `--case-insensitive`, `--keep first|last`, `--numeric`, `--reverse`, `--seed` for deterministic shuffles.\n- `scripts/wordcount.py` — word / character / line / sentence statistics. Optional `--top N` for most-frequent words, `--stopwords PATH`, `--min-length N`, `--ignore-case`, `--regex PATTERN` (default `[A-Za-z']+`).\n- `scripts/diff_text.py` — three-mode text diff using stdlib `difflib`. `--mode unified` (default), `--mode side` (custom two-column layout), `--mode html` (writes a full HTML file with red/green coloring). `--ignore-case`, `--ignore-whitespace`, `--context N`.\n- `scripts/template.py` (NEW in v0.2.0) — substitute placeholders in a text file with values from a JSON object or inline `--set key=value` overrides. Mustache (`{{name}}`), dollar (`${name}`), or percent (`%(name)s`) syntax. Filters: upper, lower, title, strip, capitalize, reverse, len, escape-html, escape-json, urlencode. Default values: `{{name ?Unknown}}`. Strict mode (`--strict`) exits 1 if any placeholder is unresolved. **No Jinja2, no `eval`.**\n- `scripts/slug.py` (NEW in v0.2.0) — turn strings into URL-safe slugs. Single string mode (`--text \"Hello World\"`) or batch mode (line-in-file -> line-out-file). Options: `--separator`, `--max-length`, `--no-lower`, `--ascii` (Unicode -> ASCII transliteration via NFKD), `--keep-dots` (useful for filenames), `--dedupe`.\n- `scripts/markdown.py` (NEW in v0.2.0) — strip Markdown to plain text, render a minimal HTML approximation, or extract structured items (headings, links, images, code blocks, list items) as JSON / JSONL / TSV. For text mode, `--link-style anchor|url|both` controls how `[text](url)` is rendered.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load entire files into memory unless an operation truly needs the whole file (full-content normalization, sort-and-write, diff). Streaming-friendly operations (`extract`, `lines --op count|head|tail`, `wordcount` for chars/lines counters) read one line at a time.\n- It does not write outside the input/output paths the caller provides.\n\n## Quick start\n\n### 1. Pull every email out of a log file\n\n```bash\npython3 scripts/extract.py app.log --kind email --unique --sort\npython3 scripts/extract.py app.log --kind email --output emails.txt --unique\n```\n\n### 2. Find every URL and tag it with the source line\n\n```bash\npython3 scripts/extract.py article.md --kind url --with-line\n```\n\n### 3. Clean up a messy OCR dump\n\n```bash\npython3 scripts/normalize.py scanned.txt clean.txt \\\n    --strip-bom --to-unix --dehyphenate --collapse-spaces \\\n    --unsmart --strip-blank --normalize-unicode NFC\n```\n\nThe transforms run in the order you list them on the command line.\n\n### 4. Redact PII before sharing a transcript\n\n```bash\npython3 scripts/redact.py transcript.txt safe.txt\n# default kinds = all\n# default placeholder = [REDACTED_{kind}_{i}]\n```\n\n```bash\n# Only redact emails and phones, give the same email the same placeholder\npython3 scripts/redact.py transcript.txt safe.txt \\\n    --kinds email,phone --keep-counts\n```\n\n```bash\n# Custom template\npython3 scripts/redact.py log.txt safe.txt \\\n    --token-template \"<<{kind}#{i}>>\"\n```\n\n```bash\n# Pad placeholder to match original length (for fixed-width layouts)\npython3 scripts/redact.py log.txt safe.txt --preserve-length\n```\n\nCredit-card matches are validated against the Luhn checksum so 16 random digits in a row don't trigger a false positive.\n\n### 5. Line utilities\n\n```bash\n# Quick file stats\npython3 scripts/lines.py haystack.txt --op count\n\n# Drop duplicates, case-insensitive\npython3 scripts/lines.py users.txt --op dedupe --case-insensitive --output unique.txt\n\n# Numeric sort (so \"100\" > \"23\" > \"7\")\npython3 scripts/lines.py scores.txt --op sort --numeric --reverse\n\n# Deterministic shuffle\npython3 scripts/lines.py prompts.txt --op shuffle --seed 42\n\n# Look at the head and tail of a multi-gig log\npython3 scripts/lines.py huge.log --op head -n 20\npython3 scripts/lines.py huge.log --op tail -n 20\n```\n\n### 6. Word counts\n\n```bash\n# Basic stats\npython3 scripts/wordcount.py essay.txt\n\n# Top words with stopwords filter\npython3 scripts/wordcount.py essay.txt --top 20 --ignore-case --stopwords stop.txt\n\n# Machine-readable output\npython3 scripts/wordcount.py essay.txt --top 10 --json > stats.json\n```\n\n### 7. Text diff\n\n```bash\n# Standard unified diff\npython3 scripts/diff_text.py before.txt after.txt\n\n# Side-by-side\npython3 scripts/diff_text.py before.txt after.txt --mode side\n\n# HTML report (colorized) for sharing\npython3 scripts/diff_text.py before.txt after.txt --mode html --output diff.html\n\n# Whitespace-insensitive compare\npython3 scripts/diff_text.py before.txt after.txt --ignore-whitespace\n```\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / one or more matches / files identical |\n| 1 | zero matches / zero redactions / files differ / empty input |\n| 2 | bad arguments / unsafe path / missing input / unknown kind / bad regex / unsupported output extension |\n\nThis 0 / 1 / 2 split is consistent across all six scripts so they slot into shell pipelines cleanly:\n\n```bash\n# Normalize, then redact, then count words in one shot\npython3 scripts/normalize.py raw.txt clean.txt --to-unix --dehyphenate \\\n  && python3 scripts/redact.py clean.txt safe.txt \\\n  && python3 scripts/wordcount.py safe.txt --top 10\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies, no `pip install`.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, etc.). The same `safe_path()` helper that powers `clean-csv-toolkit`.\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides.\n- All inputs and outputs default to UTF-8; reads fall back through `utf-8-sig`, `cp1252`, `latin-1` if needed. Writes are always UTF-8.\n- Deterministic where it matters: `shuffle --seed N` is reproducible; `extract` and `wordcount` always emit results in the same order for a given input.\n\n## Performance\n\n- `lines.py --op dedupe` processes 100,000 short lines (500 distinct) in ~0.06 s.\n- `lines.py --op sort` processes 100,000 lines in ~0.10 s.\n- `extract.py` scans the file in a single streaming pass — memory does not grow with file size.\n\n## Known limitations\n\n- The PII patterns are pragmatic heuristics, not strict RFC validators. The `email` regex accepts `user@host.tld` shapes but does not validate that `host.tld` resolves. `phone` accepts three telltale formats (`+<digits>`, `(XXX) XXX-XXXX`, `XXX-XXX-XXXX` / `XXX XXX XXXX`) so it doesn't grab IPs, dates, or credit-card numbers — but it will miss exotic local formats.\n- `credit-card` uses the Luhn checksum, but `hex-token` (and similar high-recall patterns) intentionally over-match; review the count before sharing redacted output publicly.\n- `diff_text.py --mode html` produces the standard `difflib.HtmlDiff` markup, which embeds inline styles. The file is portable but the styling is not customizable.\n\n## v0.2.0 changes\n\n- Added `scripts/template.py`: no-Jinja2 template renderer. Three placeholder syntaxes (mustache `{{x}}`, dollar `${x}`, percent `%(x)s`), pipe filters, fallback defaults, and an optional `--strict` mode for CI. **Hand-rolled regex tokenizer, no `eval`, no `subprocess`.**\n- Added `scripts/slug.py`: URL-safe slug generator. Single-string mode (prints to stdout) or batch mode (one slug per input line). Unicode-aware with optional ASCII transliteration via NFKD; `--keep-dots` for filename use; `--dedupe` for batch outputs.\n- Added `scripts/markdown.py`: three-mode Markdown processor. `text` strips all markup; `html` renders a minimal HTML approximation (headings, paragraphs, lists, blockquotes, fenced code, links, images, bold/italic/code); `extract` pulls structured items (headings, links, images, code blocks, list items) as JSON / JSONL / TSV.\n- All three new scripts share the same safe-path policy and 0 / 1 / 2 exit-code contract as the rest of the toolkit.\n\n## v0.1.0 changes\n\n- First public release of clean-text-toolkit.\n- Six scripts: `extract.py`, `normalize.py`, `redact.py`, `lines.py`, `wordcount.py`, `diff_text.py`.\n- Shared `_common.py` with `safe_path`, `read_text`, `iter_lines`, and `write_text` helpers (mirrors the design of `clean-csv-toolkit/scripts/_common.py`).\n- Bug fixed during development: initial `phone` regex was too greedy and matched IPs / ISO dates / credit-card-with-spaces; tightened to three explicit shapes (international, parenthesized, 3-3-4 dashed) that don't collide with those other patterns. Tested against a mixed-content fixture with 5 valid phones and 3 confusable non-phones.\n- Zero third-party dependencies; works on any system that ships Python 3.\n\n## Pairs well with\n\n- [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit) — same author, same design philosophy (pure stdlib, exit-code contract, safe-path policy), for structured tabular data.\n- [`openclaw-prompt-shield`](https://clawhub.ai/gopendrasharma89-tech/openclaw-prompt-shield) — pair `extract.py --kind email,url` with prompt-shield's redaction pipeline to scrub user-supplied text before passing it to an LLM.\n\n## License\n\nMIT\n\nFile v0.2.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-text-toolkit\",\n  \"version\": \"0.2.0\",\n  \"publishedAt\": 1779347797358\n}\n\nFile v0.2.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.\n\nArchive v0.1.0: 11 files, 21884 bytes\n\nFiles: LICENSE (1078b), scripts/_common.py (2147b), scripts/check_deps.sh (305b), scripts/diff_text.py (6697b), scripts/extract.py (6437b), scripts/lines.py (7186b), scripts/normalize.py (9013b), scripts/redact.py (7347b), scripts/wordcount.py (4704b), SKILL.md (10209b), _meta.json (137b)\n\nFile v0.1.0:SKILL.md\n\n---\nname: clean-text-toolkit\ndescription: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII with custom placeholders (email, phone, credit card with Luhn check, SSN, JWT, AWS keys, UUIDs), normalize text (CRLF/BOM/smart quotes/whitespace/tabs/case/Unicode NFC), line utilities (count/dedupe/sort/shuffle/head/tail), word-frequency stats with stopwords, and three-mode text diffs (unified / side-by-side / HTML). Pure Python 3 standard library, no third-party dependencies, no remote calls. Streams where possible; 100k lines deduped in under 0.1 s.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit\"}}\n---\n\n# clean-text-toolkit\n\nv0.1.0\n\nA small, honest local toolkit for the work agents end up doing constantly: read some text someone sent you, find the structured bits, clean it up, redact the secrets, and forward it downstream. Built on Python 3 standard library only. No `pandas`, no `nltk`, no pip installs, no remote calls.\n\nThis skill is the companion to [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit): that one handles structured tabular data, this one handles unstructured text.\n\n## What this skill does\n\n- `scripts/extract.py` — pull structured items out of any text file. Kinds: `url`, `email`, `phone`, `ipv4`, `ipv6`, `hashtag`, `mention`, `hex-color`, `money`, `iso-date`. Output to stdout (one-per-line or JSON), or to a `.txt` / `.json` / `.jsonl` file. Optional `--unique`, `--sort`, `--with-line` (prefix with the source line number).\n- `scripts/normalize.py` — clean up messy text. Chainable transforms applied in command-line order: `--trim`, `--collapse-spaces`, `--strip-blank`, `--to-unix`, `--to-crlf`, `--dehyphenate` (rejoin OCR/PDF hyphenated line-breaks), `--unsmart` (smart quotes / em-dashes → ASCII), `--strip-bom`, `--strip-zwsp` (zero-width spaces and joiners), `--tabs-to-spaces N`, `--spaces-to-tabs N`, `--lower` / `--upper` / `--title`, `--normalize-unicode NFC|NFD|NFKC|NFKD`.\n- `scripts/redact.py` — anonymize text by replacing PII-like patterns with placeholder tokens. Kinds: `email`, `phone`, `ipv4`, `ipv6`, `url`, `credit-card` (with Luhn validation to suppress false positives), `ssn-us`, `uuid`, `hex-token` (32+ hex chars, typical for tokens / hashes), `aws-access-key` (AKIA…), `jwt` (three base64url segments with the `eyJ` header). `--keep-counts` makes the same value always get the same placeholder; `--preserve-length` pads/truncates the placeholder to the original length.\n- `scripts/lines.py` — line-oriented utilities. `--op count | dedupe | sort | shuffle | head | tail`. Streams `count`, `head`, `tail`. `dedupe` and `sort` are O(N) memory in the number of lines, but each line is small so 1 M lines is fine on a laptop. `--case-insensitive`, `--keep first|last`, `--numeric`, `--reverse`, `--seed` for deterministic shuffles.\n- `scripts/wordcount.py` — word / character / line / sentence statistics. Optional `--top N` for most-frequent words, `--stopwords PATH`, `--min-length N`, `--ignore-case`, `--regex PATTERN` (default `[A-Za-z']+`).\n- `scripts/diff_text.py` — three-mode text diff using stdlib `difflib`. `--mode unified` (default), `--mode side` (custom two-column layout), `--mode html` (writes a full HTML file with red/green coloring). `--ignore-case`, `--ignore-whitespace`, `--context N`.\n- `scripts/check_deps.sh` — verify `python3` is available.\n\n## What this skill does not do\n\n- It does not call any LLM, web service, or remote API.\n- It does not load entire files into memory unless an operation truly needs the whole file (full-content normalization, sort-and-write, diff). Streaming-friendly operations (`extract`, `lines --op count|head|tail`, `wordcount` for chars/lines counters) read one line at a time.\n- It does not write outside the input/output paths the caller provides.\n\n## Quick start\n\n### 1. Pull every email out of a log file\n\n```bash\npython3 scripts/extract.py app.log --kind email --unique --sort\npython3 scripts/extract.py app.log --kind email --output emails.txt --unique\n```\n\n### 2. Find every URL and tag it with the source line\n\n```bash\npython3 scripts/extract.py article.md --kind url --with-line\n```\n\n### 3. Clean up a messy OCR dump\n\n```bash\npython3 scripts/normalize.py scanned.txt clean.txt \\\n    --strip-bom --to-unix --dehyphenate --collapse-spaces \\\n    --unsmart --strip-blank --normalize-unicode NFC\n```\n\nThe transforms run in the order you list them on the command line.\n\n### 4. Redact PII before sharing a transcript\n\n```bash\npython3 scripts/redact.py transcript.txt safe.txt\n# default kinds = all\n# default placeholder = [REDACTED_{kind}_{i}]\n```\n\n```bash\n# Only redact emails and phones, give the same email the same placeholder\npython3 scripts/redact.py transcript.txt safe.txt \\\n    --kinds email,phone --keep-counts\n```\n\n```bash\n# Custom template\npython3 scripts/redact.py log.txt safe.txt \\\n    --token-template \"<<{kind}#{i}>>\"\n```\n\n```bash\n# Pad placeholder to match original length (for fixed-width layouts)\npython3 scripts/redact.py log.txt safe.txt --preserve-length\n```\n\nCredit-card matches are validated against the Luhn checksum so 16 random digits in a row don't trigger a false positive.\n\n### 5. Line utilities\n\n```bash\n# Quick file stats\npython3 scripts/lines.py haystack.txt --op count\n\n# Drop duplicates, case-insensitive\npython3 scripts/lines.py users.txt --op dedupe --case-insensitive --output unique.txt\n\n# Numeric sort (so \"100\" > \"23\" > \"7\")\npython3 scripts/lines.py scores.txt --op sort --numeric --reverse\n\n# Deterministic shuffle\npython3 scripts/lines.py prompts.txt --op shuffle --seed 42\n\n# Look at the head and tail of a multi-gig log\npython3 scripts/lines.py huge.log --op head -n 20\npython3 scripts/lines.py huge.log --op tail -n 20\n```\n\n### 6. Word counts\n\n```bash\n# Basic stats\npython3 scripts/wordcount.py essay.txt\n\n# Top words with stopwords filter\npython3 scripts/wordcount.py essay.txt --top 20 --ignore-case --stopwords stop.txt\n\n# Machine-readable output\npython3 scripts/wordcount.py essay.txt --top 10 --json > stats.json\n```\n\n### 7. Text diff\n\n```bash\n# Standard unified diff\npython3 scripts/diff_text.py before.txt after.txt\n\n# Side-by-side\npython3 scripts/diff_text.py before.txt after.txt --mode side\n\n# HTML report (colorized) for sharing\npython3 scripts/diff_text.py before.txt after.txt --mode html --output diff.html\n\n# Whitespace-insensitive compare\npython3 scripts/diff_text.py before.txt after.txt --ignore-whitespace\n```\n\n## Exit codes\n\n| Code | Meaning |\n|---|---|\n| 0 | success / one or more matches / files identical |\n| 1 | zero matches / zero redactions / files differ / empty input |\n| 2 | bad arguments / unsafe path / missing input / unknown kind / bad regex / unsupported output extension |\n\nThis 0 / 1 / 2 split is consistent across all six scripts so they slot into shell pipelines cleanly:\n\n```bash\n# Normalize, then redact, then count words in one shot\npython3 scripts/normalize.py raw.txt clean.txt --to-unix --dehyphenate \\\n  && python3 scripts/redact.py clean.txt safe.txt \\\n  && python3 scripts/wordcount.py safe.txt --top 10\n```\n\n## Safety properties\n\n- Pure Python 3 standard library. No third-party dependencies, no `pip install`.\n- No `subprocess` calls. No shell invocation.\n- All file paths are validated against a strict allowlist regex that rejects shell metacharacters (`;`, `|`, `&`, `>`, `<`, `$`, `` ` ``, etc.). The same `safe_path()` helper that powers `clean-csv-toolkit`.\n- Scripts only read the input paths the caller provides and write to the output paths the caller provides.\n- All inputs and outputs default to UTF-8; reads fall back through `utf-8-sig`, `cp1252`, `latin-1` if needed. Writes are always UTF-8.\n- Deterministic where it matters: `shuffle --seed N` is reproducible; `extract` and `wordcount` always emit results in the same order for a given input.\n\n## Performance\n\n- `lines.py --op dedupe` processes 100,000 short lines (500 distinct) in ~0.06 s.\n- `lines.py --op sort` processes 100,000 lines in ~0.10 s.\n- `extract.py` scans the file in a single streaming pass — memory does not grow with file size.\n\n## Known limitations\n\n- The PII patterns are pragmatic heuristics, not strict RFC validators. The `email` regex accepts `user@host.tld` shapes but does not validate that `host.tld` resolves. `phone` accepts three telltale formats (`+<digits>`, `(XXX) XXX-XXXX`, `XXX-XXX-XXXX` / `XXX XXX XXXX`) so it doesn't grab IPs, dates, or credit-card numbers — but it will miss exotic local formats.\n- `credit-card` uses the Luhn checksum, but `hex-token` (and similar high-recall patterns) intentionally over-match; review the count before sharing redacted output publicly.\n- `diff_text.py --mode html` produces the standard `difflib.HtmlDiff` markup, which embeds inline styles. The file is portable but the styling is not customizable.\n\n## v0.1.0 changes\n\n- First public release of clean-text-toolkit.\n- Six scripts: `extract.py`, `normalize.py`, `redact.py`, `lines.py`, `wordcount.py`, `diff_text.py`.\n- Shared `_common.py` with `safe_path`, `read_text`, `iter_lines`, and `write_text` helpers (mirrors the design of `clean-csv-toolkit/scripts/_common.py`).\n- Bug fixed during development: initial `phone` regex was too greedy and matched IPs / ISO dates / credit-card-with-spaces; tightened to three explicit shapes (international, parenthesized, 3-3-4 dashed) that don't collide with those other patterns. Tested against a mixed-content fixture with 5 valid phones and 3 confusable non-phones.\n- Zero third-party dependencies; works on any system that ships Python 3.\n\n## Pairs well with\n\n- [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit) — same author, same design philosophy (pure stdlib, exit-code contract, safe-path policy), for structured tabular data.\n- [`openclaw-prompt-shield`](https://clawhub.ai/gopendrasharma89-tech/openclaw-prompt-shield) — pair `extract.py --kind email,url` with prompt-shield's redaction pipeline to scrub user-supplied text before passing it to an LLM.\n\n## License\n\nMIT\n\nFile v0.1.0:_meta.json\n\n{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-text-toolkit\",\n  \"version\": \"0.1.0\",\n  \"publishedAt\": 1779105490544\n}\n\nFile v0.1.0:LICENSE\n\nMIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE.","readmeExcerpt":"Skill: Clean Text Toolkit Owner: gopendrasharma89-tech Summary: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-... Tags: agent:0.4.0, dedupe:0.1.0, diff:0.1.0, extract:0.4.0, html:0.4.0, jinja-free:0.3.0, latest:0.4.0, links:0.4.0, markdown:0.3.0, normalize:0.1.0, pii:0.4.0, redact:0.4.0, regex:0.4.0, replace","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"python3 scripts/extract.py app.log --kind email --unique --sort\npython3 scripts/extract.py app.log --kind email --output emails.txt --unique"},{"language":"bash","snippet":"python3 scripts/extract.py article.md --kind url --with-line"},{"language":"bash","snippet":"python3 scripts/normalize.py scanned.txt clean.txt \\\n    --strip-bom --to-unix --dehyphenate --collapse-spaces \\\n    --unsmart --strip-blank --normalize-unicode NFC"},{"language":"bash","snippet":"python3 scripts/redact.py transcript.txt safe.txt\n# default kinds = all\n# default placeholder = [REDACTED_{kind}_{i}]"},{"language":"bash","snippet":"# Only redact emails and phones, give the same email the same placeholder\npython3 scripts/redact.py transcript.txt safe.txt \\\n    --kinds email,phone --keep-counts"},{"language":"bash","snippet":"# Custom template\npython3 scripts/redact.py log.txt safe.txt \\\n    --token-template \"<<{kind}#{i}>>\""}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: clean-text-toolkit\ndescription: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-with-Luhn/SSN/JWT/AWS keys/UUIDs), normalize (BOM/CRLF/smart-quotes/whitespace/tabs/case/Unicode NFC), line utilities (count/dedupe/sort/shuffle/head/tail), word-frequency stats with stopwords, three-mode text diffs (unified/side/HTML), no-Jinja2 template renderer with filters and defaults, URL-safe slug generator, and Markdown converter (strip-to-text / minimal HTML / extract headings/links/images/code/lists). Pure Python 3 standard library, no third-party dependencies, no remote calls.\nlicense: MIT\nmetadata: {\"openclaw\":{\"requires\":{\"bins\":[\"python3\"]},\"primaryEnv\":null,\"homepage\":\"https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit\"}}\n---\n\n# clean-text-toolkit\n\nv0.4.0\n\nA small, honest local toolkit for the work agents end up doing constantly: read some text someone sent you, find the structured bits, clean it up, redact the secrets, and forward it downstream. Built on Python 3 standard library only. No `pandas`, no `nltk`, no pip installs, no remote calls.\n\nThis skill is the companion to [`clean-csv-toolkit`](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit): that one handles structured tabular data, this one handles unstructured text.\n\n## What this skill does\n\n- `scripts/extract.py` — pull structured items out of any text file. Kinds: `url`, `email`, `phone`, `ipv4`, `ipv6`, `hashtag`, `mention`, `hex-color`, `money`, `iso-date`. Output to stdout (one-per-line or JSON), or to a `.txt` / `.json` / `.jsonl` file. Optional `--unique`, `--sort`, `--with-line` (prefix with the source line number).\n- `scripts/normalize.py` — clean up messy text. Chainable transforms applied in command-line order: `--trim`, `--collapse-spaces`, `--strip-blank`, `--to-unix`, `--to-crlf`, `--dehyphenate` (rejoin OCR/PDF hyphenated line-breaks), `--unsmart` (smart quotes / em-dashes → ASCII), `--strip-bom`, `--strip-zwsp` (zero-width spaces and joiners), `--tabs-to-spaces N`, `--spaces-to-tabs N`, `--lower` / `--upper` / `--title`, `--normalize-unicode NFC|NFD|NFKC|NFKD`.\n- `scripts/redact.py` — anonymize text by replacing PII-like patterns with placeholder tokens. Kinds: `email`, `phone`, `ipv4`, `ipv6`, `url`, `credit-card` (with Luhn validation to suppress false positives), `ssn-us`, `uuid`, `hex-token` (32+ hex chars, typical for tokens / hashes), `aws-access-key` (AKIA…), `jwt` (three base64url segments with the `eyJ` header). `--keep-counts` makes the same value always get the same placeholder; `--preserve-length` pads/truncates the placeholder to the original length.\n- `scripts/lines.py` — line-oriented utilities. `--op count | dedupe | sort | shuffle | head | tail`. Streams `count`, `head`, `tail`. `dedupe` and `sort` are O(N) memory in the number of lines, but each line is small so 1 M lines is fine on a laptop. `--case-insensitive`, `--keep first"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7fkwsa5knkdkkachj1p7rwr9843xts\",\n  \"slug\": \"clean-text-toolkit\",\n  \"version\": \"0.4.0\",\n  \"publishedAt\": 1780068349206\n}"},{"path":"skill-card.md","content":"## Description:\n\nLocal text cleanup and inspection toolkit that helps agents extract structured items, redact PII-like patterns, normalize text, inspect lines and word counts, compare text, render simple templates, generate slugs, and process Markdown or HTML using Python 3 standard library scripts with no remote calls.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[gopendrasharma89-tech](https://clawhub.ai/user/gopendrasharma89-tech)\n\n### License/Terms of Use:\n\nMIT\n\n## Use Case:\n\nDevelopers and agents use this skill to clean, inspect, transform, redact, and compare local text files before passing the results into downstream workflows.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: HTML sanitization and Markdown-to-HTML output may not make untrusted content safe for browser or web application use.\n\nMitigation: Use the toolkit for local text processing, and pass untrusted HTML through a maintained allowlist sanitizer before rendering it in a browser or web app.\n\nRisk: The scripts can write or overwrite caller-specified output paths.\n\nMitigation: Review output paths before execution and run commands in a working directory where overwrites are acceptable.\n\n## Reference(s):\n\n- [Clean Text Toolkit ClawHub skill page](https://clawhub.ai/gopendrasharma89-tech/skills/clean-text-toolkit)\n- [Clean Text Toolkit OpenClaw homepage](https://clawhub.ai/gopendrasharma89-tech/clean-text-toolkit)\n- [Clean CSV Toolkit companion skill](https://clawhub.ai/gopendrasharma89-tech/clean-csv-toolkit)\n- [OpenClaw Prompt Shield companion skill](https://clawhub.ai/gopendrasharma89-tech/openclaw-prompt-shield)\n\n## Skill Output:\n\n**Output Type(s):** [Text, Markdown, Code, Shell commands, Configuration]\n\n**Output Format:** [Markdown with inline shell commands and local script guidance; scripts can produce plain text, JSON, JSONL, TSV, HTML, and diff output.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Outputs are local file or stdout results from Python 3 standard library utilities; some operations write or overwrite caller-specified paths.]\n\n## Skill Version(s):\n\n0.4.0 (source: server release metadata and SKILL.md body)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."},{"path":"LICENSE","content":"MIT License\n\nCopyright (c) 2026 gopendrasharma89-tech\n\nPermission is hereby granted, free of charge, to any person obtaining a copy\nof this software and associated documentation files (the \"Software\"), to deal\nin the Software without restriction, including without limitation the rights\nto use, copy, modify, merge, publish, distribute, sublicense, and/or sell\ncopies of the Software, and to permit persons to whom the Software is\nfurnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all\ncopies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\nIMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\nFITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\nAUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\nLIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\nOUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\nSOFTWARE."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-... Skill: Clean Text Toolkit Owner: gopendrasharma89-tech Summary: Local text cleanup and inspection toolkit. Extract structured items (URLs, emails, phones, IPs, dates, hashtags, money), redact PII (email/phone/credit-card-... Tags: agent:0.4.0, dedupe:0.1.0, diff:0.1.0, extract:0.4.0, html:0.4.0, jinja-free:0.3.0, latest:0.4.0, links:0.4.0, markdown:0.3.0, normalize:0.1.0, pii:0.4.0, redact:0.4.0, regex:0.4.0, replace","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":1645,"uniquenessScore":50,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T06:50:57.039Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T06:50:57.039Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T10:50:21.594Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}