Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multipl... Skill: Pdf Owner: yang1002378395-cmyk Summary: Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multipl... Tags: chinese:1.0.0, document:1.0.0, latest:1.0.0, pdf:1.0.0 Version history: v1.0.0 | 2026-03-24T02:39:32.231Z | auto Initial release of the PDF Processor (Chinese) skill. - Provides comprehensive instructions f
Rank
62
Safety
84
Downloads
5.4k
Updated
Oct 9, 2026
Version
1.0.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 5.4K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 5.4K downloadsadoption · observed Oct 9, 2026
- Latest release
- 1.0.0release · observed Mar 24, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17627rpbb3atq3mbfabqbgtm983g9et:pdf-processor-cn- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-yang1002378395-cmyk-pdf-processor-cn/snapshot"
Documentation
CLAWHUB
41,915 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: pdf
description: Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.
license: Proprietary. LICENSE.txt has complete terms
---
# PDF Processing Guide
## Overview
This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see REFERENCE.md. If you need to fill out a PDF form, read FORMS.md and follow its instructions.
## Quick Start
```python
from pypdf import PdfReader, PdfWriter
# Read a PDF
reader = PdfReader("document.pdf")
print(f"Pages: {len(reader.pages)}")
# Extract text
text = ""
for page in reader.pages:
text += page.extract_text()
```
## Python Libraries
### pypdf - Basic Operations
#### Merge PDFs
```python
from pypdf import PdfWriter, PdfReader
writer = PdfWriter()
for pdf_file in ["doc1.pdf", "doc2.pdf", "doc3.pdf"]:
reader = PdfReader(pdf_file)
for page in reader.pages:
writer.add_page(page)
with open("merged.pdf", "wb") as output:
writer.write(output)
```
#### Split PDF
```python
reader = PdfReader("input.pdf")
for i, page in enumerate(reader.pages):
writer = PdfWriter()
writer.add_page(page)
with open(f"page_{i+1}.pdf", "wb") as output:
writer.write(output)
```
#### Extract Metadata
```python
reader = PdfReader("document.pdf")
meta = reader.metadata
print(f"Title: {meta.title}")
print(f"Author: {meta.author}")
print(f"Subject: {meta.subject}")
print(f"Creator: {meta.creator}")
```
#### Rotate Pages
```python
reader = PdfReader("input.pdf")
writer = PdfWriter()
page = reader.pages[0]
page.rotate(90) # Rotate 90 degrees clockwise
writer.add_page(page)
with open("rotated.pdf", "wb") as output:
writer.write(output)
```
### pdfplumber - Text and Table Extraction
#### Extract Text with Layout
```python
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:
for page in pdf.pages:
text = page.extract_text()
print(text)
```
#### Extract Tables
```python
with pdfplumber.open("document.pdf") as pdf:
for i, page in enumerate(pdf.pages):
tables = page.extract_tables()
for j, table in enumerate(tables):
print(f"Table {j+1} on page {i+1}:")
for row in table:
print(row)
```
#### Advanced Table Extraction
```python
import pandas as pd
with pdfplumber.open("document.pdf") as pdf:
all_tables = []
for page in pdf.pages:
tables = page.extract_tables()
for table in tables:
if table: # Check if table is not empty
df = pd.DataFrame(table[1:], columns=_meta.json
{
"ownerId": "kn7a0nk2nhyhkcwjry4knbqngn82qxtg",
"slug": "pdf-processor-cn",
"version": "1.0.0",
"publishedAt": 1774319972231
}forms.md
**CRITICAL: You MUST complete these steps in order. Do not skip ahead to writing code.**
If you need to fill out a PDF form, first check to see if the PDF has fillable form fields. Run this script from this file's directory:
`python scripts/check_fillable_fields <file.pdf>`, and depending on the result go to either the "Fillable fields" or "Non-fillable fields" and follow those instructions.
# Fillable fields
If the PDF has fillable form fields:
- Run this script from this file's directory: `python scripts/extract_form_field_info.py <input.pdf> <field_info.json>`. It will create a JSON file with a list of fields in this format:
```
[
{
"field_id": (unique ID for the field),
"page": (page number, 1-based),
"rect": ([left, bottom, right, top] bounding box in PDF coordinates, y=0 is the bottom of the page),
"type": ("text", "checkbox", "radio_group", or "choice"),
},
// Checkboxes have "checked_value" and "unchecked_value" properties:
{
"field_id": (unique ID for the field),
"page": (page number, 1-based),
"type": "checkbox",
"checked_value": (Set the field to this value to check the checkbox),
"unchecked_value": (Set the field to this value to uncheck the checkbox),
},
// Radio groups have a "radio_options" list with the possible choices.
{
"field_id": (unique ID for the field),
"page": (page number, 1-based),
"type": "radio_group",
"radio_options": [
{
"value": (set the field to this value to select this radio option),
"rect": (bounding box for the radio button for this option)
},
// Other radio options
]
},
// Multiple choice fields have a "choice_options" list with the possible choices:
{
"field_id": (unique ID for the field),
"page": (page number, 1-based),
"type": "choice",
"choice_options": [
{
"value": (set the field to this value to select this option),
"text": (display text of the option)
},
// Other choice options
],
}
]
```
- Convert the PDF to PNGs (one image for each page) with this script (run from this file's directory):
`python scripts/convert_pdf_to_images.py <file.pdf> <output_directory>`
Then analyze the images to determine the purpose of each form field (make sure to convert the bounding box PDF coordinates to image coordinates).
- Create a `field_values.json` file in this format with the values to be entered for each field:
```
[
{
"field_id": "last_name", // Must match the field_id from `extract_form_field_info.py`
"description": "The user's last name",
"page": 1, // Must match the "page" value in field_info.json
"value": "Simpson"
},
{
"field_id": "Checkbox12",
"description": "Checkbox to be checked if the user is 18 or over",
"page": 1,
"value": "/On" // If this is a checkbox, use its "checked_value" value to check it. If it's a radio button group, use one of the "value" values in "radio_options".
},
// more fields
]
reference.md
# PDF Processing Advanced Reference
This document contains advanced PDF processing features, detailed examples, and additional libraries not covered in the main skill instructions.
## pypdfium2 Library (Apache/BSD License)
### Overview
pypdfium2 is a Python binding for PDFium (Chromium's PDF library). It's excellent for fast PDF rendering, image generation, and serves as a PyMuPDF replacement.
### Render PDF to Images
```python
import pypdfium2 as pdfium
from PIL import Image
# Load PDF
pdf = pdfium.PdfDocument("document.pdf")
# Render page to image
page = pdf[0] # First page
bitmap = page.render(
scale=2.0, # Higher resolution
rotation=0 # No rotation
)
# Convert to PIL Image
img = bitmap.to_pil()
img.save("page_1.png", "PNG")
# Process multiple pages
for i, page in enumerate(pdf):
bitmap = page.render(scale=1.5)
img = bitmap.to_pil()
img.save(f"page_{i+1}.jpg", "JPEG", quality=90)
```
### Extract Text with pypdfium2
```python
import pypdfium2 as pdfium
pdf = pdfium.PdfDocument("document.pdf")
for i, page in enumerate(pdf):
text = page.get_text()
print(f"Page {i+1} text length: {len(text)} chars")
```
## JavaScript Libraries
### pdf-lib (MIT License)
pdf-lib is a powerful JavaScript library for creating and modifying PDF documents in any JavaScript environment.
#### Load and Manipulate Existing PDF
```javascript
import { PDFDocument } from 'pdf-lib';
import fs from 'fs';
async function manipulatePDF() {
// Load existing PDF
const existingPdfBytes = fs.readFileSync('input.pdf');
const pdfDoc = await PDFDocument.load(existingPdfBytes);
// Get page count
const pageCount = pdfDoc.getPageCount();
console.log(`Document has ${pageCount} pages`);
// Add new page
const newPage = pdfDoc.addPage([600, 400]);
newPage.drawText('Added by pdf-lib', {
x: 100,
y: 300,
size: 16
});
// Save modified PDF
const pdfBytes = await pdfDoc.save();
fs.writeFileSync('modified.pdf', pdfBytes);
}
```
#### Create Complex PDFs from Scratch
```javascript
import { PDFDocument, rgb, StandardFonts } from 'pdf-lib';
import fs from 'fs';
async function createPDF() {
const pdfDoc = await PDFDocument.create();
// Add fonts
const helveticaFont = await pdfDoc.embedFont(StandardFonts.Helvetica);
const helveticaBold = await pdfDoc.embedFont(StandardFonts.HelveticaBold);
// Add page
const page = pdfDoc.addPage([595, 842]); // A4 size
const { width, height } = page.getSize();
// Add text with styling
page.drawText('Invoice #12345', {
x: 50,
y: height - 50,
size: 18,
font: helveticaBold,
color: rgb(0.2, 0.2, 0.8)
});
// Add rectangle (header background)
page.drawRectangle({
x: 40,
y: height - 100,
width: width - 80,
height: 30,
color: rgb(0.9, 0.9, 0.9)
});
// Add table-like content
const items = [
['Item', 'Qty', 'Pskill-card.md
## Description: Provides guidance and helper scripts for reading, extracting, combining, splitting, modifying, filling, securing, and OCR-processing PDF files. This skill is ready for commercial/non-commercial use. ## Publisher: [yang1002378395-cmyk](https://clawhub.ai/user/yang1002378395-cmyk) ### License/Terms of Use: Proprietary ## Use Case: Developers, engineers, and document-processing agents use this skill to inspect, transform, create, secure, OCR, and fill PDF documents with Python libraries and command-line tools. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: PDF workflows can expose sensitive or unauthorized documents. Mitigation: Use the skill only on PDFs the user is authorized to process and handle generated outputs according to the document sensitivity. Risk: Repair, decryption, form filling, and overwrite-style workflows can alter document contents. Mitigation: Prefer writing outputs to new files and keep backups before making irreversible document changes. Risk: OCR and PDF-processing dependencies can introduce supply-chain or environment risk. Mitigation: Install optional dependencies in an isolated environment with reviewed and pinned versions when possible. ## Reference(s): - [PDF Processing Guide](artifact/SKILL.md) - [PDF Form Processing Guide](artifact/forms.md) - [PDF Processing Advanced Reference](artifact/reference.md) ## Skill Output: **Output Type(s):** [Text, Markdown, Code, Shell commands, Configuration, Guidance] **Output Format:** [Markdown guidance with code blocks, shell commands, and JSON examples] **Output Parameters:** [1D] **Other Properties Related to Output:** [May lead an agent to create derived PDF, image, text, spreadsheet, or JSON files when the referenced tools and scripts are run.] ## Skill Version(s): 1.0.0 (source: server release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/yang1002378395-cmyk/skills/pdf-processor-cn",
"sourceUrl": "https://clawhub.ai/yang1002378395-cmyk/skills/pdf-processor-cn",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T04:02:05.515Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-yang1002378395-cmyk-pdf-processor-cn/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-yang1002378395-cmyk-pdf-processor-cn/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T04:02:05.515Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "5.4K downloads",
"href": "https://clawhub.ai/yang1002378395-cmyk/pdf-processor-cn",
"sourceUrl": "https://clawhub.ai/yang1002378395-cmyk/pdf-processor-cn",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T04:02:05.515Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.0",
"href": "https://clawhub.ai/yang1002378395-cmyk/pdf-processor-cn",
"sourceUrl": "https://clawhub.ai/yang1002378395-cmyk/pdf-processor-cn",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-03-24T02:39:32.231Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-yang1002378395-cmyk-pdf-processor-cn/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-yang1002378395-cmyk-pdf-processor-cn/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.0",
"description": "Initial release of the PDF Processor (Chinese) skill. - Provides comprehensive instructions for common PDF operations including reading, merging, splitting, rotating, watermarking, extracting images/text/tables, password protection, and OCR. - Features Python code examples using pypdf, pdfplumber, reportlab, and command-line usage for tools like pdftotext, qpdf, pdftk, and pdfimages. - Includes important guidelines for handling subscripts/superscripts in created PDFs. - Offers references for further advanced usage and additional resources (FORMS.md, REFERENCE.md).",
"href": "https://clawhub.ai/yang1002378395-cmyk/pdf-processor-cn",
"sourceUrl": "https://clawhub.ai/yang1002378395-cmyk/pdf-processor-cn",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-03-24T02:39:32.231Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
