agentCLAWHUBUnverified

pdf-to-epub-ocr

将扫描版PDF电子书通过OCR识别转换为结构化精排的EPUB格式。当用户提到"PDF转EPUB"、"PDF转电子书"、"OCR提取PDF"、"扫描版PDF转换"、"PDF结构化处理"或上传PDF文件要求转换为电子书格式时触发。不适用于纯文本PDF(可直接提取文字的PDF不需要OCR)、图片格式转换或PDF编辑功能。 Skill: pdf-to-epub-ocr Owner: erich1566 Summary: 将扫描版PDF电子书通过OCR识别转换为结构化精排的EPUB格式。当用户提到"PDF转EPUB"、"PDF转电子书"、"OCR提取PDF"、"扫描版PDF转换"、"PDF结构化处理"或上传PDF文件要求转换为电子书格式时触发。不适用于纯文本PDF(可直接提取文字的PDF不需要OCR)、图片格式转换或PDF编辑功能。 Tags: latest:0.1.0 Version history: v0.1.0 | 2026-08-07T10:35:59.297Z | auto Initial release of PDF-to-EPUB OCR skill. - Scanned PDF ebooks can be converted to structured, reflowable EPUB files using OCR. - Suppo

OpenClaw

Rank

62

Safety

84

Downloads

2.4k

Updated

Oct 9, 2026

Version

0.1.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.4K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.4K downloadsadoption · observed Oct 9, 2026
Latest release
0.1.0release · observed Aug 7, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17amde2emhcgwty2p2w954ebn84yrw7:pdf-to-epub-ocr
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-erich1566-pdf-to-epub-ocr/snapshot"

Documentation

CLAWHUB

28,082 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: pdf-to-epub-ocr
description: 将扫描版PDF电子书通过OCR识别转换为结构化精排的EPUB格式。当用户提到"PDF转EPUB"、"PDF转电子书"、"OCR提取PDF"、"扫描版PDF转换"、"PDF结构化处理"或上传PDF文件要求转换为电子书格式时触发。不适用于纯文本PDF(可直接提取文字的PDF不需要OCR)、图片格式转换或PDF编辑功能。
---

# PDF转EPUB OCR技能

## 概述

本技能将扫描版PDF电子书通过OCR识别转换为结构化精排的EPUB格式电子书,支持封面提取、智能文本清洗、章节自动识别和元数据管理,生成适合移动设备阅读的电子书文件。

## 你的工作方式

你是一个专业的PDF到EPUB转换专家,负责将用户提供的扫描版PDF文件转换为结构化的EPUB电子书。工作流程包括:文件分析预处理、OCR文字识别、文本清洗与章节识别、EPUB结构化生成、质量验证与交付。

## Phase 1:文件分析与预处理

### 1.1 文件接收与验证
- 接收用户提供的PDF文件路径
- 验证文件是否存在且为有效的PDF格式
- 检查文件大小(建议小于100MB以避免处理超时)
- 确认PDF是否为扫描版(无文本层需要OCR)

### 1.2 元数据提取
- 使用PyPDF2提取PDF基础元数据
- 尝试获取书名、作者、出版社、出版日期等信息
- 如果元数据不完整,询问用户补充关键信息(书名、作者)

### 1.3 封面提取
- 使用pdf2image将PDF第一页转换为图片
- 转换参数:DPI=300,格式=JPEG,质量=95
- 保存为临时封面文件:`cover.jpg`

**预处理完成确认门**:向用户确认提取的元数据信息是否正确,特别是书名和作者。

## Phase 2:OCR文字识别

### 2.1 页面批处理
- 使用pdf2image将PDF逐页转换为图片(DPI=300,格式=PNG)
- 采用分页处理策略,每处理10页输出一次进度
- 大文件建议分批次处理,避免内存溢出

### 2.2 OCR识别执行
- 使用pytesseract进行OCR识别
- 语言配置:`chi_sim+eng`(简体中文+英文)
- OCR参数:`--psm 6`(假设为统一文本块)
- 保留文本坐标信息用于后续章节识别

### 2.3 识别质量控制
- 设置置信度阈值(默认0.6),低于阈值的文本标记为需要人工校对
- 统计识别准确率和处理进度
- 记录识别失败的页面供后续重试

**OCR完成确认门**:向用户报告OCR识别统计信息(总页数、识别成功数、平均置信度),询问是否继续。

## Phase 3:文本清洗与章节识别

### 3.1 文本清洗
执行以下清洗步骤,顺序无关:
1. **去除页眉页脚**:识别并删除页面顶部和底部的重复文本模式
2. **去除页码**:删除纯数字页码和"第X页"格式的内容
3. **去除水印**:识别并删除半透明或重复的水印文本
4. **规范化空白**:将连续空格、换行符、制表符统一处理
5. **OCR错误校正**:修复常见的OCR识别错误(如"0"和"O"混淆)

### 3.2 章节标题识别
使用正则表达式匹配章节标题模式:
- 中文章节:`第[一二三四五六七八九十百千]+章`、`第[0-9]+章`
- 英文章节:`Chapter [0-9]+`、`Part [0-9]+`
- 其他模式:`[0-9]+\.[0-9]+`(如1.1、2.3)

### 3.3 结构化重组
- 根据章节标题将文本分割为独立章节
- 为每个章节生成唯一ID(chapter_001, chapter_002...)
- 构建章节层级结构(支持多级标题)

**清洗完成确认门**:向用户展示识别出的章节结构,确认是否正确。

## Phase 4:EPUB结构化生成

### 4.1 EPUB框架创建
- 使用ebooklib创建EPUB2格式电子书
- 设置书名(来自PDF元数据或用户输入)
- 添加作者、出版社、语言等元数据
- 设置唯一标识符(UUID)

### 4.2 内容组织
- 将封面图片添加为EPUB封面
- 为每个章节创建独立的HTML文件
- 设置正确的MIME类型(text/html, image/jpeg)
- 添加CSS样式文件,优化阅读体验:
  - 设置合适的行高和字间距
  - 支持中文字体回退
  - 响应式布局适配移动设备

### 4.3 目录生成
- 基于章节结构生成EPUB目录
- 设置章节间的导航链接
- 添加线性阅读属性

### 4.4 文件输出
- 保存EPUB文件到workspace/output/目录
- 文件名格式:`{原文件名}_converted.epub`
- 生成转换报告(处理时间、页数、章节数、文件大小)

## Phase 5:质量验证与交付

### 5.1 文件验证
- 检查EPUB文件完整性
- 验证封面、目录、章节是否正确
- 测试在不同设备上的显示效果(如果可能)

### 5.2 交付物确认
向用户交付以下内容:
1. 转换后的EPUB文件
2. 转换报告(包含统计信息和质量评估)
3. 如有问题,提供改进建议

## 异常处理

### 常见问题处理
- **OCR识别失败**:降低DPI重新转换,或建议用户提供更清晰的PDF
- **章节识别错误**:询问用户提供正确的章节划分规则
- **元数据缺失**:提示用户手动补充关键信息
- **文件过大**:建议分批处理或使用压缩版本

### 性能优化
- 对于大文件,采用增量处理策略
- 使用多线程加速OCR识别(如果系统支持)
- 缓存中间结果,支持断点续传

## 工作流示例

### 示例1:标准PDF转换流程
用户上传扫描版PDF书籍《机器学习实战》,要求转换为EPUB格式。

**执行过程**:
1. 接收文件`machine_learning.pdf`,验证PDF格式
2. 提取元数据:书名"机器学习实战",作者"张三"
3. 提取第一页作为封面`cover.jpg`
4. 逐页OCR识别(共256页),统计识别成功率为92%
5. 清洗文本,识别出15个章节
6. 生成EPUB文件`machine_learning_converted.epub`
7. 交付文件和转换报告

### 示例2:包含复杂结构的PDF
用户上传技术文档`api_reference.pdf`,包含多级标题和代码块。

**执行过程**:
1. 预处理发现PDF包含多级结构(章、节、小节)
2. OCR识别时保留格式信息
3. 章节识别使用多级正则表达式
4. 为代码块添加特殊CSS样式保持格式
5. 生成带三级目录的EPUB文件
6. 提示用户某些代码块可能需要人工校对

## 资源目录

本技能包含以下资源目录:

### scripts/
- `pdf_to_epub_converter.py`:核心转换脚本,包含完整的PDF到EPUB转换流程
- `ocr_pr

README.md

# PDF转EPUB OCR技能

将扫描版PDF电子书通过OCR识别转换为结构化精排的EPUB格式电子书。

## 功能特点

- ✅ **智能OCR识别**: 使用Tesseract OCR引擎,支持中英文混合识别
- ✅ **封面提取**: 自动提取PDF第一页作为EPUB封面
- ✅ **文本清洗**: 去除页眉页脚、页码、水印等干扰内容
- ✅ **章节识别**: 自动识别章节标题,生成结构化目录
- ✅ **元数据管理**: 提取和写入书名、作者等元数据信息
- ✅ **移动优化**: 生成适合移动设备阅读的EPUB文件

## 系统依赖

### 必需的系统组件

1. **Tesseract OCR引擎**
   ```bash
   # Ubuntu/Debian
   sudo apt-get install tesseract-ocr tesseract-ocr-chi-sim
   
   # macOS
   brew install tesseract tesseract-lang
   
   # Windows
   # 下载安装: https://github.com/UB-Mannheim/tesseract/wiki
   ```

2. **Poppler** (PDF处理库)
   ```bash
   # Ubuntu/Debian
   sudo apt-get install poppler-utils
   
   # macOS
   brew install poppler
   
   # Windows
   # 下载安装: https://github.com/oschwartz10612/poppler-windows/releases/
   ```

## Python依赖

安装Python依赖包:
```bash
pip install -r requirements.txt
```

## 使用方法

### 基本用法

```bash
python scripts/pdf_to_epub_converter.py your_book.pdf
```

### 高级用法

```bash
# 指定输出目录
python scripts/pdf_to_epub_converter.py your_book.pdf --output-dir ./output

# 指定书名和作者
python scripts/pdf_to_epub_converter.py your_book.pdf \
    --title "我的书籍" \
    --author "作者名"
```

### 在Agent中使用

当用户提到以下内容时,此技能会自动触发:
- "PDF转EPUB"、"PDF转电子书"
- "OCR提取PDF"、"扫描版PDF转换"
- "PDF结构化处理"
- 上传PDF文件要求转换为电子书格式

## 输出说明

### 输出文件
- **位置**: `output/`目录
- **文件名**: `{原文件名}_converted.epub`
- **格式**: EPUB 2.0.1(兼容性最好)

### 转换报告
每个转换任务会生成详细报告,包括:
- 总页数和识别成功页数
- OCR平均置信度
- 识别的章节数量
- 输出文件大小
- 处理时间统计

## 工作流程

1. **文件分析**: 验证PDF格式,提取元数据
2. **封面提取**: 转换第一页为封面图片
3. **OCR识别**: 逐页进行文字识别
4. **文本清洗**: 去除噪音内容
5. **章节识别**: 自动识别章节结构
6. **EPUB生成**: 创建结构化电子书文件
7. **质量验证**: 检查文件完整性

## 质量保证

### OCR质量
- 推荐DPI: 300-400
- 语言支持: 中英文混合
- 置信度监控: 自动统计识别准确率

### EPUB质量
- 结构验证: 自动检查EPUB结构完整性
- 兼容性测试: 支持主流阅读设备
- 响应式设计: 适配不同屏幕尺寸

## 故障排除

### 常见问题

**问题**: Tesseract not found
```bash
# 解决方案: 安装Tesseract OCR引擎
sudo apt-get install tesseract-ocr tesseract-ocr-chi-sim
```

**问题**: 识别准确率低
```bash
# 解决方案: 
# 1. 提高PDF扫描质量
# 2. 调整DPI设置(推荐300-400)
# 3. 使用预处理脚本增强图像
```

**问题**: 章节识别错误
```bash
# 解决方案: 
# 1. 检查章节标题格式
# 2. 在代码中添加自定义正则表达式
# 3. 手动调整章节划分
```

## 性能参数

### 典型处理时间(基于300 DPI)
- 简单文本PDF (100页): 2-3分钟
- 复杂排版PDF (100页): 5-8分钟
- 高质量扫描PDF (100页): 3-5分钟

### 硬件建议
- **CPU**: 多核处理器(4核以上)
- **内存**: 最低4GB,推荐8GB
- **存储**: SSD硬盘(显著提升I/O性能)

## 配置选项

### OCR配置
可在`scripts/ocr_processor.py`中调整:
- `dpi`: 图片转换分辨率(默认300)
- `language`: OCR语言(默认`chi_sim+eng`)
- `psm_mode`: 页面分割模式(默认6)

### 清洗配置
可在`scripts/text_cleaner.py`中调整:
- 章节标题识别模式
- 噪音文本过滤规则
- 文本清洗策略

### EPUB配置
可在`scripts/epub_generator.py`中调整:
- CSS样式文件
- 章节HTML模板
- 元数据设置

## 参考文档

- `references/ocr_best_practices.md`: OCR识别最佳实践
- `references/chapter_patterns.md`: 章节标题识别模式库
- `references/epub_structure_guide.md`: EPUB结构说明和样式规范

## 技术支持

如遇到问题,请检查:
1. 系统依赖是否正确安装
2. Python依赖包是否完整
3. PDF文件是否有效
4. 系统资源是否充足

## 许可证

本技能仅供学习和个人使用。使用时请遵守相关版权法律法规。

## 更新日志

### v1.0.0 (2024-08-07)
- ✨ 初始版本发布
- ✅ 支持基础PDF到EPUB转换
- ✅ OCR识别和文本清洗
- ✅ 章节自动识别
- ✅ 元数据管理

_meta.json

{
  "ownerId": "kn705qkxd5q2sxd49s59dnfknd82dywx",
  "slug": "pdf-to-epub-ocr",
  "version": "0.1.0",
  "publishedAt": 1786098959297
}

references/chapter_patterns.md

# 章节标题识别正则表达式模式库

## 中文图书章节模式

### 基础章节模式
```python
# 第X章格式
r'^第[一二三四五六七八九十百千万]+章[^\n]*$'
r'^第[0-9]+章[^\n]*$'

# 第X节格式  
r'^第[一二三四五六七八九十百千万]+节[^\n]*$'
r'^第[0-9]+节[^\n]*$'

# 第X篇格式
r'^第[一二三四五六七八九十百千万]+篇[^\n]*$'
r'^第[0-9]+篇[^\n]*$'
```

### 扩展章节模式
```python
# 带副标题的章节
r'^第[一二三四五六七八九十百千万]+章[::][^\n]+$'
r'^第[0-9]+章[::][^\n]+$'

# 带括号的章节
r'^\([一二三四五六七八九十百千万]+\)[^\n]*$'
r'^\([0-9]+\)[^\n]*$'

# 中文数字章节
r'^[一二三四五六七八九十百千万]+\.[^\n]*$'
r'^[一二三四五六七八九十百千万]+、[^\n]*$'
```

## 英文图书章节模式

### 基础英文模式
```python
# Chapter格式
r'^Chapter\s+[0-9IVXLCDM]+[^\n]*$'
r'^CHAPTER\s+[0-9IVXLCDM]+[^\n]*$'

# Part格式
r'^Part\s+[0-9IVXLCDM]+[^\n]*$'
r'^PART\s+[0-9IVXLCDM]+[^\n]*$'

# Section格式
r'^Section\s+[0-9]+[^\n]*$'
r'^SECTION\s+[0-9]+[^\n]*$'
```

### 扩展英文模式
```python
# 带标题的章节
r'^Chapter\s+[0-9IVXLCDM]+:\s*[^\n]+$'
r'^Part\s+[0-9IVXLCDM]+:\s*[^\n]+$'

# 罗马数字章节
r'^[IVXLCDM]+\.\s*[^\n]+$'
r'^[IVXLCDM]+\s+[^\n]+$'

# 字母章节
r'^[A-Z]+\.\s*[^\n]+$'
r'^Appendix\s+[A-Z][^\n]*$'
```

## 数字编号模式

### 多级编号
```python
# 一级编号
r'^[0-9]+\.[^\n]*$'

# 二级编号  
r'^[0-9]+\.[0-9]+[^\n]*$'
r'^[0-9]+\.[0-9]+\.[^\n]*$'

# 三级编号
r'^[0-9]+\.[0-9]+\.[0-9]+[^\n]*$'
r'^[0-9]+\.[0-9]+\.[0-9]+\.[^\n]*$'
```

### 括号编号
```python
# 圆括号编号
r'^\([0-9]+\)[^\n]*$'
r'^\([0-9]+\.[0-9]+\)[^\n]*$'

# 方括号编号
r'^\[[0-9]+\][^\n]*$'
r'^\[[0-9]+\.[0-9]+\][^\n]*$'
```

## 学术论文模式

### 论文章节
```python
# 摘要、关键词等
r'^[摘要|关键词|Abstract|Keywords|引言|结论|参考文献][^\n]*$'

# 学术章节
r'^[0-9]+\s*[引言|文献综述|研究方法|实验结果|讨论|结论][^\n]*$'

# 图表标题
r'^图\s*[0-9]+[^\n]*$'
r'^表\s*[0-9]+[^\n]*$'
r'^Figure\s*[0-9]+[^\n]*$'
r'^Table\s*[0-9]+[^\n]*$'
```

## 技术文档模式

### 技术章节
```python
# 技术文档章节
r'^[0-9]+\s*[概述|简介|背景|目标|范围][^\n]*$'
r'^[0-9]+\s*[安装|配置|部署][^\n]*$'
r'^[0-9]+\s*[使用指南|操作说明][^\n]*$'
r'^[0-9]+\s*[故障排除|常见问题][^\n]*$'

# API文档
r'^[GET|POST|PUT|DELETE|PATCH]\s+[^\n]+$'
r'^/api/[^\n]+$'
r'^[0-9]+\.[0-9]+\s+接口[^\n]*$'
```

## 小说文学作品模式

### 小说章节
```python
# 卷篇章节
r'^第[一二三四五六七八九十百千万]+卷[^\n]*$'
r'^第[一二三四五六七八九十百千万]+部[^\n]*$'

# 回目格式(传统小说)
r'^第[一二三四五六七八九十百千万]+回[^\n]*$'
r'^[0-9]+回[^\n]*$'

# 现代小说章节
r'^[一二三四五六七八九十百千万]+、[^\n]+$'
r'^Chapter\s*[0-9]+[^\n]*$'
```

## 混合模式匹配策略

### 优先级排序
```python
# 高优先级:明确的章节标记
PRIORITY_HIGH = [
    r'^第[一二三四五六七八九十百千万]+章[^\n]*$',
    r'^Chapter\s+[0-9IVXLCDM]+[^\n]*$',
]

# 中优先级:数字编号
PRIORITY_MEDIUM = [
    r'^[0-9]+\.[0-9]+[^\n]*$',
    r'^第[0-9]+章[^\n]*$',
]

# 低优先级:可能的章节
PRIORITY_LOW = [
    r'^[0-9]+\s+[^\n]{5,30}$',
    r'^[A-Z]+\.[^\n]{5,30}$',
]
```

### 上下文验证
```python
def validate_chapter_title(line, prev_lines, next_lines):
    """
    验证章节标题的有效性
    
    Args:
        line: 当前行
        prev_lines: 前面几行
        next_lines: 后面几行
    
    Returns:
        是否为有效的章节标题
    """
    # 检查行长度(章节标题通常较短)
    if len(line) > 50:
        return False
    
    # 检查前后内容(章节标题前后通常有空白)
    if prev_lines and prev_lines[-1].strip():
        # 如果前面一行有内容,检查是否是页码等噪音
        if re.match(r'^\d+$', prev_lines[-1].strip()):
            return True
        return False
    
    # 检查后面内容(章节标题后应该有正文)
    if next_lines and not next_lines[0].strip():
        return Fa

references/epub_structure_guide.md

# EPUB文件结构说明和样式规范

## EPUB基础结构

### 目录结构
```
EPUB文件(ZIP压缩包)
├── mimetype
├── META-INF/
│   └── container.xml
└── OEBPS/
    ├── content.opf
    ├── toc.ncx
    └── chapters/
        ├── chapter_001.xhtml
        ├── chapter_002.xhtml
        ├── style/
        │   └── default.css
        └── images/
            ├── cover.jpg
            └── figure_001.jpg
```

### 关键文件说明

#### mimetype
- 必须是第一个文件且未压缩
- 内容固定为:`application/epub+zip`

#### container.xml
- 定义了OPF文件的路径
- 示例:
```xml
<?xml version="1.0"?>
<container version="1.0" xmlns="urn:oasis:names:tc:opendocument:xmlns:container">
  <rootfiles>
    <rootfile full-path="OEBPS/content.opf" media-type="application/oebps-package+xml"/>
  </rootfiles>
</container>
```

#### content.opf
- 包含书籍的元数据、清单和spine
- 描述了书籍的所有组成部分

#### toc.ncx
- 导航控制文件(EPUB2)或nav.xhtml(EPUB3)
- 定义了书籍的目录结构

## 元数据规范

### 基础元数据
```xml
<metadata xmlns:dc="http://purl.org/dc/elements/1.1/">
  <!-- 必需 -->
  <dc:title>书名</dc:title>
  <dc:identifier id="bookid">unique-id-12345</dc:identifier>
  
  <!-- 推荐 -->
  <dc:language>zh-CN</dc:language>
  <dc:creator>作者名</dc:creator>
  
  <!-- 可选 -->
  <dc:publisher>出版社</dc:publisher>
  <dc:date>2024-01-01</dc:date>
  <dc:description>书籍描述</dc:description>
  <dc:subject>主题分类</dc:subject>
  <dc:rights>版权信息</dc:rights>
</metadata>
```

### 元数据最佳实践
1. **标题**: 简洁明确,避免过长
2. **作者**: 使用真实姓名或常用笔名
3. **语言**: 使用标准的语言代码(zh-CN, en-US等)
4. **标识符**: 使用UUID或ISBN
5. **描述**: 100-500字的简介,便于搜索

## 内容组织规范

### 章节文件命名
```
chapter_001.xhtml  # 第一章
chapter_002.xhtml  # 第二章
...
chapter_999.xhtml  # 附录等
```

### 章节文件结构
```xml
<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.1//EN"
  "http://www.w3.org/TR/xhtml11/DTD/xhtml11.dtd">

<html xmlns="http://www.w3.org/1999/xhtml">
<head>
  <title>章节标题</title>
  <link rel="stylesheet" type="text/css" href="style/default.css"/>
</head>
<body>
  <h1>章节标题</h1>
  <div class="content">
    <p>段落内容...</p>
    <p>更多段落...</p>
  </div>
</body>
</html>
```

## CSS样式规范

### 基础样式框架
```css
/* 全局设置 */
* {
  margin: 0;
  padding: 0;
  box-sizing: border-box;
}

body {
  font-family: "PingFang SC", "Microsoft YaHei", "SimHei", serif;
  line-height: 1.8;
  margin: 1em;
  padding: 0;
  color: #333;
  background-color: #fff;
}
```

### 标题样式
```css
h1 {
  text-align: center;
  margin: 1em 0;
  font-size: 1.8em;
  color: #333;
  border-bottom: 2px solid #eee;
  padding-bottom: 0.5em;
  font-weight: bold;
}

h2 {
  margin: 1.5em 0 0.8em 0;
  font-size: 1.4em;
  color: #444;
  border-left: 4px solid #007bff;
  padding-left: 0.5em;
}

h3 {
  margin: 1.2em 0 0.6em 0;
  font-size: 1.2em;
  color: #555;
}
```

### 正文样式
```css
p {
  margin: 0.8em 0;
  text-align: justify;
  text-indent: 2em;
  font-size: 1em;
  line-height: 1.8;
  color: #333;
}

/* 首字下沉(可选) */
p:first-of-type::first-letter {
  font-size: 2em;
  font-weight: bold;
  float: left;
  margin-right: 0.1em;
  line-height: 1;
}
```

### 响应式设计
```css
/* 平板设备 */
@media screen and (max-width: 768px) {
  body {
    fon
Github ReposUpdated 5h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/erich1566/skills/pdf-to-epub-ocr",
      "sourceUrl": "https://clawhub.ai/erich1566/skills/pdf-to-epub-ocr",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T15:23:40.208Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-erich1566-pdf-to-epub-ocr/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-erich1566-pdf-to-epub-ocr/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T15:23:40.208Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.4K downloads",
      "href": "https://clawhub.ai/erich1566/pdf-to-epub-ocr",
      "sourceUrl": "https://clawhub.ai/erich1566/pdf-to-epub-ocr",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T15:23:40.208Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "0.1.0",
      "href": "https://clawhub.ai/erich1566/pdf-to-epub-ocr",
      "sourceUrl": "https://clawhub.ai/erich1566/pdf-to-epub-ocr",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-07T10:35:59.297Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-erich1566-pdf-to-epub-ocr/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-erich1566-pdf-to-epub-ocr/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 0.1.0",
      "description": "Initial release of PDF-to-EPUB OCR skill. - Scanned PDF ebooks can be converted to structured, reflowable EPUB files using OCR. - Supports cover image extraction, automatic chapter detection, text cleaning, and metadata management. - Handles multi-stage workflow: file validation, OCR processing, text structuring, EPUB generation, and quality checks. - Excludes pure-text PDFs and offers guidance for common issues (e.g., low quality scans, missing metadata). - Included resource directories for scripts, best practices, and templates.",
      "href": "https://clawhub.ai/erich1566/pdf-to-epub-ocr",
      "sourceUrl": "https://clawhub.ai/erich1566/pdf-to-epub-ocr",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-07T10:35:59.297Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to pdf-to-epub-ocr and adjacent AI workflows.