agentCLAWHUBUnverified

Web To Fim

将任意网页链接或本地文件一键转为结构化 Markdown,并三处存放到 Obsidian Vault、飞书云盘、腾讯 IMA 知识库。 支持的信源:(1) X/Twitter 推文、长文 Article、Thread 线程(逐字转录);(2) 微信公众号文章(带图片,IMA 服务端抓取保留排版); (3) 飞书...

OpenClaw

Rank

62

Safety

84

Downloads

2.0k

Updated

Oct 9, 2026

Version

3.7.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2K downloadsadoption · observed Oct 9, 2026
Latest release
3.7.0release · observed Jul 18, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s177q4wcvafq6fzfkhk2g3cwth83y01d:web-to-fim
  1. Install using `clawhub skill install s177q4wcvafq6fzfkhk2g3cwth83y01d:web-to-fim` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/edwardwason/web-to-fim before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-edwardwason-web-to-fim/snapshot"

Documentation

CLAWHUB

148,805 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: web-to-FIM
label: 网页内容转 Markdown/飞书/IMA
slug: web-to-fim
displayName: web-to-FIM
version: 3.7.0
summary: 将任意网页链接或本地文件一键转为结构化 Markdown,并三处存放到 Obsidian、飞书云盘、腾讯 IMA 知识库。
license: MIT-0
description: >
  将任意网页链接或本地文件一键转为结构化 Markdown,并三处存放到 Obsidian Vault、飞书云盘、腾讯 IMA 知识库。
  支持的信源:(1) X/Twitter 推文、长文 Article、Thread 线程(逐字转录);(2) 微信公众号文章(带图片,IMA 服务端抓取保留排版);
  (3) 飞书 wiki 文档(v3.7.0 改用 lark-cli docs +fetch 抓取完整内容,WebFetch 截断严重已弃用);
  (4) 小红书笔记;(5) 微博;(6) YouTube 视频;
  (7) 任意 HTML 网页(带图片,IMA 服务端抓取);(8) 本地文件:PDF、Word、PPT、Excel、图片、音频等。
  三处存放:Obsidian(本地 Markdown+frontmatter+tags)+ 飞书云盘(.md 原文件,用户身份上传立即可见)+ IMA 知识库(FIM知识库,AI 原生)。
  v3.7.0 飞书 wiki 抓取改造:从 WebFetch(截断严重,200-1000 字符)改为 lark-cli docs +fetch(完整内容,3K-25K 字符),实测 33 篇文章中 19 篇被 WebFetch 截断。
  v3.6.0 飞书存储改造:从应用身份创建在线 docx 改为 lark-cli drive +upload 上传 .md 原文件到用户云盘,用户在飞书"我的空间 > 云盘"立即可见。
  IMA 智能路由:公众号/普通网页 → import_urls 保留图片;X/Twitter/飞书/GitHub → 纯文本笔记逐字转录。
  飞书 wiki 原文链接优先转录:检查文章头部"原文链接",有则优先抓取原文内容作为最终产物存三处,IMA 只存原文链接地址自动识别。
  X 链接转录流程:飞书 wiki 中发现 X 原文链接时,用 x-tweet-fetcher 抓取完整长文(10K-15K 字符),失败时降级到飞书 wiki 内容。
  批量模式支持断点恢复,Obsidian 写入失败自动 fallback。
  工作流:自动识别 URL/文件类型 → 路由到最佳抓取工具 → 结构化 Markdown → 三处存放。
  触发词:抓网页存飞书、web to feishu、url转文档、文件转飞书、存到ima、存到obsidian、web to fim、三处存放。
  当用户明确要求将 URL 或本地文件转存为文档并存储到 Obsidian/飞书/IMA 时触发。
  Do NOT use for 创建文档内容、编辑飞书文档正文、代码开发、非文档转存任务、普通文档格式转换(非三处存放意图)。
---

# Web-to-FIM | 网页内容转 Markdown/飞书/IMA

## 🤖 AI 时代必备的信息库基础技能

在 AI 时代,无论是使用 **OpenClaw**、**Hermes Agent**,还是实践 **Obsidian + LLM** 的信息管理方法论,**一键入库、人机共用**的 AI 信息库搭建都是必备的基础设施。

**Web-to-FIM** 就是这样一个基础技能:它将任意网络内容一键转换为结构化 Markdown,并同步到:
- 📝 **Obsidian Vault** - 本地个人知识库,带 frontmatter
- 📁 **飞书云盘** - .md 原文件,用户身份上传立即可见(v3.6.0)
- 🧠 **腾讯 IMA 知识库** - AI 原生知识库(FIM知识库)

将任意网页链接或本地文件一键转为结构化 Markdown,并三处存放到 Obsidian Vault、飞书云盘、腾讯 IMA 知识库。

## 支持的信源

| 信源 | URL 特征 | 抓取方式 |
|------|---------|---------|
| X/Twitter | `x.com` / `twitter.com` | x-tweet-fetcher(逐字转录) |
| 微信公众号 | `mp.weixin.qq.com` | markitdown + 移动端 UA fallback |
| 飞书 wiki | `*.feishu.cn/wiki/` | **lark-cli docs +fetch**(v3.7.0 改造,完整内容) |
| 小红书 | `xiaohongshu.com` / `xhslink.com` | markitdown |
| 微博 | `weibo.com` | markitdown |
| YouTube | `youtube.com` / `youtu.be` | markitdown |
| 任意网页 | 其他 `http(s)://` 链接 | markitdown |

> ⚠️ **v3.7.0 飞书 wiki 抓取方式变更**:
> - **旧版本(v3.6.0 及更早)**:用 WebFetch 抓取飞书 wiki 内容,但实测会严重截断(200-1000 字符,完整文章应有 3K-25K 字符)。第十批 33 篇飞书 wiki 文章转录时,19 篇被 WebFetch 截断,导致飞书云盘上保存的内容大量丢失。
> - **新版本(v3.7.0)**:改用 `lark-cli docs +fetch --doc <url> --doc-format markdown --scope full` 抓取完整内容。lark-cli 通过用户身份认证,可读取全文,实测 33 篇全部完整获取(800-24000 字符)。
> - WebFetch 对飞书 wiki 的截断问题源于飞书 wiki 页面需 JS 渲染 + 登录认证,WebFetch 无法处理。lark-cli 通过官方 OpenAPI 直接获取结构化内容,无此问题。

### 本地文件支持

| 类型 | 扩展名 |
|------|--------|
| PDF | `.pdf` |
| Word | `.docx` / `.doc` |
| PowerPoint | `.pptx` / `.ppt` |
| Excel | `.xlsx` / `.xls` |
| 图片 | `.png` `.jpg` `.jpeg` `.gif` `.webp` |
| 音频 | `.mp3` `.wav` `.m4a` `.flac` |
| 数据 | `.csv` `.json` `.xml` |

## 输出目的地

| 目的地 | 依赖 | 说明 |
|--------|------|

README.md

[English](./README.en.md) | 中文

# web-to-FIM

> 将任意网页链接或本地文件一键转为结构化 Markdown,并三处存放到 Obsidian Vault、飞书云文档、腾讯 IMA 知识库。

## 核心功能

### 支持的信源(7 类 URL + 本地文件)

| 信源 | URL 特征 | 抓取方式 |
|------|---------|---------|
| X/Twitter | `x.com` / `twitter.com` | x-tweet-fetcher(逐字转录) |
| 微信公众号 | `mp.weixin.qq.com` | markitdown + 移动端 UA fallback |
| 飞书 wiki | `*.feishu.cn/wiki/` | **lark-cli docs +fetch**(v3.7.0,完整内容) |
| 小红书 | `xiaohongshu.com` / `xhslink.com` | markitdown |
| 微博 | `weibo.com` | markitdown |
| YouTube | `youtube.com` / `youtu.be` | markitdown |
| 任意网页 | 其他 `http(s)://` 链接 | markitdown |

本地文件支持:PDF、Word、PPT、Excel、图片、音频、CSV/JSON/XML 等。

### 三处存放

| 目的地 | 环境变量 | 说明 |
|--------|---------|------|
| Obsidian Vault | `OBSIDIAN_VAULT_PATH` | 本地 Markdown + frontmatter + tags(默认 `E:\Obsidian-Vault\00-Inbox`) |
| 飞书云文档 | `FEISHU_APP_ID` + `FEISHU_APP_SECRET` | 云端团队协作 |
| 腾讯 IMA | `IMA_OPENAPI_CLIENTID` + `IMA_OPENAPI_APIKEY` | AI 原生知识库(FIM知识库) |

## 快速开始

```bash
# Install dependencies
pip install markitdown requests python-dotenv

# Convert and save to all three destinations
python3 scripts/web_to_all.py --url "<url_or_path>"

# Convert with custom title
python3 scripts/web_to_all.py --url "<url>" --title "自定义标题"

# Save to Obsidian only
python3 scripts/web_to_all.py --url "<url>" --no-feishu --no-ima
```

## 环境变量配置

```bash
# Obsidian Vault path (optional, has default)
$env:OBSIDIAN_VAULT_PATH = "E:\Obsidian-Vault\00-Inbox"

# Feishu credentials
$env:FEISHU_APP_ID = "your_app_id"
$env:FEISHU_APP_SECRET = "your_app_secret"

# IMA credentials (v3.0 - OpenAPI v1.1.7)
$env:IMA_OPENAPI_CLIENTID = "your_client_id"
$env:IMA_OPENAPI_APIKEY = "your_api_key"

# Optional: IMA knowledge base name (default: FIM知识库)
$env:IMA_KB_NAME = "FIM知识库"
```

> 凭证必须通过环境变量配置,禁止硬编码。旧变量名 `IMA_CLIENT_ID` / `IMA_API_KEY` 仍兼容回退。

## 使用示例

```bash
# X/Twitter tweet
python3 scripts/web_to_all.py --url "https://x.com/user/status/123"

# WeChat article
python3 scripts/web_to_all.py --url "https://mp.weixin.qq.com/s/xxxxx"

# Feishu wiki
python3 scripts/web_to_all.py --url "https://xxx.feishu.cn/wiki/xxxxx"

# Local PDF file
python3 scripts/web_to_all.py --url "C:\docs\report.pdf"

# Only convert to Markdown, no storage
python3 scripts/web_to_md.py --url "<url>" --output output.md
```

## 触发词

转文档、抓网页存飞书、网页转文档、web to feishu、url转文档、文件转飞书、存到ima、存到obsidian、web to fim、三处存放。

当用户提供任意 URL 或本地文件并要求转存为文档时触发。

## 验证连接

```bash
# Verify Feishu connection
python scripts/feishu_client.py --action test

# Verify IMA connection
python scripts/ima_client.py --action test
```

## IMA 智能路由(v3.0)

| source_url 类型 | IMA 存放方式 | 原因 |
|----------------|------------|------|
| 公众号 / 普通网页 | `import_urls`(服务端抓取) | 保留图片和排版 |
| X/Twitter | 纯文本笔记(逐字转录) | 技能规则要求逐字转录 |
| 飞书 wiki | 纯文本笔记(逐字转录) | 需登录认证,IMA 无法抓取 |

## 版本历史

- **v3.7.0**:飞书 wiki 抓取改造——从 WebFetch(截断严重,200-1000 字符)改为 `lark-cli docs +fetch`(完整内容,3K-25K 字符);实测 33 篇文章全部完整获取
- **v3.6.0**:飞书存储改造——从应用身份创建在线 docx 改为 `lark-cli drive +upload` 上传 .md 原文件到用户云盘,立即可见
- **

_meta.json

{
  "ownerId": "kn75zj7vzdyvap84adxa8heyyd82f5eh",
  "slug": "web-to-fim",
  "version": "3.7.0",
  "publishedAt": 1784387155458
}

references/feishu-blocks.md

# 飞书云文档 Block 类型映射表

> 飞书 Docx API 的 block_type 枚举值。创建文档 blocks 时必须使用正确的 block_type,否则返回 400 Bad Request。

## block_type 完整映射

| block_type | 名称 | Markdown 对应 | block 字段名 |
|-----------|------|-------------|------------|
| 1 | Page(页面根块) | — | `page` |
| 2 | Text(文本) | 普通段落 | `text` |
| 3 | Heading1(一级标题) | `# ` | `heading1` |
| 4 | Heading2(二级标题) | `## ` | `heading2` |
| 5 | Heading3(三级标题) | `### ` | `heading3` |
| 6 | Heading4(四级标题) | `#### ` | `heading4` |
| 7 | Heading5(五级标题) | `##### ` | `heading5` |
| 8 | Heading6(六级标题) | `###### ` | `heading6` |
| 9 | Heading7(七级标题) | — | `heading7` |
| 10 | Heading8(八级标题) | — | `heading8` |
| 11 | Heading9(九级标题) | — | `heading9` |
| 12 | Bullet(无序列表) | `- ` 或 `* ` | `bullet` |
| 13 | Ordered(有序列表) | `1. ` | `ordered` |
| 14 | Code(代码块) | ` ``` ` | `code` |
| 15 | Quote(引用) | `> ` | `quote` |

## 常见错误

| 错误 | 正确值 | 说明 |
|------|-------|------|
| Code 用 block_type=2 | **block_type=14** | 2 是 Text,不是 Code |
| Quote 用 block_type=11 | **block_type=15** | 11 是 Heading9,不是 Quote |

## Block 结构示例

### Text(block_type=2)

```json
{
    "block_type": 2,
    "text": {
        "elements": [{"text_run": {"content": "文本内容"}}],
        "style": {}
    }
}
```

### Heading1(block_type=3)

```json
{
    "block_type": 3,
    "heading1": {
        "elements": [{"text_run": {"content": "标题内容"}}],
        "style": {}
    }
}
```

### Bullet(block_type=12)

```json
{
    "block_type": 12,
    "bullet": {
        "elements": [{"text_run": {"content": "列表项"}}],
        "style": {}
    }
}
```

### Code(block_type=14)

```json
{
    "block_type": 14,
    "code": {
        "elements": [{"text_run": {"content": "code content"}}],
        "style": {"language": 1}
    }
}
```

### Quote(block_type=15)

```json
{
    "block_type": 15,
    "quote": {
        "elements": [{"text_run": {"content": "引用内容"}}],
        "style": {}
    }
}
```

## 关键约束

### 1. elements 中不要加多余字段

```json
// ❌ 错误 — 多了 "type" 字段
{"text_run": {"type": "text_run", "content": "文本"}}

// ✅ 正确
{"text_run": {"content": "文本"}}
```

### 2. 空行不要创建空文本块

```python
# ❌ 错误 — 空行创建空块会导致 400
if not stripped:
    blocks.append({"block_type": 2, "text": {"elements": [{"text_run": {"content": ""}}]}})

# ✅ 正确 — 空行直接跳过
if not stripped:
    continue
```

### 3. 图片块需要预上传 token

图片不能直接用 URL,必须先通过飞书上传 API 获取 `file_token`,再创建图片块。当前实现暂时跳过图片。

### 4. 单次请求最多 50 个 blocks

```
POST /docx/v1/documents/{document_id}/blocks/{document_id}/children
```

**payload 中 `children` 数组最多 50 个 block**。长文章必须分批插入:

```python
for i in range(0, len(blocks), 50):
    batch = blocks[i:i+50]
    requests.post(url, headers=headers, json={"children": batch})
```

### 5. payload 中不要传 index

```json
// ❌ 错误 — index=-1 可能导致问题
{"children": blocks, "index": -1}

// ✅ 正确 — 不传 index,默认追加到末尾
{"children": blocks}
```

## API 端点

| 操作 | 方法 | URL |
|------|------|-----|
| 创建文档 | POST | `/docx/v1/documents` |
| 插入 blocks | POST | `/docx/v1/documents/{document_id}/blocks/{document_id}/children` |
| 获取文档信息 | GET | `/docx/v1/d

references/feishu-setup.md

# 飞书云文档 API 配置指南

## 获取凭证

### 步骤 1: 创建飞书应用

1. 访问 [飞书开放平台](https://open.feishu.cn/)
2. 登录后进入「开发者后台」
3. 点击「创建应用」,填写应用名称和描述
4. 获取 **App ID** 和 **App Secret**

### 步骤 2: 配置权限

在应用后台开启以下权限:

| 权限名称 | 权限说明 | 必要性 |
|---------|---------|--------|
| `docx.document:readonly` | 读取云文档 | 可选 |
| `docx.document:write` | 创建/编辑云文档 | 必须 |
| `drive:drive:folder` | 云文档文件夹操作 | 必须 |
| `sheets:Spreadsheet:readonly` | 读取表格 | 可选 |

### 步骤 3: 发布应用

1. 在「版本管理与发布」中创建版本
2. 选择「发布」并等待审核(或自建企业直接通过)

## 本地配置

### 设置环境变量

```bash
# Windows PowerShell
$env:FEISHU_APP_ID="your_app_id"
$env:FEISHU_APP_SECRET="your_app_secret"

# Linux/macOS
export FEISHU_APP_ID="your_app_id"
export FEISHU_APP_SECRET="your_app_secret"
```

### 验证配置

```python
from feishu_client import FeishuClient

client = FeishuClient()
if client.test_connection():
    print("✅ 飞书连接成功")
else:
    print("❌ 飞书连接失败,请检查凭证")
```

## 安全注意事项

⚠️ **重要提示**:
- **永远不要**将 `.env` 文件或包含真实凭证的文件提交到 Git
- 已在 `.gitignore` 中添加 `.env` 规则
- 生产环境建议使用密钥管理服务(如 AWS Secrets Manager)

## 故障排查

| 问题 | 解决方案 |
|------|---------|
| `invalid_app_id` | 检查 App ID 是否正确 |
| `app_secret` 错误 | 在飞书开放平台重置 App Secret |
| 权限不足 | 在应用后台添加所需权限并重新发布 |
| 签名校验失败 | 确保使用 HTTPS,检查时间戳是否正确 |
Github ReposUpdated 10h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/edwardwason/skills/web-to-fim",
      "sourceUrl": "https://clawhub.ai/edwardwason/skills/web-to-fim",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T20:54:54.953Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-edwardwason-web-to-fim/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edwardwason-web-to-fim/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T20:54:54.953Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2K downloads",
      "href": "https://clawhub.ai/edwardwason/web-to-fim",
      "sourceUrl": "https://clawhub.ai/edwardwason/web-to-fim",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T20:54:54.953Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "3.7.0",
      "href": "https://clawhub.ai/edwardwason/web-to-fim",
      "sourceUrl": "https://clawhub.ai/edwardwason/web-to-fim",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-07-18T15:05:55.458Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-edwardwason-web-to-fim/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edwardwason-web-to-fim/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 3.7.0",
      "description": "v3.7.0: Feishu wiki fetch via lark-cli docs +fetch (WebFetch truncation fix). Breaking change: Feishu wiki fetching migrated from WebFetch (severe truncation 200-1000 chars) to lark-cli docs +fetch (full content 3K-25K chars). Verified 33 articles all fetched completely.",
      "href": "https://clawhub.ai/edwardwason/web-to-fim",
      "sourceUrl": "https://clawhub.ai/edwardwason/web-to-fim",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-07-18T15:05:55.458Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to Web To Fim and adjacent AI workflows.