Dataify X Builder
Collect X Builder data and return results
Rank
62
Safety
84
Downloads
1.1k
Updated
Oct 11, 2026
Version
1.3.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.3.1release · observed Sep 8, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17feed8b2qc486skqmjapxmjs86bd4f:dataify-twitter-profile-by-profileurl- Install using `clawhub skill install s17feed8b2qc486skqmjapxmjs86bd4f:dataify-twitter-profile-by-profileurl` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/dataify-server/dataify-twitter-profile-by-profileurl before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-dataify-server-dataify-twitter-profile-by-profileurl/snapshot"
Documentation
CLAWHUB
66,708 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: "dataify-twitter-profile-by-profileurl"
description: "Collect an X/Twitter profile from a known profile URL. Do not use for posts, keyword search, or arbitrary X URLs."
---
# Dataify Builder Skill
Use this skill to prepare Dataify builder requests for the scraper family rooted at `twitter_profile_by-profileurl` on `x.com`.
## Quick Start
**Input:** an X/Twitter profile URL.
```bash
python3 scripts/build-dataify-request.py --tool-sign twitter_profile_by-profileurl --params-json '[{"profileurl":"https://x.com/OpenAI"}]'
```
This submits the task, waits for completion, downloads the final result, and returns it. Add `--no-wait` only when submission-only behavior is requested.
## Workflow
1. Check whether `DATAIFY_API_TOKEN` exists in the environment.
2. If the token is missing, stop and tell the user: `Dataify requires an API token. New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Once registration is complete, tell me and I'll continue the current task.`.
3. Ask the user to choose exactly one tool from the following Chinese list:
- 通过个人资料 URL采集 (twitter_profile_by-profileurl)
- 通过Twitter 用户名采集 (twitter_profile_by-username)
- 通过个人资料URL采集 (twitter_post_by-profileurl)
4. Read `references/tool-params.json` and find the chosen tool by `tool_sign` or Chinese tool name.
5. For each parameter in the chosen tool:
- If `input_mode` is `user_input`, ask the user for the value.
- If `input_mode` is `select`, present the saved options to the user.
6. Use `scripts/build-dataify-request.py` as the default cross-platform helper.
7. Use `scripts/build-dataify-request.ps1` as the Windows PowerShell helper when needed.
8. When a selectable parameter has a human-readable Chinese label, keep that label in `spider_parameters`. Do not replace it with a code such as `HK` unless the user explicitly asks for the coded value.
9. Build `spider_parameters` as a JSON array.
10. If every parameter has only one final value, build one object such as `[{"searchurl":"...","country":"Hong Kong"}]`.
11. If one or more parameters have multiple aligned values, zip them by index and build one object per row. Example: `[{"search_url":"url1","page_turning":"1","max_num":"15"},{"search_url":"url2","page_turning":"1","max_num":"15"}]`.
12. If a parameter has one value while another parameter has multiple values, reuse the single value across every generated row.
13. Set `spider_name` to `x.com`.
14. Set `spider_id` to the selected tool's `tool_sign`.
15. Always include `spider_errors=true` and `file_name={{TasksID}}`.
16. Return a curl command for `https://scraperapi.dataify.com/builder`.
## Set DATAIFY_API_TOKEN
Prefer a permanent environment-variable setup instead of setting the token only for the current terminal session.
Windows PowerShell, permanent for the current user:
```powershell
[Environment]::SetEnvironmentVariable("DATAIFY_API_TOKEN", "your_token_here", "User")
```
Then _meta.json
{
"ownerId": "kn74z5hmmwk21kw8tpphd9w21x86bkdf",
"slug": "dataify-twitter-profile-by-profileurl",
"version": "1.3.1",
"publishedAt": 1788848344315
}references/tool-params.json
[{"tool_name_cn":"个人资料URL","tool_sign":"twitter_profile_by-profileurl","spider_name":"x.com","params":[]},{"tool_name_cn":"用户名","tool_sign":"twitter_profile_by-username","spider_name":"x.com","params":[]},{"tool_name_cn":"帖子资料URL","tool_sign":"twitter_post_by-profileurl","spider_name":"x.com","params":[]}]skill-card.md
## Description: Collects an X/Twitter profile from a known profile URL via Dataify; it is not intended for posts, keyword search, or arbitrary X URLs. This skill is ready for commercial/non-commercial use. ## Publisher: [dataify-server](https://clawhub.ai/user/dataify-server) ### License/Terms of Use: MIT-0 ## Use Case: External users and developers use this skill to prepare and run Dataify Builder requests for X/Twitter profile collection, then wait for the asynchronous task and return the collected JSON result. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The package exposes broader X collection modes than the profile-URL-only skill name suggests. Mitigation: Use only the intended twitter_profile_by-profileurl workflow unless a reviewer explicitly approves another bundled Dataify collection mode for the task. Risk: The skill requires a Dataify API token and includes shell-profile setup examples that could make credential exposure persistent. Mitigation: Use a limited Dataify token, prefer session-scoped or managed secret storage where possible, never paste the token into chat or logs, and rotate it if exposed. Risk: Asynchronous task polling and broader collection settings can increase credit usage or encourage duplicate paid submissions after a timeout. Mitigation: Confirm high-volume, multi-page, or media-download scopes before execution, retain task IDs, and resume polling existing tasks instead of resubmitting them. ## Reference(s): - [Tool parameter catalog](artifact/references/tool-params.json) - [ClawHub skill page](https://clawhub.ai/dataify-server/skills/dataify-twitter-profile-by-profileurl) ## Skill Output: **Output Type(s):** [text, markdown, shell commands, configuration, API calls] **Output Format:** [Markdown with shell command examples and JSON task or result payloads] **Output Parameters:** [1D] **Other Properties Related to Output:** [By default the skill waits for task completion and returns the final collected JSON result; no-wait mode returns a submitted task_id.] ## Skill Version(s): 1.3.1 (source: server release evidence) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
SKILL.zh-CN.md
---
name: "dataify-twitter-profile-by-profileurl"
description: "为 x.com 上以 twitter_profile_by-profileurl 为根的 scraper 系列准备 Dataify builder 请求。当需要处理成功的 Dataify scraper detail 条目 twitter_profile_by-profileurl、让用户选择可用工具、读取已保存的 getToolParams 选项,并使用 DATAIFY_API_TOKEN 生成 scraperapi.dataify.com/builder curl 请求时,使用此 skill。"
---
# Dataify Builder Skill 中文版
这个 skill 用于为 `x.com` 下、以 `twitter_profile_by-profileurl` 为入口的 Dataify scraper 工具族生成 builder 请求。
## 工作流程
1. 先检查环境变量中是否存在 `DATAIFY_API_TOKEN`。
2. 如果 token 缺失,告诉用户:`Dataify 需要 API Token。新账号注册即得 50 免费积分,约可获得 6000 条试用结果,7 天有效,仅成功请求计费。注册完成后告诉我,我会继续当前任务。`。
3. 先让用户从下面的中文工具列表中明确选择一个工具:
- 通过个人资料 URL采集 (twitter_profile_by-profileurl)
- 通过Twitter 用户名采集 (twitter_profile_by-username)
- 通过个人资料URL采集 (twitter_post_by-profileurl)
4. 再读取 `references/tool-params.json`,根据 `tool_sign` 或中文工具名找到对应工具。
5. 对所选工具的每个参数分别处理:
- 如果 `input_mode` 是 `user_input`,让用户提供值。
- 如果 `input_mode` 是 `select`,把已保存的可选项展示给用户,让用户选择。
6. 默认优先使用 `scripts/build-dataify-request.py`,因为它是跨平台版本。
7. Windows 下也可以使用 `scripts/build-dataify-request.ps1`。
8. `spider_parameters` 必须是一个 JSON 数组。
9. `spider_name` 固定取 `x.com`。
10. `spider_id` 固定取用户所选工具的 `tool_sign`。
11. 始终包含 `spider_errors=true` 和 `file_name={{TasksID}}`。
## 设置 DATAIFY_API_TOKEN
推荐使用永久环境变量,而不是只在当前终端临时设置。
Windows PowerShell,当前用户永久设置:
```powershell
[Environment]::SetEnvironmentVariable("DATAIFY_API_TOKEN", "your_token_here", "User")
```
然后重新打开 PowerShell。如果当前会话也要立即生效,再执行:
```powershell
$env:DATAIFY_API_TOKEN = "your_token_here"
```
macOS 或 Linux,bash 永久设置:
```bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.bashrc
source ~/.bashrc
```
macOS 或 Linux,zsh 永久设置:
```bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.zshrc
source ~/.zshrc
```
## 脚本用法
Python:
```bash
python scripts/build-dataify-request.py --tool-sign <selected_tool_sign> --values-file values.json
```
PowerShell:
```powershell
& ".\scripts\build-dataify-request.ps1" -ToolSign "<selected_tool_sign>" -ValuesFile ".\values.json"
```
`values.json` 可以是单个对象,也可以是对象数组。
## 输出格式
最终 `curl` 命令应为:
```bash
curl -X POST 'https://scraperapi.dataify.com/builder' \
-H "Authorization: Bearer $DATAIFY_API_TOKEN" \
-H 'Content-Type: application/x-www-form-urlencoded' \
-d 'spider_name=x.com' \
-d 'spider_id=<selected_tool_sign>' \
-d 'spider_parameters=[{"param":"value"}]' \
-d 'spider_errors=true' \
-d 'file_name={{TasksID}}'
```
## 参考文件
- `references/tool-params.json` 保存了这个 skill 下所有工具及参数选项。
- `scripts/build-dataify-request.py` 是首选的跨平台实现。
- `scripts/build-dataify-request.ps1` 是 Windows PowerShell 版本。
- 如果参数没有预设选项,必须向用户要值。
- 不要假设 `spider_parameters` 永远只有一个对象;多值工具可能需要按索引生成多个对象。
- `url_example` 仅作为参考,不要默认用户就要用示例值,除非用户明确确认。
## 参数交互策略
- 当请求意图明确、只读、低风险且成本较低时,使用安全默认值直接执行。可以用一句话说明执行内容,但不要暂停等待确认。
- 只在缺少必填输入、存在会明显改变结果的歧义、大批量或多页采集、媒体下载、会明显增加积分消耗、不可逆操作,或用户明确要求查看参数时询问。
- 必须确认时,只展示会影响目标、范围、输出或成本的用户参数。优先使用一句简短说明;只有三个及以上关键值确实需要比较时才使用精简表格。
- 不要展示固定字段、空的可选字段、未修改的默认值、凭据或内部实现参数,例如引擎选择、响应格式开关、偏移量、spider ID 和文件名模板。
- 默AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/dataify-server/skills/dataify-twitter-profile-by-profileurl",
"sourceUrl": "https://clawhub.ai/dataify-server/skills/dataify-twitter-profile-by-profileurl",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T12:05:45.056Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-dataify-server-dataify-twitter-profile-by-profileurl/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-dataify-server-dataify-twitter-profile-by-profileurl/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T12:05:45.056Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.1K downloads",
"href": "https://clawhub.ai/dataify-server/dataify-twitter-profile-by-profileurl",
"sourceUrl": "https://clawhub.ai/dataify-server/dataify-twitter-profile-by-profileurl",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T12:05:45.056Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.3.1",
"href": "https://clawhub.ai/dataify-server/dataify-twitter-profile-by-profileurl",
"sourceUrl": "https://clawhub.ai/dataify-server/dataify-twitter-profile-by-profileurl",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-08T06:19:04.315Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-dataify-server-dataify-twitter-profile-by-profileurl/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-dataify-server-dataify-twitter-profile-by-profileurl/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.3.1",
"description": "Fix natural-language usage failures: validate required targets and URLs, preserve catalog references, default Amazon region safely, normalize Google News links, and improve UTF-8 error output.",
"href": "https://clawhub.ai/dataify-server/dataify-twitter-profile-by-profileurl",
"sourceUrl": "https://clawhub.ai/dataify-server/dataify-twitter-profile-by-profileurl",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-08T06:19:04.315Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
