AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
Crawler Summary
Automated daily AI research engine powered by CrewAI & serverless Nvidia T4 GPUs. <div align="center"> <img src="https://capsule-render.vercel.app/api?type=waving&height=240&color=0:0f0c29,50:302b63,100:20BEFF&text=SIFT%20AI&fontColor=ffffff&fontSize=78&fontAlignY=38&desc=The%20best%20article%20of%20the%20day%2C%20picked%20by%20AI%20%C2%B7%20Zero%20servers%20%C2%B7%20Zero%20API%20bills&descAlignY=60&descSize=17&animation=fadeIn" alt="Sift AI" width="100%" /> <a href="https://github.com/bytebymanas Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.
Freshness
Last checked 10/9/2026
Best For
Sift-AI is best for crewai, multi-agent workflows where OpenClaw compatibility matters.
Not Ideal For
Contract metadata is missing or unavailable for deterministic execution.
Evidence Sources Checked
editorial-content, GITHUB REPOS, runtime-metrics, public facts pack
Automated daily AI research engine powered by CrewAI & serverless Nvidia T4 GPUs. <div align="center"> <img src="https://capsule-render.vercel.app/api?type=waving&height=240&color=0:0f0c29,50:302b63,100:20BEFF&text=SIFT%20AI&fontColor=ffffff&fontSize=78&fontAlignY=38&desc=The%20best%20article%20of%20the%20day%2C%20picked%20by%20AI%20%C2%B7%20Zero%20servers%20%C2%B7%20Zero%20API%20bills&descAlignY=60&descSize=17&animation=fadeIn" alt="Sift AI" width="100%" /> <a href="https://github.com/bytebymanas
Public facts
4
Change events
1
Artifacts
0
Freshness
Oct 9, 2026
Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.
Trust score
Unknown
Compatibility
OpenClaw
Freshness
Oct 9, 2026
Vendor
Bytebymanas
Artifacts
0
Benchmarks
0
Last release
Unpublished
Key links, install path, and a quick operational read before the deeper crawl record.
Summary
Capability contract not published. No trust telemetry is available yet. Last updated 10/9/2026.
Setup snapshot
Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.
Vendor
Bytebymanas
Protocol compatibility
OpenClaw
Handshake status
UNKNOWN
Crawlable docs
6 indexed pages on the official domain
Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.
Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.
Extracted files
0
Examples
6
Snippets
0
Languages
python
bash
git clone https://github.com/bytebymanas/sift-ai.git cd sift-ai
json
{
"id": "YOUR_KAGGLE_USERNAME/sift-ai-daily-runner",
"title": "Sift AI Daily Runner",
"code_file": "sift-ai.ipynb",
"language": "python",
"kernel_type": "notebook",
"is_private": true,
"enable_gpu": true,
"enable_internet": true
}bash
export KAGGLE_USERNAME="your_username" export KAGGLE_KEY="your_api_key" kaggle kernels push -p .
python
USER_PREFERRED_TOPICS = [
"latest technology inventions",
"politics",
"artificial intelligence",
"business",
"technology",
"economy",
"world news",
"AI companies",
]python
TRUSTED_SOURCES = {
"reuters", "associated press", "ap", "bbc", "the guardian", "bloomberg",
"the new york times", "the washington post", "npr", "the wall street journal",
"the hindu", "the indian express", "hindustan times", "livemint",
"business standard", "press trust of india", "pti",
}python
EDITORIAL_FEEDS = {
"The Hindu": "https://www.thehindu.com/opinion/editorial/feeder/default.rss",
"The Hindu (Op-Ed)": "https://www.thehindu.com/opinion/op-ed/feeder/default.rss",
"Indian Express (Editorials)": "https://indianexpress.com/section/opinion/editorials/feed/",
"Indian Express (Columns)": "https://indianexpress.com/section/opinion/columns/feed/",
"Hindustan Times (Editorials)": "https://www.hindustantimes.com/feeds/rss/editorials/rssfeed.xml",
}Full documentation captured from public sources, including the complete README when available.
Docs source
GITHUB REPOS
Editorial quality
ready
Automated daily AI research engine powered by CrewAI & serverless Nvidia T4 GPUs. <div align="center"> <img src="https://capsule-render.vercel.app/api?type=waving&height=240&color=0:0f0c29,50:302b63,100:20BEFF&text=SIFT%20AI&fontColor=ffffff&fontSize=78&fontAlignY=38&desc=The%20best%20article%20of%20the%20day%2C%20picked%20by%20AI%20%C2%B7%20Zero%20servers%20%C2%B7%20Zero%20API%20bills&descAlignY=60&descSize=17&animation=fadeIn" alt="Sift AI" width="100%" /> <a href="https://github.com/bytebymanas
How it works · The agents · Setup · Customize · FAQ
</div> <br/>Most news is noise. Sift AI reads the day's candidates and hands you exactly one article worth your time, complete with key takeaways and a word of the day to grow your vocabulary.
Every morning a GitHub Actions workflow launches a Kaggle notebook. Inside it, a small crew of AI agents running open-source Qwen models collects fresh stories, picks the strongest one, summarizes it, chooses a useful English word, and emails everything to you as a clean PDF.
No server stays on. No LLM API is billed. Nobody has to press a button.
<table> <tr> <td align="center" width="25%"><h3>06:30</h3><sub>Fires daily on schedule</sub></td> <td align="center" width="25%"><h3>1 Article</h3><sub>Chosen from ~20 candidates</sub></td> <td align="center" width="25%"><h3>4 Agents</h3><sub>Select · Summarize · Vocab</sub></td> <td align="center" width="25%"><h3>$0</h3><sub>Inference cost</sub></td> </tr> </table> <br/>| Step | What happens |
| :---: | :--- |
| 1 | GitHub Actions fires on the cron schedule (or a repository_dispatch event) and pushes the notebook to Kaggle with the CLI |
| 2 | Kaggle starts a T4 GPU kernel. The notebook installs Ollama and pulls the two Qwen models |
| 3 | It queries NewsData for each of your topics (globally and India-scoped) and reads editorial RSS feeds, then scrapes and cleans every candidate |
| 4 | The Senior News Editor agent picks the one article that best rewards a reader's time |
| 5 | Two lighter agents extract takeaways and choose a word of the day |
| 6 | A fixed HTML template is rendered to PDF with WeasyPrint and sent through Gmail |
Two models, four agents. The bigger model does the one job that needs real judgment. The smaller one handles tasks on a single, already-clean input.
| Agent | Model | Job | | :--- | :--- | :--- | | Senior News Editor | Qwen 2.5 · 7B Instruct | Compares all candidates and picks exactly one, or declares that none qualify | | Content Summarizer | Qwen 3 · 1.7B | Extracts 3 to 5 specific takeaways, never inventing a claim | | Vocabulary Word Selector | Qwen 3 · 1.7B | Picks one common, everyday English word, ignoring the article's topic | | English Vocabulary Coach | Qwen 3 · 1.7B | Explains that word with a meaning, example and practice scenario. It never sees the article |
<details> <summary><b>How the editor chooses</b></summary> <br/>The editor scores candidates in this order of priority:
Hard filters remove listicles, clickbait, press releases and vague "experts say" sourcing. Opinion and news are judged equally, with no built-in favorite.
</details> <details> <summary><b>What happens on a bad news day</b></summary> <br/>The pipeline never forces a weak pick and never fails silently.
| Attempt | Minimum length | Behavior | | :---: | :---: | :--- | | 1 | 220 words | Strict bar. The editor may reject everything | | 2 | 150 words | Looser bar, topic match no longer a reason to reject | | 3 | 120 words | Last chance. Picks the best acceptable candidate | | Fallback | n/a | No LLM call. Picks the best known-reliable, longest candidate |
The article is scraped once. Retries only re-run the cheap selection step.
</details> <br/>git clone https://github.com/bytebymanas/sift-ai.git
cd sift-ai
Edit kernel-metadata.json and replace the username:
{
"id": "YOUR_KAGGLE_USERNAME/sift-ai-daily-runner",
"title": "Sift AI Daily Runner",
"code_file": "sift-ai.ipynb",
"language": "python",
"kernel_type": "notebook",
"is_private": true,
"enable_gpu": true,
"enable_internet": true
}
Secrets live in Kaggle, never in this repo. Push the notebook once (step 4) so it exists in your account, then open it on Kaggle and go to Add-ons → Secrets. Add these four, with the names matching exactly, and make sure each one is attached to the notebook:
| Secret name | What to put in it |
| :--- | :--- |
| NEWSDATA_API_KEY | Your API key from newsdata.io |
| GMAIL_ADDRESS | The Gmail address the brief is sent from |
| GMAIL_APP_PASSWORD | A Google app password for that sender account (not your normal password) |
| RECIPIENT_EMAIL | The address that receives the brief. It can be the same as the sender |
GMAIL_APP_PASSWORD secretInternet access must be on for the notebook (the NewsData API, RSS feeds, article scraping and email all need it). The
"enable_internet": trueline inkernel-metadata.jsonhandles this when you push through the CLI.
In your repository go to Settings → Secrets and variables → Actions and add KAGGLE_USERNAME and KAGGLE_KEY. Get the key from your Kaggle account settings by creating a new API token. The workflow uses these to push the notebook.
export KAGGLE_USERNAME="your_username"
export KAGGLE_KEY="your_api_key"
kaggle kernels push -p .
From here the scheduled workflow takes over, and a fresh brief arrives every morning.
<br/>Open sift-ai.ipynb and edit the configuration cell near the top. Commit and push, and the next scheduled run uses your changes.
Topics you care about (USER_PREFERRED_TOPICS). These are a preference, not a filter. A great article outside your topics still beats a mediocre one inside them.
USER_PREFERRED_TOPICS = [
"latest technology inventions",
"politics",
"artificial intelligence",
"business",
"technology",
"economy",
"world news",
"AI companies",
]
Trusted sources (TRUSTED_SOURCES). Publications listed here get a credibility boost when the editor compares candidates. This is a soft signal, not a hard filter, so niche topics never end up with an empty pool. Names must be lowercase and match the publication name NewsData reports.
TRUSTED_SOURCES = {
"reuters", "associated press", "ap", "bbc", "the guardian", "bloomberg",
"the new york times", "the washington post", "npr", "the wall street journal",
"the hindu", "the indian express", "hindustan times", "livemint",
"business standard", "press trust of india", "pti",
}
Editorial and opinion feeds (EDITORIAL_FEEDS). These RSS feeds are read directly, because news APIs mostly index wire news rather than opinion sections. Add any publication that offers an RSS feed.
EDITORIAL_FEEDS = {
"The Hindu": "https://www.thehindu.com/opinion/editorial/feeder/default.rss",
"The Hindu (Op-Ed)": "https://www.thehindu.com/opinion/op-ed/feeder/default.rss",
"Indian Express (Editorials)": "https://indianexpress.com/section/opinion/editorials/feed/",
"Indian Express (Columns)": "https://indianexpress.com/section/opinion/columns/feed/",
"Hindustan Times (Editorials)": "https://www.hindustantimes.com/feeds/rss/editorials/rssfeed.xml",
}
<br/>
| Layer | Tooling |
| :--- | :--- |
| Orchestration | GitHub Actions (cron, repository_dispatch) |
| Compute | Kaggle GPU kernels · Nvidia T4 · Python 3.10 |
| LLM inference | Ollama · Qwen 2.5 7B Instruct · Qwen 3 1.7B |
| Agents | CrewAI |
| Sources | NewsData API · RSS via feedparser |
| Scraping | requests · BeautifulSoup |
| Output | HTML template · WeasyPrint (PDF) · Gmail SMTP |
The first version drove Google Colab through Playwright. It worked until the DOM changed. Sift AI now talks to Kaggle directly through its official CLI.
| | Legacy (Playwright + Colab) | Sift AI (Kaggle API) | | :--- | :--- | :--- | | Execution | Simulated browser clicks | Direct CLI dispatch | | Authentication | Session cookies, UI logins | Scoped API key | | Reliability | Breaks on DOM changes and captchas | Deterministic cloud execution | | Overhead | Browser boot, DOM parsing | Lightweight API payload |
</details> <details> <summary><b>Models judge, code handles the text</b></summary> <br/>Scraping and cleaning are deterministic, so the article reaches you exactly as published. Models only do what they are good at: comparing candidates, extracting takeaways and teaching a word.
</details> <details> <summary><b>Open-weight models, not paid APIs</b></summary> <br/>Qwen served through Ollama means no per-token billing and no vendor lock-in.
</details> <br/>sift-ai/
├── .github/
│ └── workflows/
│ └── daily_colab_trigger.yml # Cron + dispatch automation
├── assets/ # README graphics
├── sift-ai.ipynb # Full pipeline: collect, select, summarize, PDF, email
├── kernel-metadata.json # Kaggle GPU execution spec
└── README.md
<br/>
Built by Manas
<sub>Sift through the noise. Keep the signal.</sub>
<img src="https://capsule-render.vercel.app/api?type=waving&height=120&color=0:20BEFF,50:302b63,100:0f0c29§ion=footer" alt="" width="100%" /> </div>Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.
Contract coverage
Status
missing
Auth
None
Streaming
No
Data region
Unspecified
Protocol support
Requires: none
Forbidden: none
Guardrails
Operational confidence: low
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/contract"
curl -s "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/trust"
Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.
Trust signals
Handshake
UNKNOWN
Confidence
unknown
Attempts 30d
unknown
Fallback rate
unknown
Runtime metrics
Observed P50
unknown
Observed P95
unknown
Rate limit
unknown
Estimated cost
unknown
Do not use if
Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.
Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
The Frontend for Agents & Generative UI. React + Angular
Contract JSON
{
"contractStatus": "missing",
"authModes": [],
"requires": [],
"forbidden": [],
"supportsMcp": false,
"supportsA2a": false,
"supportsStreaming": false,
"inputSchemaRef": null,
"outputSchemaRef": null,
"dataRegion": null,
"contractUpdatedAt": null,
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Invocation Guide
{
"preferredApi": {
"snapshotUrl": "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/snapshot",
"contractUrl": "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/contract",
"trustUrl": "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/trust"
},
"curlExamples": [
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/snapshot\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/contract\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/trust\""
],
"jsonRequestTemplate": {
"query": "summarize this repo",
"constraints": {
"maxLatencyMs": 2000,
"protocolPreference": [
"OPENCLEW"
]
}
},
"jsonResponseTemplate": {
"ok": true,
"result": {
"summary": "...",
"confidence": 0.9
},
"meta": {
"source": "GITHUB_REPOS",
"generatedAt": "2026-10-09T23:05:02.672Z"
}
},
"retryPolicy": {
"maxAttempts": 3,
"backoffMs": [
500,
1500,
3500
],
"retryableConditions": [
"HTTP_429",
"HTTP_503",
"NETWORK_TIMEOUT"
]
}
}Trust JSON
{
"status": "unavailable",
"handshakeStatus": "UNKNOWN",
"verificationFreshnessHours": null,
"reputationScore": null,
"p95LatencyMs": null,
"successRate30d": null,
"fallbackRate": null,
"attempts30d": null,
"trustUpdatedAt": null,
"trustConfidence": "unknown",
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Capability Matrix
{
"rows": [
{
"key": "OPENCLEW",
"type": "protocol",
"support": "unknown",
"confidenceSource": "profile",
"notes": "Listed on profile"
},
{
"key": "crewai",
"type": "capability",
"support": "supported",
"confidenceSource": "profile",
"notes": "Declared in agent profile metadata"
},
{
"key": "multi-agent",
"type": "capability",
"support": "supported",
"confidenceSource": "profile",
"notes": "Declared in agent profile metadata"
}
],
"flattenedTokens": "protocol:OPENCLEW|unknown|profile capability:crewai|supported|profile capability:multi-agent|supported|profile"
}Facts JSON
[
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Bytebymanas",
"href": "https://github.com/bytebymanas/Sift-AI",
"sourceUrl": "https://github.com/bytebymanas/Sift-AI",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T01:11:42.805Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T01:11:42.805Z",
"isPublic": true
},
{
"factKey": "docs_crawl",
"category": "integration",
"label": "Crawlable docs",
"value": "6 indexed pages on the official domain",
"href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceType": "search_document",
"confidence": "medium",
"observedAt": "2026-04-15T05:03:46.393Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/crewai-bytebymanas-sift-ai/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
]Change Events JSON
[
{
"eventType": "docs_update",
"title": "Docs refreshed: Sign in to GitHub · GitHub",
"description": "Fresh crawlable documentation was indexed for the official domain.",
"href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
"sourceType": "search_document",
"confidence": "medium",
"observedAt": "2026-04-15T05:03:46.393Z",
"isPublic": true
}
]Sponsored
Ads related to Sift-AI and adjacent AI workflows.