prompt-archaeology
Excavate forgotten solutions, code snippets, and decisions from past conversation sessions. Use when the user is re-solving a problem you've likely solved before, hunting for a lost snippet, or wants to mine session history for buried knowledge instead of starting from scratch. Skill: prompt-archaeology Owner: voronindenis5 Summary: Excavate forgotten solutions, code snippets, and decisions from past conversation sessions. Use when the user is re-solving a problem you've likely solved before, hunting for a lost snippet, or wants to mine session history for buried knowledge instead of starting from scratch. Tags: latest:0.1.1 Version history: v0.1.1 | 2026-08-11T11:58:40.748Z | auto Version
Rank
62
Safety
84
Downloads
2.5k
Updated
Oct 9, 2026
Version
0.1.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 2.5K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 2.5K downloadsadoption · observed Oct 9, 2026
- Latest release
- 0.1.1release · observed Aug 11, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17b6amkd3wzqgg640v03a9r1n83gxs1:prompt-archaeology- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-prompt-archaeology/snapshot"
Documentation
CLAWHUB
95,624 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: prompt-archaeology
description: "Excavate forgotten solutions, code snippets, and decisions from past conversation sessions. Use when the user is re-solving a problem you've likely solved before, hunting for a lost snippet, or wants to mine session history for buried knowledge instead of starting from scratch."
version: 1.0.0
author: Denis Voronin
license: MIT
metadata:
hermes:
tags: [sessions, history, search, knowledge-mining, archaeology, recovery, hermes-agent]
related_skills: []
---
# Prompt Archaeology
## Overview
**Prompt Archaeology** is the practice of excavating your own conversation history instead of re-solving problems from scratch. Every AI session is a stratum — a sedimented layer of debugging, decision-making, and discovery. Over time, valuable artifacts sink below the surface: a one-liner that fixed a gnarly race condition, a config that satisfied a finicky build, the exact incantation that convinced a model to behave. Most agents never dig for these. They re-derive, re-guess, and re-fail.
This skill turns that history into a quarryable resource. It bundles:
- **Search strategies** — keyword, semantic-adjacent, temporal, and structural queries tuned for session transcripts.
- **Relevance scoring** — a transparent, composable ranking that surfaces the one session that actually matters.
- **Knowledge extraction patterns** — recipes for pulling *decisions* and *solutions* out of a wall of chat, not just matching text.
- **Deduplication** — collapse near-duplicate fixes across sessions into a single canonical answer.
- **`excavate.py`** — a standalone Python script that crawls session logs and markdown files, ranks them, and prints the buried artifacts.
The metaphor is deliberate. An archaeologist does not grep the desert for "pottery" and ship the first hit. They survey, triangulate, carefully extract, and catalog. This skill teaches the agent to do the same with its own past.
## When to Use
- **The user is about to re-solve a known problem.** They describe a bug or task and you have a flicker of "we've done this before." Excavate before answering.
- **"Didn't we figure out...?" / "What did we land on?"** — retrieve the prior decision and its rationale, not just the outcome.
- **Hunting for a lost code snippet, config value, or command** that worked months ago.
- **Onboarding to a codebase you've touched before** — pull the architectural decisions out of old sessions.
- **Avoiding repeated dead ends** — find the approaches that were *rejected* and why, so you don't walk back into them.
- **Writing postmortems or ADRs** from scattered session evidence.
### Don't use for
- Fresh problems with no prior history — there's nothing to excavate; solve forward.
- When you already hold the answer in active context — don't pad the turn with a search.
- Sensitive retrieval across other users' private profiles unless explicitly authorized.
## The Excavation Workflow
A dig has five phases. Skipping any phase degradREADME.md
# Prompt Archaeology
> An AI agent skill for excavating forgotten solutions, code snippets, and decisions from past conversation sessions — instead of re-solving problems from scratch.
[](LICENSE)
[](https://www.python.org/downloads/)
Every AI session is a stratum — a sedimented layer of debugging, decision-making, and discovery. Over time, valuable artifacts sink below the surface: a one-liner that fixed a gnarly race condition, a config that satisfied a finicky build, the exact incantation that convinced a model to behave. Most agents never dig for these. **Prompt Archaeology** turns that history into a quarryable resource.
## What it gives you
- **Search strategies** — keyword, semantic-adjacent, temporal, and structural queries tuned for session transcripts.
- **Relevance scoring** — a transparent, composable ranking that surfaces the one session that actually matters (not just the one that mentions the term most).
- **Knowledge extraction patterns** — recipes for pulling *decisions* and *solutions* out of a wall of chat, not just matching text.
- **Deduplication** — collapse near-duplicate fixes across sessions into a single canonical answer.
- **`excavate.py`** — a standalone Python script that crawls session logs and markdown files, ranks them, and prints the buried artifacts. **Zero third-party dependencies** (stdlib only).
## The workflow in one breath
**Survey** what artifact you want → **Locate** it with multiple query passes → **Score** the finds with the composite ranker → **Extract** the minimal artifact (with citation) → **Deduplicate** before reporting.
## Quick start
```bash
git clone https://github.com/voronindenis5/prompt-archaeology.git
cd prompt-archaeology
# Basic keyword dig over a directory of session logs (.md/.txt/.json/.jsonl)
python3 scripts/excavate.py dig ./my-sessions --query "kafka consumer rebalance"
# Multiple terms, top 5, with per-signal score breakdown
python3 scripts/excavate.py dig ./my-sessions --query "rebalance retry backoff" --top 5 --explain
# Date window + dedup near-identical results
python3 scripts/excavate.py dig ./my-sessions \
--query "connection pool exhaustion" \
--after 2024-01-01 --before 2024-06-01 --dedup
# Dump just the extracted code blocks across matches
python3 scripts/excavate.py dig ./my-sessions --query "ffmpeg concatenate" --extract code
```
### Index once, query many
For large corpora, build an index and query it repeatedly:
```bash
python3 scripts/excavate.py index ./my-sessions --out sessions.idx
python3 scripts/excavate.py query sessions.idx --query "oauth refresh token" --top 3 --explain
```
### Programmatic use
```python
from excavate import ArchaeologyIndex
idx = ArchaeologyIndex()
idx.scan("./my-sessions")
for hit in idx.search("kafka rebalance", top=5, explain=True):
print(hit.score, hit.path)
print(hit.extraction_meta.json
{
"ownerId": "kn75wwn4x6djaf28jbykeamazd81gtdp",
"slug": "prompt-archaeology",
"version": "0.1.1",
"publishedAt": 1786449520748
}references/cli-reference.md
# `excavate.py` CLI Reference `excavate.py` is a stdlib-only Python script for excavating session logs and markdown files. It supports three subcommands: `dig`, `index`, and `query`. ## Global behavior - **No third-party dependencies.** Runs on Python 3.8+ with the standard library only. - **Reads** `.md`, `.txt`, `.json`, and `.jsonl` files (see [File formats](#file-formats)). - **Writes** nothing unless `--out` is given (index subcommand). - **Exits non-zero** on argument errors; zero on successful search (even with no hits). ## Subcommands ### `dig` — search a directory directly ```bash python3 scripts/excavate.py dig <directory> --query <terms> [options] ``` Scans `<directory>` recursively, scores every file against the query, prints the top matches. Use `dig` for one-off searches; use `index` + `query` for repeated searches over the same corpus. **Options:** | Flag | Type | Default | Description | |---|---|---|---| | `--query`, `-q` | str (required) | — | Search terms. Multiple terms are AND'd within a file (all must appear). Use `--query "a b c"` or repeat `--query` for OR semantics is not supported; pass a single string. | | `--top`, `-n` | int | 5 | Number of top results to print. | | `--after` | date (ISO) | — | Only include files modified on/after this date (`YYYY-MM-DD`). | | `--before` | date (ISO) | — | Only include files modified on/before this date (`YYYY-MM-DD`). | | `--not` | str | — | Exclude files containing this term. Repeatable. | | `--explain` | flag | off | Print per-signal score breakdown for each result. | | `--extract` | `code` \| `all` \| `none` | `none` | Print extracted code blocks (`code`), full extraction (`all`), or just scores (`none`). | | `--dedup` | flag | off | Collapse near-duplicate results into clusters. | | `--dedup-threshold` | float | 0.85 | Jaccard threshold for near-dup (0–1). | | `--keep-variants` | flag | off | With `--dedup`, report all version/path variants instead of collapsing. | **Examples:** ```bash # Basic dig python3 scripts/excavate.py dig ./sessions --query "kafka rebalance" # Top 10 with score breakdown python3 scripts/excavate.py dig ./sessions --query "pool exhaustion" --top 10 --explain # Date window, exclude sidekiq mentions, dedup python3 scripts/excavate.py dig ./sessions \ --query "redis cache" \ --after 2024-01-01 --before 2024-06-30 \ --not sidekiq \ --dedup # Extract code blocks only python3 scripts/excavate.py dig ./sessions --query "ffmpeg concat" --extract code ``` ### `index` — build a reusable index ```bash python3 scripts/excavate.py index <directory> --out <index-file> [options] ``` Scans `<directory>` once and serializes the `ArchaeologyIndex` to `<index-file>` (pickle format). Subsequent `query` calls load the index instead of re-scanning. **Options:** | Flag | Type | Default | Description | |---|---|---|---| | `--out`, `-o` | path (required) | — | Output index file path. | | `--after` / `--before` | date | — | Pre-filter by mtime at index time
references/deduplication.md
# Deduplication
The same fix often appears in three sessions: the first attempt, the retry, and the "oh and also" follow-up. Deduplication collapses these into a single canonical answer before you report.
`excavate.py --dedup` runs three layers of dedup in order: exact → near-dup → variant.
## Layer 1: Exact dedup
**Rule:** identical fenced code blocks collapse to one.
Two sessions with byte-identical ` ```bash ... ``` ` blocks are the same artifact. Keep the higher-scoring session as the canonical instance; drop the others from the result set but note them as duplicates.
```
sessions/fix-v1.md score 0.71 [canonical]
≡ sessions/fix-v1-retry.md (exact dup of fix-v1.md)
≡ sessions/fix-v1-followup.md (exact dup of fix-v1.md)
```
Exact dedup is cheap (a hash) and catches the easy cases.
## Layer 2: Near-dup detection
**Rule:** normalized text similarity > 0.85 merges into a cluster.
Two sessions that describe the same fix in slightly different words are the same artifact. Near-dup uses **normalized token overlap**:
1. Lowercase, strip punctuation, remove stopwords.
2. Tokenize on whitespace.
3. Compute Jaccard similarity: `|A ∩ B| / |A ∪ B|`.
4. If Jaccard > 0.85, merge.
Keep the highest-scoring session in the cluster as canonical.
```
sessions/2024-03-12-a.md score 0.82 [canonical]
≈ sessions/2024-03-13-b.md (near-dup, Jaccard 0.91)
≈ sessions/2024-03-14-c.md (near-dup, Jaccard 0.88)
```
### Tuning the threshold
- **0.85 (default):** conservative. Only clearly-the-same fixes merge.
- **0.75:** aggressive. Catches more dups but risks merging distinct fixes that share boilerplate (e.g., two different nginx fixes with similar config structure).
- **0.95:** paranoid. Almost only exact dups merge.
Set via `--dedup-threshold` on the CLI, or `DEDUP_THRESHOLD` in code.
### What gets normalized away
- Case (`Error` ≡ `error`)
- Punctuation (`failed.` ≡ `failed`)
- Stopwords (`the connection pool` ≡ `connection pool`)
- Whitespace runs
### What does NOT get normalized away
- Numbers and identifiers (`pool_size=20` ≢ `pool_size=50`)
- Code structure (two fenced blocks with different commands stay distinct)
- File paths
## Layer 3: Variant detection
**Rule:** if two snippets differ only in a version number, path, or date, treat as the same artifact and note the latest variant.
```
sessions/2023-09-01.md: pip install fastapi==0.68.0
sessions/2024-03-12.md: pip install fastapi==0.110.0
```
These are the same artifact ("install fastapi") with a version variant. Collapse to one, note the latest (0.110.0).
Variant detection works by:
1. Masking version-like tokens (`\d+\.\d+\.\d++`), date-like tokens, and path-like tokens.
2. Re-running near-dup (Layer 2) on the masked text.
3. If they merge, they're variants.
### When variants matter
Sometimes the variant *is* the artifact — e.g., "the exact version that worked with Python 3.9." Variant detection has a `--keep-variants` fAionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/voronindenis5/skills/prompt-archaeology",
"sourceUrl": "https://clawhub.ai/voronindenis5/skills/prompt-archaeology",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T13:43:44.391Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-prompt-archaeology/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-prompt-archaeology/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T13:43:44.391Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "2.5K downloads",
"href": "https://clawhub.ai/voronindenis5/prompt-archaeology",
"sourceUrl": "https://clawhub.ai/voronindenis5/prompt-archaeology",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T13:43:44.391Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.1.1",
"href": "https://clawhub.ai/voronindenis5/prompt-archaeology",
"sourceUrl": "https://clawhub.ai/voronindenis5/prompt-archaeology",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-11T11:58:40.748Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-prompt-archaeology/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-prompt-archaeology/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.1.1",
"description": "Version 0.1.1 - Removed the sample file: skill-card.md - No other functional or documentation changes detected.",
"href": "https://clawhub.ai/voronindenis5/prompt-archaeology",
"sourceUrl": "https://clawhub.ai/voronindenis5/prompt-archaeology",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-11T11:58:40.748Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
