{"id":"7d992139-125c-439a-a822-ef3602d6c30d","entityType":"agent","slug":"clawhub-globalcaos-memory-bench-pioneer","name":"TinkerClaw Memory Bench","canonicalUrl":"https://www.xpersona.co/agent/clawhub-globalcaos-memory-bench-pioneer","canonicalPath":"/agent/clawhub-globalcaos-memory-bench-pioneer","generatedAt":"2026-10-11T08:41:46.320Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:14:48.633Z","emptyReason":null},"description":"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine, and excerpts are redacted even on the local path. The optional OpenAI judge is opt-in, prints the exact request body it would send, redacts secrets first, requires typed consent, and cannot be switched on by an unattended run. Submitting results is a separate confirmed step that validates the report against the full schema and previews every field in it, and identifies you only if you pass --contributor. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw. See Permissions, Data Flow & Consent.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17324vfeqe0ptzp84z2z9vttx883zdg:memory-bench-pioneer","sourceUrl":"https://clawhub.ai/globalcaos/memory-bench-pioneer","homepage":"https://clawhub.ai/globalcaos/skills/memory-bench-pioneer","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/globalcaos/memory-bench-pioneer","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/globalcaos/skills/memory-bench-pioneer","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":61,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"TinkerClaw Memory Bench technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:14:48.633Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:14:48.633Z","emptyReason":null},"stars":null,"forks":null,"downloads":1136,"packageName":null,"latestVersion":"2.1.3","tractionLabel":"1.1K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:14:48.568Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T06:14:48.633Z","lastCrawledAt":"2026-10-11T06:14:48.568Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T06:14:48.568Z","lastVerifiedAt":null,"highlights":[{"version":"2.1.3","createdAt":"2026-09-09T09:04:13.102Z","changelog":"Exact request-body preview before OpenAI consent; schema-validated public report; no unattended bypass.","fileCount":10,"zipByteSize":39517},{"version":"2.1.2","createdAt":"2026-09-07T10:15:25.757Z","changelog":"Fixes two real defects the audit found, not just wording. submit.sh interpolated the report path directly into python3 -c source in five places, so a crafted report filename could execute arbitrary Python; every site now passes the path as sys.argv. collect.py probed the Node version through os.popen (a shell); replaced with shutil.which plus a non-shell subprocess.","fileCount":8,"zipByteSize":28360},{"version":"2.1.1","createdAt":"2026-09-07T10:02:17.474Z","changelog":"Summary now matches the code: benchmarking runs locally by default and no memory content leaves the machine; the OpenAI judge is opt-in with a printed data-flow notice, secret redaction, typed consent and no unattended use; submission is a separate confirmed step and is anonymous unless --contributor is passed.","fileCount":8,"zipByteSize":28069},{"version":"2.1.0","createdAt":"2026-09-07T09:41:05.870Z","changelog":"Security and disclosure pass. Adds a Permissions, Data Flow & Consent section declaring every capability, what data is touched, where it is written and what leaves the machine. Documents the off-switch and requires explicit opt-in for privacy-affecting or destructive actions. Removes documentation claims the shipped code did not implement. No functionality removed.","fileCount":8,"zipByteSize":28156},{"version":"2.0.1","createdAt":"2026-06-06T16:51:22.217Z","changelog":"TinkerClaw rebrand + funnel to github.com/globalcaos/tinkerclaw","fileCount":7,"zipByteSize":19267},{"version":"2.0.0","createdAt":"2026-02-16T22:57:09.924Z","changelog":"A-grade research methodology: LLM-as-judge (GPT-4o-mini), 30-query standard test set (4 types × 3 difficulties), nDCG/MAP/MRR/RAR with 95% bootstrap CIs, ablation (with/without spreading activation), Cohen's kappa inter-rater reliability, longitudinal tracking, algorithm versioning, 11 unit tests. Anonymized — your memories never leave your machine.","fileCount":7,"zipByteSize":19226}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17324vfeqe0ptzp84z2z9vttx883zdg:memory-bench-pioneer","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17324vfeqe0ptzp84z2z9vttx883zdg:memory-bench-pioneer` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/globalcaos/memory-bench-pioneer before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T08:41:46.314Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:14:48.633Z","emptyReason":null},"readme":"Skill: TinkerClaw Memory Bench\n\nOwner: globalcaos\n\nSummary: Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine, and excerpts are redacted even on the local path. The optional OpenAI judge is opt-in, prints the exact request body it would send, redacts secrets first, requires typed consent, and cannot be switched on by an unattended run. Submitting results is a separate confirmed step that validates the report against the full schema and previews every field in it, and identifies you only if you pass --contributor. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw. See Permissions, Data Flow & Consent.\n\nTags: latest:2.1.3\n\nVersion history:\n\nv2.1.3 | 2026-09-09T09:04:13.102Z | user\n\nExact request-body preview before OpenAI consent; schema-validated public report; no unattended bypass.\n\nv2.1.2 | 2026-09-07T10:15:25.757Z | user\n\nFixes two real defects the audit found, not just wording. submit.sh interpolated the report path directly into python3 -c source in five places, so a crafted report filename could execute arbitrary Python; every site now passes the path as sys.argv. collect.py probed the Node version through os.popen (a shell); replaced with shutil.which plus a non-shell subprocess.\n\nv2.1.1 | 2026-09-07T10:02:17.474Z | user\n\nSummary now matches the code: benchmarking runs locally by default and no memory content leaves the machine; the OpenAI judge is opt-in with a printed data-flow notice, secret redaction, typed consent and no unattended use; submission is a separate confirmed step and is anonymous unless --contributor is passed.\n\nv2.1.0 | 2026-09-07T09:41:05.870Z | user\n\nSecurity and disclosure pass. Adds a Permissions, Data Flow & Consent section declaring every capability, what data is touched, where it is written and what leaves the machine. Documents the off-switch and requires explicit opt-in for privacy-affecting or destructive actions. Removes documentation claims the shipped code did not implement. No functionality removed.\n\nv2.0.1 | 2026-06-06T16:51:22.217Z | user\n\nTinkerClaw rebrand + funnel to github.com/globalcaos/tinkerclaw\n\nv2.0.0 | 2026-02-16T22:57:09.924Z | user\n\nA-grade research methodology: LLM-as-judge (GPT-4o-mini), 30-query standard test set (4 types × 3 difficulties), nDCG/MAP/MRR/RAR with 95% bootstrap CIs, ablation (with/without spreading activation), Cohen's kappa inter-rater reliability, longitudinal tracking, algorithm versioning, 11 unit tests. Anonymized — your memories never leave your machine.\n\nArchive index:\n\nArchive v2.1.3: 10 files, 39517 bytes\n\nFiles: _meta.json (139b), scripts/collect.py (17523b), scripts/rate.py (27466b), scripts/report_schema.py (11130b), scripts/submit.sh (9337b), scripts/test_metrics.py (4009b), scripts/test_privacy.py (15761b), scripts/testset.json (6202b), skill-card.md (2509b), SKILL.md (16013b)\n\nFile v2.1.3:SKILL.md\n\n---\nname: memory-bench-pioneer\nversion: 2.1.3\ndescription: \"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine, and excerpts are redacted even on the local path. The optional OpenAI judge is opt-in, prints the exact request body it would send, redacts secrets first, requires typed consent, and cannot be switched on by an unattended run. Submitting results is a separate confirmed step that validates the report against the full schema and previews every field in it, and identifies you only if you pass --contributor. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw. See Permissions, Data Flow & Consent.\"\nmetadata:\n  {\n    \"openclaw\":\n      {\n        \"emoji\": \"🧠\",\n        \"requires\": { \"bins\": [\"python3\"] },\n        \"notes\":\n          {\n            \"security\": \"Benchmarks a LIVE memory database, so disclosure matters more than usual. DEFAULT PATH IS LOCAL: the default judge (--judge local) sends retrieved excerpts only to a local embedding server on 127.0.0.1; nothing leaves the machine, and those excerpts are redacted too because that server keeps its own logs. Two actions can send data out and BOTH are opt-in, confirmed, and IMPOSSIBLE to trigger unattended: --judge openai transmits the benchmark query plus up to 300 redacted characters of each RETRIEVED MEMORY to api.openai.com, printing the exact JSON request body first and then requiring a typed confirmation at a terminal; and scripts/submit.sh opens a PUBLIC GitHub pull request containing the statistics report, after validating it against the complete schema and printing every field, the commit message and the PR body (--dry-run to preview, typed confirmation to publish). --yes-send-to-openai and --yes skip only the keystroke; both still require a terminal, so an agent running this on its own cannot transmit or publish. Attribution is anonymous unless you pass --contributor; token/cost totals are excluded unless you pass --include-token-stats. rate.py WRITES a retrieval_log table into the database you point it at — use --db on a copy to sandbox it. No daemon, no cron, nothing runs on its own. See the Permissions, Data Flow & Consent section.\"\n          },\n      },\n  }\n---\n\n# Memory Bench\n\n> One of dozens of skills and plugins in **[TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — a self-improving OpenClaw fork that's been running 24/7 for months.\n\nEveryone has an opinion about whether their agent's memory is any good. Almost nobody has a number.\n\nThis produces the number — nDCG, MAP, MRR, Precision@5, each with a 95% bootstrap confidence interval, plus an ablation that isolates what spreading activation actually contributes. Then, if you want, it contributes your (anonymous) results to the ENGRAM and CORTEX research papers, where a few dozen real deployments beat any amount of arguing.\n\n**Part of [TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — real-time token tracking, self-improving crons, persistent cognitive memory. This is one piece of that stack; the repo has dozens more.\n\n👉 **https://github.com/globalcaos/tinkerclaw**\n\n_Clone it. Fork it. Break it. Make it yours._\n\n## Three-Step Pipeline\n\nStep 1 measures, step 2 summarises, step 3 publishes. Steps 1 and 2 are local. **Step 3 is the only one that uploads anything, and it asks first.**\n\n### 1. Assess Retrieval Quality\n\nRun the standard test set (30 queries across 4 types × 3 difficulty levels):\n\n```bash\n# Local judge — the default. Nothing leaves your machine.\npython3 scripts/rate.py --queries 30 --judge local --ablation\n\n# Stronger judge, but it TRANSMITS retrieved memory excerpts to OpenAI.\n# You will be asked to type 'send' before the first request.\npython3 scripts/rate.py --queries 30 --judge openai --ablation\n\n# Benchmark a copy instead of your live database\npython3 scripts/rate.py --db /tmp/memory-copy.db --judge local\n\n# Custom test set\npython3 scripts/rate.py --testset path/to/queries.json --judge local\n```\n\n**What it measures:**\n\n- **RAR** (Recall Accuracy Ratio), **MRR** (Mean Reciprocal Rank)\n- **nDCG@5**, **MAP@5**, **Precision@5**, **Hit Rate**\n- All metrics include **95% bootstrap confidence intervals**\n- **Ablation**: runs with AND without spreading activation to isolate its contribution\n\n**Judge methods — this is the privacy decision in this skill:**\n\n| Judge | What it sees | Where it goes | Cost |\n| --- | --- | --- | --- |\n| `local` **(default)** | query + 300 **redacted** chars of each result | `http://127.0.0.1:8900/embed` on your own machine | free |\n| `openai` | query + 300 **redacted** chars of each result | `api.openai.com`, gpt-4o-mini | ~$0.01/run |\n\nThe `openai` judge is more discriminating and independent of the retrieval system, which is why the research protocol prefers it. It is also the option that puts pieces of your memories in someone else's logs. Both are legitimate; pick deliberately, and see the consent section below for exactly what is sent.\n\n**It writes to your database.** `rate.py` creates (or extends) a `retrieval_log` table in the database you point it at and inserts one row per benchmark query: the benchmark query text, the ratings, the metrics. Your memory content is never written there. `collect.py` reads that table later. If you would rather not touch your live DB, run both scripts with `--db` against a copy.\n\n**Standard test set** (`scripts/testset.json`): 30 queries stratified across semantic/episodic/procedural/strategic types and easy/medium/hard difficulty. As of 2.1.0 the queries are phrased to probe the operational and technical side of your memory — configuration, procedures, incidents, architecture decisions — rather than family, health, contacts or calendar. All deployments run the same queries, so results are comparable across sites; reports produced with the pre-2.1.0 set are not directly comparable with these.\n\nNote the honest limit: the queries steer *what is asked*, not *what your memory system returns*. Retrieval searches whatever database you point it at. If your memory holds things you would not want an external judge to rate, use `--judge local` (the default), or benchmark a filtered copy with `--db`.\n\n### 2. Collect Statistics\n\n```bash\n# Anonymous — the default\npython3 scripts/collect.py --days 14 --output /tmp/memory-bench-report.json\n\n# Attributed to you (your username goes in the report, and it becomes public if you submit)\npython3 scripts/collect.py --days 14 --contributor YOUR_GITHUB_USER --output /tmp/memory-bench-report.json\n```\n\n**What goes in the report:** memory counts, type/age distributions, strength and importance histograms, association graph size, hierarchy levels, consolidation run counts, embedding coverage, retrieval metrics from `retrieval_log` (RAR/MRR/nDCG/MAP, judge method, ablation config), the algorithm version as a short git SHA, and coarse system info (OS, CPU architecture, Python version, Node version). Instance ID is a random UUID.\n\n**What never goes in it:** memory content, benchmark queries, file paths, hostnames, environment variables.\n\n**Opt-in extras, off unless you ask:**\n\n- `--contributor NAME` — your username. Without it the report says `anonymous`.\n- `--include-token-stats` — total tokens and USD spend from your OpenClaw usage files.\n\n`collect.py` uploads nothing. It writes a JSON file and tells you what is in it. Read that file before step 3.\n\n### 3. Submit as PR — the step that publishes\n\n```bash\n# ALWAYS do this first: shows every field that would become public, uploads nothing\nscripts/submit.sh /tmp/memory-bench-report.json --dry-run\n\n# Real submission — prints the same preview, then asks you to type 'publish'\nscripts/submit.sh /tmp/memory-bench-report.json YOUR_GITHUB_USERNAME\n```\n\nThis forks `globalcaos/clawdbot-moltbot-openclaw`, pushes a branch to **your** fork, and opens a **public pull request** containing the report file. A merged PR is public and permanent. Requires the `gh` CLI, authenticated.\n\nBefore any of that, the report is validated against the complete schema and every field in it is printed, along with the commit message and the PR body that would be created. An unknown field anywhere — a key a future collector adds, or one you edited in — is a hard refusal rather than a note, because an unknown field is exactly the case nobody has reviewed.\n\nIt refuses to run unattended: with no terminal to confirm on, it exits rather than publishing. `--yes` skips only the typed word — it still requires a terminal, so it is not a way for an agent to publish on your behalf.\n\n## Permissions, Data Flow & Consent\n\nShort version: steps 1 and 2 are local; step 3 publishes, and so does the optional OpenAI judge. Both ask first. Longer version, because you should not have to take that on trust:\n\n**What it needs, and why.**\n\n| Capability | Why | Scope |\n| --- | --- | --- |\n| Read your memory DB | Counts, histograms, and running the benchmark queries | `~/.openclaw/workspace/db/{memory,cognitive_memory,jarvis}.db`, or `--db` |\n| **Write** your memory DB | `rate.py` logs one `retrieval_log` row per benchmark query | Same DB (or the copy you pass to `--db`); no other table is touched |\n| File write | `<db_dir>/.memory-bench-instance-id` (random UUID, so repeat reports group together) and the `--output` report path | Two files, both of which you can delete |\n| Read usage files | Token/cost totals — **only** with `--include-token-stats` | `~/.openclaw/workspace/memory/*-usage.json` |\n| Local shell exec | `git log` for the algorithm SHA, `node --version` for system info, `git`/`gh` in `submit.sh` | Fixed commands |\n| Network to localhost | Local judge embeddings | `http://127.0.0.1:8900/embed` — stays on the machine |\n| Network to OpenAI | **Opt-in.** Query + up to 300 redacted chars per retrieved memory | `api.openai.com`, only with `--judge openai` after confirmation |\n| Network to GitHub | **Opt-in.** Uploads the report and opens a public PR | `github.com`, only in `submit.sh` after confirmation |\n| Credentials | `OPENAI_API_KEY` (env or `--api-key`) only when you choose the OpenAI judge; `gh`'s existing login in `submit.sh` | Read at call time, never stored, never written to the report |\n| Scheduling | **None.** No daemon, no cron, no install hook. Nothing runs unless you run it | — |\n\n**About that redaction.** Every excerpt — external judge, local judge, and the lines printed to your terminal — goes through one function that truncates to 300 characters and replaces email addresses, phone-shaped numbers, API-key-shaped strings, long hex blobs, home directory paths and URL credentials with placeholders. Localhost is not the same thing as this process: the embedding server keeps its own logs, and terminal output ends up in agent transcripts, so neither gets raw memory content either.\n\nThat is pattern matching, not comprehension — it will not catch a secret written in prose. Treat it as a seatbelt, not a force field. If the content is genuinely sensitive, the answer is `--judge local`, not better regexes.\n\nOne consequence worth knowing: since 2.1.3 the local judge embeds the redacted excerpt rather than the raw one, so local-judge scores can shift slightly against a pre-2.1.3 run if your memories contain redactable patterns. It also means both judges now rate identical text, which is what makes the Cohen's κ between them meaningful.\n\n**And you see the request itself.** With `--judge openai`, the script runs one retrieval locally, builds the actual first request from it, and prints that JSON body — endpoint, model, prompt, redacted excerpt and all — before asking. The same function builds the preview and the outgoing request, so what you approve is what is sent.\n\n**Consent, concretely.** Two actions can leave your machine, and neither happens by accident:\n\n```bash\n# 1. External judging — prints the exact request body, then waits\npython3 scripts/rate.py --judge openai        # asks you to type 'send'\npython3 scripts/rate.py --judge openai --yes-send-to-openai   # skips the keystroke, still needs a terminal\n\n# 2. Publishing — validates, prints every field that becomes public, then waits\nscripts/submit.sh report.json --dry-run       # preview only, uploads nothing\nscripts/submit.sh report.json                 # asks you to type 'publish'\n```\n\nBoth refuse outright when there is no terminal to confirm on. `--yes-send-to-openai` and `--yes` remove the keystroke, never the terminal, so an agent running this unattended can neither transmit your memories nor publish on your behalf.\n\n**Turning it off, and undoing it:**\n\n```bash\n# never transmit: just use the default judge and skip step 3\npython3 scripts/rate.py --judge local\n\n# benchmark a throwaway copy instead of your live memory\ncp ~/.openclaw/workspace/db/memory.db /tmp/bench.db\npython3 scripts/rate.py --db /tmp/bench.db --judge local\n\n# forget this installation ever ran\nrm ~/.openclaw/workspace/db/.memory-bench-instance-id\nsqlite3 ~/.openclaw/workspace/db/memory.db 'DROP TABLE retrieval_log;'\n```\n\nDeleting the instance ID file makes your next report a new anonymous instance; it also breaks the longitudinal link, which is the trade.\n\n**Read it before you run it.** Four scripts, all plain text, none of them long. `rate.py` is the only one that can talk to a third party and `submit.sh` is the only one that can publish — both are worth the two minutes. `report_schema.py` is the list of what is publishable, and you can run it against any report on its own.\n\n## Validation Protocol\n\nFor peer-review-ready data, contributors should:\n\n1. Run `rate.py --ablation` over the full N=30 test set\n2. Use `--judge openai` if you are comfortable with the data flow above — it agrees better with human raters, and the script reports Cohen's κ between the two judges so you can see the gap on your own data. Local-judge submissions are still welcome and are marked as such in the report\n3. Collect at least 2 reports from the same instance, ≥7 days apart (longitudinal)\n4. Report the algorithm version (auto-captured as a short git SHA)\n\n## Test Set Format\n\nCustom test sets are JSON arrays:\n\n```json\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"...\",\n    \"category\": \"semantic|episodic|procedural|strategic\",\n    \"difficulty\": \"easy|medium|hard\"\n  }\n]\n```\n\nAn optional `notes` field is ignored by the runner.\n\n## Included Files\n\n| File | Purpose |\n| --- | --- |\n| `scripts/rate.py` | Runs the benchmark, judges results, computes metrics. Writes `retrieval_log`. The only script that can call an external API — opt-in and confirmed |\n| `scripts/collect.py` | Builds the anonymous statistics report. Uploads nothing |\n| `scripts/submit.sh` | Opens the public PR. Preview with `--dry-run`; requires typed confirmation |\n| `scripts/report_schema.py` | The publishable report schema. Validates and prints every field; run it on any report by itself |\n| `scripts/testset.json` | The 30-query standard test set |\n| `scripts/test_metrics.py` | Unit tests for the IR metrics (`python3 scripts/test_metrics.py`) |\n| `scripts/test_privacy.py` | Tests for the disclosure gates: redaction, consent, schema (`python3 scripts/test_privacy.py`) |\n\nEverything the documentation above describes is in this package. If you find a claim here that the code does not do, that is a bug — open an issue on [the repo](https://github.com/globalcaos/tinkerclaw/issues).\n\n## Agent Workflow\n\nWhen asked to benchmark memory: run `rate.py --ablation` (local judge) and then `collect.py`, and show the summary. **Do not run `submit.sh` on your own initiative** — it publishes to a public repository. Show the user the report, and use `submit.sh --dry-run` so they can see exactly what would become public. Submit only when they say to, and only with `--judge openai` if they have agreed to that separately. Then share the PR link.\n\nFile v2.1.3:_meta.json\n\n{\n  \"ownerId\": \"kn7623hrcwt6rg73a67xw3wyx580asdw\",\n  \"slug\": \"memory-bench-pioneer\",\n  \"version\": \"2.1.3\",\n  \"publishedAt\": 1788944653102\n}\n\nFile v2.1.3:scripts/testset.json\n\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"What preferences are recorded for the development environment setup?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests basic semantic recall of preference-type memories\"\n  },\n  {\n    \"id\": \"T02\",\n    \"query\": \"What happened during the last consolidation cycle?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests temporal/episodic retrieval\"\n  },\n  {\n    \"id\": \"T03\",\n    \"query\": \"How should I handle sensitive data when sharing between agents?\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests procedural knowledge retrieval\"\n  },\n  {\n    \"id\": \"T04\",\n    \"query\": \"recurring patterns in recorded project planning failures\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract/strategic retrieval; phrased to avoid keyword matching\"\n  },\n  {\n    \"id\": \"T05\",\n    \"query\": \"What tools were configured for audio processing?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests factual tool/config recall\"\n  },\n  {\n    \"id\": \"T06\",\n    \"query\": \"what happened, and in what order, during the most recent incident\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests ordered event reconstruction without keyword match\"\n  },\n  {\n    \"id\": \"T07\",\n    \"query\": \"steps to deploy a new version safely\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests multi-step procedural recall\"\n  },\n  {\n    \"id\": \"T08\",\n    \"query\": \"Which external services or integrations are configured?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity-type memory retrieval over configured services\"\n  },\n  {\n    \"id\": \"T09\",\n    \"query\": \"lessons learned from debugging difficult issues\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract lesson retrieval across multiple episodes\"\n  },\n  {\n    \"id\": \"T10\",\n    \"query\": \"What was decided about rate limits or usage quotas?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests topical episodic recall\"\n  },\n  {\n    \"id\": \"T11\",\n    \"query\": \"privacy rules for handling user data\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests policy/procedural recall\"\n  },\n  {\n    \"id\": \"T12\",\n    \"query\": \"connections between automation projects and hardware\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests multi-hop associative retrieval\"\n  },\n  {\n    \"id\": \"T13\",\n    \"query\": \"which scheduled or recurring jobs have run recently\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recurring-job pattern retrieval\"\n  },\n  {\n    \"id\": \"T14\",\n    \"query\": \"error handling best practices\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests general procedural knowledge\"\n  },\n  {\n    \"id\": \"T15\",\n    \"query\": \"which design proposals were recorded for the retrieval pipeline\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent episodic recall with abstract query\"\n  },\n  {\n    \"id\": \"T16\",\n    \"query\": \"tradeoffs between local and cloud solutions\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests comparative/analytical retrieval\"\n  },\n  {\n    \"id\": \"T17\",\n    \"query\": \"which components make up the system and what each one does\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests component/architecture relationship recall\"\n  },\n  {\n    \"id\": \"T18\",\n    \"query\": \"What failed and why in previous attempts?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests failure/negative experience retrieval\"\n  },\n  {\n    \"id\": \"T19\",\n    \"query\": \"configuration steps for new integrations\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests setup/config procedural recall\"\n  },\n  {\n    \"id\": \"T20\",\n    \"query\": \"long-term technical objectives recorded for the memory system\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests goal-tracking retrieval across time\"\n  },\n  {\n    \"id\": \"T21\",\n    \"query\": \"formatting and output conventions that are preferred\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests convention/preference recall\"\n  },\n  {\n    \"id\": \"T22\",\n    \"query\": \"What did the last test or verification run report as failing?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests failure-report episodic recall\"\n  },\n  {\n    \"id\": \"T23\",\n    \"query\": \"how to resolve merge conflicts in the codebase\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests specific technical procedural recall\"\n  },\n  {\n    \"id\": \"T24\",\n    \"query\": \"evolution of the memory architecture over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal strategic retrieval\"\n  },\n  {\n    \"id\": \"T25\",\n    \"query\": \"the most recent benchmark runs and when they happened\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests temporal recall of recent runs\"\n  },\n  {\n    \"id\": \"T26\",\n    \"query\": \"Which file formats or codecs are documented as supported?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests documented-capability recall\"\n  },\n  {\n    \"id\": \"T27\",\n    \"query\": \"backup and recovery procedures\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests disaster recovery procedural recall\"\n  },\n  {\n    \"id\": \"T28\",\n    \"query\": \"how tradeoffs between retrieval accuracy and latency were weighed over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal tradeoff-reasoning retrieval\"\n  },\n  {\n    \"id\": \"T29\",\n    \"query\": \"What was the outcome of the last research task?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent task outcome retrieval\"\n  },\n  {\n    \"id\": \"T30\",\n    \"query\": \"communication style preferences for different channels\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests nuanced preference retrieval\"\n  }\n]\n\nFile v2.1.3:skill-card.md\n\n## Description:\n\nTinkerClaw Memory Bench benchmarks an OpenClaw or TinkerClaw memory database locally by default, producing retrieval-quality metrics with confidence intervals and optional, consent-gated OpenAI judging and public report submission.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[globalcaos](https://clawhub.ai/user/globalcaos)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and agent operators use this skill to measure retrieval quality for a live or copied OpenClaw or TinkerClaw memory database and collect aggregate benchmark reports. It supports local default runs, optional external LLM judging, and optional public submission after preview and confirmation.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Benchmarking can touch sensitive local memory data from a live memory database.\n\nMitigation: Use the default local judge and run against a copied or filtered database when contents are sensitive.\n\nRisk: The optional OpenAI judge can send redacted benchmark queries and retrieved memory excerpts to OpenAI.\n\nMitigation: Use it only after reviewing the printed request preview and confirming that redacted excerpts may leave the machine; prefer OPENAI_API_KEY over command-line keys.\n\nRisk: Submitting results opens a public GitHub pull request with aggregate report metadata.\n\nMitigation: Run submit.sh with --dry-run first, review every field in the schema-validated preview, and publish only after explicit confirmation.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/globalcaos/skills/memory-bench-pioneer)\n- [Publisher profile](https://clawhub.ai/user/globalcaos)\n- [TinkerClaw project link from skill documentation](https://github.com/globalcaos/tinkerclaw)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with shell commands and JSON report files]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Produces local benchmark metrics and aggregate reports; external judging and public submission are optional and confirmation-gated.]\n\n## Skill Version(s):\n\n2.1.3 (source: frontmatter and server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v2.1.2: 8 files, 28360 bytes\n\nFiles: _meta.json (139b), scripts/collect.py (17523b), scripts/rate.py (24005b), scripts/submit.sh (9408b), scripts/test_metrics.py (4009b), scripts/testset.json (6202b), skill-card.md (2732b), SKILL.md (13585b)\n\nFile v2.1.2:SKILL.md\n\n---\nname: memory-bench-pioneer\nversion: 2.1.2\ndescription: \"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine. The optional OpenAI judge is opt-in, prints exactly what it would send, redacts secrets first, requires typed consent, and refuses to run unattended. Submitting results is a separate confirmed step that previews every public field, and identifies you only if you pass --contributor. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw. See Permissions, Data Flow & Consent.\"\nmetadata:\n  {\n    \"openclaw\":\n      {\n        \"emoji\": \"🧠\",\n        \"requires\": { \"bins\": [\"python3\"] },\n        \"notes\":\n          {\n            \"security\": \"Benchmarks a LIVE memory database, so disclosure matters more than usual. DEFAULT PATH IS LOCAL: the default judge (--judge local) sends retrieved excerpts only to a local embedding server on 127.0.0.1; nothing leaves the machine. Two actions can send data out and BOTH are opt-in and confirmed: --judge openai transmits the benchmark query plus up to 300 redacted characters of each RETRIEVED MEMORY to api.openai.com (typed confirmation, or --yes-send-to-openai), and scripts/submit.sh opens a PUBLIC GitHub pull request containing the statistics report (--dry-run to preview, typed confirmation to publish, refuses non-interactively). Attribution is anonymous unless you pass --contributor; token/cost totals are excluded unless you pass --include-token-stats. rate.py WRITES a retrieval_log table into the database you point it at — use --db on a copy to sandbox it. No daemon, no cron, nothing runs on its own. See the Permissions, Data Flow & Consent section.\"\n          },\n      },\n  }\n---\n\n# Memory Bench\n\n> One of dozens of skills and plugins in **[TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — a self-improving OpenClaw fork that's been running 24/7 for months.\n\nEveryone has an opinion about whether their agent's memory is any good. Almost nobody has a number.\n\nThis produces the number — nDCG, MAP, MRR, Precision@5, each with a 95% bootstrap confidence interval, plus an ablation that isolates what spreading activation actually contributes. Then, if you want, it contributes your (anonymous) results to the ENGRAM and CORTEX research papers, where a few dozen real deployments beat any amount of arguing.\n\n**Part of [TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — real-time token tracking, self-improving crons, persistent cognitive memory. This is one piece of that stack; the repo has dozens more.\n\n👉 **https://github.com/globalcaos/tinkerclaw**\n\n_Clone it. Fork it. Break it. Make it yours._\n\n## Three-Step Pipeline\n\nStep 1 measures, step 2 summarises, step 3 publishes. Steps 1 and 2 are local. **Step 3 is the only one that uploads anything, and it asks first.**\n\n### 1. Assess Retrieval Quality\n\nRun the standard test set (30 queries across 4 types × 3 difficulty levels):\n\n```bash\n# Local judge — the default. Nothing leaves your machine.\npython3 scripts/rate.py --queries 30 --judge local --ablation\n\n# Stronger judge, but it TRANSMITS retrieved memory excerpts to OpenAI.\n# You will be asked to type 'send' before the first request.\npython3 scripts/rate.py --queries 30 --judge openai --ablation\n\n# Benchmark a copy instead of your live database\npython3 scripts/rate.py --db /tmp/memory-copy.db --judge local\n\n# Custom test set\npython3 scripts/rate.py --testset path/to/queries.json --judge local\n```\n\n**What it measures:**\n\n- **RAR** (Recall Accuracy Ratio), **MRR** (Mean Reciprocal Rank)\n- **nDCG@5**, **MAP@5**, **Precision@5**, **Hit Rate**\n- All metrics include **95% bootstrap confidence intervals**\n- **Ablation**: runs with AND without spreading activation to isolate its contribution\n\n**Judge methods — this is the privacy decision in this skill:**\n\n| Judge | What it sees | Where it goes | Cost |\n| --- | --- | --- | --- |\n| `local` **(default)** | query + 300 chars of each result | `http://127.0.0.1:8900/embed` on your own machine | free |\n| `openai` | query + 300 **redacted** chars of each result | `api.openai.com`, gpt-4o-mini | ~$0.01/run |\n\nThe `openai` judge is more discriminating and independent of the retrieval system, which is why the research protocol prefers it. It is also the option that puts pieces of your memories in someone else's logs. Both are legitimate; pick deliberately, and see the consent section below for exactly what is sent.\n\n**It writes to your database.** `rate.py` creates (or extends) a `retrieval_log` table in the database you point it at and inserts one row per benchmark query: the benchmark query text, the ratings, the metrics. Your memory content is never written there. `collect.py` reads that table later. If you would rather not touch your live DB, run both scripts with `--db` against a copy.\n\n**Standard test set** (`scripts/testset.json`): 30 queries stratified across semantic/episodic/procedural/strategic types and easy/medium/hard difficulty. As of 2.1.0 the queries are phrased to probe the operational and technical side of your memory — configuration, procedures, incidents, architecture decisions — rather than family, health, contacts or calendar. All deployments run the same queries, so results are comparable across sites; reports produced with the pre-2.1.0 set are not directly comparable with these.\n\nNote the honest limit: the queries steer *what is asked*, not *what your memory system returns*. Retrieval searches whatever database you point it at. If your memory holds things you would not want an external judge to rate, use `--judge local` (the default), or benchmark a filtered copy with `--db`.\n\n### 2. Collect Statistics\n\n```bash\n# Anonymous — the default\npython3 scripts/collect.py --days 14 --output /tmp/memory-bench-report.json\n\n# Attributed to you (your username goes in the report, and it becomes public if you submit)\npython3 scripts/collect.py --days 14 --contributor YOUR_GITHUB_USER --output /tmp/memory-bench-report.json\n```\n\n**What goes in the report:** memory counts, type/age distributions, strength and importance histograms, association graph size, hierarchy levels, consolidation run counts, embedding coverage, retrieval metrics from `retrieval_log` (RAR/MRR/nDCG/MAP, judge method, ablation config), the algorithm version as a short git SHA, and coarse system info (OS, CPU architecture, Python version, Node version). Instance ID is a random UUID.\n\n**What never goes in it:** memory content, benchmark queries, file paths, hostnames, environment variables.\n\n**Opt-in extras, off unless you ask:**\n\n- `--contributor NAME` — your username. Without it the report says `anonymous`.\n- `--include-token-stats` — total tokens and USD spend from your OpenClaw usage files.\n\n`collect.py` uploads nothing. It writes a JSON file and tells you what is in it. Read that file before step 3.\n\n### 3. Submit as PR — the step that publishes\n\n```bash\n# ALWAYS do this first: shows every field that would become public, uploads nothing\nscripts/submit.sh /tmp/memory-bench-report.json --dry-run\n\n# Real submission — prints the same preview, then asks you to type 'publish'\nscripts/submit.sh /tmp/memory-bench-report.json YOUR_GITHUB_USERNAME\n```\n\nThis forks `globalcaos/clawdbot-moltbot-openclaw`, pushes a branch to **your** fork, and opens a **public pull request** containing the report file. A merged PR is public and permanent. Requires the `gh` CLI, authenticated.\n\nIt refuses to run unattended: with no terminal to confirm on, it exits rather than publishing. `--yes` is available for scripted use and means you accept the upload.\n\n## Permissions, Data Flow & Consent\n\nShort version: steps 1 and 2 are local; step 3 publishes, and so does the optional OpenAI judge. Both ask first. Longer version, because you should not have to take that on trust:\n\n**What it needs, and why.**\n\n| Capability | Why | Scope |\n| --- | --- | --- |\n| Read your memory DB | Counts, histograms, and running the benchmark queries | `~/.openclaw/workspace/db/{memory,cognitive_memory,jarvis}.db`, or `--db` |\n| **Write** your memory DB | `rate.py` logs one `retrieval_log` row per benchmark query | Same DB (or the copy you pass to `--db`); no other table is touched |\n| File write | `<db_dir>/.memory-bench-instance-id` (random UUID, so repeat reports group together) and the `--output` report path | Two files, both of which you can delete |\n| Read usage files | Token/cost totals — **only** with `--include-token-stats` | `~/.openclaw/workspace/memory/*-usage.json` |\n| Local shell exec | `git log` for the algorithm SHA, `node --version` for system info, `git`/`gh` in `submit.sh` | Fixed commands |\n| Network to localhost | Local judge embeddings | `http://127.0.0.1:8900/embed` — stays on the machine |\n| Network to OpenAI | **Opt-in.** Query + up to 300 redacted chars per retrieved memory | `api.openai.com`, only with `--judge openai` after confirmation |\n| Network to GitHub | **Opt-in.** Uploads the report and opens a public PR | `github.com`, only in `submit.sh` after confirmation |\n| Credentials | `OPENAI_API_KEY` (env or `--api-key`) only when you choose the OpenAI judge; `gh`'s existing login in `submit.sh` | Read at call time, never stored, never written to the report |\n| Scheduling | **None.** No daemon, no cron, no install hook. Nothing runs unless you run it | — |\n\n**About that redaction.** Before an excerpt goes to OpenAI, `rate.py` replaces email addresses, phone-shaped numbers, API-key-shaped strings, long hex blobs, home directory paths and URL credentials with placeholders. That is pattern matching, not comprehension — it will not catch a secret written in prose. Treat it as a seatbelt, not a force field. If the content is genuinely sensitive, the answer is `--judge local`, not better regexes.\n\n**Consent, concretely.** Two actions can leave your machine, and neither happens by accident:\n\n```bash\n# 1. External judging — prints exactly what will be sent, then waits\npython3 scripts/rate.py --judge openai        # asks you to type 'send'\npython3 scripts/rate.py --judge openai --yes-send-to-openai   # scripted consent\n\n# 2. Publishing — prints every field that becomes public, then waits\nscripts/submit.sh report.json --dry-run       # preview only, uploads nothing\nscripts/submit.sh report.json                 # asks you to type 'publish'\n```\n\nBoth refuse outright when there is no terminal to confirm on, so an agent running this unattended cannot publish on your behalf.\n\n**Turning it off, and undoing it:**\n\n```bash\n# never transmit: just use the default judge and skip step 3\npython3 scripts/rate.py --judge local\n\n# benchmark a throwaway copy instead of your live memory\ncp ~/.openclaw/workspace/db/memory.db /tmp/bench.db\npython3 scripts/rate.py --db /tmp/bench.db --judge local\n\n# forget this installation ever ran\nrm ~/.openclaw/workspace/db/.memory-bench-instance-id\nsqlite3 ~/.openclaw/workspace/db/memory.db 'DROP TABLE retrieval_log;'\n```\n\nDeleting the instance ID file makes your next report a new anonymous instance; it also breaks the longitudinal link, which is the trade.\n\n**Read it before you run it.** Three scripts, all plain text, none of them long. `rate.py` is the only one that can talk to a third party and `submit.sh` is the only one that can publish — both are worth the two minutes.\n\n## Validation Protocol\n\nFor peer-review-ready data, contributors should:\n\n1. Run `rate.py --ablation` over the full N=30 test set\n2. Use `--judge openai` if you are comfortable with the data flow above — it agrees better with human raters, and the script reports Cohen's κ between the two judges so you can see the gap on your own data. Local-judge submissions are still welcome and are marked as such in the report\n3. Collect at least 2 reports from the same instance, ≥7 days apart (longitudinal)\n4. Report the algorithm version (auto-captured as a short git SHA)\n\n## Test Set Format\n\nCustom test sets are JSON arrays:\n\n```json\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"...\",\n    \"category\": \"semantic|episodic|procedural|strategic\",\n    \"difficulty\": \"easy|medium|hard\"\n  }\n]\n```\n\nAn optional `notes` field is ignored by the runner.\n\n## Included Files\n\n| File | Purpose |\n| --- | --- |\n| `scripts/rate.py` | Runs the benchmark, judges results, computes metrics. Writes `retrieval_log`. The only script that can call an external API — opt-in and confirmed |\n| `scripts/collect.py` | Builds the anonymous statistics report. Uploads nothing |\n| `scripts/submit.sh` | Opens the public PR. Preview with `--dry-run`; requires typed confirmation |\n| `scripts/testset.json` | The 30-query standard test set |\n| `scripts/test_metrics.py` | Unit tests for the IR metrics (`python3 scripts/test_metrics.py`) |\n\nEverything the documentation above describes is in this package. If you find a claim here that the code does not do, that is a bug — open an issue on [the repo](https://github.com/globalcaos/tinkerclaw/issues).\n\n## Agent Workflow\n\nWhen asked to benchmark memory: run `rate.py --ablation` (local judge) and then `collect.py`, and show the summary. **Do not run `submit.sh` on your own initiative** — it publishes to a public repository. Show the user the report, and use `submit.sh --dry-run` so they can see exactly what would become public. Submit only when they say to, and only with `--judge openai` if they have agreed to that separately. Then share the PR link.\n\nFile v2.1.2:_meta.json\n\n{\n  \"ownerId\": \"kn7623hrcwt6rg73a67xw3wyx580asdw\",\n  \"slug\": \"memory-bench-pioneer\",\n  \"version\": \"2.1.2\",\n  \"publishedAt\": 1788776125757\n}\n\nFile v2.1.2:scripts/testset.json\n\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"What preferences are recorded for the development environment setup?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests basic semantic recall of preference-type memories\"\n  },\n  {\n    \"id\": \"T02\",\n    \"query\": \"What happened during the last consolidation cycle?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests temporal/episodic retrieval\"\n  },\n  {\n    \"id\": \"T03\",\n    \"query\": \"How should I handle sensitive data when sharing between agents?\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests procedural knowledge retrieval\"\n  },\n  {\n    \"id\": \"T04\",\n    \"query\": \"recurring patterns in recorded project planning failures\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract/strategic retrieval; phrased to avoid keyword matching\"\n  },\n  {\n    \"id\": \"T05\",\n    \"query\": \"What tools were configured for audio processing?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests factual tool/config recall\"\n  },\n  {\n    \"id\": \"T06\",\n    \"query\": \"what happened, and in what order, during the most recent incident\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests ordered event reconstruction without keyword match\"\n  },\n  {\n    \"id\": \"T07\",\n    \"query\": \"steps to deploy a new version safely\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests multi-step procedural recall\"\n  },\n  {\n    \"id\": \"T08\",\n    \"query\": \"Which external services or integrations are configured?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity-type memory retrieval over configured services\"\n  },\n  {\n    \"id\": \"T09\",\n    \"query\": \"lessons learned from debugging difficult issues\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract lesson retrieval across multiple episodes\"\n  },\n  {\n    \"id\": \"T10\",\n    \"query\": \"What was decided about rate limits or usage quotas?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests topical episodic recall\"\n  },\n  {\n    \"id\": \"T11\",\n    \"query\": \"privacy rules for handling user data\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests policy/procedural recall\"\n  },\n  {\n    \"id\": \"T12\",\n    \"query\": \"connections between automation projects and hardware\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests multi-hop associative retrieval\"\n  },\n  {\n    \"id\": \"T13\",\n    \"query\": \"which scheduled or recurring jobs have run recently\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recurring-job pattern retrieval\"\n  },\n  {\n    \"id\": \"T14\",\n    \"query\": \"error handling best practices\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests general procedural knowledge\"\n  },\n  {\n    \"id\": \"T15\",\n    \"query\": \"which design proposals were recorded for the retrieval pipeline\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent episodic recall with abstract query\"\n  },\n  {\n    \"id\": \"T16\",\n    \"query\": \"tradeoffs between local and cloud solutions\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests comparative/analytical retrieval\"\n  },\n  {\n    \"id\": \"T17\",\n    \"query\": \"which components make up the system and what each one does\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests component/architecture relationship recall\"\n  },\n  {\n    \"id\": \"T18\",\n    \"query\": \"What failed and why in previous attempts?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests failure/negative experience retrieval\"\n  },\n  {\n    \"id\": \"T19\",\n    \"query\": \"configuration steps for new integrations\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests setup/config procedural recall\"\n  },\n  {\n    \"id\": \"T20\",\n    \"query\": \"long-term technical objectives recorded for the memory system\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests goal-tracking retrieval across time\"\n  },\n  {\n    \"id\": \"T21\",\n    \"query\": \"formatting and output conventions that are preferred\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests convention/preference recall\"\n  },\n  {\n    \"id\": \"T22\",\n    \"query\": \"What did the last test or verification run report as failing?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests failure-report episodic recall\"\n  },\n  {\n    \"id\": \"T23\",\n    \"query\": \"how to resolve merge conflicts in the codebase\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests specific technical procedural recall\"\n  },\n  {\n    \"id\": \"T24\",\n    \"query\": \"evolution of the memory architecture over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal strategic retrieval\"\n  },\n  {\n    \"id\": \"T25\",\n    \"query\": \"the most recent benchmark runs and when they happened\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests temporal recall of recent runs\"\n  },\n  {\n    \"id\": \"T26\",\n    \"query\": \"Which file formats or codecs are documented as supported?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests documented-capability recall\"\n  },\n  {\n    \"id\": \"T27\",\n    \"query\": \"backup and recovery procedures\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests disaster recovery procedural recall\"\n  },\n  {\n    \"id\": \"T28\",\n    \"query\": \"how tradeoffs between retrieval accuracy and latency were weighed over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal tradeoff-reasoning retrieval\"\n  },\n  {\n    \"id\": \"T29\",\n    \"query\": \"What was the outcome of the last research task?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent task outcome retrieval\"\n  },\n  {\n    \"id\": \"T30\",\n    \"query\": \"communication style preferences for different channels\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests nuanced preference retrieval\"\n  }\n]\n\nFile v2.1.2:skill-card.md\n\n## Description:\n\nBenchmarks an agent's live memory system with retrieval metrics, confidence intervals, and ablation analysis while defaulting to local execution.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[globalcaos](https://clawhub.ai/user/globalcaos)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and external OpenClaw or TinkerClaw users use this skill to measure retrieval quality, collect aggregate memory statistics, and optionally prepare a public benchmark submission.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Optional OpenAI judging can expose benchmark queries and retrieved memory excerpts to an external API.\n\nMitigation: Use the default local judge for sensitive memories, or review the data flow and provide explicit consent only after deciding external judging is acceptable.\n\nRisk: Public GitHub submission can publish aggregate report data and attribution when a contributor name is supplied.\n\nMitigation: Inspect the complete JSON report, run the dry-run preview first, omit the contributor for anonymous reports, and publish only after explicit user approval.\n\nRisk: Benchmarking writes a retrieval_log table into the database selected for the run.\n\nMitigation: Run against a copied or filtered database with --db when the live memory database should not be modified.\n\nRisk: The built-in preview may not be a complete substitute for reviewing the full report payload.\n\nMitigation: Open and inspect the complete JSON report before any public submission.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/globalcaos/skills/memory-bench-pioneer)\n- [TinkerClaw repository](https://github.com/globalcaos/tinkerclaw)\n- [ENGRAM context compaction paper](https://github.com/globalcaos/clawdbot-moltbot-openclaw/blob/main/docs/papers/context-compaction.md)\n- [CORTEX agent memory paper](https://github.com/globalcaos/clawdbot-moltbot-openclaw/blob/main/docs/papers/agent-memory.md)\n\n## Skill Output:\n\n**Output Type(s):** [Text, Markdown, Shell commands, JSON]\n\n**Output Format:** [Markdown guidance with shell commands and JSON report files]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Writes benchmark rows to retrieval_log and can produce a JSON report; optional OpenAI judging and GitHub pull request submission require explicit consent.]\n\n## Skill Version(s):\n\n2.1.2 (source: SKILL.md frontmatter and server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v2.1.1: 8 files, 28069 bytes\n\nFiles: _meta.json (139b), scripts/collect.py (16952b), scripts/rate.py (24005b), scripts/submit.sh (9293b), scripts/test_metrics.py (4009b), scripts/testset.json (6202b), skill-card.md (2441b), SKILL.md (13585b)\n\nFile v2.1.1:SKILL.md\n\n---\nname: memory-bench-pioneer\nversion: 2.1.1\ndescription: \"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine. The optional OpenAI judge is opt-in, prints exactly what it would send, redacts secrets first, requires typed consent, and refuses to run unattended. Submitting results is a separate confirmed step that previews every public field, and identifies you only if you pass --contributor. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw. See Permissions, Data Flow & Consent.\"\nmetadata:\n  {\n    \"openclaw\":\n      {\n        \"emoji\": \"🧠\",\n        \"requires\": { \"bins\": [\"python3\"] },\n        \"notes\":\n          {\n            \"security\": \"Benchmarks a LIVE memory database, so disclosure matters more than usual. DEFAULT PATH IS LOCAL: the default judge (--judge local) sends retrieved excerpts only to a local embedding server on 127.0.0.1; nothing leaves the machine. Two actions can send data out and BOTH are opt-in and confirmed: --judge openai transmits the benchmark query plus up to 300 redacted characters of each RETRIEVED MEMORY to api.openai.com (typed confirmation, or --yes-send-to-openai), and scripts/submit.sh opens a PUBLIC GitHub pull request containing the statistics report (--dry-run to preview, typed confirmation to publish, refuses non-interactively). Attribution is anonymous unless you pass --contributor; token/cost totals are excluded unless you pass --include-token-stats. rate.py WRITES a retrieval_log table into the database you point it at — use --db on a copy to sandbox it. No daemon, no cron, nothing runs on its own. See the Permissions, Data Flow & Consent section.\"\n          },\n      },\n  }\n---\n\n# Memory Bench\n\n> One of dozens of skills and plugins in **[TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — a self-improving OpenClaw fork that's been running 24/7 for months.\n\nEveryone has an opinion about whether their agent's memory is any good. Almost nobody has a number.\n\nThis produces the number — nDCG, MAP, MRR, Precision@5, each with a 95% bootstrap confidence interval, plus an ablation that isolates what spreading activation actually contributes. Then, if you want, it contributes your (anonymous) results to the ENGRAM and CORTEX research papers, where a few dozen real deployments beat any amount of arguing.\n\n**Part of [TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — real-time token tracking, self-improving crons, persistent cognitive memory. This is one piece of that stack; the repo has dozens more.\n\n👉 **https://github.com/globalcaos/tinkerclaw**\n\n_Clone it. Fork it. Break it. Make it yours._\n\n## Three-Step Pipeline\n\nStep 1 measures, step 2 summarises, step 3 publishes. Steps 1 and 2 are local. **Step 3 is the only one that uploads anything, and it asks first.**\n\n### 1. Assess Retrieval Quality\n\nRun the standard test set (30 queries across 4 types × 3 difficulty levels):\n\n```bash\n# Local judge — the default. Nothing leaves your machine.\npython3 scripts/rate.py --queries 30 --judge local --ablation\n\n# Stronger judge, but it TRANSMITS retrieved memory excerpts to OpenAI.\n# You will be asked to type 'send' before the first request.\npython3 scripts/rate.py --queries 30 --judge openai --ablation\n\n# Benchmark a copy instead of your live database\npython3 scripts/rate.py --db /tmp/memory-copy.db --judge local\n\n# Custom test set\npython3 scripts/rate.py --testset path/to/queries.json --judge local\n```\n\n**What it measures:**\n\n- **RAR** (Recall Accuracy Ratio), **MRR** (Mean Reciprocal Rank)\n- **nDCG@5**, **MAP@5**, **Precision@5**, **Hit Rate**\n- All metrics include **95% bootstrap confidence intervals**\n- **Ablation**: runs with AND without spreading activation to isolate its contribution\n\n**Judge methods — this is the privacy decision in this skill:**\n\n| Judge | What it sees | Where it goes | Cost |\n| --- | --- | --- | --- |\n| `local` **(default)** | query + 300 chars of each result | `http://127.0.0.1:8900/embed` on your own machine | free |\n| `openai` | query + 300 **redacted** chars of each result | `api.openai.com`, gpt-4o-mini | ~$0.01/run |\n\nThe `openai` judge is more discriminating and independent of the retrieval system, which is why the research protocol prefers it. It is also the option that puts pieces of your memories in someone else's logs. Both are legitimate; pick deliberately, and see the consent section below for exactly what is sent.\n\n**It writes to your database.** `rate.py` creates (or extends) a `retrieval_log` table in the database you point it at and inserts one row per benchmark query: the benchmark query text, the ratings, the metrics. Your memory content is never written there. `collect.py` reads that table later. If you would rather not touch your live DB, run both scripts with `--db` against a copy.\n\n**Standard test set** (`scripts/testset.json`): 30 queries stratified across semantic/episodic/procedural/strategic types and easy/medium/hard difficulty. As of 2.1.0 the queries are phrased to probe the operational and technical side of your memory — configuration, procedures, incidents, architecture decisions — rather than family, health, contacts or calendar. All deployments run the same queries, so results are comparable across sites; reports produced with the pre-2.1.0 set are not directly comparable with these.\n\nNote the honest limit: the queries steer *what is asked*, not *what your memory system returns*. Retrieval searches whatever database you point it at. If your memory holds things you would not want an external judge to rate, use `--judge local` (the default), or benchmark a filtered copy with `--db`.\n\n### 2. Collect Statistics\n\n```bash\n# Anonymous — the default\npython3 scripts/collect.py --days 14 --output /tmp/memory-bench-report.json\n\n# Attributed to you (your username goes in the report, and it becomes public if you submit)\npython3 scripts/collect.py --days 14 --contributor YOUR_GITHUB_USER --output /tmp/memory-bench-report.json\n```\n\n**What goes in the report:** memory counts, type/age distributions, strength and importance histograms, association graph size, hierarchy levels, consolidation run counts, embedding coverage, retrieval metrics from `retrieval_log` (RAR/MRR/nDCG/MAP, judge method, ablation config), the algorithm version as a short git SHA, and coarse system info (OS, CPU architecture, Python version, Node version). Instance ID is a random UUID.\n\n**What never goes in it:** memory content, benchmark queries, file paths, hostnames, environment variables.\n\n**Opt-in extras, off unless you ask:**\n\n- `--contributor NAME` — your username. Without it the report says `anonymous`.\n- `--include-token-stats` — total tokens and USD spend from your OpenClaw usage files.\n\n`collect.py` uploads nothing. It writes a JSON file and tells you what is in it. Read that file before step 3.\n\n### 3. Submit as PR — the step that publishes\n\n```bash\n# ALWAYS do this first: shows every field that would become public, uploads nothing\nscripts/submit.sh /tmp/memory-bench-report.json --dry-run\n\n# Real submission — prints the same preview, then asks you to type 'publish'\nscripts/submit.sh /tmp/memory-bench-report.json YOUR_GITHUB_USERNAME\n```\n\nThis forks `globalcaos/clawdbot-moltbot-openclaw`, pushes a branch to **your** fork, and opens a **public pull request** containing the report file. A merged PR is public and permanent. Requires the `gh` CLI, authenticated.\n\nIt refuses to run unattended: with no terminal to confirm on, it exits rather than publishing. `--yes` is available for scripted use and means you accept the upload.\n\n## Permissions, Data Flow & Consent\n\nShort version: steps 1 and 2 are local; step 3 publishes, and so does the optional OpenAI judge. Both ask first. Longer version, because you should not have to take that on trust:\n\n**What it needs, and why.**\n\n| Capability | Why | Scope |\n| --- | --- | --- |\n| Read your memory DB | Counts, histograms, and running the benchmark queries | `~/.openclaw/workspace/db/{memory,cognitive_memory,jarvis}.db`, or `--db` |\n| **Write** your memory DB | `rate.py` logs one `retrieval_log` row per benchmark query | Same DB (or the copy you pass to `--db`); no other table is touched |\n| File write | `<db_dir>/.memory-bench-instance-id` (random UUID, so repeat reports group together) and the `--output` report path | Two files, both of which you can delete |\n| Read usage files | Token/cost totals — **only** with `--include-token-stats` | `~/.openclaw/workspace/memory/*-usage.json` |\n| Local shell exec | `git log` for the algorithm SHA, `node --version` for system info, `git`/`gh` in `submit.sh` | Fixed commands |\n| Network to localhost | Local judge embeddings | `http://127.0.0.1:8900/embed` — stays on the machine |\n| Network to OpenAI | **Opt-in.** Query + up to 300 redacted chars per retrieved memory | `api.openai.com`, only with `--judge openai` after confirmation |\n| Network to GitHub | **Opt-in.** Uploads the report and opens a public PR | `github.com`, only in `submit.sh` after confirmation |\n| Credentials | `OPENAI_API_KEY` (env or `--api-key`) only when you choose the OpenAI judge; `gh`'s existing login in `submit.sh` | Read at call time, never stored, never written to the report |\n| Scheduling | **None.** No daemon, no cron, no install hook. Nothing runs unless you run it | — |\n\n**About that redaction.** Before an excerpt goes to OpenAI, `rate.py` replaces email addresses, phone-shaped numbers, API-key-shaped strings, long hex blobs, home directory paths and URL credentials with placeholders. That is pattern matching, not comprehension — it will not catch a secret written in prose. Treat it as a seatbelt, not a force field. If the content is genuinely sensitive, the answer is `--judge local`, not better regexes.\n\n**Consent, concretely.** Two actions can leave your machine, and neither happens by accident:\n\n```bash\n# 1. External judging — prints exactly what will be sent, then waits\npython3 scripts/rate.py --judge openai        # asks you to type 'send'\npython3 scripts/rate.py --judge openai --yes-send-to-openai   # scripted consent\n\n# 2. Publishing — prints every field that becomes public, then waits\nscripts/submit.sh report.json --dry-run       # preview only, uploads nothing\nscripts/submit.sh report.json                 # asks you to type 'publish'\n```\n\nBoth refuse outright when there is no terminal to confirm on, so an agent running this unattended cannot publish on your behalf.\n\n**Turning it off, and undoing it:**\n\n```bash\n# never transmit: just use the default judge and skip step 3\npython3 scripts/rate.py --judge local\n\n# benchmark a throwaway copy instead of your live memory\ncp ~/.openclaw/workspace/db/memory.db /tmp/bench.db\npython3 scripts/rate.py --db /tmp/bench.db --judge local\n\n# forget this installation ever ran\nrm ~/.openclaw/workspace/db/.memory-bench-instance-id\nsqlite3 ~/.openclaw/workspace/db/memory.db 'DROP TABLE retrieval_log;'\n```\n\nDeleting the instance ID file makes your next report a new anonymous instance; it also breaks the longitudinal link, which is the trade.\n\n**Read it before you run it.** Three scripts, all plain text, none of them long. `rate.py` is the only one that can talk to a third party and `submit.sh` is the only one that can publish — both are worth the two minutes.\n\n## Validation Protocol\n\nFor peer-review-ready data, contributors should:\n\n1. Run `rate.py --ablation` over the full N=30 test set\n2. Use `--judge openai` if you are comfortable with the data flow above — it agrees better with human raters, and the script reports Cohen's κ between the two judges so you can see the gap on your own data. Local-judge submissions are still welcome and are marked as such in the report\n3. Collect at least 2 reports from the same instance, ≥7 days apart (longitudinal)\n4. Report the algorithm version (auto-captured as a short git SHA)\n\n## Test Set Format\n\nCustom test sets are JSON arrays:\n\n```json\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"...\",\n    \"category\": \"semantic|episodic|procedural|strategic\",\n    \"difficulty\": \"easy|medium|hard\"\n  }\n]\n```\n\nAn optional `notes` field is ignored by the runner.\n\n## Included Files\n\n| File | Purpose |\n| --- | --- |\n| `scripts/rate.py` | Runs the benchmark, judges results, computes metrics. Writes `retrieval_log`. The only script that can call an external API — opt-in and confirmed |\n| `scripts/collect.py` | Builds the anonymous statistics report. Uploads nothing |\n| `scripts/submit.sh` | Opens the public PR. Preview with `--dry-run`; requires typed confirmation |\n| `scripts/testset.json` | The 30-query standard test set |\n| `scripts/test_metrics.py` | Unit tests for the IR metrics (`python3 scripts/test_metrics.py`) |\n\nEverything the documentation above describes is in this package. If you find a claim here that the code does not do, that is a bug — open an issue on [the repo](https://github.com/globalcaos/tinkerclaw/issues).\n\n## Agent Workflow\n\nWhen asked to benchmark memory: run `rate.py --ablation` (local judge) and then `collect.py`, and show the summary. **Do not run `submit.sh` on your own initiative** — it publishes to a public repository. Show the user the report, and use `submit.sh --dry-run` so they can see exactly what would become public. Submit only when they say to, and only with `--judge openai` if they have agreed to that separately. Then share the PR link.\n\nFile v2.1.1:_meta.json\n\n{\n  \"ownerId\": \"kn7623hrcwt6rg73a67xw3wyx580asdw\",\n  \"slug\": \"memory-bench-pioneer\",\n  \"version\": \"2.1.1\",\n  \"publishedAt\": 1788775337474\n}\n\nFile v2.1.1:scripts/testset.json\n\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"What preferences are recorded for the development environment setup?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests basic semantic recall of preference-type memories\"\n  },\n  {\n    \"id\": \"T02\",\n    \"query\": \"What happened during the last consolidation cycle?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests temporal/episodic retrieval\"\n  },\n  {\n    \"id\": \"T03\",\n    \"query\": \"How should I handle sensitive data when sharing between agents?\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests procedural knowledge retrieval\"\n  },\n  {\n    \"id\": \"T04\",\n    \"query\": \"recurring patterns in recorded project planning failures\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract/strategic retrieval; phrased to avoid keyword matching\"\n  },\n  {\n    \"id\": \"T05\",\n    \"query\": \"What tools were configured for audio processing?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests factual tool/config recall\"\n  },\n  {\n    \"id\": \"T06\",\n    \"query\": \"what happened, and in what order, during the most recent incident\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests ordered event reconstruction without keyword match\"\n  },\n  {\n    \"id\": \"T07\",\n    \"query\": \"steps to deploy a new version safely\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests multi-step procedural recall\"\n  },\n  {\n    \"id\": \"T08\",\n    \"query\": \"Which external services or integrations are configured?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity-type memory retrieval over configured services\"\n  },\n  {\n    \"id\": \"T09\",\n    \"query\": \"lessons learned from debugging difficult issues\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract lesson retrieval across multiple episodes\"\n  },\n  {\n    \"id\": \"T10\",\n    \"query\": \"What was decided about rate limits or usage quotas?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests topical episodic recall\"\n  },\n  {\n    \"id\": \"T11\",\n    \"query\": \"privacy rules for handling user data\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests policy/procedural recall\"\n  },\n  {\n    \"id\": \"T12\",\n    \"query\": \"connections between automation projects and hardware\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests multi-hop associative retrieval\"\n  },\n  {\n    \"id\": \"T13\",\n    \"query\": \"which scheduled or recurring jobs have run recently\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recurring-job pattern retrieval\"\n  },\n  {\n    \"id\": \"T14\",\n    \"query\": \"error handling best practices\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests general procedural knowledge\"\n  },\n  {\n    \"id\": \"T15\",\n    \"query\": \"which design proposals were recorded for the retrieval pipeline\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent episodic recall with abstract query\"\n  },\n  {\n    \"id\": \"T16\",\n    \"query\": \"tradeoffs between local and cloud solutions\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests comparative/analytical retrieval\"\n  },\n  {\n    \"id\": \"T17\",\n    \"query\": \"which components make up the system and what each one does\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests component/architecture relationship recall\"\n  },\n  {\n    \"id\": \"T18\",\n    \"query\": \"What failed and why in previous attempts?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests failure/negative experience retrieval\"\n  },\n  {\n    \"id\": \"T19\",\n    \"query\": \"configuration steps for new integrations\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests setup/config procedural recall\"\n  },\n  {\n    \"id\": \"T20\",\n    \"query\": \"long-term technical objectives recorded for the memory system\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests goal-tracking retrieval across time\"\n  },\n  {\n    \"id\": \"T21\",\n    \"query\": \"formatting and output conventions that are preferred\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests convention/preference recall\"\n  },\n  {\n    \"id\": \"T22\",\n    \"query\": \"What did the last test or verification run report as failing?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests failure-report episodic recall\"\n  },\n  {\n    \"id\": \"T23\",\n    \"query\": \"how to resolve merge conflicts in the codebase\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests specific technical procedural recall\"\n  },\n  {\n    \"id\": \"T24\",\n    \"query\": \"evolution of the memory architecture over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal strategic retrieval\"\n  },\n  {\n    \"id\": \"T25\",\n    \"query\": \"the most recent benchmark runs and when they happened\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests temporal recall of recent runs\"\n  },\n  {\n    \"id\": \"T26\",\n    \"query\": \"Which file formats or codecs are documented as supported?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests documented-capability recall\"\n  },\n  {\n    \"id\": \"T27\",\n    \"query\": \"backup and recovery procedures\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests disaster recovery procedural recall\"\n  },\n  {\n    \"id\": \"T28\",\n    \"query\": \"how tradeoffs between retrieval accuracy and latency were weighed over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal tradeoff-reasoning retrieval\"\n  },\n  {\n    \"id\": \"T29\",\n    \"query\": \"What was the outcome of the last research task?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent task outcome retrieval\"\n  },\n  {\n    \"id\": \"T30\",\n    \"query\": \"communication style preferences for different channels\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests nuanced preference retrieval\"\n  }\n]\n\nFile v2.1.1:skill-card.md\n\n## Description:\n\nTinkerClaw Memory Bench benchmarks an agent's live memory system locally by default, producing retrieval-quality metrics with confidence intervals and an optional, separately confirmed public submission flow.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[globalcaos](https://clawhub.ai/user/globalcaos)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and engineers use this skill to benchmark a TinkerClaw or OpenClaw memory database, measure retrieval quality with standard IR metrics, and generate an aggregate JSON report. It can also help users preview and optionally submit benchmark results as a public GitHub pull request.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The public submission step can publish more report data than the preview fully summarizes.\n\nMitigation: Run submit.sh with --dry-run first and inspect the full JSON report before publishing; avoid --yes unless the full report has already been reviewed.\n\nRisk: Benchmarking a live memory database may expose sensitive retrieved excerpts if the optional OpenAI judge is used.\n\nMitigation: Use the default local judge or benchmark a copied or filtered database with --db; choose --judge openai only when sending redacted excerpts outside the machine is acceptable.\n\nRisk: rate.py writes retrieval_log benchmark rows into the database being tested.\n\nMitigation: Run against a database copy with --db when isolation is needed, or remove the retrieval_log table after testing.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/globalcaos/skills/memory-bench-pioneer)\n- [TinkerClaw repository](https://github.com/globalcaos/tinkerclaw)\n\n## Skill Output:\n\n**Output Type(s):** [Analysis, JSON, Shell commands, Guidance]\n\n**Output Format:** [Markdown guidance with shell commands and generated JSON benchmark reports]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May write retrieval metrics to a SQLite database and optional report files; public submission requires separate confirmation.]\n\n## Skill Version(s):\n\n2.1.1 (source: server release evidence and SKILL.md frontmatter)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v2.1.0: 8 files, 28156 bytes\n\nFiles: _meta.json (139b), scripts/collect.py (16952b), scripts/rate.py (24005b), scripts/submit.sh (9293b), scripts/test_metrics.py (4009b), scripts/testset.json (6202b), skill-card.md (2699b), SKILL.md (13438b)\n\nFile v2.1.0:SKILL.md\n\n---\nname: memory-bench-pioneer\nversion: 2.1.0\ndescription: \"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Runs a peer-review-grade evaluation suite (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablation studies) against your live memory system and submits anonymized results to the ENGRAM/CORTEX research papers. Your data stays private; only aggregate stats leave. Works with agent-memory-ultimate. For the bold few who believe AI memory should be measured, not guessed at. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw.\"\nmetadata:\n  {\n    \"openclaw\":\n      {\n        \"emoji\": \"🧠\",\n        \"requires\": { \"bins\": [\"python3\"] },\n        \"notes\":\n          {\n            \"security\": \"Benchmarks a LIVE memory database, so disclosure matters more than usual. DEFAULT PATH IS LOCAL: the default judge (--judge local) sends retrieved excerpts only to a local embedding server on 127.0.0.1; nothing leaves the machine. Two actions can send data out and BOTH are opt-in and confirmed: --judge openai transmits the benchmark query plus up to 300 redacted characters of each RETRIEVED MEMORY to api.openai.com (typed confirmation, or --yes-send-to-openai), and scripts/submit.sh opens a PUBLIC GitHub pull request containing the statistics report (--dry-run to preview, typed confirmation to publish, refuses non-interactively). Attribution is anonymous unless you pass --contributor; token/cost totals are excluded unless you pass --include-token-stats. rate.py WRITES a retrieval_log table into the database you point it at — use --db on a copy to sandbox it. No daemon, no cron, nothing runs on its own. See the Permissions, Data Flow & Consent section.\"\n          },\n      },\n  }\n---\n\n# Memory Bench\n\n> One of dozens of skills and plugins in **[TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — a self-improving OpenClaw fork that's been running 24/7 for months.\n\nEveryone has an opinion about whether their agent's memory is any good. Almost nobody has a number.\n\nThis produces the number — nDCG, MAP, MRR, Precision@5, each with a 95% bootstrap confidence interval, plus an ablation that isolates what spreading activation actually contributes. Then, if you want, it contributes your (anonymous) results to the ENGRAM and CORTEX research papers, where a few dozen real deployments beat any amount of arguing.\n\n**Part of [TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — real-time token tracking, self-improving crons, persistent cognitive memory. This is one piece of that stack; the repo has dozens more.\n\n👉 **https://github.com/globalcaos/tinkerclaw**\n\n_Clone it. Fork it. Break it. Make it yours._\n\n## Three-Step Pipeline\n\nStep 1 measures, step 2 summarises, step 3 publishes. Steps 1 and 2 are local. **Step 3 is the only one that uploads anything, and it asks first.**\n\n### 1. Assess Retrieval Quality\n\nRun the standard test set (30 queries across 4 types × 3 difficulty levels):\n\n```bash\n# Local judge — the default. Nothing leaves your machine.\npython3 scripts/rate.py --queries 30 --judge local --ablation\n\n# Stronger judge, but it TRANSMITS retrieved memory excerpts to OpenAI.\n# You will be asked to type 'send' before the first request.\npython3 scripts/rate.py --queries 30 --judge openai --ablation\n\n# Benchmark a copy instead of your live database\npython3 scripts/rate.py --db /tmp/memory-copy.db --judge local\n\n# Custom test set\npython3 scripts/rate.py --testset path/to/queries.json --judge local\n```\n\n**What it measures:**\n\n- **RAR** (Recall Accuracy Ratio), **MRR** (Mean Reciprocal Rank)\n- **nDCG@5**, **MAP@5**, **Precision@5**, **Hit Rate**\n- All metrics include **95% bootstrap confidence intervals**\n- **Ablation**: runs with AND without spreading activation to isolate its contribution\n\n**Judge methods — this is the privacy decision in this skill:**\n\n| Judge | What it sees | Where it goes | Cost |\n| --- | --- | --- | --- |\n| `local` **(default)** | query + 300 chars of each result | `http://127.0.0.1:8900/embed` on your own machine | free |\n| `openai` | query + 300 **redacted** chars of each result | `api.openai.com`, gpt-4o-mini | ~$0.01/run |\n\nThe `openai` judge is more discriminating and independent of the retrieval system, which is why the research protocol prefers it. It is also the option that puts pieces of your memories in someone else's logs. Both are legitimate; pick deliberately, and see the consent section below for exactly what is sent.\n\n**It writes to your database.** `rate.py` creates (or extends) a `retrieval_log` table in the database you point it at and inserts one row per benchmark query: the benchmark query text, the ratings, the metrics. Your memory content is never written there. `collect.py` reads that table later. If you would rather not touch your live DB, run both scripts with `--db` against a copy.\n\n**Standard test set** (`scripts/testset.json`): 30 queries stratified across semantic/episodic/procedural/strategic types and easy/medium/hard difficulty. As of 2.1.0 the queries are phrased to probe the operational and technical side of your memory — configuration, procedures, incidents, architecture decisions — rather than family, health, contacts or calendar. All deployments run the same queries, so results are comparable across sites; reports produced with the pre-2.1.0 set are not directly comparable with these.\n\nNote the honest limit: the queries steer *what is asked*, not *what your memory system returns*. Retrieval searches whatever database you point it at. If your memory holds things you would not want an external judge to rate, use `--judge local` (the default), or benchmark a filtered copy with `--db`.\n\n### 2. Collect Statistics\n\n```bash\n# Anonymous — the default\npython3 scripts/collect.py --days 14 --output /tmp/memory-bench-report.json\n\n# Attributed to you (your username goes in the report, and it becomes public if you submit)\npython3 scripts/collect.py --days 14 --contributor YOUR_GITHUB_USER --output /tmp/memory-bench-report.json\n```\n\n**What goes in the report:** memory counts, type/age distributions, strength and importance histograms, association graph size, hierarchy levels, consolidation run counts, embedding coverage, retrieval metrics from `retrieval_log` (RAR/MRR/nDCG/MAP, judge method, ablation config), the algorithm version as a short git SHA, and coarse system info (OS, CPU architecture, Python version, Node version). Instance ID is a random UUID.\n\n**What never goes in it:** memory content, benchmark queries, file paths, hostnames, environment variables.\n\n**Opt-in extras, off unless you ask:**\n\n- `--contributor NAME` — your username. Without it the report says `anonymous`.\n- `--include-token-stats` — total tokens and USD spend from your OpenClaw usage files.\n\n`collect.py` uploads nothing. It writes a JSON file and tells you what is in it. Read that file before step 3.\n\n### 3. Submit as PR — the step that publishes\n\n```bash\n# ALWAYS do this first: shows every field that would become public, uploads nothing\nscripts/submit.sh /tmp/memory-bench-report.json --dry-run\n\n# Real submission — prints the same preview, then asks you to type 'publish'\nscripts/submit.sh /tmp/memory-bench-report.json YOUR_GITHUB_USERNAME\n```\n\nThis forks `globalcaos/clawdbot-moltbot-openclaw`, pushes a branch to **your** fork, and opens a **public pull request** containing the report file. A merged PR is public and permanent. Requires the `gh` CLI, authenticated.\n\nIt refuses to run unattended: with no terminal to confirm on, it exits rather than publishing. `--yes` is available for scripted use and means you accept the upload.\n\n## Permissions, Data Flow & Consent\n\nShort version: steps 1 and 2 are local; step 3 publishes, and so does the optional OpenAI judge. Both ask first. Longer version, because you should not have to take that on trust:\n\n**What it needs, and why.**\n\n| Capability | Why | Scope |\n| --- | --- | --- |\n| Read your memory DB | Counts, histograms, and running the benchmark queries | `~/.openclaw/workspace/db/{memory,cognitive_memory,jarvis}.db`, or `--db` |\n| **Write** your memory DB | `rate.py` logs one `retrieval_log` row per benchmark query | Same DB (or the copy you pass to `--db`); no other table is touched |\n| File write | `<db_dir>/.memory-bench-instance-id` (random UUID, so repeat reports group together) and the `--output` report path | Two files, both of which you can delete |\n| Read usage files | Token/cost totals — **only** with `--include-token-stats` | `~/.openclaw/workspace/memory/*-usage.json` |\n| Local shell exec | `git log` for the algorithm SHA, `node --version` for system info, `git`/`gh` in `submit.sh` | Fixed commands |\n| Network to localhost | Local judge embeddings | `http://127.0.0.1:8900/embed` — stays on the machine |\n| Network to OpenAI | **Opt-in.** Query + up to 300 redacted chars per retrieved memory | `api.openai.com`, only with `--judge openai` after confirmation |\n| Network to GitHub | **Opt-in.** Uploads the report and opens a public PR | `github.com`, only in `submit.sh` after confirmation |\n| Credentials | `OPENAI_API_KEY` (env or `--api-key`) only when you choose the OpenAI judge; `gh`'s existing login in `submit.sh` | Read at call time, never stored, never written to the report |\n| Scheduling | **None.** No daemon, no cron, no install hook. Nothing runs unless you run it | — |\n\n**About that redaction.** Before an excerpt goes to OpenAI, `rate.py` replaces email addresses, phone-shaped numbers, API-key-shaped strings, long hex blobs, home directory paths and URL credentials with placeholders. That is pattern matching, not comprehension — it will not catch a secret written in prose. Treat it as a seatbelt, not a force field. If the content is genuinely sensitive, the answer is `--judge local`, not better regexes.\n\n**Consent, concretely.** Two actions can leave your machine, and neither happens by accident:\n\n```bash\n# 1. External judging — prints exactly what will be sent, then waits\npython3 scripts/rate.py --judge openai        # asks you to type 'send'\npython3 scripts/rate.py --judge openai --yes-send-to-openai   # scripted consent\n\n# 2. Publishing — prints every field that becomes public, then waits\nscripts/submit.sh report.json --dry-run       # preview only, uploads nothing\nscripts/submit.sh report.json                 # asks you to type 'publish'\n```\n\nBoth refuse outright when there is no terminal to confirm on, so an agent running this unattended cannot publish on your behalf.\n\n**Turning it off, and undoing it:**\n\n```bash\n# never transmit: just use the default judge and skip step 3\npython3 scripts/rate.py --judge local\n\n# benchmark a throwaway copy instead of your live memory\ncp ~/.openclaw/workspace/db/memory.db /tmp/bench.db\npython3 scripts/rate.py --db /tmp/bench.db --judge local\n\n# forget this installation ever ran\nrm ~/.openclaw/workspace/db/.memory-bench-instance-id\nsqlite3 ~/.openclaw/workspace/db/memory.db 'DROP TABLE retrieval_log;'\n```\n\nDeleting the instance ID file makes your next report a new anonymous instance; it also breaks the longitudinal link, which is the trade.\n\n**Read it before you run it.** Three scripts, all plain text, none of them long. `rate.py` is the only one that can talk to a third party and `submit.sh` is the only one that can publish — both are worth the two minutes.\n\n## Validation Protocol\n\nFor peer-review-ready data, contributors should:\n\n1. Run `rate.py --ablation` over the full N=30 test set\n2. Use `--judge openai` if you are comfortable with the data flow above — it agrees better with human raters, and the script reports Cohen's κ between the two judges so you can see the gap on your own data. Local-judge submissions are still welcome and are marked as such in the report\n3. Collect at least 2 reports from the same instance, ≥7 days apart (longitudinal)\n4. Report the algorithm version (auto-captured as a short git SHA)\n\n## Test Set Format\n\nCustom test sets are JSON arrays:\n\n```json\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"...\",\n    \"category\": \"semantic|episodic|procedural|strategic\",\n    \"difficulty\": \"easy|medium|hard\"\n  }\n]\n```\n\nAn optional `notes` field is ignored by the runner.\n\n## Included Files\n\n| File | Purpose |\n| --- | --- |\n| `scripts/rate.py` | Runs the benchmark, judges results, computes metrics. Writes `retrieval_log`. The only script that can call an external API — opt-in and confirmed |\n| `scripts/collect.py` | Builds the anonymous statistics report. Uploads nothing |\n| `scripts/submit.sh` | Opens the public PR. Preview with `--dry-run`; requires typed confirmation |\n| `scripts/testset.json` | The 30-query standard test set |\n| `scripts/test_metrics.py` | Unit tests for the IR metrics (`python3 scripts/test_metrics.py`) |\n\nEverything the documentation above describes is in this package. If you find a claim here that the code does not do, that is a bug — open an issue on [the repo](https://github.com/globalcaos/tinkerclaw/issues).\n\n## Agent Workflow\n\nWhen asked to benchmark memory: run `rate.py --ablation` (local judge) and then `collect.py`, and show the summary. **Do not run `submit.sh` on your own initiative** — it publishes to a public repository. Show the user the report, and use `submit.sh --dry-run` so they can see exactly what would become public. Submit only when they say to, and only with `--judge openai` if they have agreed to that separately. Then share the PR link.\n\nFile v2.1.0:_meta.json\n\n{\n  \"ownerId\": \"kn7623hrcwt6rg73a67xw3wyx580asdw\",\n  \"slug\": \"memory-bench-pioneer\",\n  \"version\": \"2.1.0\",\n  \"publishedAt\": 1788774065870\n}\n\nFile v2.1.0:scripts/testset.json\n\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"What preferences are recorded for the development environment setup?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests basic semantic recall of preference-type memories\"\n  },\n  {\n    \"id\": \"T02\",\n    \"query\": \"What happened during the last consolidation cycle?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests temporal/episodic retrieval\"\n  },\n  {\n    \"id\": \"T03\",\n    \"query\": \"How should I handle sensitive data when sharing between agents?\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests procedural knowledge retrieval\"\n  },\n  {\n    \"id\": \"T04\",\n    \"query\": \"recurring patterns in recorded project planning failures\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract/strategic retrieval; phrased to avoid keyword matching\"\n  },\n  {\n    \"id\": \"T05\",\n    \"query\": \"What tools were configured for audio processing?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests factual tool/config recall\"\n  },\n  {\n    \"id\": \"T06\",\n    \"query\": \"what happened, and in what order, during the most recent incident\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests ordered event reconstruction without keyword match\"\n  },\n  {\n    \"id\": \"T07\",\n    \"query\": \"steps to deploy a new version safely\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests multi-step procedural recall\"\n  },\n  {\n    \"id\": \"T08\",\n    \"query\": \"Which external services or integrations are configured?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity-type memory retrieval over configured services\"\n  },\n  {\n    \"id\": \"T09\",\n    \"query\": \"lessons learned from debugging difficult issues\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract lesson retrieval across multiple episodes\"\n  },\n  {\n    \"id\": \"T10\",\n    \"query\": \"What was decided about rate limits or usage quotas?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests topical episodic recall\"\n  },\n  {\n    \"id\": \"T11\",\n    \"query\": \"privacy rules for handling user data\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests policy/procedural recall\"\n  },\n  {\n    \"id\": \"T12\",\n    \"query\": \"connections between automation projects and hardware\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests multi-hop associative retrieval\"\n  },\n  {\n    \"id\": \"T13\",\n    \"query\": \"which scheduled or recurring jobs have run recently\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recurring-job pattern retrieval\"\n  },\n  {\n    \"id\": \"T14\",\n    \"query\": \"error handling best practices\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests general procedural knowledge\"\n  },\n  {\n    \"id\": \"T15\",\n    \"query\": \"which design proposals were recorded for the retrieval pipeline\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent episodic recall with abstract query\"\n  },\n  {\n    \"id\": \"T16\",\n    \"query\": \"tradeoffs between local and cloud solutions\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests comparative/analytical retrieval\"\n  },\n  {\n    \"id\": \"T17\",\n    \"query\": \"which components make up the system and what each one does\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests component/architecture relationship recall\"\n  },\n  {\n    \"id\": \"T18\",\n    \"query\": \"What failed and why in previous attempts?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests failure/negative experience retrieval\"\n  },\n  {\n    \"id\": \"T19\",\n    \"query\": \"configuration steps for new integrations\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests setup/config procedural recall\"\n  },\n  {\n    \"id\": \"T20\",\n    \"query\": \"long-term technical objectives recorded for the memory system\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests goal-tracking retrieval across time\"\n  },\n  {\n    \"id\": \"T21\",\n    \"query\": \"formatting and output conventions that are preferred\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests convention/preference recall\"\n  },\n  {\n    \"id\": \"T22\",\n    \"query\": \"What did the last test or verification run report as failing?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests failure-report episodic recall\"\n  },\n  {\n    \"id\": \"T23\",\n    \"query\": \"how to resolve merge conflicts in the codebase\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests specific technical procedural recall\"\n  },\n  {\n    \"id\": \"T24\",\n    \"query\": \"evolution of the memory architecture over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal strategic retrieval\"\n  },\n  {\n    \"id\": \"T25\",\n    \"query\": \"the most recent benchmark runs and when they happened\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests temporal recall of recent runs\"\n  },\n  {\n    \"id\": \"T26\",\n    \"query\": \"Which file formats or codecs are documented as supported?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests documented-capability recall\"\n  },\n  {\n    \"id\": \"T27\",\n    \"query\": \"backup and recovery procedures\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests disaster recovery procedural recall\"\n  },\n  {\n    \"id\": \"T28\",\n    \"query\": \"how tradeoffs between retrieval accuracy and latency were weighed over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal tradeoff-reasoning retrieval\"\n  },\n  {\n    \"id\": \"T29\",\n    \"query\": \"What was the outcome of the last research task?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent task outcome retrieval\"\n  },\n  {\n    \"id\": \"T30\",\n    \"query\": \"communication style preferences for different channels\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests nuanced preference retrieval\"\n  }\n]\n\nFile v2.1.0:skill-card.md\n\n## Description:\n\nBenchmarks an agent memory system with standardized retrieval queries, relevance judging, IR metrics, confidence intervals, ablations, and optional anonymous result collection for public research submission.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[globalcaos](https://clawhub.ai/user/globalcaos)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and engineers use this skill to measure retrieval quality in OpenClaw or TinkerClaw-style memory systems, compare local and external judge results, and produce JSON benchmark reports. It is intended for users who can review database and publication side effects before running the scripts.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The benchmark reads a live memory database and rate.py writes retrieval_log rows to the database selected by the user.\n\nMitigation: Run against a copied or filtered database with --db when benchmarking sensitive or production memory stores.\n\nRisk: The optional OpenAI judge sends benchmark queries and up to 300 redacted characters from retrieved memories to api.openai.com.\n\nMitigation: Use the default local judge for sensitive memories, and choose --judge openai only after reviewing the preview and consenting to the data flow.\n\nRisk: The public submission path can publish statistics through a GitHub pull request and evidence.security flags enough privacy and code-execution risk to require careful review.\n\nMitigation: Inspect the generated JSON and use submit.sh --dry-run first; do not run submit.sh on untrusted or hand-edited reports until the identified submission-path issues are fixed.\n\n## Reference(s):\n\n- [ClawHub Skill Page](https://clawhub.ai/globalcaos/skills/memory-bench-pioneer)\n- [TinkerClaw Repository](https://github.com/globalcaos/tinkerclaw)\n- [TinkerClaw Issues](https://github.com/globalcaos/tinkerclaw/issues)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with shell commands and JSON benchmark report files]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Produces local retrieval metrics, optional ablation summaries, and optional public submission previews; external judging and publication require explicit user action.]\n\n## Skill Version(s):\n\n2.1.0 (source: server release evidence and frontmatter)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v2.0.1: 7 files, 19267 bytes\n\nFiles: _meta.json (139b), scripts/collect.py (15569b), scripts/rate.py (19413b), scripts/submit.sh (6110b), scripts/test_metrics.py (4009b), scripts/testset.json (5975b), SKILL.md (3378b)\n\nFile v2.0.1:SKILL.md\n\n---\nname: memory-bench-pioneer\ndescription: \"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Runs a peer-review-grade evaluation suite (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablation studies) against your live memory system and submits anonymized results to the ENGRAM/CORTEX research papers. Your data stays private; only aggregate stats leave. Works with agent-memory-ultimate. For the bold few who believe AI memory should be measured, not guessed at. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw.\"\n---\n\n# Memory Bench\n\nCollect, assess, and submit anonymized memory system statistics for the ENGRAM and CORTEX research papers.\n\n## Three-Step Pipeline\n\n### 1. Assess Retrieval Quality\n\nRun the standard test set (30 queries across 4 types × 3 difficulty levels) with LLM-as-judge:\n\n```bash\n# Full assessment with GPT-4o-mini judge + ablation (recommended)\npython3 scripts/rate.py --queries 30 --judge openai --ablation\n\n# Without OpenAI key: local embedding judge (weaker, marked in output)\npython3 scripts/rate.py --queries 30 --judge local --ablation\n\n# Custom test set\npython3 scripts/rate.py --testset path/to/queries.json --judge openai\n```\n\n**What it measures:**\n\n- **RAR** (Recall Accuracy Ratio), **MRR** (Mean Reciprocal Rank)\n- **nDCG@5**, **MAP@5**, **Precision@5**, **Hit Rate**\n- All metrics include **95% bootstrap confidence intervals**\n- **Ablation**: runs with AND without spreading activation to isolate its contribution\n\n**Judge methods:**\n\n- `openai` — GPT-4o-mini rates each (query, result) pair 1-5. Independent from retrieval system. ~$0.01 per run.\n- `local` — Embedding cosine similarity. Weaker, marked as such in output. Zero cost.\n\n**Standard test set** (`scripts/testset.json`): 30 queries stratified across semantic/episodic/procedural/strategic types and easy/medium/hard difficulty. No lexical overlap with stored memories. All deployments run the same queries for cross-site comparability.\n\n### 2. Collect Statistics\n\n```bash\npython3 scripts/collect.py --contributor GITHUB_USER --days 14 --output /tmp/memory-bench-report.json\n```\n\n**Collected (anonymized):** Memory counts/types/ages, strength/importance histograms, association graph size, hierarchy levels, consolidation history, retrieval metrics (RAR/MRR/nDCG/MAP with CIs), ablation results, judge method, algorithm version, embedding coverage. Instance ID is a random UUID (not reversible).\n\n**Never collected:** Memory content, queries, file paths, usernames, hostnames.\n\n### 3. Submit as PR\n\n```bash\nscripts/submit.sh /tmp/memory-bench-report.json GITHUB_USERNAME\n```\n\nForks, branches, places report, updates INDEX.json, opens PR. Requires `gh` CLI.\n\n## Validation Protocol\n\nFor peer-review-ready data, contributors should:\n\n1. Run `rate.py --ablation --judge openai` (minimum N=30 queries)\n2. Collect at least 2 reports from the same instance, ≥7 days apart (longitudinal)\n3. Report the algorithm version (auto-captured from git)\n\n## Test Set Format\n\nCustom test sets are JSON arrays:\n\n```json\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"...\",\n    \"category\": \"semantic|episodic|procedural|strategic\",\n    \"difficulty\": \"easy|medium|hard\"\n  }\n]\n```\n\n## Agent Workflow\n\nWhen asked to submit benchmarks: run `rate.py --ablation --judge openai`, then `collect.py`, review summary, then `submit.sh`. Share the PR link.\n\nFile v2.0.1:_meta.json\n\n{\n  \"ownerId\": \"kn7623hrcwt6rg73a67xw3wyx580asdw\",\n  \"slug\": \"memory-bench-pioneer\",\n  \"version\": \"2.0.1\",\n  \"publishedAt\": 1780764682217\n}\n\nFile v2.0.1:scripts/testset.json\n\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"What preferences does the user have for their home setup?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests basic semantic recall of preference-type memories\"\n  },\n  {\n    \"id\": \"T02\",\n    \"query\": \"What happened during the last consolidation cycle?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests temporal/episodic retrieval\"\n  },\n  {\n    \"id\": \"T03\",\n    \"query\": \"How should I handle sensitive data when sharing between agents?\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests procedural knowledge retrieval\"\n  },\n  {\n    \"id\": \"T04\",\n    \"query\": \"recurring patterns in project planning failures\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract/strategic memory retrieval with no lexical overlap\"\n  },\n  {\n    \"id\": \"T05\",\n    \"query\": \"What tools were configured for audio processing?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests factual tool/config recall\"\n  },\n  {\n    \"id\": \"T06\",\n    \"query\": \"emotional context of recent family discussions\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests retrieval requiring semantic understanding, no keyword match\"\n  },\n  {\n    \"id\": \"T07\",\n    \"query\": \"steps to deploy a new version safely\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests multi-step procedural recall\"\n  },\n  {\n    \"id\": \"T08\",\n    \"query\": \"Which contacts are in the professional network?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity-type memory retrieval\"\n  },\n  {\n    \"id\": \"T09\",\n    \"query\": \"lessons learned from debugging difficult issues\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract lesson retrieval across multiple episodes\"\n  },\n  {\n    \"id\": \"T10\",\n    \"query\": \"What was discussed about budget and spending limits?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests topical episodic recall\"\n  },\n  {\n    \"id\": \"T11\",\n    \"query\": \"privacy rules for handling user data\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests policy/procedural recall\"\n  },\n  {\n    \"id\": \"T12\",\n    \"query\": \"connections between automation projects and hardware\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests multi-hop associative retrieval\"\n  },\n  {\n    \"id\": \"T13\",\n    \"query\": \"morning routine and daily schedule patterns\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests habitual pattern retrieval\"\n  },\n  {\n    \"id\": \"T14\",\n    \"query\": \"error handling best practices\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests general procedural knowledge\"\n  },\n  {\n    \"id\": \"T15\",\n    \"query\": \"What creative ideas have been proposed recently?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent episodic recall with abstract query\"\n  },\n  {\n    \"id\": \"T16\",\n    \"query\": \"tradeoffs between local and cloud solutions\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests comparative/analytical retrieval\"\n  },\n  {\n    \"id\": \"T17\",\n    \"query\": \"who are the family members and their roles\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity relationship recall\"\n  },\n  {\n    \"id\": \"T18\",\n    \"query\": \"What failed and why in previous attempts?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests failure/negative experience retrieval\"\n  },\n  {\n    \"id\": \"T19\",\n    \"query\": \"configuration steps for new integrations\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests setup/config procedural recall\"\n  },\n  {\n    \"id\": \"T20\",\n    \"query\": \"long-term goals and their current progress\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests goal-tracking retrieval across time\"\n  },\n  {\n    \"id\": \"T21\",\n    \"query\": \"dietary preferences and health considerations\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests personal preference recall\"\n  },\n  {\n    \"id\": \"T22\",\n    \"query\": \"What security incidents or vulnerabilities were found?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests security-related episodic recall\"\n  },\n  {\n    \"id\": \"T23\",\n    \"query\": \"how to resolve merge conflicts in the codebase\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests specific technical procedural recall\"\n  },\n  {\n    \"id\": \"T24\",\n    \"query\": \"evolution of the memory architecture over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal strategic retrieval\"\n  },\n  {\n    \"id\": \"T25\",\n    \"query\": \"meetings and appointments scheduled this month\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests calendar/temporal recall\"\n  },\n  {\n    \"id\": \"T26\",\n    \"query\": \"What music or media does the user enjoy?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests lifestyle preference recall\"\n  },\n  {\n    \"id\": \"T27\",\n    \"query\": \"backup and recovery procedures\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests disaster recovery procedural recall\"\n  },\n  {\n    \"id\": \"T28\",\n    \"query\": \"relationship dynamics between team members\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests interpersonal/social graph retrieval\"\n  },\n  {\n    \"id\": \"T29\",\n    \"query\": \"What was the outcome of the last research task?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent task outcome retrieval\"\n  },\n  {\n    \"id\": \"T30\",\n    \"query\": \"communication style preferences for different channels\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests nuanced preference retrieval\"\n  }\n]\n\nArchive v2.0.0: 7 files, 19226 bytes\n\nFiles: scripts/collect.py (15569b), scripts/rate.py (19413b), scripts/submit.sh (6110b), scripts/test_metrics.py (4009b), scripts/testset.json (5975b), SKILL.md (3310b), _meta.json (139b)\n\nFile v2.0.0:SKILL.md\n\n---\nname: memory-bench-pioneer\ndescription: \"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Runs a peer-review-grade evaluation suite (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablation studies) against your live memory system and submits anonymized results to the ENGRAM/CORTEX research papers. Your data stays private; only aggregate stats leave. Works with agent-memory-ultimate. For the bold few who believe AI memory should be measured, not guessed at.\"\n---\n\n# Memory Bench\n\nCollect, assess, and submit anonymized memory system statistics for the ENGRAM and CORTEX research papers.\n\n## Three-Step Pipeline\n\n### 1. Assess Retrieval Quality\n\nRun the standard test set (30 queries across 4 types × 3 difficulty levels) with LLM-as-judge:\n\n```bash\n# Full assessment with GPT-4o-mini judge + ablation (recommended)\npython3 scripts/rate.py --queries 30 --judge openai --ablation\n\n# Without OpenAI key: local embedding judge (weaker, marked in output)\npython3 scripts/rate.py --queries 30 --judge local --ablation\n\n# Custom test set\npython3 scripts/rate.py --testset path/to/queries.json --judge openai\n```\n\n**What it measures:**\n\n- **RAR** (Recall Accuracy Ratio), **MRR** (Mean Reciprocal Rank)\n- **nDCG@5**, **MAP@5**, **Precision@5**, **Hit Rate**\n- All metrics include **95% bootstrap confidence intervals**\n- **Ablation**: runs with AND without spreading activation to isolate its contribution\n\n**Judge methods:**\n\n- `openai` — GPT-4o-mini rates each (query, result) pair 1-5. Independent from retrieval system. ~$0.01 per run.\n- `local` — Embedding cosine similarity. Weaker, marked as such in output. Zero cost.\n\n**Standard test set** (`scripts/testset.json`): 30 queries stratified across semantic/episodic/procedural/strategic types and easy/medium/hard difficulty. No lexical overlap with stored memories. All deployments run the same queries for cross-site comparability.\n\n### 2. Collect Statistics\n\n```bash\npython3 scripts/collect.py --contributor GITHUB_USER --days 14 --output /tmp/memory-bench-report.json\n```\n\n**Collected (anonymized):** Memory counts/types/ages, strength/importance histograms, association graph size, hierarchy levels, consolidation history, retrieval metrics (RAR/MRR/nDCG/MAP with CIs), ablation results, judge method, algorithm version, embedding coverage. Instance ID is a random UUID (not reversible).\n\n**Never collected:** Memory content, queries, file paths, usernames, hostnames.\n\n### 3. Submit as PR\n\n```bash\nscripts/submit.sh /tmp/memory-bench-report.json GITHUB_USERNAME\n```\n\nForks, branches, places report, updates INDEX.json, opens PR. Requires `gh` CLI.\n\n## Validation Protocol\n\nFor peer-review-ready data, contributors should:\n\n1. Run `rate.py --ablation --judge openai` (minimum N=30 queries)\n2. Collect at least 2 reports from the same instance, ≥7 days apart (longitudinal)\n3. Report the algorithm version (auto-captured from git)\n\n## Test Set Format\n\nCustom test sets are JSON arrays:\n\n```json\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"...\",\n    \"category\": \"semantic|episodic|procedural|strategic\",\n    \"difficulty\": \"easy|medium|hard\"\n  }\n]\n```\n\n## Agent Workflow\n\nWhen asked to submit benchmarks: run `rate.py --ablation --judge openai`, then `collect.py`, review summary, then `submit.sh`. Share the PR link.\n\nFile v2.0.0:_meta.json\n\n{\n  \"ownerId\": \"kn7623hrcwt6rg73a67xw3wyx580asdw\",\n  \"slug\": \"memory-bench-pioneer\",\n  \"version\": \"2.0.0\",\n  \"publishedAt\": 1771282629924\n}\n\nFile v2.0.0:scripts/testset.json\n\n[\n  {\n    \"id\": \"T01\",\n    \"query\": \"What preferences does the user have for their home setup?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests basic semantic recall of preference-type memories\"\n  },\n  {\n    \"id\": \"T02\",\n    \"query\": \"What happened during the last consolidation cycle?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests temporal/episodic retrieval\"\n  },\n  {\n    \"id\": \"T03\",\n    \"query\": \"How should I handle sensitive data when sharing between agents?\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests procedural knowledge retrieval\"\n  },\n  {\n    \"id\": \"T04\",\n    \"query\": \"recurring patterns in project planning failures\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract/strategic memory retrieval with no lexical overlap\"\n  },\n  {\n    \"id\": \"T05\",\n    \"query\": \"What tools were configured for audio processing?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests factual tool/config recall\"\n  },\n  {\n    \"id\": \"T06\",\n    \"query\": \"emotional context of recent family discussions\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests retrieval requiring semantic understanding, no keyword match\"\n  },\n  {\n    \"id\": \"T07\",\n    \"query\": \"steps to deploy a new version safely\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests multi-step procedural recall\"\n  },\n  {\n    \"id\": \"T08\",\n    \"query\": \"Which contacts are in the professional network?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity-type memory retrieval\"\n  },\n  {\n    \"id\": \"T09\",\n    \"query\": \"lessons learned from debugging difficult issues\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract lesson retrieval across multiple episodes\"\n  },\n  {\n    \"id\": \"T10\",\n    \"query\": \"What was discussed about budget and spending limits?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests topical episodic recall\"\n  },\n  {\n    \"id\": \"T11\",\n    \"query\": \"privacy rules for handling user data\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests policy/procedural recall\"\n  },\n  {\n    \"id\": \"T12\",\n    \"query\": \"connections between automation projects and hardware\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests multi-hop associative retrieval\"\n  },\n  {\n    \"id\": \"T13\",\n    \"query\": \"morning routine and daily schedule patterns\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests habitual pattern retrieval\"\n  },\n  {\n    \"id\": \"T14\",\n    \"query\": \"error handling best practices\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests general procedural knowledge\"\n  },\n  {\n    \"id\": \"T15\",\n    \"query\": \"What creative ideas have been proposed recently?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent episodic recall with abstract query\"\n  },\n  {\n    \"id\": \"T16\",\n    \"query\": \"tradeoffs between local and cloud solutions\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests comparative/analytical retrieval\"\n  },\n  {\n    \"id\": \"T17\",\n    \"query\": \"who are the family members and their roles\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity relationship recall\"\n  },\n  {\n    \"id\": \"T18\",\n    \"query\": \"What failed and why in previous attempts?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests failure/negative experience retrieval\"\n  },\n  {\n    \"id\": \"T19\",\n    \"query\": \"configuration steps for new integrations\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests setup/config procedural recall\"\n  },\n  {\n    \"id\": \"T20\",\n    \"query\": \"long-term goals and their current progress\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests goal-tracking retrieval across time\"\n  },\n  {\n    \"id\": \"T21\",\n    \"query\": \"dietary preferences and health considerations\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests personal preference recall\"\n  },\n  {\n    \"id\": \"T22\",\n    \"query\": \"What security incidents or vulnerabilities were found?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests security-related episodic recall\"\n  },\n  {\n    \"id\": \"T23\",\n    \"query\": \"how to resolve merge conflicts in the codebase\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests specific technical procedural recall\"\n  },\n  {\n    \"id\": \"T24\",\n    \"query\": \"evolution of the memory architecture over time\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests longitudinal strategic retrieval\"\n  },\n  {\n    \"id\": \"T25\",\n    \"query\": \"meetings and appointments scheduled this month\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests calendar/temporal recall\"\n  },\n  {\n    \"id\": \"T26\",\n    \"query\": \"What music or media does the user enjoy?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests lifestyle preference recall\"\n  },\n  {\n    \"id\": \"T27\",\n    \"query\": \"backup and recovery procedures\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests disaster recovery procedural recall\"\n  },\n  {\n    \"id\": \"T28\",\n    \"query\": \"relationship dynamics between team members\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests interpersonal/social graph retrieval\"\n  },\n  {\n    \"id\": \"T29\",\n    \"query\": \"What was the outcome of the last research task?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recent task outcome retrieval\"\n  },\n  {\n    \"id\": \"T30\",\n    \"query\": \"communication style preferences for different channels\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests nuanced preference retrieval\"\n  }\n]","readmeExcerpt":"Skill: TinkerClaw Memory Bench Owner: globalcaos Summary: Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine, and excerpts are redacted even on the local path. The optional OpenAI judge is opt-in","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"# Local judge — the default. Nothing leaves your machine.\npython3 scripts/rate.py --queries 30 --judge local --ablation\n\n# Stronger judge, but it TRANSMITS retrieved memory excerpts to OpenAI.\n# You will be asked to type 'send' before the first request.\npython3 scripts/rate.py --queries 30 --judge openai --ablation\n\n# Benchmark a copy instead of your live database\npython3 scripts/rate.py --db /tmp/memory-copy.db --judge local\n\n# Custom test set\npython3 scripts/rate.py --testset path/to/queries.json --judge local"},{"language":"bash","snippet":"# Anonymous — the default\npython3 scripts/collect.py --days 14 --output /tmp/memory-bench-report.json\n\n# Attributed to you (your username goes in the report, and it becomes public if you submit)\npython3 scripts/collect.py --days 14 --contributor YOUR_GITHUB_USER --output /tmp/memory-bench-report.json"},{"language":"bash","snippet":"# ALWAYS do this first: shows every field that would become public, uploads nothing\nscripts/submit.sh /tmp/memory-bench-report.json --dry-run\n\n# Real submission — prints the same preview, then asks you to type 'publish'\nscripts/submit.sh /tmp/memory-bench-report.json YOUR_GITHUB_USERNAME"},{"language":"bash","snippet":"# 1. External judging — prints the exact request body, then waits\npython3 scripts/rate.py --judge openai        # asks you to type 'send'\npython3 scripts/rate.py --judge openai --yes-send-to-openai   # skips the keystroke, still needs a terminal\n\n# 2. Publishing — validates, prints every field that becomes public, then waits\nscripts/submit.sh report.json --dry-run       # preview only, uploads nothing\nscripts/submit.sh report.json                 # asks you to type 'publish'"},{"language":"bash","snippet":"# never transmit: just use the default judge and skip step 3\npython3 scripts/rate.py --judge local\n\n# benchmark a throwaway copy instead of your live memory\ncp ~/.openclaw/workspace/db/memory.db /tmp/bench.db\npython3 scripts/rate.py --db /tmp/bench.db --judge local\n\n# forget this installation ever ran\nrm ~/.openclaw/workspace/db/.memory-bench-instance-id\nsqlite3 ~/.openclaw/workspace/db/memory.db 'DROP TABLE retrieval_log;'"},{"language":"json","snippet":"[\n  {\n    \"id\": \"T01\",\n    \"query\": \"...\",\n    \"category\": \"semantic|episodic|procedural|strategic\",\n    \"difficulty\": \"easy|medium|hard\"\n  }\n]"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: memory-bench-pioneer\nversion: 2.1.3\ndescription: \"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine, and excerpts are redacted even on the local path. The optional OpenAI judge is opt-in, prints the exact request body it would send, redacts secrets first, requires typed consent, and cannot be switched on by an unattended run. Submitting results is a separate confirmed step that validates the report against the full schema and previews every field in it, and identifies you only if you pass --contributor. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw. See Permissions, Data Flow & Consent.\"\nmetadata:\n  {\n    \"openclaw\":\n      {\n        \"emoji\": \"🧠\",\n        \"requires\": { \"bins\": [\"python3\"] },\n        \"notes\":\n          {\n            \"security\": \"Benchmarks a LIVE memory database, so disclosure matters more than usual. DEFAULT PATH IS LOCAL: the default judge (--judge local) sends retrieved excerpts only to a local embedding server on 127.0.0.1; nothing leaves the machine, and those excerpts are redacted too because that server keeps its own logs. Two actions can send data out and BOTH are opt-in, confirmed, and IMPOSSIBLE to trigger unattended: --judge openai transmits the benchmark query plus up to 300 redacted characters of each RETRIEVED MEMORY to api.openai.com, printing the exact JSON request body first and then requiring a typed confirmation at a terminal; and scripts/submit.sh opens a PUBLIC GitHub pull request containing the statistics report, after validating it against the complete schema and printing every field, the commit message and the PR body (--dry-run to preview, typed confirmation to publish). --yes-send-to-openai and --yes skip only the keystroke; both still require a terminal, so an agent running this on its own cannot transmit or publish. Attribution is anonymous unless you pass --contributor; token/cost totals are excluded unless you pass --include-token-stats. rate.py WRITES a retrieval_log table into the database you point it at — use --db on a copy to sandbox it. No daemon, no cron, nothing runs on its own. See the Permissions, Data Flow & Consent section.\"\n          },\n      },\n  }\n---\n\n# Memory Bench\n\n> One of dozens of skills and plugins in **[TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — a self-improving OpenClaw fork that's been running 24/7 for months.\n\nEveryone has an opinion about whether their agent's memory is any good. Almost nobody has a number.\n\nThis produces the number — nDCG, MAP, MRR, Precision@5, each with a 95% bootstrap confidence interval, plus an ablation that isolates what spreading activation actually contributes. Then, if you want, it contributes your (anonymous) results to the ENGRAM and CORTEX research papers, where a few dozen real d"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7623hrcwt6rg73a67xw3wyx580asdw\",\n  \"slug\": \"memory-bench-pioneer\",\n  \"version\": \"2.1.3\",\n  \"publishedAt\": 1788944653102\n}"},{"path":"scripts/testset.json","content":"[\n  {\n    \"id\": \"T01\",\n    \"query\": \"What preferences are recorded for the development environment setup?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests basic semantic recall of preference-type memories\"\n  },\n  {\n    \"id\": \"T02\",\n    \"query\": \"What happened during the last consolidation cycle?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests temporal/episodic retrieval\"\n  },\n  {\n    \"id\": \"T03\",\n    \"query\": \"How should I handle sensitive data when sharing between agents?\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests procedural knowledge retrieval\"\n  },\n  {\n    \"id\": \"T04\",\n    \"query\": \"recurring patterns in recorded project planning failures\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract/strategic retrieval; phrased to avoid keyword matching\"\n  },\n  {\n    \"id\": \"T05\",\n    \"query\": \"What tools were configured for audio processing?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests factual tool/config recall\"\n  },\n  {\n    \"id\": \"T06\",\n    \"query\": \"what happened, and in what order, during the most recent incident\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests ordered event reconstruction without keyword match\"\n  },\n  {\n    \"id\": \"T07\",\n    \"query\": \"steps to deploy a new version safely\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests multi-step procedural recall\"\n  },\n  {\n    \"id\": \"T08\",\n    \"query\": \"Which external services or integrations are configured?\",\n    \"category\": \"semantic\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests entity-type memory retrieval over configured services\"\n  },\n  {\n    \"id\": \"T09\",\n    \"query\": \"lessons learned from debugging difficult issues\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests abstract lesson retrieval across multiple episodes\"\n  },\n  {\n    \"id\": \"T10\",\n    \"query\": \"What was decided about rate limits or usage quotas?\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests topical episodic recall\"\n  },\n  {\n    \"id\": \"T11\",\n    \"query\": \"privacy rules for handling user data\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"easy\",\n    \"notes\": \"Tests policy/procedural recall\"\n  },\n  {\n    \"id\": \"T12\",\n    \"query\": \"connections between automation projects and hardware\",\n    \"category\": \"strategic\",\n    \"difficulty\": \"hard\",\n    \"notes\": \"Tests multi-hop associative retrieval\"\n  },\n  {\n    \"id\": \"T13\",\n    \"query\": \"which scheduled or recurring jobs have run recently\",\n    \"category\": \"episodic\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests recurring-job pattern retrieval\"\n  },\n  {\n    \"id\": \"T14\",\n    \"query\": \"error handling best practices\",\n    \"category\": \"procedural\",\n    \"difficulty\": \"medium\",\n    \"notes\": \"Tests general procedural knowledge\"\n  },\n  {\n    \"id\": \"T15\",\n    \"query\": \"which design proposals were recorded for the retrieval"},{"path":"skill-card.md","content":"## Description:\n\nTinkerClaw Memory Bench benchmarks an OpenClaw or TinkerClaw memory database locally by default, producing retrieval-quality metrics with confidence intervals and optional, consent-gated OpenAI judging and public report submission.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[globalcaos](https://clawhub.ai/user/globalcaos)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and agent operators use this skill to measure retrieval quality for a live or copied OpenClaw or TinkerClaw memory database and collect aggregate benchmark reports. It supports local default runs, optional external LLM judging, and optional public submission after preview and confirmation.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Benchmarking can touch sensitive local memory data from a live memory database.\n\nMitigation: Use the default local judge and run against a copied or filtered database when contents are sensitive.\n\nRisk: The optional OpenAI judge can send redacted benchmark queries and retrieved memory excerpts to OpenAI.\n\nMitigation: Use it only after reviewing the printed request preview and confirming that redacted excerpts may leave the machine; prefer OPENAI_API_KEY over command-line keys.\n\nRisk: Submitting results opens a public GitHub pull request with aggregate report metadata.\n\nMitigation: Run submit.sh with --dry-run first, review every field in the schema-validated preview, and publish only after explicit confirmation.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/globalcaos/skills/memory-bench-pioneer)\n- [Publisher profile](https://clawhub.ai/user/globalcaos)\n- [TinkerClaw project link from skill documentation](https://github.com/globalcaos/tinkerclaw)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with shell commands and JSON report files]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Produces local benchmark metrics and aggregate reports; external judging and public submission are optional and confirmation-gated.]\n\n## Skill Version(s):\n\n2.1.3 (source: frontmatter and server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1805,"uniquenessScore":44,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T06:14:48.633Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T06:14:48.633Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T08:41:46.320Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}