agentCLAWHUBUnverified

TinkerClaw Memory Bench

Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine, and excerpts are redacted even on the local path. The optional OpenAI judge is opt-in, prints the exact request body it would send, redacts secrets first, requires typed consent, and cannot be switched on by an unattended run. Submitting results is a separate confirmed step that validates the report against the full schema and previews every field in it, and identifies you only if you pass --contributor. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw. See Permissions, Data Flow & Consent.

OpenClaw

Rank

62

Safety

84

Downloads

1.1k

Updated

Oct 11, 2026

Version

2.1.3

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.1K downloadsadoption · observed Oct 11, 2026
Latest release
2.1.3release · observed Sep 9, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17324vfeqe0ptzp84z2z9vttx883zdg:memory-bench-pioneer
  1. Install using `clawhub skill install s17324vfeqe0ptzp84z2z9vttx883zdg:memory-bench-pioneer` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/globalcaos/memory-bench-pioneer before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/snapshot"

Documentation

CLAWHUB

115,875 characters of source documentation, loaded on request.

Extracted files

4 files captured from the source.

SKILL.md

---
name: memory-bench-pioneer
version: 2.1.3
description: "Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine, and excerpts are redacted even on the local path. The optional OpenAI judge is opt-in, prints the exact request body it would send, redacts secrets first, requires typed consent, and cannot be switched on by an unattended run. Submitting results is a separate confirmed step that validates the report against the full schema and previews every field in it, and identifies you only if you pass --contributor. Built for the TinkerClaw fork — github.com/globalcaos/tinkerclaw. See Permissions, Data Flow & Consent."
metadata:
  {
    "openclaw":
      {
        "emoji": "🧠",
        "requires": { "bins": ["python3"] },
        "notes":
          {
            "security": "Benchmarks a LIVE memory database, so disclosure matters more than usual. DEFAULT PATH IS LOCAL: the default judge (--judge local) sends retrieved excerpts only to a local embedding server on 127.0.0.1; nothing leaves the machine, and those excerpts are redacted too because that server keeps its own logs. Two actions can send data out and BOTH are opt-in, confirmed, and IMPOSSIBLE to trigger unattended: --judge openai transmits the benchmark query plus up to 300 redacted characters of each RETRIEVED MEMORY to api.openai.com, printing the exact JSON request body first and then requiring a typed confirmation at a terminal; and scripts/submit.sh opens a PUBLIC GitHub pull request containing the statistics report, after validating it against the complete schema and printing every field, the commit message and the PR body (--dry-run to preview, typed confirmation to publish). --yes-send-to-openai and --yes skip only the keystroke; both still require a terminal, so an agent running this on its own cannot transmit or publish. Attribution is anonymous unless you pass --contributor; token/cost totals are excluded unless you pass --include-token-stats. rate.py WRITES a retrieval_log table into the database you point it at — use --db on a copy to sandbox it. No daemon, no cron, nothing runs on its own. See the Permissions, Data Flow & Consent section."
          },
      },
  }
---

# Memory Bench

> One of dozens of skills and plugins in **[TinkerClaw](https://github.com/globalcaos/tinkerclaw)** — a self-improving OpenClaw fork that's been running 24/7 for months.

Everyone has an opinion about whether their agent's memory is any good. Almost nobody has a number.

This produces the number — nDCG, MAP, MRR, Precision@5, each with a 95% bootstrap confidence interval, plus an ablation that isolates what spreading activation actually contributes. Then, if you want, it contributes your (anonymous) results to the ENGRAM and CORTEX research papers, where a few dozen real d

_meta.json

{
  "ownerId": "kn7623hrcwt6rg73a67xw3wyx580asdw",
  "slug": "memory-bench-pioneer",
  "version": "2.1.3",
  "publishedAt": 1788944653102
}

scripts/testset.json

[
  {
    "id": "T01",
    "query": "What preferences are recorded for the development environment setup?",
    "category": "semantic",
    "difficulty": "easy",
    "notes": "Tests basic semantic recall of preference-type memories"
  },
  {
    "id": "T02",
    "query": "What happened during the last consolidation cycle?",
    "category": "episodic",
    "difficulty": "medium",
    "notes": "Tests temporal/episodic retrieval"
  },
  {
    "id": "T03",
    "query": "How should I handle sensitive data when sharing between agents?",
    "category": "procedural",
    "difficulty": "medium",
    "notes": "Tests procedural knowledge retrieval"
  },
  {
    "id": "T04",
    "query": "recurring patterns in recorded project planning failures",
    "category": "strategic",
    "difficulty": "hard",
    "notes": "Tests abstract/strategic retrieval; phrased to avoid keyword matching"
  },
  {
    "id": "T05",
    "query": "What tools were configured for audio processing?",
    "category": "semantic",
    "difficulty": "easy",
    "notes": "Tests factual tool/config recall"
  },
  {
    "id": "T06",
    "query": "what happened, and in what order, during the most recent incident",
    "category": "episodic",
    "difficulty": "hard",
    "notes": "Tests ordered event reconstruction without keyword match"
  },
  {
    "id": "T07",
    "query": "steps to deploy a new version safely",
    "category": "procedural",
    "difficulty": "medium",
    "notes": "Tests multi-step procedural recall"
  },
  {
    "id": "T08",
    "query": "Which external services or integrations are configured?",
    "category": "semantic",
    "difficulty": "easy",
    "notes": "Tests entity-type memory retrieval over configured services"
  },
  {
    "id": "T09",
    "query": "lessons learned from debugging difficult issues",
    "category": "strategic",
    "difficulty": "hard",
    "notes": "Tests abstract lesson retrieval across multiple episodes"
  },
  {
    "id": "T10",
    "query": "What was decided about rate limits or usage quotas?",
    "category": "episodic",
    "difficulty": "medium",
    "notes": "Tests topical episodic recall"
  },
  {
    "id": "T11",
    "query": "privacy rules for handling user data",
    "category": "procedural",
    "difficulty": "easy",
    "notes": "Tests policy/procedural recall"
  },
  {
    "id": "T12",
    "query": "connections between automation projects and hardware",
    "category": "strategic",
    "difficulty": "hard",
    "notes": "Tests multi-hop associative retrieval"
  },
  {
    "id": "T13",
    "query": "which scheduled or recurring jobs have run recently",
    "category": "episodic",
    "difficulty": "medium",
    "notes": "Tests recurring-job pattern retrieval"
  },
  {
    "id": "T14",
    "query": "error handling best practices",
    "category": "procedural",
    "difficulty": "medium",
    "notes": "Tests general procedural knowledge"
  },
  {
    "id": "T15",
    "query": "which design proposals were recorded for the retrieval

skill-card.md

## Description:

TinkerClaw Memory Bench benchmarks an OpenClaw or TinkerClaw memory database locally by default, producing retrieval-quality metrics with confidence intervals and optional, consent-gated OpenAI judging and public report submission.

This skill is ready for commercial/non-commercial use.

## Publisher:

[globalcaos](https://clawhub.ai/user/globalcaos)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and agent operators use this skill to measure retrieval quality for a live or copied OpenClaw or TinkerClaw memory database and collect aggregate benchmark reports. It supports local default runs, optional external LLM judging, and optional public submission after preview and confirmation.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: Benchmarking can touch sensitive local memory data from a live memory database.

Mitigation: Use the default local judge and run against a copied or filtered database when contents are sensitive.

Risk: The optional OpenAI judge can send redacted benchmark queries and retrieved memory excerpts to OpenAI.

Mitigation: Use it only after reviewing the printed request preview and confirming that redacted excerpts may leave the machine; prefer OPENAI_API_KEY over command-line keys.

Risk: Submitting results opens a public GitHub pull request with aggregate report metadata.

Mitigation: Run submit.sh with --dry-run first, review every field in the schema-validated preview, and publish only after explicit confirmation.

## Reference(s):

- [ClawHub skill page](https://clawhub.ai/globalcaos/skills/memory-bench-pioneer)
- [Publisher profile](https://clawhub.ai/user/globalcaos)
- [TinkerClaw project link from skill documentation](https://github.com/globalcaos/tinkerclaw)

## Skill Output:

**Output Type(s):** [text, markdown, shell commands, configuration, guidance]

**Output Format:** [Markdown guidance with shell commands and JSON report files]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Produces local benchmark metrics and aggregate reports; external judging and public submission are optional and confirmation-gated.]

## Skill Version(s):

2.1.3 (source: frontmatter and server release evidence)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 2d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/globalcaos/skills/memory-bench-pioneer",
      "sourceUrl": "https://clawhub.ai/globalcaos/skills/memory-bench-pioneer",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T06:14:48.633Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T06:14:48.633Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.1K downloads",
      "href": "https://clawhub.ai/globalcaos/memory-bench-pioneer",
      "sourceUrl": "https://clawhub.ai/globalcaos/memory-bench-pioneer",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T06:14:48.633Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "2.1.3",
      "href": "https://clawhub.ai/globalcaos/memory-bench-pioneer",
      "sourceUrl": "https://clawhub.ai/globalcaos/memory-bench-pioneer",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-09T09:04:13.102Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-globalcaos-memory-bench-pioneer/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 2.1.3",
      "description": "Exact request-body preview before OpenAI consent; schema-validated public report; no unattended bypass.",
      "href": "https://clawhub.ai/globalcaos/memory-bench-pioneer",
      "sourceUrl": "https://clawhub.ai/globalcaos/memory-bench-pioneer",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-09T09:04:13.102Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to TinkerClaw Memory Bench and adjacent AI workflows.