adversarial-code-loop
BUILD → REVIEW → (FIX → VERIFY)^N → ARBITER on isolated git branches. Git-native: each loop runs on its own branch, changes are committed, reviews inspect git diffs. Skill: adversarial-code-loop Owner: chpomob Summary: BUILD → REVIEW → (FIX → VERIFY)^N → ARBITER on isolated git branches. Git-native: each loop runs on its own branch, changes are committed, reviews inspect git diffs. Tags: latest:0.1.0 Version history: v0.1.0 | 2026-08-03T18:12:20.116Z | auto Adversarial Code Loop v4 is a major update making the workflow fully git-native. - Each code-review loop runs on its own iso
Rank
62
Safety
84
Downloads
2.6k
Updated
Oct 9, 2026
Version
0.1.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 2.6K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 2.6K downloadsadoption · observed Oct 9, 2026
- Latest release
- 0.1.0release · observed Aug 3, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17435m3chty5jmw4jhpkyhnb58brn8g:adversarial-code-loop- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-chpomob-adversarial-code-loop/snapshot"
Documentation
CLAWHUB
75,803 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: adversarial-code-loop
description: "BUILD → REVIEW → (FIX → VERIFY)^N → ARBITER on isolated git branches. Git-native: each loop runs on its own branch, changes are committed, reviews inspect git diffs."
version: 4.2.0
author: Hermes Agent
license: 0BSD
platforms: [linux, macos]
metadata:
hermes:
tags: [adversarial, code-review, multi-model, sequential, loop, persona, git]
related_skills: [adversarial-code-review, triangle-code-review, claude-tmux-wrapper]
---
# Adversarial Code Loop v4
**BUILD → REVIEW → (FIX → VERIFY)^N → ARBITER.** A sequential pipeline where one model
writes code, another critiques the git diff, the first fixes, the second validates, and
an optional arbiter resolves the last disagreement. Every loop runs on its own git
branch; each BUILD/FIX is a commit; reviews inspect real git diffs; the result is squash-
merged into the parent branch (or marked `[REJECTED]`).
> **Rule: the orchestrator never writes code directly.** This skill delegates code to
> DEV/FIXER agents (codex, claude-tmux, pi). The orchestrator writes the spec, launches
> the pipeline, and interprets the results. Never use `patch`/`write`/`bash` to edit code
> inside a task covered by this skill — always go through the DEV role. If no DEV agent is
> configured explicitly, use `pi` with the current model.
Based on Multi-Persona adversarial debate (Smit et al., ICML 2024): each role gets a
distinct persona, which improves quality even when both roles share the same model.
**When to use:** code that must be **reviewed by another model** before delivery
(breaking the echo chamber), critical code (security, auth, money), and well-scoped
multi-file refactors (up to ~15 files with a structured spec). Not for simple questions,
trivial 1-file changes, or open-ended design exploration.
## Installation
Requires the `adversarial-common` sibling repo (shared engine). One-line install:
curl -fsSL https://raw.githubusercontent.com/chpomob/adversarial-code-loop/main/scripts/install.sh | bash
or, from an existing checkout:
bash scripts/install.sh
Both place adversarial-code-loop and adversarial-common side by side under `~/.hermes/skills` (override the target with `$1` or `$HERMES_HOME`).
## Overview — what's new in v4
v4 is **git-native**. Where v3 wrote files directly to the worktree and reviewed a stdin
concatenation of file contents, v4 isolates every loop on a dedicated branch and reviews
real diffs.
| Concern | v3 | v4 |
|---------|----|----|
| Isolation | none — writes to live worktree | dedicated branch `loop/<feature>/<N>` |
| Review input | concatenated file contents (stdin) | `git diff <branch-point>..HEAD` |
| BUILD/FIX output | prose/JSON the orchestrator extracts | files committed by the model |
| Recovery on failure | manual file salvage | `git reset`/`git checkout` to restore |
| Merge | manual `git add -A` | squash-merge into parent branch |
| Rejection | exit code only | `[REJECTED]` marker commit + branch preserved |
| Resume | README.md
# adversarial-code-loop **BUILD → REVIEW → (FIX → VERIFY)^N → ARBITER.** A git-native adversarial development pipeline where one model writes code, another critiques the real `git diff`, the first fixes, the second validates, and an optional arbiter resolves deadlocks. For Hermes Agent, Claude Code, Codex, or any LLM CLI. ## How it works Every loop runs on an isolated git branch (`loop/<feature>/<N>`): ``` PHASE 0 ──→ GIT SETUP (branch, stash, identity, gitignore) PHASE 1 ──→ BUILD (DEV model writes code, commits) PHASE 2 ──→ REVIEW (CRITIC model inspects `git diff <branch>..HEAD`) PHASE 3 ──→ FIX (DEV addresses findings, commits) PHASE 4 ──→ VERIFY (CRITIC checks each finding resolved) └── loop 3-4 until APPROVED or max-loops PHASE 5 ──→ ARBITER (resolves last dispute, optional) MERGE ──→ squash-merge into parent, or [REJECTED] marker ``` ## Comparison | Feature | adversarial-code-loop | claude-wizard | opencode-spec-kit | |---------|----------------------|---------------|-------------------| | Git-native (reviews real diffs) | ✅ | ❌ | ❌ | | Multi-model (Codex DEV + Claude REVIEW) | ✅ | ❌ Single model | ❌ | | Per-step plan mode | ❌ (manual) | ❌ | ❌ | | Resume on interrupt | ✅ `--resume` from `state.json` | ❌ | ❌ | | Build/test gates | ✅ `--build-cmd` / `--test-cmd` | ❌ | ❌ | ## Quick start ```bash python3 scripts/adversarial_loop.py \ --spec /path/to/spec.md \ --workdir /path/to/project \ --dev-cmd "codex exec --sandbox workspace-write" \ --review-cmd "pi -p --provider zai --model glm-5.2 --thinking high" ``` See `SKILL.md` for full CLI reference and 30+ validated pitfalls. ## Dependencies - Python ≥ 3.11 - Git ≥ 2.5 - A DEV CLI (codex, pi, claude-tmux, …) - A REVIEW CLI (pi, claude-tmux, …) Uses `adversarial-common` as the shared engine. ## License 0BSD — see [LICENSE](LICENSE).
_meta.json
{
"ownerId": "kn7e26az9x7m8bgwfwg90q1wkh8bsqw0",
"slug": "adversarial-code-loop",
"version": "0.1.0",
"publishedAt": 1785780740116
}references/batch-splitting-strategy.md
# Batch Splitting Strategy for Adversarial Dev Loops How to go from an adversarial-code-review report (17 findings) to fixed code via adversarial dev loops, without stepping on your own feet. ## The Pipeline 1. **adversarial-code-review** → finds bugs, classifies by severity, cross-validates 2. **Group findings into batches** by file dependency (see below) 3. **Write a spec per batch** covering the fixes, files, and test requirements 4. **Run adversarial dev loops sequentially** — NEVER parallel on overlapping files 5. **Commit after each batch** before starting the next ## Why Sequential? The DEV/FIXER role writes files to the workdir. Two simultaneous loops that touch the same file will overwrite each other. The review batches may touch disjoint files but the dev loop writes what the spec asks for — and specs often overlap on shared files (e.g. `routes_api.py`, `gui.py`). **Rule:** only parallelize when `git diff --stat` between batches shows zero file overlap. In practice, this almost never happens — just run sequentially. ## Batch Grouping Rules 1. **Tightly coupled bugs → same batch.** If fixing bug A enables bug B (e.g. callback fix enables refresh-loop fix), batch them together. 2. **Same file modified → same batch if possible.** Two bugs in `routes_api.py` should be one batch. 3. **Independent fixes → separate batches.** Config atomicity (routes_api.py) and version sync (__init__.py) can be separate if they don't touch the same function. 4. **Size: 3-6 bugs per batch max.** More than 6 and the spec gets too long; the DEV may skip items. ## Validated Example (pz-save-manager, 2026-06-16) 17 adversarial-review findings, 10 fixed across 3 batches: | Batch | Bugs | Files | Cycles | Files touched | |-------|------|-------|--------|---------------| | 1 | A2/B1, B2, B3, C1, C2 | 5 | 1 | gui.py, routes_api.py, watcher.py, index.html, test_gui.py | | 2 | B5, B6, B7, B8 | 4 | 1 | routes_api.py, backup.py, index.html, __init__.py | | 3 | A4 | 1 | 2 | gui.py | Batch 1 and 2 both touched `routes_api.py` and `index.html` → must be sequential. Batch 3 touched `gui.py` (also touched by batch 1) → must be sequential. **Total: ~30 min for 3 batches, 4 cycles, all APPROVED, Codex DEV + Claude Opus REVIEW.** ## Pitfalls - **Don't batch security hardening with functional fixes.** The review's A1 (loopback gate) is a cross-cutting concern touching every route — it deserves its own batch with careful testing. - **The adversarial-code-review cross-review was one-directional** (A reviewed B only). Single-reviewer findings (A1-A9) have lower confidence. Prioritize cross-validated (A2/B1) and consensus (B2-B8) findings first. - **Disputed findings (A7/B4) need human adjudication** before entering a dev loop. Don't automate a fix for something the reviewers disagreed on.
references/bug-fix-after-review-workflow.md
# Bug Fix After Review — Workflow Pattern When a code review produces 5+ findings across multiple modules, organize fixes into small adversarial-loop-sized specs rather than a single large spec. ## The Pattern 1. **Taxonomy pass**: Group findings by module/severity. Each group becomes one spec (e.g. F1a = wifi_csi data races, F1b = ble_rssi data races, etc.). 2. **Batch by risk**: - Lot 1 (HIGH/CRITICAL): data races, memory safety, init bugs - Lot 2 (MEDIUM): dead code, edge cases, missing init - Lot 3 (LOW): cosmetic, performance, style 3. **Each spec targets 1-3 files max** and adds zero or very few tests. The review already validated the tests — the fix just needs to compile and pass. 4. **Launch order**: most critical first, simplest first. This maximizes the chance that every fix gets done even if quota runs out. ## Example: OmniSense Bug Fix Session ``` Lot 1 — Data races (4 specs, 4 files total) F1a wifi_csi atomics F1b ble_rssi atomics F1c subghz atomics F1d fusion volatile -> _Atomic Lot 2 — Logic bugs (3 specs, 4 files) F2a config SD init in setup() F2b VHCI stale-response race F2c call fusion_set_band_snr() Lot 3 — Edge cases (2 specs, 2 files) F3a millis() wraparound F3b csi_doppler min subcarriers ``` ## When FIX Times Out If the FIX phase times out (common with Codex sandbox at 300s on integration specs): 1. Kill the process 2. `git diff --stat` — verify the BUILD wrote the core changes 3. Read `02_review.json` — check if the findings are critical 4. Apply critical fixes manually 5. `make all && pio run` — if it compiles and tests pass, commit 6. Document skipped findings as technical debt If Codex wrote files during BUILD (sandbox mode), they're always recoverable via `git diff --stat` even if FIX never ran.
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/chpomob/skills/adversarial-code-loop",
"sourceUrl": "https://clawhub.ai/chpomob/skills/adversarial-code-loop",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T13:35:10.043Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-chpomob-adversarial-code-loop/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-chpomob-adversarial-code-loop/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T13:35:10.043Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "2.6K downloads",
"href": "https://clawhub.ai/chpomob/adversarial-code-loop",
"sourceUrl": "https://clawhub.ai/chpomob/adversarial-code-loop",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T13:35:10.043Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.1.0",
"href": "https://clawhub.ai/chpomob/adversarial-code-loop",
"sourceUrl": "https://clawhub.ai/chpomob/adversarial-code-loop",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-03T18:12:20.116Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-chpomob-adversarial-code-loop/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-chpomob-adversarial-code-loop/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.1.0",
"description": "Adversarial Code Loop v4 is a major update making the workflow fully git-native. - Each code-review loop runs on its own isolated git branch; changes are committed and reviewed via real git diffs, not concatenated files. - All phases (BUILD, REVIEW, FIX, VERIFY, ARBITER) are implemented as shell-invocable steps, with phase responsibilities clarified. - Adds optional gates: custom build/test commands can now block progress at key stages. - Robust resume support: interrupted sessions can be continued with a single flag. - All merge/reject/arbitrate outcomes are precisely tracked by commit/message and machine-readable artifacts. - CLI flags and environment variable handling explicitly defined and streamlined for better configuration.",
"href": "https://clawhub.ai/chpomob/adversarial-code-loop",
"sourceUrl": "https://clawhub.ai/chpomob/adversarial-code-loop",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-03T18:12:20.116Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
