{"id":"7a38d732-9bd6-4d8a-b94c-6a0caa1b0756","entityType":"agent","slug":"clawhub-psyb0t-audiolla","name":"Audiolla","canonicalUrl":"https://www.xpersona.co/agent/clawhub-psyb0t-audiolla","canonicalPath":"/agent/clawhub-psyb0t-audiolla","generatedAt":"2026-10-10T22:47:38.643Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T20:27:38.304Z","emptyReason":null},"description":"Connect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files. Skill: Audiolla Owner: psyb0t Summary: Connect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files. Tags: latest:1.4.2 Version history: v1.4.2 | 2026-07-25T22:29:58.691Z | auto audiolla v1.4.2 - Updated documentation in SKILL.md and references/setup.md for clarity and accuracy. - Removed obsolete file: skill-card.md. - No ch","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.3K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s17fq93tmpky791n7516jcn08n83sfn2:audiolla","sourceUrl":"https://clawhub.ai/psyb0t/audiolla","homepage":"https://clawhub.ai/psyb0t/skills/audiolla","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/psyb0t/audiolla","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/psyb0t/skills/audiolla","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":62,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Connect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files. Skill"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T20:27:38.304Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T20:27:38.304Z","emptyReason":null},"stars":null,"forks":null,"downloads":1266,"packageName":null,"latestVersion":"1.4.2","tractionLabel":"1.3K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T20:27:38.304Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T20:27:38.304Z","lastCrawledAt":"2026-10-10T20:27:38.304Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T20:27:38.304Z","lastVerifiedAt":null,"highlights":[{"version":"1.4.2","createdAt":"2026-07-25T22:29:58.691Z","changelog":"audiolla v1.4.2 - Updated documentation in SKILL.md and references/setup.md for clarity and accuracy. - Removed obsolete file: skill-card.md. - No changes to core functionality or endpoint references. - Minor content cleanup and improved setup instructions.","fileCount":4,"zipByteSize":28594},{"version":"1.4.1","createdAt":"2026-07-25T00:17:42.795Z","changelog":"- Minor documentation update: the skill-card.md file was removed. - SKILL.md refined, but capabilities and usage remain unchanged. - No behavior or API changes to the skill itself.","fileCount":4,"zipByteSize":27450},{"version":"1.4.0","createdAt":"2026-06-08T19:45:32.528Z","changelog":"**audiolla 1.4.0** — Major update to input/output modes and supported features. - Audio endpoints now use JSON-only I/O: all audio-processing routes accept JSON inputs; inline bytes responses are removed. - File handling now requires either a pre-staged `file_path` or a remote `file_url` (when fetch mode is enabled). Removed multipart-input except for direct file uploads. - Output is returned as a JSON object describing result locations (`output_path`, `output_url`); inline audio output is no longer supported. - Added native support for text-to-music and SFX generation with five generation engines (Stable Audio Open, MusicGen, Riffusion, AudioLDM2). - All documentation updated to reflect v1.0.0 API structure, stricter I/O, and expanded music/SFX generation capabilities. - Removed deprecated `skill-card.md`.","fileCount":4,"zipByteSize":27440},{"version":"1.3.0","createdAt":"2026-06-07T17:08:28.853Z","changelog":"**Major capability and documentation expansion** - Vastly expanded supported features, including AudioSet tagging, CLAP embeddings, audio repair, advanced DSP, vocal enhancement, AI audio restoration, speaker diarization, zero-shot audio classification, async jobs, and more. - Added and updated endpoints for modern workflows: curated server-side presets, ad-hoc pipelines, multiband compression, transient shaping, de-essing, sidechain ducking, convolution reverb, harmonic/percussive separation, beat slicing, pitch correction, BPM/key matching, noise reduction, voice activity detection, polyphonic audio-to-MIDI, and more. - New authoritative endpoint catalog reference (`GET /v1/catalog`), detailed discovery/documentation of endpoints, and workflow examples. - Updated constraints: do not use for general audio-processing unless explicitly asked for audiolla; do not access secrets from files autonomously; clarified stance on which jobs are supported. - Documentation now emphasizes curated/pipeline workflows, async job capabilities, and output handling improvements. - Removed redundant/legacy documentation file (`skill-card.md`) for clarity.","fileCount":4,"zipByteSize":25873},{"version":"1.2.0","createdAt":"2026-06-01T17:09:31.513Z","changelog":"**Major feature update: Expanded audio/MIDI features and effects.** - Added support for audio silence detection/trimming, PNG spectrogram/waveform images, 8-mode video visualization, and Chromaprint audio fingerprinting. - Introduced general effects chain processor (full pedalboard catalog), and silence detection/FFmpeg engine. - Enhanced MIR analysis: beat grid, onset detection, melody extraction, and structural segmentation are now available. - Added MIDI composition (from JSON), MIDI inspection, transformation (transpose/quantize/tempo/filter), and MIDI-to-audio rendering capabilities. - Updated documentation to reflect newly supported engines and endpoints. - Removed obsolete skill-card.md file.","fileCount":4,"zipByteSize":20471},{"version":"1.1.0","createdAt":"2026-05-31T14:44:19.709Z","changelog":"- Adds support for multiple audio I/O modes: files can now be input as multipart uploads, staged file paths, or remote URLs (when operator enables AUDIOLLA_FETCH_MODE). - Supports three output modes for processed audio: inline bytes, server staging, or direct PUT to a presigned URL. - Clarifies that audiolla only fetches from or uploads to remote URLs when AUDIOLLA_FETCH_MODE is enabled; requests will fail with \"URL fetch/upload is disabled\" if not permitted. - Updated documentation to describe all input/output modes and stricter guidance on environment variable requirements. - No changes to core processing features or required dependencies.","fileCount":4,"zipByteSize":14511},{"version":"1.0.0","createdAt":"2026-05-31T12:16:16.966Z","changelog":"audiolla 1.0.0 – Initial release - Provides an HTTP/MCP client for user-deployed audiolla servers, enabling music stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization. - Only operates when the user explicitly names audiolla AND provides AUDIOLLA_URL. - Supports Demucs for stem separation, matchering/pedalboard for mastering, MIR feature extraction (BPM, key, LUFS, etc.), and robust DSP workflows via SoX. - Read-only processing: files are handled by the user’s server and never sent to external services. - Requires curl, a running audiolla Docker instance, and user-supplied connection/auth settings. - Provides clear error handling, engine management, and documentation of typical use cases and limitations.","fileCount":4,"zipByteSize":11781}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17fq93tmpky791n7516jcn08n83sfn2:audiolla","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T22:47:38.637Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-psyb0t-audiolla/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T20:27:38.304Z","emptyReason":null},"readme":"Skill: Audiolla\n\nOwner: psyb0t\n\nSummary: Connect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files.\n\nTags: latest:1.4.2\n\nVersion history:\n\nv1.4.2 | 2026-07-25T22:29:58.691Z | auto\n\naudiolla v1.4.2\n\n- Updated documentation in SKILL.md and references/setup.md for clarity and accuracy.\n- Removed obsolete file: skill-card.md.\n- No changes to core functionality or endpoint references.\n- Minor content cleanup and improved setup instructions.\n\nv1.4.1 | 2026-07-25T00:17:42.795Z | auto\n\n- Minor documentation update: the skill-card.md file was removed.\n- SKILL.md refined, but capabilities and usage remain unchanged.\n- No behavior or API changes to the skill itself.\n\nv1.4.0 | 2026-06-08T19:45:32.528Z | auto\n\n**audiolla 1.4.0** — Major update to input/output modes and supported features.\n\n- Audio endpoints now use JSON-only I/O: all audio-processing routes accept JSON inputs; inline bytes responses are removed.\n- File handling now requires either a pre-staged `file_path` or a remote `file_url` (when fetch mode is enabled). Removed multipart-input except for direct file uploads.\n- Output is returned as a JSON object describing result locations (`output_path`, `output_url`); inline audio output is no longer supported.\n- Added native support for text-to-music and SFX generation with five generation engines (Stable Audio Open, MusicGen, Riffusion, AudioLDM2).\n- All documentation updated to reflect v1.0.0 API structure, stricter I/O, and expanded music/SFX generation capabilities.\n- Removed deprecated `skill-card.md`.\n\nv1.3.0 | 2026-06-07T17:08:28.853Z | auto\n\n**Major capability and documentation expansion**\n\n- Vastly expanded supported features, including AudioSet tagging, CLAP embeddings, audio repair, advanced DSP, vocal enhancement, AI audio restoration, speaker diarization, zero-shot audio classification, async jobs, and more.\n- Added and updated endpoints for modern workflows: curated server-side presets, ad-hoc pipelines, multiband compression, transient shaping, de-essing, sidechain ducking, convolution reverb, harmonic/percussive separation, beat slicing, pitch correction, BPM/key matching, noise reduction, voice activity detection, polyphonic audio-to-MIDI, and more.\n- New authoritative endpoint catalog reference (`GET /v1/catalog`), detailed discovery/documentation of endpoints, and workflow examples.\n- Updated constraints: do not use for general audio-processing unless explicitly asked for audiolla; do not access secrets from files autonomously; clarified stance on which jobs are supported.\n- Documentation now emphasizes curated/pipeline workflows, async job capabilities, and output handling improvements.\n- Removed redundant/legacy documentation file (`skill-card.md`) for clarity.\n\nv1.2.0 | 2026-06-01T17:09:31.513Z | auto\n\n**Major feature update: Expanded audio/MIDI features and effects.**\n\n- Added support for audio silence detection/trimming, PNG spectrogram/waveform images, 8-mode video visualization, and Chromaprint audio fingerprinting.\n- Introduced general effects chain processor (full pedalboard catalog), and silence detection/FFmpeg engine.\n- Enhanced MIR analysis: beat grid, onset detection, melody extraction, and structural segmentation are now available.\n- Added MIDI composition (from JSON), MIDI inspection, transformation (transpose/quantize/tempo/filter), and MIDI-to-audio rendering capabilities.\n- Updated documentation to reflect newly supported engines and endpoints.\n- Removed obsolete skill-card.md file.\n\nv1.1.0 | 2026-05-31T14:44:19.709Z | auto\n\n- Adds support for multiple audio I/O modes: files can now be input as multipart uploads, staged file paths, or remote URLs (when operator enables AUDIOLLA_FETCH_MODE).\n- Supports three output modes for processed audio: inline bytes, server staging, or direct PUT to a presigned URL.\n- Clarifies that audiolla only fetches from or uploads to remote URLs when AUDIOLLA_FETCH_MODE is enabled; requests will fail with \"URL fetch/upload is disabled\" if not permitted.\n- Updated documentation to describe all input/output modes and stricter guidance on environment variable requirements.\n- No changes to core processing features or required dependencies.\n\nv1.0.0 | 2026-05-31T12:16:16.966Z | auto\n\naudiolla 1.0.0 – Initial release\n\n- Provides an HTTP/MCP client for user-deployed audiolla servers, enabling music stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization.\n- Only operates when the user explicitly names audiolla AND provides AUDIOLLA_URL.\n- Supports Demucs for stem separation, matchering/pedalboard for mastering, MIR feature extraction (BPM, key, LUFS, etc.), and robust DSP workflows via SoX.\n- Read-only processing: files are handled by the user’s server and never sent to external services.\n- Requires curl, a running audiolla Docker instance, and user-supplied connection/auth settings.\n- Provides clear error handling, engine management, and documentation of typical use cases and limitations.\n\nArchive index:\n\nArchive v1.4.2: 4 files, 28594 bytes\n\nFiles: references/setup.md (12248b), skill-card.md (2233b), SKILL.md (66641b), _meta.json (127b)\n\nFile v1.4.2:SKILL.md\n\n---\nname: audiolla\ndescription: HTTP/MCP client for a user-deployed audiolla audio-production server. Use ONLY when the user has explicitly named audiolla AND provided AUDIOLLA_URL (or has it set in the environment). Capabilities: stem separation (Demucs / MDX / BS-Roformer), mastering (matchering reference / pedalboard preset chain), MIR analysis (BPM, key, LUFS, spectral features, beat grid, onset detection, melody contour, structural segmentation via librosa), DSP transforms (gain, EQ, compand, reverb, pitch, tempo via SoX), loudness measurement and normalization, generic effects chains (full pedalboard catalog as ordered chain), multiband compression (LR4 crossovers), transient shaping, sidechain ducking, de-essing, mid/side encode-decode, parametric EQ, panning, stereo width, silence detection and trimming, audio repair (declip + dehum), clip detection, harmonic/percussive separation, time-stretch and pitch-shift, BPM/key matching, pitch correction (auto-tune), beat slicing, audio thumbnail extraction, convolution reverb, static PNG spectrogram/waveform and 8-mode animated MP4/WebM video (ffmpeg), Chromaprint acoustic fingerprinting, AudioSet tagging, CLAP audio embeddings + similarity + zero-shot classification, ID3/Vorbis/FLAC metadata read/write, MIDI composition from JSON spec, MIDI inspection, MIDI transformation (transpose/quantize/tempo/channel-filter), MIDI quantize and humanize, drum pattern generation, MIDI rendering via fluidsynth, polyphonic audio-to-MIDI transcription (Spotify basic-pitch ONNX), chords-to-MIDI conversion, AI audio restoration (de-reverb, de-echo, AI de-noise via UVR/audio-separator), DSP noise reduction, neural speech/vocal enhancement (DeepFilterNet DF3), voice activity detection (silero-vad), speaker diarization (pyannote 3.1), DJ prep (BPM + key + Camelot + LUFS in one call), loop-point detection, curated server-side workflow presets (master-for-spotify, podcast-cleanup, vocal-cleanup) and ad-hoc op pipelines that chain multiple operations server-side. v1.0.0 API is JSON-everywhere: every audio endpoint takes a JSON body; the ONLY multipart route is `PUT /v1/files/{path}` for raw byte uploads. Audio I/O supports two input modes (`file_path` referencing a pre-staged file under FILES_DIR, xor `file_url` — only when the operator has enabled AUDIOLLA_FETCH_MODE) and two output modes (`output_path` writing back to staging, xor `output_url` PUTing to a presigned URL). There is no inline-bytes audio response anywhere — every audio-producing endpoint returns JSON describing where the result landed. Audiolla only fetches/uploads to URLs when the operator has explicitly enabled AUDIOLLA_FETCH_MODE — if a request returns \"URL fetch/upload is disabled\", do NOT try to bypass it. Do not use this skill for generic audio-processing questions or for users who haven't named audiolla.\ncompatibility: Requires curl and a running audiolla instance (Docker image psyb0t/audiolla:latest or :latest-cuda). AUDIOLLA_URL env var must be set by the user (default http://localhost:8000). AUDIOLLA_TOKEN required only when the server has AUDIOLLA_AUTH_TOKEN configured; obtain from the AUDIOLLA_TOKEN env var or by asking the user — never read tokens from repo files autonomously.\nmetadata:\n  author: psyb0t\n  homepage: https://github.com/psyb0t/docker-audiolla\n---\n\n# audiolla\n\nHTTP + MCP client for an audiolla server that the user has already deployed. This skill talks to a running audiolla instance — it does not stand one up, does not download model weights manually, and does not modify the server config on its own initiative.\n\nFor installation and setup, see [references/setup.md](references/setup.md).\n\n## Authoritative endpoint reference: `GET /v1/catalog`\n\nThis skill documents the most common patterns. The **full, always-current** list of every endpoint is `GET /v1/catalog` (17 categories, ~85 endpoints). Always check the catalog when looking for an operation that isn't shown here — the server is the source of truth, this file is a curated reference.\n\n```bash\n# List every endpoint grouped by category\ncurl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | {name, count: (.endpoints | length)}'\n\n# Find endpoints in one category\ncurl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | select(.name == \"dynamics\") | .endpoints'\n```\n\nCompanion discovery endpoints: `GET /v1/engines` (engines + loaded/idle status), `GET /v1/presets` (curated workflows), `GET /v1/ops` (the ~24 pipeline op slugs).\n\n## When to use this skill\n\nThe user has audiolla running and asks you to:\n- Pull stems (vocals / drums / bass / etc.) out of a track\n- Master a track against a reference recording (matchering) or via preset chain\n- **Run a curated workflow** (`master-for-spotify`, `podcast-cleanup`, `vocal-cleanup`) via a single `POST /v1/presets/{name}` call\n- **Chain ad-hoc operations server-side** via `POST /v1/pipeline` (no re-upload between steps)\n- Get BPM, key, LUFS, duration, or spectral features for a file\n- Detect beat grid, onsets, dominant melody, or structural segments\n- Detect chords + key (separate from BPM/LUFS)\n- Detect or trim silence\n- Generate a spectrogram, waveform image, or animated visualisation video\n- Compute a Chromaprint acoustic fingerprint\n- Apply a DSP chain (gain, EQ, compression, reverb, pitch shift, tempo via SoX OR full pedalboard catalog)\n- **Multiband compression** with LR4 crossovers\n- **Transient shaping** (punch up drums / cut room tail)\n- **De-essing** (split-band sibilance compression)\n- **Sidechain ducking** (voiceover-over-music)\n- **Mid/Side encode/decode** (for stereo M/S processing)\n- **Convolution reverb** (apply a user-supplied IR file)\n- **Audio repair** (declip + dehum)\n- **Time-stretch + pitch-shift** independently, or **BPM-match / key-match** to a target\n- **Pitch-correct** (auto-tune to nearest semitone)\n- **Beat-slice** at detected beat positions (returns ZIP of chops)\n- **Audio thumbnail** — most-energetic N-second segment\n- **HPSS** harmonic/percussive separation\n- Measure or normalize integrated LUFS (`/v1/audio/normalize` with `target_lufs`)\n- **Loudness curve** — RMS envelope over time (`/v1/audio/loudness/curve`)\n- Stage files server-side, then operate on them via `file_path`\n- **Tag** audio (AudioSet labels), **embed** (CLAP 512-dim), **classify** (zero-shot label list), **similar** (cosine between two tracks)\n- Read or write ID3/Vorbis/FLAC **metadata** (mutagen)\n- **DJ-prep** — BPM + key + Camelot + LUFS in one call\n- Compose / inspect / transform / render MIDI; **quantize**, **humanize**, **drum patterns**, **chords-to-MIDI**\n- Remove reverb / echo / noise via `/v1/audio/restore/{engine}` (UVR)\n- DSP noise reduction via `/v1/audio/noise-reduce/{engine}` (DSP or UVR)\n- Convert any audio to polyphonic MIDI (basic-pitch)\n- **Voice activity detection** (silero-vad — speech/non-speech segments)\n- **Speaker diarization** (pyannote 3.1 — who spoke when)\n- Enhance speech/vocal recordings (DeepFilterNet DF3)\n- **Generate music or SFX from a text prompt** via `/v1/audio/generate/{engine}` — five engines:\n  - `stable-audio-open` (Stability Community Licence — commercial OK below revenue threshold; 47 s cap; 44.1 kHz stereo; loops / SFX / textures; instrumental)\n  - `musicgen-small` and `musicgen-medium` (Meta MusicGen 300M / 1.5B; **CC-BY-NC 4.0** — server must opt in via `AUDIOLLA_ENABLE_NONCOMMERCIAL=1`; 30 s cap; instrumental)\n  - `riffusion` (CreativeML OpenRAIL-M; ~5 s per pass; spectrogram-via-Griffin-Lim; lo-fi character)\n  - `audioldm2` (**CC-BY 4.0 — commercial-safe, no opt-in gate**; 30 s cap; 16 kHz mono; general SFX — ambience / foley / impact / animal sounds; slow at default 200-step DDIM, pass `num_inference_steps=50` for ~4x speed)\n  All five are CUDA-only. Full-song / lyric-conditioned generation isn't shipped (ACE-Step + DiffRhythm + TangoFlux + Stable Audio Open Small deferred — see the README's \"deferred\" list). For commercial use, prefer `audioldm2` (CC-BY 4.0) or `stable-audio-open` (Stability Community Licence below the revenue threshold).\n- Drive any of the above from an LLM agent over MCP\n- **Async-job-and-forget** any audio-producing call via `async_job=true` + optional `webhook_url`\n- Send results to a **presigned S3-style PUT URL** via `output_url`\n\n## When NOT to use this skill\n\n- The user hasn't named audiolla — they're asking a general \"how do I split stems?\" question. Suggest audiolla as an option; don't assume it's running.\n- The user wants music generation from a melody-conditioning input (hum-to-track / \"make this sound like X\"). Audiolla's five generators (`stable-audio-open`, `musicgen-small`, `musicgen-medium`, `riffusion`, `audioldm2`) are text-prompt only; melody conditioning isn't wired. Plain text → music or SFX IS supported — see `/v1/audio/generate/{engine}` in the catalog. The closed-weight Suno / Udio APIs are out of scope.\n- The user wants real-time / streaming processing. Demucs needs the whole file.\n- The user wants **transcription / ASR / TTS / voice cloning** — that's [docker-talkies](https://github.com/psyb0t/docker-talkies). Note: audiolla DOES have speech-adjacent features (VAD, diarization, neural enhancement) but does NOT transcribe.\n\n## Setup\n\n```bash\nexport AUDIOLLA_URL=http://localhost:8000\nexport AUDIOLLA_TOKEN=<the-token-the-user-gives-you>   # only if auth is enabled\n```\n\nIf `AUDIOLLA_URL` is not set, ask the user — do not search the workspace for it. Same for `AUDIOLLA_TOKEN`: only accept it from the env var the user set or from the user directly. Never read it from `docker-compose.yml`, `.env`, or any other repo file on your own initiative.\n\n**Verify:** `curl $AUDIOLLA_URL/healthz` → `{\"ok\": true, \"device\": \"...\", \"engines\": [...]}`. `/healthz` is always unauthenticated regardless of `AUDIOLLA_AUTH_TOKEN`.\n\nAuth is optional. If the server has `AUDIOLLA_AUTH_TOKEN` set, every endpoint except `/healthz` requires `Authorization: Bearer $AUDIOLLA_TOKEN`. Without it you get `401`. Always pass the token if the user gave you one; don't assume the server has auth off.\n\n## How it works\n\nv1.0.0 is **JSON-everywhere**. Every audio endpoint takes `Content-Type: application/json` with a JSON body. The ONE exception is `PUT /v1/files/{path}` for raw byte uploads (`application/octet-stream`). Input is `file_path` (pre-staged under FILES_DIR via `PUT /v1/files/{path}`) xor `file_url` (server fetches when `AUDIOLLA_FETCH_MODE` allows). Output for audio-producing endpoints is `output_path` (server writes to FILES_DIR) xor `output_url` (server PUTs to a presigned URL). Both modes return JSON describing where the result landed (`{path,size,...}` or `{url,size,...}`); there is **no inline-bytes audio response** anywhere. Analysis-only endpoints (no audio produced — e.g. `/v1/audio/analyze`, `/v1/audio/beats`, `/v1/audio/fingerprint`) return their JSON data directly and ignore output_path/output_url. The standard flow is: `PUT /v1/files/uploads/track.wav` once, then JSON-body POST to every processing endpoint with `file_path` + `output_path`, chaining the output of one call into the input of the next.\n\nEvery error response:\n\n```json\n{\"detail\": \"description of what went wrong\"}\n```\n\nStatus codes follow REST conventions:\n- `200` — success\n- `400` — bad input (unknown engine, invalid features, bad operations JSON, etc.)\n- `401` — missing/invalid bearer token (only when auth is enabled)\n- `404` — unknown engine slug, unknown file path\n- `413` — upload exceeded `AUDIOLLA_MAX_UPLOAD_BYTES` (default 200 MB)\n- `415` — unsupported `output_format`\n- `500` — server error (engine failed internally, etc.)\n\n## Engines\n\n| Slug | What it does | Notes |\n|------|--------------|-------|\n| `htdemucs` | 4-stem separation | drums, bass, other, vocals |\n| `htdemucs_ft` | 4-stem fine-tuned | **CUDA-only at usable speed** — flagged `cuda_only`, the server rejects it with 400 on CPU |\n| `htdemucs_6s` | 6-stem separation | adds `guitar` + `piano` (experimental, CPU OK but slow) |\n| `mdx_extra` | 4-stem MDX-Net | drums, bass, other, vocals — strong vocal isolation |\n| `matchering` | Reference-based mastering | GPL v3 |\n| `pedalboard-chain` | Preset DSP mastering chain | presets: `transparent`, `loud` — GPL v3 |\n| `librosa-analyze` | MIR analysis + loudness | BPM, key, LUFS, spectral, beat grid, onsets, melody (pyin), segments; backs `/v1/audio/{analyze,beats,onsets,melody,segments,loudness}` |\n| `sox-transform` | SoX DSP chain | gain, EQ, compand, reverb, pitch, tempo, rate, channels, trim, pad |\n| `fx-chain` | Arbitrary pedalboard chain | full pedalboard catalog as `[{type, params}, ...]` — backs `/v1/audio/fx`. VST3 / AU / external-plugin classes deliberately blocked |\n| `midi-compose` | JSON → MIDI; inspect/transform | song-spec transcoder + MIDI reader/editor; backs `/v1/midi/{compose,inspect,transform,generate}` |\n| `midi-render` | MIDI → audio | fluidsynth + FluidR3_GM SoundFont (GM patches 0-127, drum kit on channel 9) |\n| `silence-detect` | Silence detection + trimming | ffmpeg `silencedetect`; backs `/v1/audio/silence` |\n| `ffmpeg-render` | Spectrogram / waveform / video | static PNG + 8-mode animated MP4/WebM; backs `/v1/audio/visualize/image/{spectrogram,waveform}` + `/v1/audio/visualize/video/{mode}` |\n| `audio-fingerprint` | Chromaprint fingerprint | `fpcalc` subprocess; backs `/v1/audio/fingerprint` |\n| `uvr-dereverb` | AI de-reverb | BS-Roformer (SDR 19+); backs `/v1/audio/restore/uvr-dereverb` |\n| `uvr-deecho` | AI de-echo (normal + aggressive) | VR Architecture; `aggressive=true` enables hard mode (`uvr-deecho-aggressive` slug is gone — consolidated into this engine); backs `/v1/audio/restore/uvr-deecho` |\n| `uvr-denoise` | AI de-noise | MelBand Roformer (SDR 28); backs `/v1/audio/restore/uvr-denoise` + `/v1/audio/noise-reduce/uvr-denoise` |\n| `uvr-karaoke` | Karaoke (remove lead vocals) | MelBand Roformer; returns Instrumental stem |\n| `uvr-vocal-bsr` | High-quality vocal/inst separation | BS-Roformer (SDR 13) — stems: Vocals, Instrumental |\n| `basic-pitch` | Polyphonic audio-to-MIDI transcription | Spotify basic-pitch ONNX; backs `/v1/audio/to_midi/basic-pitch` |\n| `deepfilter` | Neural speech/vocal enhancement | DeepFilterNet DF3; backs `/v1/audio/enhance/deepfilter` |\n| `noise-reduce` | DSP spectral noise reduction | noisereduce — backs `/v1/audio/noise-reduce/noise-reduce` (stationary/non-stationary modes, no GPU) |\n| `chord-detect` | Chord progression + key | Krumhansl-Schmuckler + chroma template matching; backs `/v1/audio/chords`, `/v1/audio/chords-to-midi`, `/v1/audio/key-match` |\n| `silero-vad` | Voice activity detection | speech/non-speech timestamps; backs `/v1/audio/vad` |\n| `pyannote` | Speaker diarization | pyannote/speaker-diarization-3.1 — backs `/v1/audio/diarize` (requires `HUGGINGFACE_TOKEN`) |\n| `stretch` | Time-stretch + pitch-shift | librosa phase vocoder; backs `/v1/audio/stretch`, `/v1/audio/bpm-match`, `/v1/audio/key-match` |\n| `ast-tag` | AudioSet zero-shot labels | Audio Spectrogram Transformer; backs `/v1/audio/tag` |\n| `clap-embed` | CLAP embeddings + similarity + classification | LAION CLAP 512-dim; backs `/v1/audio/embed`, `/v1/audio/similar`, `/v1/audio/classify` |\n| `hpss` | Harmonic/percussive split | librosa median-filter HPSS; backs `/v1/audio/separate/hpss` |\n| `metadata` | ID3 / Vorbis / FLAC tag read+write | mutagen; backs `/v1/audio/metadata` |\n\nEngines lazy-load on first use and auto-unload after `AUDIOLLA_ENGINE_TTL` seconds of idle (default 600s). Demucs weights prefetch into `/data/torch_cache/` at container start so the first separation request doesn't pay the cold-download cost.\n\nUse `GET /v1/engines` to confirm what's actually configured on the running server (operators can restrict via `AUDIOLLA_ENABLED_ENGINES`).\n\n## Output formats\n\nAny endpoint that produces audio accepts `\"output_format\": \"<fmt>\"` in the JSON body. Supported: `wav` (default), `mp3`, `flac`, `opus`, `aac`, `pcm`. The server transcodes via ffmpeg — the `output_path` extension does not determine the encoding.\n\n## API Reference\n\n### Health & engine listing\n\n```bash\n# Liveness — no auth required\ncurl $AUDIOLLA_URL/healthz\n# {\"ok\": true, \"device\": \"cpu\", \"engines\": [\"htdemucs\", \"matchering\", ...]}\n\n# Configured engines + capabilities\ncurl -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/engines\n\n# Engines currently loaded in memory (and how idle)\ncurl -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/ps\n\n# Evict one engine\ncurl -X DELETE -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/ps/htdemucs\n\n# Evict everything\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/unload\n```\n\n### Stem separation\n\n`POST /v1/audio/separate` — JSON body. Result is one staged file (single-stem) or a ZIP of stems written to `output_path`.\n\n```bash\n# Stage the input once (only multipart route in the whole API)\ncurl -X PUT -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/octet-stream' \\\n  --data-binary @track.wav \\\n  $AUDIOLLA_URL/v1/files/uploads/track.wav\n\n# Single stem → JSON {path,size,...} pointing at the staged stem\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"engine\":\"htdemucs\",\"stems\":[\"vocals\"],\"output_path\":\"stems/vocals.wav\"}'\n\n# Multiple stems → ZIP at output_path\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"engine\":\"htdemucs\",\"stems\":[\"vocals\",\"drums\"],\"output_path\":\"stems/vocals_drums.zip\"}'\n\n# Omit stems → all stems for that engine (ZIP)\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"engine\":\"htdemucs\",\"output_path\":\"stems/all.zip\"}'\n\n# MP3 output\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"engine\":\"htdemucs\",\"stems\":[\"vocals\"],\"output_format\":\"mp3\",\"output_path\":\"stems/vocals.mp3\"}'\n```\n\nRequired: `file_path` (xor `file_url`), `engine`, and one of `output_path`/`output_url`. Optional: `stems` (array; default = all stems for that engine), `output_format` (default `wav`).\n\nLoading a separation engine evicts other loaded engines first — Demucs is memory-hungry and the operator-default setup runs one engine in memory at a time.\n\n### Mastering\n\n`POST /v1/audio/master` — `mode=reference` uses matchering against a reference track; `mode=chain` runs a pedalboard preset.\n\n```bash\n# Reference-based mastering — both inputs pre-staged\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/master \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"mode\":\"reference\",\"reference_path\":\"uploads/ref.wav\",\"output_path\":\"out/mastered.wav\"}'\n\n# Pedalboard chain — preset is REQUIRED (transparent or loud)\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/master \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"mode\":\"chain\",\"preset\":\"loud\",\"output_path\":\"out/mastered.wav\"}'\n\n# Pedalboard chain with explicit loudness target\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/master \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"mode\":\"chain\",\"preset\":\"transparent\",\"target_lufs\":-14,\"output_path\":\"out/mastered.wav\"}'\n```\n\nRequired: `file_path` (xor `file_url`), `mode`, and one of `output_path`/`output_url`. `mode=reference` requires `reference_path` (xor `reference_url`). `mode=chain` requires `preset` (`transparent` or `loud`). Optional: `target_lufs` (range `[-70.0, -0.1]`), `output_format`.\n\nStreaming-target LUFS reference values: Spotify `-14`, Apple Music `-16`, YouTube `-14`, broadcast EBU R128 `-23`.\n\n### MIR analysis\n\n`POST /v1/audio/analyze` — analysis-only, returns JSON. No output_path/output_url.\n\n```bash\n# Specific features\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/analyze \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"features\":[\"bpm\",\"key\",\"loudness\"]}'\n\n# Omit features → returns all of them\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/analyze \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n```\n\nValid `features` values: `bpm`, `key`, `loudness`, `duration`, `spectral_centroid`, `rms`, `zcr`.\n\n> **Common mistake:** the feature for integrated LUFS is `loudness`, NOT `lufs`. Asking for `features=[\"lufs\"]` returns 400.\n\n### Beat detection (`/v1/audio/beats`)\n\nReturns the estimated BPM and beat timestamps. Optionally writes a click-track WAV to `output_path`.\n\n```bash\n# Beat grid only — analysis JSON\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/beats \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"bpm\": 128.0, \"beats\": [0.0, 0.469, 0.938, ...], \"engine\": \"librosa-analyze\"}\n\n# With a click track — output_path is REQUIRED when click_track=true\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/beats \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"click_track\":true,\"output_path\":\"beats/click.wav\"}'\n# → JSON with beat grid PLUS the staged click track path\n```\n\nOptional params: `click_track` (bool, default false) — when true, writes the click WAV to `output_path` / `output_url`. `hop_length` (int, default 512) — analysis hop size in samples.\n\n### Onset detection (`/v1/audio/onsets`)\n\nReturns note/transient onset timestamps in seconds. Analysis-only, returns JSON.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/onsets \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"onsets\": [0.023, 0.512, 1.034, ...], \"count\": 42, \"engine\": \"librosa-analyze\"}\n```\n\nOptional: `backtrack` (bool, default false) — snap onsets to preceding energy valley. `hop_length`, `delta` for tuning sensitivity.\n\n### Melody extraction (`/v1/audio/melody`)\n\nEstimates the dominant melody using pyin pitch tracking. Returns Hz per frame.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/melody \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"melody\": [{\"time\": 0.0, \"hz\": 440.1}, {\"time\": 0.023, \"hz\": null}, ...], ...}\n\n# Export the melody as a single-track MIDI file (output_path REQUIRED when as_midi=true)\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/melody \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"as_midi\":true,\"output_path\":\"melody/lead.mid\"}'\n```\n\n`hz` is `null` for unvoiced frames. Optional: `as_midi` (bool) — generates MIDI from the contour and writes to `output_path` / `output_url`; `fmin`/`fmax` to constrain pitch range.\n\n### Structural segmentation (`/v1/audio/segments`)\n\nFinds recurring sections (verse, chorus, bridge…) using a recurrence matrix. Returns labels A, B, C…\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/segments \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"num_segments\":6}'\n# {\"segments\": [{\"label\":\"A\",\"start_sec\":0.0,\"end_sec\":32.5},\n#               {\"label\":\"B\",\"start_sec\":32.5,\"end_sec\":65.0}, ...]}\n```\n\nOptional: `num_segments` (int, default 6, valid range [2, 32]). Short inputs (fewer beats than `num_segments`) return a single `A` span with a `note` field explaining the fallback.\n\n### Silence detection and trimming (`/v1/audio/silence`)\n\nFinds silent gaps via ffmpeg `silencedetect`. Optionally trims them.\n\n```bash\n# Detect only — analysis JSON, no audio produced\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/silence \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"threshold_db\":-30,\"min_duration_sec\":1.0}'\n# {\"silent_ranges\": [...], \"non_silent_ranges\": [...], \"duration\": 215.3}\n\n# Trim all silence → trimmed audio staged\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/silence \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"threshold_db\":-30,\"min_duration_sec\":0.5,\"trim_mode\":\"all\",\"output_path\":\"proc/trimmed.wav\"}'\n\n# Trim only edges\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/silence \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"threshold_db\":-40,\"min_duration_sec\":0.3,\"trim_mode\":\"edges\",\"output_path\":\"proc/trimmed.wav\"}'\n```\n\n`threshold_db` must be ≤ 0. `trim_mode`: `edges` (leading + trailing only), `all` (every detected gap). Without `trim_mode`, response is JSON only — no audio produced. With `trim_mode` set, `output_path` (or `output_url`) is required and the response JSON points at the trimmed file.\n\n### Spectrogram (`/v1/audio/visualize/image/spectrogram`)\n\nStatic PNG spectrogram via ffmpeg `showspectrumpic`. PNG is written to `output_path` / `output_url`.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/visualize/image/spectrogram \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"width\":1280,\"height\":720,\"output_path\":\"viz/spec.png\"}'\n```\n\nOptional: `width`, `height` (64–8192, defaults 1920×1080), `color` (default `intensity`), `scale` (default `log`).\n\n### Waveform (`/v1/audio/visualize/image/waveform`)\n\nStatic PNG waveform via ffmpeg `showwavespic`. PNG is written to `output_path` / `output_url`.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/visualize/image/waveform \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"width\":1920,\"height\":240,\"output_path\":\"viz/wave.png\"}'\n```\n\nOptional: `width`, `height` (64–8192, defaults 1920×320), `color` (default `lime`).\n\n### Animated visualisation (`/v1/audio/visualize/video/{mode}`)\n\nAnimated MP4 or WebM video from one of 8 ffmpeg filter modes. Video is written to `output_path` / `output_url`.\n\n```bash\n# `mode` is in the URL path\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/visualize/video/spectrum \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"width\":1280,\"height\":720,\"fps\":30,\"container\":\"mp4\",\"output_path\":\"viz/spectrum.mp4\"}'\n```\n\n`mode` options (URL path segment): `spectrum` (scrolling FFT), `waves` (oscilloscope), `cqt` (constant-Q transform), `freqs` (bar-graph), `volume` (VU meter), `vectorscope` (stereo X/Y), `phasemeter`, `histogram`. `container`: `mp4` (default) or `webm`. `fps` 1–120.\n\n### Acoustic fingerprint (`/v1/audio/fingerprint`)\n\nChromaprint fingerprint via `fpcalc`. The base64 string is AcoustID-compatible. Analysis-only — no output_path.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/fingerprint \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"duration\": 215.34, \"fingerprint\": \"AQADtEqRRIuQ...\"}\n\n# Include the raw integer array\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/fingerprint \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"return_raw\":true}'\n# adds \"fingerprint_raw\": [12345, 67890, ...]\n```\n\nOptional: `analyze_seconds` (default 120 — AcoustID standard; pass 0 to fingerprint the whole file), `return_raw` (bool).\n\n### DSP transform chain\n\n`POST /v1/audio/transform` — applies an array of SoX operations in order.\n\n```bash\n# Pitch shift up 2 semitones, then add reverb\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/transform \\\n  -d '{\n    \"file_path\":\"uploads/track.wav\",\n    \"operations\":[\n      {\"op\":\"pitch\",\"params\":{\"n_semitones\":2}},\n      {\"op\":\"reverb\",\"params\":{\"reverberance\":50,\"room_scale\":80}}\n    ],\n    \"output_format\":\"wav\",\n    \"output_path\":\"out/transformed.wav\"\n  }'\n\n# Trim first 30s, pad 2s silence at end, gain -3dB\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/transform \\\n  -d '{\n    \"file_path\":\"uploads/track.wav\",\n    \"operations\":[\n      {\"op\":\"trim\",\"params\":{\"start_time\":0,\"end_time\":30}},\n      {\"op\":\"pad\",\"params\":{\"end_duration\":2}},\n      {\"op\":\"gain\",\"params\":{\"db\":-3}}\n    ],\n    \"output_path\":\"out/trimmed.wav\"\n  }'\n```\n\n`operations` is a JSON array of `{\"op\": \"<name>\", \"params\": {...}}`. Order matters — ops apply left-to-right.\n\n**Ops and their params:**\n\n| op | required params | optional params | what it does |\n|----|-----------------|-----------------|--------------|\n| `gain` | `db` (float) | | gain in dB |\n| `equalizer` | `frequency`, `gain_db` | `width_q` (default 1.0) | peaking EQ |\n| `compand` | | `attack_time`, `decay_time`, `soft_knee_db`, `tf_points` ([[in_db, out_db], ...]) | dynamic range compression |\n| `reverb` | | `reverberance` (0-100, default 50), `pre_delay_ms` (default 0), `room_scale` (default 100) | reverb |\n| `pitch` | `n_semitones` (float) | | pitch shift in **semitones**, not cents |\n| `tempo` | `factor` (float) | | tempo factor (1.5 = 1.5x faster, 0.5 = half speed) |\n| `rate` | `samplerate` (int) | | resample |\n| `channels` | `n_channels` (int) | | mix to N channels |\n| `trim` | `start_time` (float, sec) | `end_time` (float, sec; null = end of file) | trim |\n| `pad` | | `start_duration`, `end_duration` (both floats, sec) | pad silence |\n\nUnknown ops return 400 with the valid list.\n\n### Loudness\n\n`POST /v1/audio/loudness` — analysis-only. Returns integrated LUFS as JSON. Use `/v1/audio/normalize` (separate endpoint) for actual normalization.\n\n```bash\n# Measure\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/loudness \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"loudness_lufs\": -16.3}\n\n# Normalize to -14 LUFS (streaming target). Result is staged audio.\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/normalize \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"target_lufs\":-14,\"output_path\":\"out/normalized.wav\"}'\n# → {\"path\":\"out/normalized.wav\",\"size\":...,\"measured_lufs\":-16.3,\"target_lufs\":-14,...}\n```\n\n`target_lufs` must be in `[-70.0, -0.1]` — outside that range returns 400 (anything closer to 0 will clip catastrophically; anything below -70 silences the audio).\n\n### Effects chain (`/v1/audio/fx`)\n\nArbitrary pedalboard effect chain — full catalog. Different from `/v1/audio/master` (which runs presets).\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/fx \\\n  -d '{\n    \"file_path\":\"uploads/track.wav\",\n    \"effects\":[\n      {\"type\":\"Compressor\",\"params\":{\"threshold_db\":-18,\"ratio\":4.0}},\n      {\"type\":\"Reverb\",\"params\":{\"room_size\":0.5,\"wet_level\":0.3}},\n      {\"type\":\"PitchShift\",\"params\":{\"semitones\":2}},\n      {\"type\":\"Gain\",\"params\":{\"gain_db\":-3}}\n    ],\n    \"output_path\":\"out/fx.wav\"\n  }'\n```\n\nAllowed `type` values: `Compressor`, `Limiter`, `NoiseGate`, `Gain`, `Clipping`, `Distortion`, `Bitcrush`, `Reverb`, `Chorus`, `Delay`, `Phaser`, `PitchShift`, `HighShelfFilter`, `LowShelfFilter`, `PeakFilter`, `HighpassFilter`, `LowpassFilter`, `LadderFilter`, `IIRFilter`, `GSMFullRateCompressor`, `MP3Compressor`, `Resample`, `Invert`, `Convolution`.\n\n`VST3Plugin`, `AudioUnitPlugin`, `ExternalPlugin` are deliberately blocked — they load arbitrary native code from arbitrary filesystem paths. Server returns 400 if asked.\n\n### MIDI composition (`/v1/midi/compose`)\n\nTranscode a JSON song spec to a Standard MIDI File. **No AI runs server-side** — your agent writes the spec, audiolla turns it into MIDI bytes staged at `output_path`.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/compose \\\n  -d '{\n    \"output_path\":\"midi/song.mid\",\n    \"spec\":{\n      \"tempo_bpm\": 120,\n      \"time_signature\": [4, 4],\n      \"key_signature\": \"C\",\n      \"tracks\": [\n        {\"name\":\"Lead\",\"program\":0,\"channel\":0,\"notes\":[\n          {\"pitch\":60,\"start_beats\":0.0,\"duration_beats\":0.5,\"velocity\":100},\n          {\"pitch\":64,\"start_beats\":0.5,\"duration_beats\":0.5,\"velocity\":100},\n          {\"pitch\":67,\"start_beats\":1.0,\"duration_beats\":0.5,\"velocity\":100}\n        ]},\n        {\"name\":\"Drums\",\"program\":0,\"channel\":9,\"notes\":[\n          {\"pitch\":36,\"start_beats\":0.0,\"duration_beats\":0.1,\"velocity\":110}\n        ]}\n      ]\n    }\n  }'\n```\n\nSpec fields (inside the `spec` object):\n\n| Field | Type | Default | Notes |\n|-------|------|---------|-------|\n| `tempo_bpm` | float | 120 | 1.0 ≤ bpm ≤ 999.0 |\n| `time_signature` | `[num, den]` | `[4, 4]` | denominator must be 1/2/4/8/16/32 |\n| `key_signature` | string | none | `\"C\"`, `\"Am\"`, `\"F#\"`, `\"Bbm\"` — letter [+ #/b] [+ m for minor] |\n| `ticks_per_beat` | int | 480 | 24 ≤ tpb ≤ 1920 |\n| `tracks[].name` | string | none | optional, writes a `track_name` meta event |\n| `tracks[].program` | int 0-127 | 0 | General MIDI program (Acoustic Grand Piano = 0, Distortion Guitar = 30, Synth Brass 1 = 62, etc.) |\n| `tracks[].channel` | int 0-15 | 0 | **Channel 9 is the GM drum channel** — pitch maps to drum kit, not piano |\n| `tracks[].volume` | int 0-127 | 100 | MIDI CC#7 — initial volume |\n| `tracks[].pan` | int 0-127 | 64 | MIDI CC#10 — initial pan (64 = centre) |\n| `tracks[].notes[].pitch` | int 0-127 | required | 60 = middle C |\n| `tracks[].notes[].start_beats` | float ≥ 0 | 0 | beat-based absolute position |\n| `tracks[].notes[].duration_beats` | float > 0 | required | must be > 1/64 beat (≈ a 256th note) |\n| `tracks[].notes[].velocity` | int 1-127 | 100 | |\n\nGM drum kit reference for channel 9: 35 acoustic bass drum, 36 kick, 38 snare, 39 hand clap, 40 electric snare, 42 closed hi-hat, 46 open hi-hat, 49 crash, 51 ride, 57 crash 2.\n\nSpec validation is fail-loud — bad pitch / negative duration / unknown program returns a 400 with the offending path in the message (e.g. `tracks[1].notes[3].pitch must be in [0, 127], got 200`).\n\nOne of `output_path` / `output_url` is required — the staged MIDI is then referenced via `file_path` on any subsequent MIDI call.\n\n### MIDI inspection (`/v1/midi/inspect`)\n\nRead the structure of any Standard MIDI File. Analysis-only, returns JSON.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/inspect \\\n  -d '{\"file_path\":\"midi/song.mid\"}'\n# {\n#   \"type\": 1, \"ticks_per_beat\": 480, \"length_seconds\": 16.0,\n#   \"tempo_changes\": [{\"tick\": 0, \"bpm\": 120.0}],\n#   \"time_signatures\": [{\"tick\": 0, \"numerator\": 4, \"denominator\": 4}],\n#   \"tracks\": [\n#     {\"index\": 1, \"name\": \"Lead\", \"note_on_count\": 32,\n#      \"channels\": [0], \"programs\": [0], \"length_beats\": 8.0},\n#     ...\n#   ],\n#   \"track_count\": 3, \"size_bytes\": 1024\n# }\n```\n\nNon-MIDI input returns 400 with `\"MThd\"` mentioned in the detail.\n\n### MIDI transformation (`/v1/midi/transform`)\n\nModify an existing MIDI file. Result is staged at `output_path` / `output_url`.\n\n```bash\n# Transpose all non-drum tracks up an octave\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"transpose_semitones\":12,\"output_path\":\"midi/transposed.mid\"}'\n\n# Override tempo to 140 BPM\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"tempo_bpm\":140,\"output_path\":\"midi/fast.mid\"}'\n\n# Drop the drum channel\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"drop_channels\":[9],\"output_path\":\"midi/no-drums.mid\"}'\n\n# Keep only channels 0 and 1\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"keep_channels\":[0,1],\"output_path\":\"midi/two-ch.mid\"}'\n\n# Quantize to 1/16th notes (0.25 beats)\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"quantize\":0.25,\"output_path\":\"midi/quantized.mid\"}'\n```\n\nTransform params (all optional — omit for a no-op):\n\n| Param | Type | Notes |\n|-------|------|-------|\n| `transpose_semitones` | int ±48 | Shifts all non-drum (non-ch9) pitches. Out-of-range notes after shift are dropped (not clipped). |\n| `tempo_bpm` | float 1–999 | Replaces all `set_tempo` events. |\n| `quantize` | float > 0 | Beat grid in beats (0.25 = 1/16th at 4/4). Snaps note starts; note-off shifts by the same delta to preserve duration. |\n| `keep_channels` | int array (0–15) | Whitelist — drop all other channels. Mutually exclusive with `drop_channels`. |\n| `drop_channels` | int array (0–15) | Blacklist — drop only these channels. Mutually exclusive with `keep_channels`. |\n\nSupplying both `keep_channels` and `drop_channels` returns 400.\n\n### MIDI rendering (`/v1/midi/render`)\n\nSynthesise MIDI to audio via fluidsynth. Default SoundFont is FluidR3_GM (bundled in the prod image). Override per-request with a staged `.sf2`.\n\n```bash\n# Render a staged MIDI to staged audio\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/render \\\n  -d '{\"file_path\":\"midi/song.mid\",\"output_format\":\"wav\",\"output_path\":\"audio/song.wav\"}'\n\n# Render with a custom SoundFont (stage it first)\ncurl -X PUT -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/octet-stream' \\\n  --data-binary @my.sf2 \\\n  $AUDIOLLA_URL/v1/files/sf/orchestral.sf2\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/render \\\n  -d '{\"file_path\":\"midi/song.mid\",\"soundfont_path\":\"sf/orchestral.sf2\",\"output_format\":\"flac\",\"gain\":0.3,\"samplerate\":48000,\"output_path\":\"audio/orch.flac\"}'\n```\n\n`gain` range `[0.0, 5.0]` — default `0.5` is calibrated to avoid clipping on percussive MIDI. `samplerate` must be 22050 / 44100 / 48000 / 88200 / 96000.\n\n### MIDI generate (`/v1/midi/generate`)\n\nOne-shot compose + render. Body has the same `spec` field as `/v1/midi/compose` plus audio knobs (`output_format`, `soundfont_path`, `gain`, `samplerate`). Result audio is staged at `output_path` / `output_url`.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/generate \\\n  -d '{\n    \"output_format\":\"wav\",\n    \"output_path\":\"songs/v1.wav\",\n    \"spec\":{\"tempo_bpm\":120,\"tracks\":[{\"channel\":0,\"notes\":[\n      {\"pitch\":60,\"start_beats\":0,\"duration_beats\":1,\"velocity\":100}\n    ]}]}\n  }'\n```\n\n### File staging\n\nA simple server-side file store under `/v1/files`. **This is the only multipart-ish route in the API** — the body is raw bytes (`application/octet-stream`). Plain CRUD: upload, list, download, delete. Once a file is staged, every audio endpoint references it by relative path via the `file_path` field in its JSON body.\n\n```bash\n# Upload (path can have subdirectories: uploads/bands/myband/track.wav)\ncurl -X PUT -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/octet-stream' \\\n  --data-binary @track.wav \\\n  $AUDIOLLA_URL/v1/files/uploads/mytrack.wav\n\n# Use the staged path on any audio call\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/mytrack.wav\",\"engine\":\"htdemucs\",\"stems\":[\"vocals\"],\"output_path\":\"stems/mytrack-vocals.wav\"}'\n# → {\"path\":\"stems/mytrack-vocals.wav\",\"size\":...,\"engine\":\"htdemucs\",\"stem\":\"vocals\",\"output_format\":\"wav\"}\n\n# List\ncurl -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/files\n\n# Download (raw bytes — Content-Type matches the stored file)\ncurl -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  $AUDIOLLA_URL/v1/files/uploads/mytrack.wav -o copy.wav\n\n# Delete\ncurl -X DELETE -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  $AUDIOLLA_URL/v1/files/uploads/mytrack.wav\n```\n\nPath traversal (`..`, leading `/`, etc.) is rejected with 400. Symlinks are not followed. Size cap is `AUDIOLLA_MAX_UPLOAD_BYTES`.\n\n### Input and output modes (every audio endpoint)\n\nEvery audio endpoint accepts exactly one of two input forms — supplying zero or both returns 400:\n\n- `file_path` — relative path under FILES_DIR (pre-staged via `PUT /v1/files/{path}`)\n- `file_url` — remote URL the server fetches (subject to the `AUDIOLLA_FETCH_MODE` policy — see below)\n\nAudio-producing endpoints (separate, master, transform, normalize, fx, restore, enhance, visualize, midi compose/transform/render/generate, melody-as-midi, beats-with-click-track, etc.) require exactly one of:\n\n- `output_path` — server writes the result to `FILES_DIR / <path>`; response is JSON `{path, size, ...}`\n- `output_url` — server PUTs the result to a presigned URL; response is JSON `{url, size, ...}`\n\n`output_path` and `output_url` are mutually exclusive — supplying both is 400. Supplying neither is 400 too (no inline-bytes audio response exists in v1.0.0) — except when `async_job=true`, which auto-stages to `jobs/{job_id}.{ext}` if neither is set.\n\nAnalysis-only endpoints (`/v1/audio/analyze`, `/v1/audio/onsets`, `/v1/audio/fingerprint`, `/v1/audio/loudness`, beats without `click_track`, silence without `trim_mode`, etc.) ignore `output_path` / `output_url` — they return their JSON data directly.\n\nThe master endpoint additionally accepts `reference_path` xor `reference_url` for the reference track in `mode=reference` — same exactly-one-of rule.\n\n### Remote URLs (file_url / output_url)\n\nThe server-side URL fetch is **disabled by default**. To enable it, the operator sets:\n\n```\nAUDIOLLA_FETCH_MODE = disabled | allowlist | denylist     (default: disabled)\nAUDIOLLA_FETCH_HOSTS = comma-separated host patterns       (required when mode=allowlist)\nAUDIOLLA_FETCH_SCHEMES = https,http                        (default: https only)\nAUDIOLLA_FETCH_TIMEOUT = 30s                               (per fetch/upload)\nAUDIOLLA_FETCH_ALLOW_PRIVATE = false                       (allow private/loopback IPs)\nAUDIOLLA_FETCH_MAX_REDIRECTS = 5\n```\n\nHost patterns are exact match (`bucket.s3.amazonaws.com`) or single-wildcard subdomain (`*.s3.amazonaws.com`, matches any `<x>.s3.amazonaws.com` but NOT `s3.amazonaws.com` itself).\n\nAlways-on protections regardless of mode:\n- DNS-resolved private / loopback / link-local / metadata-service IPs (`169.254.169.254`) rejected unless `AUDIOLLA_FETCH_ALLOW_PRIVATE=true`\n- Only schemes in `AUDIOLLA_FETCH_SCHEMES` accepted; `file://`, `gopher://`, etc. always rejected\n- Each redirect's `Location` re-validated through the full policy before following\n- Body streamed; abort if it exceeds `AUDIOLLA_MAX_UPLOAD_BYTES`\n\nIf you're scripting and the server returns `URL fetch/upload is disabled` (400), tell the user — don't try to bypass it. The operator chose `disabled` for a reason.\n\nExample — fetch from S3, master, PUT to a presigned URL:\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/master \\\n  -d '{\n    \"file_url\":\"https://my-bucket.s3.amazonaws.com/track.wav\",\n    \"mode\":\"chain\",\n    \"preset\":\"loud\",\n    \"output_url\":\"https://my-bucket.s3.amazonaws.com/mastered.wav?X-Amz-Signature=...\"\n  }'\n# → {\"url\":\"...\",\"size\":...,\"engine\":\"pedalboard-chain\",\"mode\":\"chain\",\"output_format\":\"wav\"}\n```\n\n## MCP\n\naudiolla exposes a Model Context Protocol server at `/v1/mcp` using the streamable HTTP transport. Same auth as REST — pass `Authorization: Bearer $AUDIOLLA_TOKEN`.\n\nThe MCP contract mostly mirrors REST: every audio tool requires exactly one of `file_path` or `file_url` for input (same `AUDIOLLA_FETCH_MODE` policy as REST), and most audio-producing tools require exactly one of `output_path` or `output_url` for output, returning either `{path, size, output_format}` or `{url, size, output_format}`. There are narrow exceptions where the MCP transport still returns inline base64 instead of staging — see the `beats`/`melody`/`silence` rows below.\n\nThe table below is a curated subset — the server exposes ~78 MCP tools total (one per REST endpoint, roughly), including ops like `trim`, `mix`, `concat`, `fade`, `reverse`, `loop`, `speed`, `convert`, `pan`, `eq`, `key_match`, `bpm_match`, `stereo_width`, `mid_side`, `sidechain_duck`, `multiband_compress`, `transient_shaper`, `convolution_reverb`, `deess`, `dj_prep`, `pitch_correct`, `repair_audio`, `find_loop_point`, `chords`, `chords_to_midi`, `vad`, `diarize`, `stretch`, `tag`, `embed`, `classify`, `similar`, `hpss`, `noise_reduce`, `detect_clipping`, `slice_at_beats`, `audio_thumbnail`, `stereo_field`, `midi_quantize`, `midi_humanize`, `drum_pattern`, `generate_music`, `list_presets`, `describe_preset`, `run_preset`, `list_ops`, `run_pipeline_tool`, `audio_metadata`, `info`, and job control (`list_jobs`, `get_job`, `cancel_job`). Each mirrors its REST counterpart's params 1:1 (see the REST reference above for the full param list of each) — connect an MCP client and call `list_engines`/introspect the tool list, or check `GET /v1/catalog`, for the authoritative current set.\n\n| Tool | Inputs | Output |\n|------|--------|--------|\n| `list_engines` | — | engine catalog with `loaded` flag |\n| `separate` | `engine`, `stems`, `file_path` or `file_url`, `output_paths: {stem: path}` or `output_urls: {stem: url}` (one of the two maps is required — no single `output_path`/`output_url`) | `{staged_stems: {stem: {path, size}}, output_format}` or `{uploaded_stems: {stem: {url, size}}, output_format}` |\n| `master` | `mode`, `file_path` or `file_url`, `reference_path` or `reference_url` (mode=reference), `preset` (mode=chain), `target_lufs`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `analyze` | `file_path` or `file_url`, `features` | librosa feature dict |\n| `beats` | `file_path` or `file_url`, `click_track`, `output_format` | `{tempo_bpm, beats, beat_count, duration}` — with `click_track=true`, adds `click_track_base64` (inline base64, NOT staged — no `output_path`/`output_url` param on this tool) |\n| `onsets` | `file_path` or `file_url` | `{onsets, ...}` |\n| `melody` | `file_path` or `file_url`, `fmin`, `fmax`, `as_midi` | pitch contour — with `as_midi=true`, adds `midi_base64` (inline base64, NOT staged — no `output_path`/`output_url` param on this tool) |\n| `segments` | `file_path` or `file_url`, `num_segments` (default 6) | `{segments: [{label, start_sec, end_sec}, ...]}` |\n| `silence` | `file_path` or `file_url`, `threshold_db`, `min_duration_sec`, `trim_mode`, `output_format` | `{silent_ranges, non_silent_ranges, duration, ...}` — with `trim_mode` set, adds `trimmed_audio_base64` (inline base64, NOT staged — no `output_path`/`output_url` param on this tool) |\n| `visualize` | `file_path` or `file_url`, `engine`, `mode` (`spectrogram`/`waveform` for static PNG, or `spectrum`/`waves`/`cqt`/`freqs`/`volume`/`vectorscope`/`phasemeter`/`histogram` for animated video), `width`, `height`, `color`, `scale`, `fps`, `container`, `output_path` or `output_url` | `{path, size}` or `{url, size}` — replaces the separate `spectrogram`/`waveform` REST endpoints as one tool with a `mode` switch |\n| `fingerprint` | `file_path` or `file_url`, `analyze_seconds`, `return_raw` | `{duration, fingerprint, fingerprint_raw?}` |\n| `transform` | `operations`, `file_path` or `file_url`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `loudness` | `file_path` or `file_url` | `{loudness_lufs}` — measurement only |\n| `normalize` | `file_path` or `file_url`, `target_lufs`, `output_format`, `output_path` or `output_url` | `{path or url, size, measured_lufs, target_lufs}` |\n| `fx` | `effects`, `file_path` or `file_url`, `output_format`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `midi_compose` | `spec` (song JSON), `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `midi_inspect` | `file_path` or `file_url` (MIDI) | `{type, ticks_per_beat, tempo_changes, tracks, ...}` |\n| `midi_transform` | `file_path` or `file_url` (MIDI), `transpose_semitones`, `tempo_bpm`, `quantize`, `keep_channels`, `drop_channels`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `midi_render` | `file_path` or `file_url` (MIDI), `soundfont_path`, `gain`, `samplerate`, `output_format`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `midi_generate` | `spec`, `soundfont_path`, `gain`, `samplerate`, `output_format`, `output_path` or `output_url` | `{path, size, midi_size}` or `{url, size, midi_size}` |\n| `restore` | `file_path` or `file_url`, `engine` (default `uvr-dereverb`; also `uvr-deecho`, `uvr-denoise`), `aggressive` (uvr-deecho only), `output_format`, `output_url` only (no `output_path` on this tool — inline base64 is not returned either; `output_url` is required) | `{url, size, engine, aggressive, output_format}` |\n| `denoise` | `file_path` or `file_url`, `engine`, `output_format`, `output_path` or `output_url` | `{path, size, engine, output_format}` or `{url, size, engine, output_format}` — thin shim over `restore`/`noise_reduce` |\n| `audio_to_midi` | `file_path` or `file_url`, `engine`, `onset_threshold`, `frame_threshold`, `minimum_note_lengt\n\nFile v1.4.2:_meta.json\n\n{\n  \"ownerId\": \"kn79dhvmpjng4rp2jjk8k0v5xx80ccbk\",\n  \"slug\": \"audiolla\",\n  \"version\": \"1.4.2\",\n  \"publishedAt\": 1785018598691\n}\n\nFile v1.4.2:references/setup.md\n\n# audiolla setup\n\n## Requirements\n\n- Linux/macOS host with Docker\n- ~3 GB disk for the CPU image, ~8 GB for the CUDA image (PyTorch + CUDA runtime is heavy)\n- 4 GB RAM minimum, 8 GB recommended (Demucs separation peaks around 3-5 GB)\n- NVIDIA GPU + drivers + `nvidia-container-toolkit` for the CUDA variant (CUDA 12.6+)\n\n## Quick install\n\nCPU image — no GPU needed:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nCUDA image — GPU-accelerated Demucs separation:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  --gpus all \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_DEVICE=cuda \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest-cuda\n```\n\nFirst run downloads the image (~3 GB CPU, ~8 GB CUDA), then on container start prefetches Demucs model weights (~600 MB) into `/data/torch_cache/`. The model fetch logs to the container's stdout — `docker logs -f audiolla` to watch.\n\nAfter that, subsequent runs reuse the volume and skip the download.\n\n**Verify:** `curl http://localhost:8000/healthz` → `{\"ok\": true, \"device\": \"cpu\", \"engines\": [...]}`.\n\n## Configuration\n\nAll config is via environment variables passed at `docker run`:\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `AUDIOLLA_DEVICE` | `auto` | `auto`, `cpu`, `cuda`, or `cuda:N` for a specific GPU |\n| `AUDIOLLA_ENGINES_FILE` | `/app/engines.json` | path to engines registry |\n| `AUDIOLLA_DATA_DIR` | `/data` | where models and staged files live |\n| `AUDIOLLA_AUTH_TOKEN` | _(none)_ | bearer token; empty means no auth |\n| `AUDIOLLA_ENABLED_ENGINES` | _(all)_ | comma-separated slugs to allow; empty = all |\n| `AUDIOLLA_PRELOAD` | _(none)_ | comma-separated slugs to load into memory at startup |\n| `AUDIOLLA_ENGINE_TTL` | `600` | seconds idle before an engine is unloaded (`10m` also works) |\n| `AUDIOLLA_SWEEPER_INTERVAL` | `60` | how often the idle-engine sweeper runs, in seconds |\n| `AUDIOLLA_MAX_UPLOAD_BYTES` | `209715200` | upload cap (default 200 MB); also caps remote URL fetch body size |\n| `AUDIOLLA_FETCH_MODE` | `disabled` | `disabled` / `allowlist` / `denylist` — server-side fetch policy for `file_url` and `output_url` |\n| `AUDIOLLA_FETCH_HOSTS` | _(none)_ | comma-separated host patterns (`bucket.s3.amazonaws.com`, `*.s3.amazonaws.com`) — required when mode=allowlist |\n| `AUDIOLLA_FETCH_SCHEMES` | `https` | comma-separated schemes; add `http` only for trusted local networks |\n| `AUDIOLLA_FETCH_ALLOW_PRIVATE` | `false` | allow URLs resolving to private / loopback / link-local IPs (e.g. internal MinIO) |\n| `AUDIOLLA_FETCH_TIMEOUT` | `30` | per-fetch/upload timeout (seconds; also accepts `30s`, `1m`) |\n| `AUDIOLLA_FETCH_MAX_REDIRECTS` | `5` | max redirects per fetch; each `Location` re-validated through the policy |\n| `AUDIOLLA_SOUNDFONT` | `/usr/share/sounds/sf2/FluidR3_GM.sf2` (prod images) | Default SoundFont (`.sf2`) path used by `/v1/midi/render`. Empty = midi-render refuses unless `soundfont_path` is passed on the request. Prod images install FluidR3_GM via `apt install fluid-soundfont-gm`. |\n| `AUDIOLLA_UVR_MODELS_DIR` | `/data/uvr_models` | directory holding UVR (`audio-separator`) `.ckpt`/`.pth` model files for `uvr-dereverb`, `uvr-deecho`, `uvr-denoise`, `uvr-karaoke`, `uvr-vocal-bsr` — not bundled in the image |\n| `AUDIOLLA_ENABLE_NONCOMMERCIAL` | `false` | opt-in gate (`1`/`true`/`yes`/`on`) required to use CC-BY-NC-licensed `musicgen-small` / `musicgen-medium` weights |\n| `AUDIOLLA_LOAD_TIMEOUT` | `300` | seconds allowed for an engine's cold load before it's treated as failed (also accepts Go-style durations) |\n| `AUDIOLLA_JOB_TTL` | `3600` | seconds a completed/failed async job's record is retained before eviction |\n| `AUDIOLLA_JOB_MAX_CONCURRENT` | `8` | parsed at startup for future async-job concurrency limiting; not yet enforced anywhere in the current build |\n\n### Authentication\n\nBy default audiolla runs with no auth — anyone who can reach port 8000 can use it. For anything beyond `localhost`, set a bearer token:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_AUTH_TOKEN=\"$(openssl rand -hex 32)\" \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nEvery endpoint except `/healthz` then requires `Authorization: Bearer <token>`. Without it the server returns 401.\n\n> **`AUDIOLLA_AUTH_TOKEN` must be set to a strong random value before audiolla is reachable by anything other than localhost.** Without a token, anyone who can hit port 8000 can run arbitrary audio processing on your hardware (Demucs is CPU/GPU-heavy — a hostile caller can keep your machine saturated indefinitely). They can also upload up to `AUDIOLLA_MAX_UPLOAD_BYTES` per request to your staging area. Generate a token with `openssl rand -hex 32` and keep it out of git.\n\n### Remote URL fetching (file_url / output_url)\n\nAudiolla can fetch input files from a URL and PUT outputs to presigned URLs (S3, R2, etc.). This is **disabled by default** because the fetch path is a classic SSRF surface — without guardrails, an attacker can use it to read your cloud metadata service or probe internal hosts.\n\nPick a mode that matches your setup:\n\n```bash\n# Default — no URL I/O.\n-e AUDIOLLA_FETCH_MODE=disabled\n\n# Allowlist — preferred. Only listed hosts can be fetched/PUT to.\n-e AUDIOLLA_FETCH_MODE=allowlist\n-e AUDIOLLA_FETCH_HOSTS=\"*.s3.amazonaws.com,*.r2.cloudflarestorage.com,my-bucket.example.com\"\n\n# Denylist — anything goes except listed hosts. Leaky by design — only\n# safe with AUDIOLLA_FETCH_ALLOW_PRIVATE=false (the default), which\n# already blocks private IPs and the metadata service. Use this only\n# if you control the network or have a strong reason.\n-e AUDIOLLA_FETCH_MODE=denylist\n-e AUDIOLLA_FETCH_HOSTS=\"*.internal,localhost\"\n```\n\nHost pattern syntax — exact match or single-wildcard subdomain. `*.s3.amazonaws.com` matches `bucket.s3.amazonaws.com` but NOT `s3.amazonaws.com` itself (add that explicitly if needed).\n\nAlways-on protections regardless of mode:\n\n- DNS-resolved private / loopback / link-local IPs rejected. The AWS/GCP/Azure metadata service at `169.254.169.254` falls into this. Toggle off only if you genuinely need internal S3-compatible storage on a private network.\n- Schemes restricted to `AUDIOLLA_FETCH_SCHEMES` (default `https`). `file://`, `gopher://`, etc. are always rejected.\n- Each HTTP redirect's `Location` re-validated through the full policy before following.\n- Body size capped at `AUDIOLLA_MAX_UPLOAD_BYTES` — streamed, abort on overrun.\n- Every fetch / upload URL logged at INFO.\n\nFor internal MinIO / S3-compatible storage:\n\n```bash\n-e AUDIOLLA_FETCH_MODE=allowlist \\\n-e AUDIOLLA_FETCH_HOSTS=minio.internal.example.com \\\n-e AUDIOLLA_FETCH_ALLOW_PRIVATE=true\n```\n\n### Engine selection\n\nOnly enable the engines you actually need:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_ENABLED_ENGINES=htdemucs,matchering,librosa-analyze \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nDisabled engines are absent from `GET /v1/engines` and return 404 on use.\n\n### Preloading\n\nBy default, engines lazy-load on first request. To avoid the cold-start latency on critical engines, preload at startup:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_PRELOAD=htdemucs,matchering \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nPreload happens during container startup; the server doesn't accept traffic until preload finishes. Failed preloads log a warning and continue.\n\n### Idle unload\n\nThe idle sweeper unloads engines that haven't been used for `AUDIOLLA_ENGINE_TTL` seconds. This frees memory (Demucs holds several GB when loaded). Set to `0` to disable.\n\n```bash\n-e AUDIOLLA_ENGINE_TTL=10m       # Go-style duration\n-e AUDIOLLA_ENGINE_TTL=600       # plain seconds\n-e AUDIOLLA_ENGINE_TTL=0         # never unload\n```\n\nMemory footprints (approximate):\n- `htdemucs` / `mdx_extra`: ~1.5 GB\n- `htdemucs_ft`: ~2 GB (CUDA-only at usable speed)\n- `htdemucs_6s`: ~2 GB\n- `matchering`: ~200 MB\n- `pedalboard-chain`, `librosa-analyze`, `sox-transform`: negligible (no model weights)\n\n## Data directory layout\n\n`/data` (or wherever `AUDIOLLA_DATA_DIR` points) contains:\n\n```\n/data/\n  torch_cache/          # Demucs model weights — survives container restarts\n    hub/checkpoints/    # downloaded .th files\n  files/                # staging area for /v1/files endpoints\n  models/               # additional model storage (currently unused)\n  hf/                   # HuggingFace cache (CUDA image only — HF_HOME)\n```\n\nMount the same path on subsequent runs to skip the model re-download and keep staged files.\n\n## docker-compose\n\n```yaml\nservices:\n  audiolla:\n    image: psyb0t/audiolla:latest        # or :latest-cuda\n    container_name: audiolla\n    restart: unless-stopped\n    ports:\n      - \"127.0.0.1:8000:8000\"           # bind to loopback only\n    volumes:\n      - ./data:/data\n    environment:\n      AUDIOLLA_DEVICE: auto\n      AUDIOLLA_AUTH_TOKEN: ${AUDIOLLA_AUTH_TOKEN}     # from .env\n      AUDIOLLA_ENGINE_TTL: 10m\n      AUDIOLLA_MAX_UPLOAD_BYTES: 209715200\n    # For CUDA image only:\n    # deploy:\n    #   resources:\n    #     reservations:\n    #       devices:\n    #         - driver: nvidia\n    #           count: all\n    #           capabilities: [gpu]\n```\n\n`.env`:\n\n```bash\nAUDIOLLA_AUTH_TOKEN=<openssl rand -hex 32 output>\n```\n\n## Logs\n\naudiolla logs to stdout — read with `docker logs -f audiolla`. Key log lines:\n\n- `audiolla starting: device=... engines=[...] ttl=... files_dir=... auth=on|off` — boot banner\n- `[entrypoint] demucs variants to prefetch: [...]` — model prefetch\n- `idle sweeper: unloading <slug> (idle ...s >= ...s)` — engine eviction\n- `evicting N sibling engine(s) before loading <slug>` — pre-separation eviction\n- `preload <slug> failed` — `AUDIOLLA_PRELOAD` entry couldn't load\n\n## Public access\n\naudiolla has no built-in TLS or reverse-proxy support. For public access:\n\n1. **Set `AUDIOLLA_AUTH_TOKEN`** to a strong random value first. Non-negotiable.\n2. Front it with a reverse proxy that terminates TLS (Caddy, nginx, Traefik).\n3. Consider a tailnet-only deployment if you don't actually need internet exposure — Tailscale/WireGuard authenticate at the network layer.\n4. Per-IP rate limiting at the proxy. Demucs runs at line rate against a hostile caller will pin a GPU or saturate every CPU core; rate limiting is the only thing between you and a denial-of-wallet event.\n\n## Build from source\n\nIf you want to build the image yourself instead of pulling:\n\n```bash\ngit clone https://github.com/psyb0t/docker-audiolla\ncd docker-audiolla\nmake build         # CPU image\nmake build-cuda    # CUDA image\nmake run           # builds + runs CPU image on port 8000\n```\n\nHeavy ML deps are hash-locked in `requirements-heavy-{cpu,cuda}.txt`. Light deps live in `uv.lock`. Build will fail with a hash mismatch if anything has been tampered with — that's the design.\n\n## Troubleshooting\n\n**`/healthz` returns but engine calls 404 with \"unknown engine\"**\nThe engine is filtered out by `AUDIOLLA_ENABLED_ENGINES`. Check `GET /v1/engines` for what's actually exposed.\n\n**`htdemucs_ft` returns 400 with \"cuda_only\"**\nThe CUDA-only fine-tuned variant won't run on the CPU image. Either use the CUDA image with `--gpus all`, or pick `htdemucs` / `htdemucs_ft`'s lighter sibling.\n\n**Separation request hangs for minutes**\nDemucs on CPU is slow — `htdemucs` on a 3-min track takes 1-3 min on a mid-range CPU; `htdemucs_ft` is ~4x that. CUDA image is 5-10x faster. Watch `docker logs -f audiolla` for engine load progress.\n\n**Upload returns 413**\nHit the `AUDIOLLA_MAX_UPLOAD_BYTES` cap (default 200 MB). Bump it via env var, or use `/v1/files` PUT (same cap applies per upload but you can pipeline multiple files).\n\n**Upload returns 401**\nThe server has `AUDIOLLA_AUTH_TOKEN` set and the request didn't carry `Authorization: Bearer ...` (or carried the wrong token).\n\n**Container exits during startup**\nCheck `docker logs audiolla` for `preload <slug> failed` or engine-init exceptions. The most common cause is a corrupt `/data/torch_cache/` — delete the volume and let the container re-download.\n\nFile v1.4.2:skill-card.md\n\n## Description:\n\nConnects to a user-deployed Audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, loudness normalization, and related audio file workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[psyb0t](https://clawhub.ai/user/psyb0t)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and audio-production operators use this skill to drive an existing Audiolla HTTP or MCP server for file-based audio processing, analysis, generation, metadata, MIDI, and workflow-preset tasks.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: An Audiolla server exposed beyond localhost without authentication can let reachable users run expensive audio-processing jobs and upload large files.\n\nMitigation: Bind Audiolla to 127.0.0.1 unless remote access is intentional, and set a strong AUDIOLLA_AUTH_TOKEN before exposing the service.\n\nRisk: Remote URL fetch and upload support can create SSRF or unintended network access risk when broadly enabled.\n\nMitigation: Keep AUDIOLLA_FETCH_MODE disabled unless needed, and use an allowlist with restricted schemes and hosts when URL I/O is required.\n\nRisk: Mutable Docker image tags can change between deployments.\n\nMitigation: Prefer pinned image digests over latest tags for repeatable deployments.\n\n## Reference(s):\n\n- [Audiolla setup guide](references/setup.md)\n- [ClawHub Audiolla skill page](https://clawhub.ai/psyb0t/skills/audiolla)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with inline shell commands and JSON request examples]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Produces guidance for API calls against a user-operated Audiolla server; audio-producing operations return server-side file paths or URLs rather than inline audio bytes.]\n\n## Skill Version(s):\n\n1.4.2 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.4.1: 4 files, 27450 bytes\n\nFiles: references/setup.md (11452b), skill-card.md (2614b), SKILL.md (65384b), _meta.json (127b)\n\nFile v1.4.1:SKILL.md\n\n---\nname: audiolla\ndescription: HTTP/MCP client for a user-deployed audiolla audio-production server. Use ONLY when the user has explicitly named audiolla AND provided AUDIOLLA_URL (or has it set in the environment). Capabilities: stem separation (Demucs / MDX / BS-Roformer), mastering (matchering reference / pedalboard preset chain), MIR analysis (BPM, key, LUFS, spectral features, beat grid, onset detection, melody contour, structural segmentation via librosa), DSP transforms (gain, EQ, compand, reverb, pitch, tempo via SoX), loudness measurement and normalization, generic effects chains (full pedalboard catalog as ordered chain), multiband compression (LR4 crossovers), transient shaping, sidechain ducking, de-essing, mid/side encode-decode, parametric EQ, panning, stereo width, silence detection and trimming, audio repair (declip + dehum), clip detection, harmonic/percussive separation, time-stretch and pitch-shift, BPM/key matching, pitch correction (auto-tune), beat slicing, audio thumbnail extraction, convolution reverb, static PNG spectrogram/waveform and 8-mode animated MP4/WebM video (ffmpeg), Chromaprint acoustic fingerprinting, AudioSet tagging, CLAP audio embeddings + similarity + zero-shot classification, ID3/Vorbis/FLAC metadata read/write, MIDI composition from JSON spec, MIDI inspection, MIDI transformation (transpose/quantize/tempo/channel-filter), MIDI quantize and humanize, drum pattern generation, MIDI rendering via fluidsynth, polyphonic audio-to-MIDI transcription (Spotify basic-pitch ONNX), chords-to-MIDI conversion, AI audio restoration (de-reverb, de-echo, AI de-noise via UVR/audio-separator), DSP noise reduction, neural speech/vocal enhancement (DeepFilterNet DF3), voice activity detection (silero-vad), speaker diarization (pyannote 3.1), DJ prep (BPM + key + Camelot + LUFS in one call), loop-point detection, curated server-side workflow presets (master-for-spotify, podcast-cleanup, vocal-cleanup) and ad-hoc op pipelines that chain multiple operations server-side. v1.0.0 API is JSON-everywhere: every audio endpoint takes a JSON body; the ONLY multipart route is `PUT /v1/files/{path}` for raw byte uploads. Audio I/O supports two input modes (`file_path` referencing a pre-staged file under FILES_DIR, xor `file_url` — only when the operator has enabled AUDIOLLA_FETCH_MODE) and two output modes (`output_path` writing back to staging, xor `output_url` PUTing to a presigned URL). There is no inline-bytes audio response anywhere — every audio-producing endpoint returns JSON describing where the result landed. Audiolla only fetches/uploads to URLs when the operator has explicitly enabled AUDIOLLA_FETCH_MODE — if a request returns \"URL fetch/upload is disabled\", do NOT try to bypass it. Do not use this skill for generic audio-processing questions or for users who haven't named audiolla.\ncompatibility: Requires curl and a running audiolla instance (Docker image psyb0t/audiolla:latest or :latest-cuda). AUDIOLLA_URL env var must be set by the user (default http://localhost:8000). AUDIOLLA_TOKEN required only when the server has AUDIOLLA_AUTH_TOKEN configured; obtain from the AUDIOLLA_TOKEN env var or by asking the user — never read tokens from repo files autonomously.\nmetadata:\n  author: psyb0t\n  homepage: https://github.com/psyb0t/docker-audiolla\n---\n\n# audiolla\n\nHTTP + MCP client for an audiolla server that the user has already deployed. This skill talks to a running audiolla instance — it does not stand one up, does not download model weights manually, and does not modify the server config on its own initiative.\n\nFor installation and setup, see [references/setup.md](references/setup.md).\n\n## Authoritative endpoint reference: `GET /v1/catalog`\n\nThis skill documents the most common patterns. The **full, always-current** list of every endpoint is `GET /v1/catalog` (17 categories, ~85 endpoints). Always check the catalog when looking for an operation that isn't shown here — the server is the source of truth, this file is a curated reference.\n\n```bash\n# List every endpoint grouped by category\ncurl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | {name, count: (.endpoints | length)}'\n\n# Find endpoints in one category\ncurl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | select(.name == \"dynamics\") | .endpoints'\n```\n\nCompanion discovery endpoints: `GET /v1/engines` (engines + loaded/idle status), `GET /v1/presets` (curated workflows), `GET /v1/ops` (the ~24 pipeline op slugs).\n\n## When to use this skill\n\nThe user has audiolla running and asks you to:\n- Pull stems (vocals / drums / bass / etc.) out of a track\n- Master a track against a reference recording (matchering) or via preset chain\n- **Run a curated workflow** (`master-for-spotify`, `podcast-cleanup`, `vocal-cleanup`) via a single `POST /v1/presets/{name}` call\n- **Chain ad-hoc operations server-side** via `POST /v1/pipeline` (no re-upload between steps)\n- Get BPM, key, LUFS, duration, or spectral features for a file\n- Detect beat grid, onsets, dominant melody, or structural segments\n- Detect chords + key (separate from BPM/LUFS)\n- Detect or trim silence\n- Generate a spectrogram, waveform image, or animated visualisation video\n- Compute a Chromaprint acoustic fingerprint\n- Apply a DSP chain (gain, EQ, compression, reverb, pitch shift, tempo via SoX OR full pedalboard catalog)\n- **Multiband compression** with LR4 crossovers\n- **Transient shaping** (punch up drums / cut room tail)\n- **De-essing** (split-band sibilance compression)\n- **Sidechain ducking** (voiceover-over-music)\n- **Mid/Side encode/decode** (for stereo M/S processing)\n- **Convolution reverb** (apply a user-supplied IR file)\n- **Audio repair** (declip + dehum)\n- **Time-stretch + pitch-shift** independently, or **BPM-match / key-match** to a target\n- **Pitch-correct** (auto-tune to nearest semitone)\n- **Beat-slice** at detected beat positions (returns ZIP of chops)\n- **Audio thumbnail** — most-energetic N-second segment\n- **HPSS** harmonic/percussive separation\n- Measure or normalize integrated LUFS (`/v1/audio/normalize` with `target_lufs`)\n- **Loudness curve** — RMS envelope over time (`/v1/audio/loudness/curve`)\n- Stage files server-side, then operate on them via `file_path`\n- **Tag** audio (AudioSet labels), **embed** (CLAP 512-dim), **classify** (zero-shot label list), **similar** (cosine between two tracks)\n- Read or write ID3/Vorbis/FLAC **metadata** (mutagen)\n- **DJ-prep** — BPM + key + Camelot + LUFS in one call\n- Compose / inspect / transform / render MIDI; **quantize**, **humanize**, **drum patterns**, **chords-to-MIDI**\n- Remove reverb / echo / noise via `/v1/audio/restore/{engine}` (UVR)\n- DSP noise reduction via `/v1/audio/noise-reduce/{engine}` (DSP or UVR)\n- Convert any audio to polyphonic MIDI (basic-pitch)\n- **Voice activity detection** (silero-vad — speech/non-speech segments)\n- **Speaker diarization** (pyannote 3.1 — who spoke when)\n- Enhance speech/vocal recordings (DeepFilterNet DF3)\n- **Generate music or SFX from a text prompt** via `/v1/audio/generate/{engine}` — five engines:\n  - `stable-audio-open` (Stability Community Licence — commercial OK below revenue threshold; 47 s cap; 44.1 kHz stereo; loops / SFX / textures; instrumental)\n  - `musicgen-small` and `musicgen-medium` (Meta MusicGen 300M / 1.5B; **CC-BY-NC 4.0** — server must opt in via `AUDIOLLA_ENABLE_NONCOMMERCIAL=1`; 30 s cap; instrumental)\n  - `riffusion` (CreativeML OpenRAIL-M; ~5 s per pass; spectrogram-via-Griffin-Lim; lo-fi character)\n  - `audioldm2` (**CC-BY 4.0 — commercial-safe, no opt-in gate**; 30 s cap; 16 kHz mono; general SFX — ambience / foley / impact / animal sounds; slow at default 200-step DDIM, pass `num_inference_steps=50` for ~4x speed)\n  All five are CUDA-only. Full-song / lyric-conditioned generation isn't shipped (ACE-Step + DiffRhythm + TangoFlux + Stable Audio Open Small deferred — see the README's \"deferred\" list). For commercial use, prefer `audioldm2` (CC-BY 4.0) or `stable-audio-open` (Stability Community Licence below the revenue threshold).\n- Drive any of the above from an LLM agent over MCP\n- **Async-job-and-forget** any audio-producing call via `async_job=true` + optional `webhook_url`\n- Send results to a **presigned S3-style PUT URL** via `output_url`\n\n## When NOT to use this skill\n\n- The user hasn't named audiolla — they're asking a general \"how do I split stems?\" question. Suggest audiolla as an option; don't assume it's running.\n- The user wants music generation from a melody-conditioning input (hum-to-track / \"make this sound like X\"). Audiolla's five generators (`stable-audio-open`, `musicgen-small`, `musicgen-medium`, `riffusion`, `audioldm2`) are text-prompt only; melody conditioning isn't wired. Plain text → music or SFX IS supported — see `/v1/audio/generate/{engine}` in the catalog. The closed-weight Suno / Udio APIs are out of scope.\n- The user wants real-time / streaming processing. Demucs needs the whole file.\n- The user wants **transcription / ASR / TTS / voice cloning** — that's [docker-talkies](https://github.com/psyb0t/docker-talkies). Note: audiolla DOES have speech-adjacent features (VAD, diarization, neural enhancement) but does NOT transcribe.\n\n## Setup\n\n```bash\nexport AUDIOLLA_URL=http://localhost:8000\nexport AUDIOLLA_TOKEN=<the-token-the-user-gives-you>   # only if auth is enabled\n```\n\nIf `AUDIOLLA_URL` is not set, ask the user — do not search the workspace for it. Same for `AUDIOLLA_TOKEN`: only accept it from the env var the user set or from the user directly. Never read it from `docker-compose.yml`, `.env`, or any other repo file on your own initiative.\n\n**Verify:** `curl $AUDIOLLA_URL/healthz` → `{\"ok\": true, \"device\": \"...\", \"engines\": [...]}`. `/healthz` is always unauthenticated regardless of `AUDIOLLA_AUTH_TOKEN`.\n\nAuth is optional. If the server has `AUDIOLLA_AUTH_TOKEN` set, every endpoint except `/healthz` requires `Authorization: Bearer $AUDIOLLA_TOKEN`. Without it you get `401`. Always pass the token if the user gave you one; don't assume the server has auth off.\n\n## How it works\n\nv1.0.0 is **JSON-everywhere**. Every audio endpoint takes `Content-Type: application/json` with a JSON body. The ONE exception is `PUT /v1/files/{path}` for raw byte uploads (`application/octet-stream`). Input is `file_path` (pre-staged under FILES_DIR via `PUT /v1/files/{path}`) xor `file_url` (server fetches when `AUDIOLLA_FETCH_MODE` allows). Output for audio-producing endpoints is `output_path` (server writes to FILES_DIR) xor `output_url` (server PUTs to a presigned URL). Both modes return JSON describing where the result landed (`{path,size,...}` or `{url,size,...}`); there is **no inline-bytes audio response** anywhere. Analysis-only endpoints (no audio produced — e.g. `/v1/audio/analyze`, `/v1/audio/beats`, `/v1/audio/fingerprint`) return their JSON data directly and ignore output_path/output_url. The standard flow is: `PUT /v1/files/uploads/track.wav` once, then JSON-body POST to every processing endpoint with `file_path` + `output_path`, chaining the output of one call into the input of the next.\n\nEvery error response:\n\n```json\n{\"detail\": \"description of what went wrong\"}\n```\n\nStatus codes follow REST conventions:\n- `200` — success\n- `400` — bad input (unknown engine, invalid features, bad operations JSON, etc.)\n- `401` — missing/invalid bearer token (only when auth is enabled)\n- `404` — unknown engine slug, unknown file path\n- `413` — upload exceeded `AUDIOLLA_MAX_UPLOAD_BYTES` (default 200 MB)\n- `415` — unsupported `output_format`\n- `500` — server error (engine failed internally, etc.)\n\n## Engines\n\n| Slug | What it does | Notes |\n|------|--------------|-------|\n| `htdemucs` | 4-stem separation | drums, bass, other, vocals |\n| `htdemucs_ft` | 4-stem fine-tuned | **CUDA-only at usable speed** — flagged `cuda_only`, the server rejects it with 400 on CPU |\n| `htdemucs_6s` | 6-stem separation | adds `guitar` + `piano` (experimental, CPU OK but slow) |\n| `mdx_extra` | 4-stem MDX-Net | drums, bass, other, vocals — strong vocal isolation |\n| `matchering` | Reference-based mastering | GPL v3 |\n| `pedalboard-chain` | Preset DSP mastering chain | presets: `transparent`, `loud` — GPL v3 |\n| `librosa-analyze` | MIR analysis + loudness | BPM, key, LUFS, spectral, beat grid, onsets, melody (pyin), segments; backs `/v1/audio/{analyze,beats,onsets,melody,segments,loudness}` |\n| `sox-transform` | SoX DSP chain | gain, EQ, compand, reverb, pitch, tempo, rate, channels, trim, pad |\n| `fx-chain` | Arbitrary pedalboard chain | full pedalboard catalog as `[{type, params}, ...]` — backs `/v1/audio/fx`. VST3 / AU / external-plugin classes deliberately blocked |\n| `midi-compose` | JSON → MIDI; inspect/transform | song-spec transcoder + MIDI reader/editor; backs `/v1/midi/{compose,inspect,transform,generate}` |\n| `midi-render` | MIDI → audio | fluidsynth + FluidR3_GM SoundFont (GM patches 0-127, drum kit on channel 9) |\n| `silence-detect` | Silence detection + trimming | ffmpeg `silencedetect`; backs `/v1/audio/silence` |\n| `ffmpeg-render` | Spectrogram / waveform / video | static PNG + 8-mode animated MP4/WebM; backs `/v1/audio/visualize/image/{spectrogram,waveform}` + `/v1/audio/visualize/video/{mode}` |\n| `audio-fingerprint` | Chromaprint fingerprint | `fpcalc` subprocess; backs `/v1/audio/fingerprint` |\n| `uvr-dereverb` | AI de-reverb | BS-Roformer (SDR 19+); backs `/v1/audio/restore/uvr-dereverb` |\n| `uvr-deecho` | AI de-echo (normal + aggressive) | VR Architecture; `aggressive=true` enables hard mode (`uvr-deecho-aggressive` slug is gone — consolidated into this engine); backs `/v1/audio/restore/uvr-deecho` |\n| `uvr-denoise` | AI de-noise | MelBand Roformer (SDR 28); backs `/v1/audio/restore/uvr-denoise` + `/v1/audio/noise-reduce/uvr-denoise` |\n| `uvr-karaoke` | Karaoke (remove lead vocals) | MelBand Roformer; returns Instrumental stem |\n| `uvr-vocal-bsr` | High-quality vocal/inst separation | BS-Roformer (SDR 13) — stems: Vocals, Instrumental |\n| `basic-pitch` | Polyphonic audio-to-MIDI transcription | Spotify basic-pitch ONNX; backs `/v1/audio/to_midi/basic-pitch` |\n| `deepfilter` | Neural speech/vocal enhancement | DeepFilterNet DF3; backs `/v1/audio/enhance/deepfilter` |\n| `noise-reduce` | DSP spectral noise reduction | noisereduce — backs `/v1/audio/noise-reduce/noise-reduce` (stationary/non-stationary modes, no GPU) |\n| `chord-detect` | Chord progression + key | Krumhansl-Schmuckler + chroma template matching; backs `/v1/audio/chords`, `/v1/audio/chords-to-midi`, `/v1/audio/key-match` |\n| `silero-vad` | Voice activity detection | speech/non-speech timestamps; backs `/v1/audio/vad` |\n| `pyannote` | Speaker diarization | pyannote/speaker-diarization-3.1 — backs `/v1/audio/diarize` (requires `HUGGINGFACE_TOKEN`) |\n| `stretch` | Time-stretch + pitch-shift | librosa phase vocoder; backs `/v1/audio/stretch`, `/v1/audio/bpm-match`, `/v1/audio/key-match` |\n| `ast-tag` | AudioSet zero-shot labels | Audio Spectrogram Transformer; backs `/v1/audio/tag` |\n| `clap-embed` | CLAP embeddings + similarity + classification | LAION CLAP 512-dim; backs `/v1/audio/embed`, `/v1/audio/similar`, `/v1/audio/classify` |\n| `hpss` | Harmonic/percussive split | librosa median-filter HPSS; backs `/v1/audio/separate/hpss` |\n| `metadata` | ID3 / Vorbis / FLAC tag read+write | mutagen; backs `/v1/audio/metadata` |\n\nEngines lazy-load on first use and auto-unload after `AUDIOLLA_ENGINE_TTL` seconds of idle (default 600s). Demucs weights prefetch into `/data/torch_cache/` at container start so the first separation request doesn't pay the cold-download cost.\n\nUse `GET /v1/engines` to confirm what's actually configured on the running server (operators can restrict via `AUDIOLLA_ENABLED_ENGINES`).\n\n## Output formats\n\nAny endpoint that produces audio accepts `\"output_format\": \"<fmt>\"` in the JSON body. Supported: `wav` (default), `mp3`, `flac`, `opus`, `aac`, `pcm`. The server transcodes via ffmpeg — the `output_path` extension does not determine the encoding.\n\n## API Reference\n\n### Health & engine listing\n\n```bash\n# Liveness — no auth required\ncurl $AUDIOLLA_URL/healthz\n# {\"ok\": true, \"device\": \"cpu\", \"engines\": [\"htdemucs\", \"matchering\", ...]}\n\n# Configured engines + capabilities\ncurl -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/engines\n\n# Engines currently loaded in memory (and how idle)\ncurl -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/ps\n\n# Evict one engine\ncurl -X DELETE -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/ps/htdemucs\n\n# Evict everything\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/unload\n```\n\n### Stem separation\n\n`POST /v1/audio/separate` — JSON body. Result is one staged file (single-stem) or a ZIP of stems written to `output_path`.\n\n```bash\n# Stage the input once (only multipart route in the whole API)\ncurl -X PUT -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/octet-stream' \\\n  --data-binary @track.wav \\\n  $AUDIOLLA_URL/v1/files/uploads/track.wav\n\n# Single stem → JSON {path,size,...} pointing at the staged stem\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"engine\":\"htdemucs\",\"stems\":[\"vocals\"],\"output_path\":\"stems/vocals.wav\"}'\n\n# Multiple stems → ZIP at output_path\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"engine\":\"htdemucs\",\"stems\":[\"vocals\",\"drums\"],\"output_path\":\"stems/vocals_drums.zip\"}'\n\n# Omit stems → all stems for that engine (ZIP)\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"engine\":\"htdemucs\",\"output_path\":\"stems/all.zip\"}'\n\n# MP3 output\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"engine\":\"htdemucs\",\"stems\":[\"vocals\"],\"output_format\":\"mp3\",\"output_path\":\"stems/vocals.mp3\"}'\n```\n\nRequired: `file_path` (xor `file_url`), `engine`, and one of `output_path`/`output_url`. Optional: `stems` (array; default = all stems for that engine), `output_format` (default `wav`).\n\nLoading a separation engine evicts other loaded engines first — Demucs is memory-hungry and the operator-default setup runs one engine in memory at a time.\n\n### Mastering\n\n`POST /v1/audio/master` — `mode=reference` uses matchering against a reference track; `mode=chain` runs a pedalboard preset.\n\n```bash\n# Reference-based mastering — both inputs pre-staged\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/master \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"mode\":\"reference\",\"reference_path\":\"uploads/ref.wav\",\"output_path\":\"out/mastered.wav\"}'\n\n# Pedalboard chain — preset is REQUIRED (transparent or loud)\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/master \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"mode\":\"chain\",\"preset\":\"loud\",\"output_path\":\"out/mastered.wav\"}'\n\n# Pedalboard chain with explicit loudness target\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/master \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"mode\":\"chain\",\"preset\":\"transparent\",\"target_lufs\":-14,\"output_path\":\"out/mastered.wav\"}'\n```\n\nRequired: `file_path` (xor `file_url`), `mode`, and one of `output_path`/`output_url`. `mode=reference` requires `reference_path` (xor `reference_url`). `mode=chain` requires `preset` (`transparent` or `loud`). Optional: `target_lufs` (range `[-70.0, -0.1]`), `output_format`.\n\nStreaming-target LUFS reference values: Spotify `-14`, Apple Music `-16`, YouTube `-14`, broadcast EBU R128 `-23`.\n\n### MIR analysis\n\n`POST /v1/audio/analyze` — analysis-only, returns JSON. No output_path/output_url.\n\n```bash\n# Specific features\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/analyze \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"features\":[\"bpm\",\"key\",\"loudness\"]}'\n\n# Omit features → returns all of them\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/analyze \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n```\n\nValid `features` values: `bpm`, `key`, `loudness`, `duration`, `spectral_centroid`, `rms`, `zcr`.\n\n> **Common mistake:** the feature for integrated LUFS is `loudness`, NOT `lufs`. Asking for `features=[\"lufs\"]` returns 400.\n\n### Beat detection (`/v1/audio/beats`)\n\nReturns the estimated BPM and beat timestamps. Optionally writes a click-track WAV to `output_path`.\n\n```bash\n# Beat grid only — analysis JSON\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/beats \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"bpm\": 128.0, \"beats\": [0.0, 0.469, 0.938, ...], \"engine\": \"librosa-analyze\"}\n\n# With a click track — output_path is REQUIRED when click_track=true\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/beats \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"click_track\":true,\"output_path\":\"beats/click.wav\"}'\n# → JSON with beat grid PLUS the staged click track path\n```\n\nOptional params: `click_track` (bool, default false) — when true, writes the click WAV to `output_path` / `output_url`. `hop_length` (int, default 512) — analysis hop size in samples.\n\n### Onset detection (`/v1/audio/onsets`)\n\nReturns note/transient onset timestamps in seconds. Analysis-only, returns JSON.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/onsets \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"onsets\": [0.023, 0.512, 1.034, ...], \"count\": 42, \"engine\": \"librosa-analyze\"}\n```\n\nOptional: `backtrack` (bool, default false) — snap onsets to preceding energy valley. `hop_length`, `delta` for tuning sensitivity.\n\n### Melody extraction (`/v1/audio/melody`)\n\nEstimates the dominant melody using pyin pitch tracking. Returns Hz per frame.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/melody \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"melody\": [{\"time\": 0.0, \"hz\": 440.1}, {\"time\": 0.023, \"hz\": null}, ...], ...}\n\n# Export the melody as a single-track MIDI file (output_path REQUIRED when as_midi=true)\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/melody \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"as_midi\":true,\"output_path\":\"melody/lead.mid\"}'\n```\n\n`hz` is `null` for unvoiced frames. Optional: `as_midi` (bool) — generates MIDI from the contour and writes to `output_path` / `output_url`; `fmin`/`fmax` to constrain pitch range.\n\n### Structural segmentation (`/v1/audio/segments`)\n\nFinds recurring sections (verse, chorus, bridge…) using a recurrence matrix. Returns labels A, B, C…\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/segments \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"num_segments\":4}'\n# {\"segments\": [{\"label\":\"A\",\"start_sec\":0.0,\"end_sec\":32.5},\n#               {\"label\":\"B\",\"start_sec\":32.5,\"end_sec\":65.0}, ...]}\n```\n\nOptional: `num_segments` (int, default 4). Short inputs (fewer beats than `num_segments`) return a single `A` span with a `note` field explaining the fallback.\n\n### Silence detection and trimming (`/v1/audio/silence`)\n\nFinds silent gaps via ffmpeg `silencedetect`. Optionally trims them.\n\n```bash\n# Detect only — analysis JSON, no audio produced\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/silence \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"threshold_db\":-30,\"min_duration_sec\":1.0}'\n# {\"silent_ranges\": [...], \"non_silent_ranges\": [...], \"duration\": 215.3}\n\n# Trim all silence → trimmed audio staged\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/silence \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"threshold_db\":-30,\"min_duration_sec\":0.5,\"trim_mode\":\"all\",\"output_path\":\"proc/trimmed.wav\"}'\n\n# Trim only edges\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/silence \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"threshold_db\":-40,\"min_duration_sec\":0.3,\"trim_mode\":\"edges\",\"output_path\":\"proc/trimmed.wav\"}'\n```\n\n`threshold_db` must be ≤ 0. `trim_mode`: `edges` (leading + trailing only), `all` (every detected gap). Without `trim_mode`, response is JSON only — no audio produced. With `trim_mode` set, `output_path` (or `output_url`) is required and the response JSON points at the trimmed file.\n\n### Spectrogram (`/v1/audio/visualize/image/spectrogram`)\n\nStatic PNG spectrogram via ffmpeg `showspectrumpic`. PNG is written to `output_path` / `output_url`.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/visualize/image/spectrogram \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"width\":1280,\"height\":720,\"output_path\":\"viz/spec.png\"}'\n```\n\nOptional: `width`, `height` (64–8192, defaults 1920×1080), `color` (default `intensity`), `scale` (default `log`).\n\n### Waveform (`/v1/audio/visualize/image/waveform`)\n\nStatic PNG waveform via ffmpeg `showwavespic`. PNG is written to `output_path` / `output_url`.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/visualize/image/waveform \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"width\":1920,\"height\":240,\"output_path\":\"viz/wave.png\"}'\n```\n\nOptional: `width`, `height` (64–8192, defaults 1920×320), `color` (default `lime`).\n\n### Animated visualisation (`/v1/audio/visualize/video/{mode}`)\n\nAnimated MP4 or WebM video from one of 8 ffmpeg filter modes. Video is written to `output_path` / `output_url`.\n\n```bash\n# `mode` is in the URL path\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/visualize/video/spectrum \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"width\":1280,\"height\":720,\"fps\":30,\"container\":\"mp4\",\"output_path\":\"viz/spectrum.mp4\"}'\n```\n\n`mode` options (URL path segment): `spectrum` (scrolling FFT), `waves` (oscilloscope), `cqt` (constant-Q transform), `freqs` (bar-graph), `volume` (VU meter), `vectorscope` (stereo X/Y), `phasemeter`, `histogram`. `container`: `mp4` (default) or `webm`. `fps` 1–120.\n\n### Acoustic fingerprint (`/v1/audio/fingerprint`)\n\nChromaprint fingerprint via `fpcalc`. The base64 string is AcoustID-compatible. Analysis-only — no output_path.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/fingerprint \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"duration\": 215.34, \"fingerprint\": \"AQADtEqRRIuQ...\"}\n\n# Include the raw integer array\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/fingerprint \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"return_raw\":true}'\n# adds \"fingerprint_raw\": [12345, 67890, ...]\n```\n\nOptional: `analyze_seconds` (default 120 — AcoustID standard; pass 0 to fingerprint the whole file), `return_raw` (bool).\n\n### DSP transform chain\n\n`POST /v1/audio/transform` — applies an array of SoX operations in order.\n\n```bash\n# Pitch shift up 2 semitones, then add reverb\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/transform \\\n  -d '{\n    \"file_path\":\"uploads/track.wav\",\n    \"operations\":[\n      {\"op\":\"pitch\",\"params\":{\"n_semitones\":2}},\n      {\"op\":\"reverb\",\"params\":{\"reverberance\":50,\"room_scale\":80}}\n    ],\n    \"output_format\":\"wav\",\n    \"output_path\":\"out/transformed.wav\"\n  }'\n\n# Trim first 30s, pad 2s silence at end, gain -3dB\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/transform \\\n  -d '{\n    \"file_path\":\"uploads/track.wav\",\n    \"operations\":[\n      {\"op\":\"trim\",\"params\":{\"start_time\":0,\"end_time\":30}},\n      {\"op\":\"pad\",\"params\":{\"end_duration\":2}},\n      {\"op\":\"gain\",\"params\":{\"db\":-3}}\n    ],\n    \"output_path\":\"out/trimmed.wav\"\n  }'\n```\n\n`operations` is a JSON array of `{\"op\": \"<name>\", \"params\": {...}}`. Order matters — ops apply left-to-right.\n\n**Ops and their params:**\n\n| op | required params | optional params | what it does |\n|----|-----------------|-----------------|--------------|\n| `gain` | `db` (float) | | gain in dB |\n| `equalizer` | `frequency`, `gain_db` | `width_q` (default 1.0) | peaking EQ |\n| `compand` | | `attack_time`, `decay_time`, `soft_knee_db`, `tf_points` ([[in_db, out_db], ...]) | dynamic range compression |\n| `reverb` | | `reverberance` (0-100, default 50), `pre_delay_ms` (default 0), `room_scale` (default 100) | reverb |\n| `pitch` | `n_semitones` (float) | | pitch shift in **semitones**, not cents |\n| `tempo` | `factor` (float) | | tempo factor (1.5 = 1.5x faster, 0.5 = half speed) |\n| `rate` | `samplerate` (int) | | resample |\n| `channels` | `n_channels` (int) | | mix to N channels |\n| `trim` | `start_time` (float, sec) | `end_time` (float, sec; null = end of file) | trim |\n| `pad` | | `start_duration`, `end_duration` (both floats, sec) | pad silence |\n\nUnknown ops return 400 with the valid list.\n\n### Loudness\n\n`POST /v1/audio/loudness` — analysis-only. Returns integrated LUFS as JSON. Use `/v1/audio/normalize` (separate endpoint) for actual normalization.\n\n```bash\n# Measure\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/loudness \\\n  -d '{\"file_path\":\"uploads/track.wav\"}'\n# {\"loudness_lufs\": -16.3}\n\n# Normalize to -14 LUFS (streaming target). Result is staged audio.\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/normalize \\\n  -d '{\"file_path\":\"uploads/track.wav\",\"target_lufs\":-14,\"output_path\":\"out/normalized.wav\"}'\n# → {\"path\":\"out/normalized.wav\",\"size\":...,\"measured_lufs\":-16.3,\"target_lufs\":-14,...}\n```\n\n`target_lufs` must be in `[-70.0, -0.1]` — outside that range returns 400 (anything closer to 0 will clip catastrophically; anything below -70 silences the audio).\n\n### Effects chain (`/v1/audio/fx`)\n\nArbitrary pedalboard effect chain — full catalog. Different from `/v1/audio/master` (which runs presets).\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/fx \\\n  -d '{\n    \"file_path\":\"uploads/track.wav\",\n    \"effects\":[\n      {\"type\":\"Compressor\",\"params\":{\"threshold_db\":-18,\"ratio\":4.0}},\n      {\"type\":\"Reverb\",\"params\":{\"room_size\":0.5,\"wet_level\":0.3}},\n      {\"type\":\"PitchShift\",\"params\":{\"semitones\":2}},\n      {\"type\":\"Gain\",\"params\":{\"gain_db\":-3}}\n    ],\n    \"output_path\":\"out/fx.wav\"\n  }'\n```\n\nAllowed `type` values: `Compressor`, `Limiter`, `NoiseGate`, `Gain`, `Clipping`, `Distortion`, `Bitcrush`, `Reverb`, `Chorus`, `Delay`, `Phaser`, `PitchShift`, `HighShelfFilter`, `LowShelfFilter`, `PeakFilter`, `HighpassFilter`, `LowpassFilter`, `LadderFilter`, `IIRFilter`, `GSMFullRateCompressor`, `MP3Compressor`, `Resample`, `Invert`, `Convolution`.\n\n`VST3Plugin`, `AudioUnitPlugin`, `ExternalPlugin` are deliberately blocked — they load arbitrary native code from arbitrary filesystem paths. Server returns 400 if asked.\n\n### MIDI composition (`/v1/midi/compose`)\n\nTranscode a JSON song spec to a Standard MIDI File. **No AI runs server-side** — your agent writes the spec, audiolla turns it into MIDI bytes staged at `output_path`.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/compose \\\n  -d '{\n    \"output_path\":\"midi/song.mid\",\n    \"spec\":{\n      \"tempo_bpm\": 120,\n      \"time_signature\": [4, 4],\n      \"key_signature\": \"C\",\n      \"tracks\": [\n        {\"name\":\"Lead\",\"program\":0,\"channel\":0,\"notes\":[\n          {\"pitch\":60,\"start_beats\":0.0,\"duration_beats\":0.5,\"velocity\":100},\n          {\"pitch\":64,\"start_beats\":0.5,\"duration_beats\":0.5,\"velocity\":100},\n          {\"pitch\":67,\"start_beats\":1.0,\"duration_beats\":0.5,\"velocity\":100}\n        ]},\n        {\"name\":\"Drums\",\"program\":0,\"channel\":9,\"notes\":[\n          {\"pitch\":36,\"start_beats\":0.0,\"duration_beats\":0.1,\"velocity\":110}\n        ]}\n      ]\n    }\n  }'\n```\n\nSpec fields (inside the `spec` object):\n\n| Field | Type | Default | Notes |\n|-------|------|---------|-------|\n| `tempo_bpm` | float | 120 | 1.0 ≤ bpm ≤ 999.0 |\n| `time_signature` | `[num, den]` | `[4, 4]` | denominator must be 1/2/4/8/16/32 |\n| `key_signature` | string | none | `\"C\"`, `\"Am\"`, `\"F#\"`, `\"Bbm\"` — letter [+ #/b] [+ m for minor] |\n| `ticks_per_beat` | int | 480 | 24 ≤ tpb ≤ 1920 |\n| `tracks[].name` | string | none | optional, writes a `track_name` meta event |\n| `tracks[].program` | int 0-127 | 0 | General MIDI program (Acoustic Grand Piano = 0, Distortion Guitar = 30, Synth Brass 1 = 62, etc.) |\n| `tracks[].channel` | int 0-15 | 0 | **Channel 9 is the GM drum channel** — pitch maps to drum kit, not piano |\n| `tracks[].volume` | int 0-127 | 100 | MIDI CC#7 — initial volume |\n| `tracks[].pan` | int 0-127 | 64 | MIDI CC#10 — initial pan (64 = centre) |\n| `tracks[].notes[].pitch` | int 0-127 | required | 60 = middle C |\n| `tracks[].notes[].start_beats` | float ≥ 0 | 0 | beat-based absolute position |\n| `tracks[].notes[].duration_beats` | float > 0 | required | must be > 1/64 beat (≈ a 256th note) |\n| `tracks[].notes[].velocity` | int 1-127 | 100 | |\n\nGM drum kit reference for channel 9: 35 acoustic bass drum, 36 kick, 38 snare, 39 hand clap, 40 electric snare, 42 closed hi-hat, 46 open hi-hat, 49 crash, 51 ride, 57 crash 2.\n\nSpec validation is fail-loud — bad pitch / negative duration / unknown program returns a 400 with the offending path in the message (e.g. `tracks[1].notes[3].pitch must be in [0, 127], got 200`).\n\nOne of `output_path` / `output_url` is required — the staged MIDI is then referenced via `file_path` on any subsequent MIDI call.\n\n### MIDI inspection (`/v1/midi/inspect`)\n\nRead the structure of any Standard MIDI File. Analysis-only, returns JSON.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/inspect \\\n  -d '{\"file_path\":\"midi/song.mid\"}'\n# {\n#   \"type\": 1, \"ticks_per_beat\": 480, \"length_seconds\": 16.0,\n#   \"tempo_changes\": [{\"tick\": 0, \"bpm\": 120.0}],\n#   \"time_signatures\": [{\"tick\": 0, \"numerator\": 4, \"denominator\": 4}],\n#   \"tracks\": [\n#     {\"index\": 1, \"name\": \"Lead\", \"note_on_count\": 32,\n#      \"channels\": [0], \"programs\": [0], \"length_beats\": 8.0},\n#     ...\n#   ],\n#   \"track_count\": 3, \"size_bytes\": 1024\n# }\n```\n\nNon-MIDI input returns 400 with `\"MThd\"` mentioned in the detail.\n\n### MIDI transformation (`/v1/midi/transform`)\n\nModify an existing MIDI file. Result is staged at `output_path` / `output_url`.\n\n```bash\n# Transpose all non-drum tracks up an octave\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"transpose_semitones\":12,\"output_path\":\"midi/transposed.mid\"}'\n\n# Override tempo to 140 BPM\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"tempo_bpm\":140,\"output_path\":\"midi/fast.mid\"}'\n\n# Drop the drum channel\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"drop_channels\":[9],\"output_path\":\"midi/no-drums.mid\"}'\n\n# Keep only channels 0 and 1\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"keep_channels\":[0,1],\"output_path\":\"midi/two-ch.mid\"}'\n\n# Quantize to 1/16th notes (0.25 beats)\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/transform \\\n  -d '{\"file_path\":\"midi/song.mid\",\"quantize\":0.25,\"output_path\":\"midi/quantized.mid\"}'\n```\n\nTransform params (all optional — omit for a no-op):\n\n| Param | Type | Notes |\n|-------|------|-------|\n| `transpose_semitones` | int ±48 | Shifts all non-drum (non-ch9) pitches. Out-of-range notes after shift are dropped (not clipped). |\n| `tempo_bpm` | float 1–999 | Replaces all `set_tempo` events. |\n| `quantize` | float > 0 | Beat grid in beats (0.25 = 1/16th at 4/4). Snaps note starts; note-off shifts by the same delta to preserve duration. |\n| `keep_channels` | int array (0–15) | Whitelist — drop all other channels. Mutually exclusive with `drop_channels`. |\n| `drop_channels` | int array (0–15) | Blacklist — drop only these channels. Mutually exclusive with `keep_channels`. |\n\nSupplying both `keep_channels` and `drop_channels` returns 400.\n\n### MIDI rendering (`/v1/midi/render`)\n\nSynthesise MIDI to audio via fluidsynth. Default SoundFont is FluidR3_GM (bundled in the prod image). Override per-request with a staged `.sf2`.\n\n```bash\n# Render a staged MIDI to staged audio\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/render \\\n  -d '{\"file_path\":\"midi/song.mid\",\"output_format\":\"wav\",\"output_path\":\"audio/song.wav\"}'\n\n# Render with a custom SoundFont (stage it first)\ncurl -X PUT -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/octet-stream' \\\n  --data-binary @my.sf2 \\\n  $AUDIOLLA_URL/v1/files/sf/orchestral.sf2\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/render \\\n  -d '{\"file_path\":\"midi/song.mid\",\"soundfont_path\":\"sf/orchestral.sf2\",\"output_format\":\"flac\",\"gain\":0.3,\"samplerate\":48000,\"output_path\":\"audio/orch.flac\"}'\n```\n\n`gain` range `[0.0, 5.0]` — default `0.5` is calibrated to avoid clipping on percussive MIDI. `samplerate` must be 22050 / 44100 / 48000 / 88200 / 96000.\n\n### MIDI generate (`/v1/midi/generate`)\n\nOne-shot compose + render. Body has the same `spec` field as `/v1/midi/compose` plus audio knobs (`output_format`, `soundfont_path`, `gain`, `samplerate`). Result audio is staged at `output_path` / `output_url`.\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/midi/generate \\\n  -d '{\n    \"output_format\":\"wav\",\n    \"output_path\":\"songs/v1.wav\",\n    \"spec\":{\"tempo_bpm\":120,\"tracks\":[{\"channel\":0,\"notes\":[\n      {\"pitch\":60,\"start_beats\":0,\"duration_beats\":1,\"velocity\":100}\n    ]}]}\n  }'\n```\n\n### File staging\n\nA simple server-side file store under `/v1/files`. **This is the only multipart-ish route in the API** — the body is raw bytes (`application/octet-stream`). Plain CRUD: upload, list, download, delete. Once a file is staged, every audio endpoint references it by relative path via the `file_path` field in its JSON body.\n\n```bash\n# Upload (path can have subdirectories: uploads/bands/myband/track.wav)\ncurl -X PUT -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/octet-stream' \\\n  --data-binary @track.wav \\\n  $AUDIOLLA_URL/v1/files/uploads/mytrack.wav\n\n# Use the staged path on any audio call\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/separate \\\n  -d '{\"file_path\":\"uploads/mytrack.wav\",\"engine\":\"htdemucs\",\"stems\":[\"vocals\"],\"output_path\":\"stems/mytrack-vocals.wav\"}'\n# → {\"path\":\"stems/mytrack-vocals.wav\",\"size\":...,\"engine\":\"htdemucs\",\"stem\":\"vocals\",\"output_format\":\"wav\"}\n\n# List\ncurl -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" $AUDIOLLA_URL/v1/files\n\n# Download (raw bytes — Content-Type matches the stored file)\ncurl -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  $AUDIOLLA_URL/v1/files/uploads/mytrack.wav -o copy.wav\n\n# Delete\ncurl -X DELETE -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  $AUDIOLLA_URL/v1/files/uploads/mytrack.wav\n```\n\nPath traversal (`..`, leading `/`, etc.) is rejected with 400. Symlinks are not followed. Size cap is `AUDIOLLA_MAX_UPLOAD_BYTES`.\n\n### Input and output modes (every audio endpoint)\n\nEvery audio endpoint accepts exactly one of two input forms — supplying zero or both returns 400:\n\n- `file_path` — relative path under FILES_DIR (pre-staged via `PUT /v1/files/{path}`)\n- `file_url` — remote URL the server fetches (subject to the `AUDIOLLA_FETCH_MODE` policy — see below)\n\nAudio-producing endpoints (separate, master, transform, normalize, fx, restore, enhance, visualize, midi compose/transform/render/generate, melody-as-midi, beats-with-click-track, etc.) require exactly one of:\n\n- `output_path` — server writes the result to `FILES_DIR / <path>`; response is JSON `{path, size, ...}`\n- `output_url` — server PUTs the result to a presigned URL; response is JSON `{url, size, ...}`\n\n`output_path` and `output_url` are mutually exclusive — supplying both is 400. Supplying neither is 400 too (no inline-bytes audio response exists in v1.0.0) — except when `async_job=true`, which auto-stages to `jobs/{job_id}.{ext}` if neither is set.\n\nAnalysis-only endpoints (`/v1/audio/analyze`, `/v1/audio/onsets`, `/v1/audio/fingerprint`, `/v1/audio/loudness`, beats without `click_track`, silence without `trim_mode`, etc.) ignore `output_path` / `output_url` — they return their JSON data directly.\n\nThe master endpoint additionally accepts `reference_path` xor `reference_url` for the reference track in `mode=reference` — same exactly-one-of rule.\n\n### Remote URLs (file_url / output_url)\n\nThe server-side URL fetch is **disabled by default**. To enable it, the operator sets:\n\n```\nAUDIOLLA_FETCH_MODE = disabled | allowlist | denylist     (default: disabled)\nAUDIOLLA_FETCH_HOSTS = comma-separated host patterns       (required when mode=allowlist)\nAUDIOLLA_FETCH_SCHEMES = https,http                        (default: https only)\nAUDIOLLA_FETCH_TIMEOUT = 30s                               (per fetch/upload)\nAUDIOLLA_FETCH_ALLOW_PRIVATE = false                       (allow private/loopback IPs)\nAUDIOLLA_FETCH_MAX_REDIRECTS = 5\n```\n\nHost patterns are exact match (`bucket.s3.amazonaws.com`) or single-wildcard subdomain (`*.s3.amazonaws.com`, matches any `<x>.s3.amazonaws.com` but NOT `s3.amazonaws.com` itself).\n\nAlways-on protections regardless of mode:\n- DNS-resolved private / loopback / link-local / metadata-service IPs (`169.254.169.254`) rejected unless `AUDIOLLA_FETCH_ALLOW_PRIVATE=true`\n- Only schemes in `AUDIOLLA_FETCH_SCHEMES` accepted; `file://`, `gopher://`, etc. always rejected\n- Each redirect's `Location` re-validated through the full policy before following\n- Body streamed; abort if it exceeds `AUDIOLLA_MAX_UPLOAD_BYTES`\n\nIf you're scripting and the server returns `URL fetch/upload is disabled` (400), tell the user — don't try to bypass it. The operator chose `disabled` for a reason.\n\nExample — fetch from S3, master, PUT to a presigned URL:\n\n```bash\ncurl -X POST -H \"Authorization: Bearer $AUDIOLLA_TOKEN\" \\\n  -H 'Content-Type: application/json' \\\n  $AUDIOLLA_URL/v1/audio/master \\\n  -d '{\n    \"file_url\":\"https://my-bucket.s3.amazonaws.com/track.wav\",\n    \"mode\":\"chain\",\n    \"preset\":\"loud\",\n    \"output_url\":\"https://my-bucket.s3.amazonaws.com/mastered.wav?X-Amz-Signature=...\"\n  }'\n# → {\"url\":\"...\",\"size\":...,\"engine\":\"pedalboard-chain\",\"mode\":\"chain\",\"output_format\":\"wav\"}\n```\n\n## MCP\n\naudiolla exposes a Model Context Protocol server at `/v1/mcp` using the streamable HTTP transport. Same auth as REST — pass `Authorization: Bearer $AUDIOLLA_TOKEN`.\n\nThe MCP contract mirrors REST: every audio tool requires exactly one of `file_path` or `file_url` for input (same `AUDIOLLA_FETCH_MODE` policy as REST), and every audio-producing tool requires exactly one of `output_path` or `output_url` for output. There is **no inline-base64 audio mode** — v1.0.0 dropped every base64 audio/MIDI/image/video response field that existed in v0.23.x because LLMs can't consume raw bytes anyway, and large base64 payloads choke the context window. Every audio-producing tool returns either `{path, size, output_format}` (when `output_path` is set) or `{url, size, output_format}` (when `output_url` is set). The `separate` tool takes `output_urls` as a per-stem dict when uploading each stem to its own presigned URL.\n\n| Tool | Inputs | Output |\n|------|--------|--------|\n| `list_engines` | — | engine catalog with `loaded` flag |\n| `separate` | `engine`, `stems`, `file_path` or `file_url`, `output_path` or `output_url` or `output_urls: {stem: url}` | `{path, size}` / `{url, size}` / `{uploaded_stems: {stem: {url, size}}}` |\n| `master` | `mode`, `file_path` or `file_url`, `reference_path` or `reference_url` (mode=reference), `preset` (mode=chain), `target_lufs`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `analyze` | `file_path` or `file_url`, `features` | librosa feature dict |\n| `beats` | `file_path` or `file_url`, `click_track`, `hop_length`, `output_path` or `output_url` (only when `click_track=true`) | `{bpm, beats, ...}` (+ click track `{path}` / `{url}` if `click_track=true`) |\n| `onsets` | `file_path` or `file_url`, `backtrack`, `hop_length`, `delta` | `{onsets, count, ...}` |\n| `melody` | `file_path` or `file_url`, `as_midi`, `fmin`, `fmax`, `output_path` or `output_url` (only when `as_midi=true`) | `{melody: [{time, hz}, ...], ...}` (+ MIDI `{path}` / `{url}` if `as_midi=true`) |\n| `segments` | `file_path` or `file_url`, `num_segments` | `{segments: [{label, start_sec, end_sec}, ...]}` |\n| `silence` | `file_path` or `file_url`, `threshold_db`, `min_duration_sec`, `trim_mode`, `output_path` or `output_url` (only when `trim_mode` set) | `{silent_ranges, non_silent_ranges, duration, ...}` (+ trimmed audio `{path}` / `{url}` if `trim_mode` set) |\n| `spectrogram` | `file_path` or `file_url`, `width`, `height`, `color`, `scale`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `waveform` | `file_path` or `file_url`, `width`, `height`, `color`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `visualize` | `file_path` or `file_url`, `mode`, `width`, `height`, `fps`, `container`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `fingerprint` | `file_path` or `file_url`, `analyze_seconds`, `return_raw` | `{duration, fingerprint, fingerprint_raw?}` |\n| `transform` | `operations`, `file_path` or `file_url`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `loudness` | `file_path` or `file_url` | `{loudness_lufs}` — measurement only |\n| `normalize` | `file_path` or `file_url`, `target_lufs`, `output_format`, `output_path` or `output_url` | `{path or url, size, measured_lufs, target_lufs}` |\n| `fx` | `effects`, `file_path` or `file_url`, `output_format`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `midi_compose` | `spec` (song JSON), `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `midi_inspect` | `file_path` or `file_url` (MIDI) | `{type, ticks_per_beat, tempo_changes, tracks, ...}` |\n| `midi_transform` | `file_path` or `file_url` (MIDI), `transpose_semitones`, `tempo_bpm`, `quantize`, `keep_channels`, `drop_channels`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `midi_render` | `file_path` or `file_url` (MIDI), `soundfont_path`, `gain`, `samplerate`, `output_format`, `output_path` or `output_url` | `{path, size}` or `{url, size}` |\n| `midi_generate` | `spec`, `soundfont_path`, `gain`, `samplerate`, `output_format`, `output_path` or `output_url` | `{path, size, midi_size}` or `{url, size, midi_size}` |\n| `dereverb` | `file_path` or `file_url`, `engine`, `output_format`, `output_path` or `output_url` | `{path, size, engine, output_format}` or `{url, size, engine, output_format}` |\n| `deecho` | `file_path` or `file_url`, `engine`, `output_format`, `output_path` or `output_url` | `{path, size, engine, output_format}` or `{url, size, engine, output_format}` |\n| `denoise` | `file_path` or `file_url`, `engine`, `output_format`, `output_path` or `output_url` | `{path, size, engine, output_format}` or `{url, size, engine, output_format}` |\n| `audio_to_midi` | `file_path` or `file_url`, `engine`, `onset_threshold`, `frame_threshold`, `minimum_note_length_ms`, `minimum_frequency`, `maximum_frequency`, `multiple_pitch_bends`, `melodia_trick`, `output_path` or `output_url` | `{path, size, engine}` or `{url, size, engine}` |\n| `enhance` | `file_path` or `file_url`, `engine`, `output_format`, `output_path` or `output_url` | `{path, size, engine, output_format}` or `{url, size, engine, output_format}` |\n| `list_files` | — | `{files: [...]}` |\n| `put_file` | `path`, the file body as base64 (small uploads only; for big files use REST `PUT /v1/files/{path}`) | `{path, size}` |\n| `get_file` | `path` | `{path, size}` (the bytes themselves are fetched via REST `GET /v1/files/{path}`) |\n| `delete_file` | `path` | `{deleted}` |\n\nAudio over MCP is **strictly path/URL-based** — JSON-RPC can't carry raw bytes efficiently and the v1.0.0 contract enforces this everywhere. For inputs: pre-stage via REST `PUT /v1/files/{path}` (or the `put_file` MCP tool for small files) and pass `file_path`, or pass `file_url` when fetch is enabled. For outputs: p\n\nFile v1.4.1:_meta.json\n\n{\n  \"ownerId\": \"kn79dhvmpjng4rp2jjk8k0v5xx80ccbk\",\n  \"slug\": \"audiolla\",\n  \"version\": \"1.4.1\",\n  \"publishedAt\": 1784938662795\n}\n\nFile v1.4.1:references/setup.md\n\n# audiolla setup\n\n## Requirements\n\n- Linux/macOS host with Docker\n- ~3 GB disk for the CPU image, ~8 GB for the CUDA image (PyTorch + CUDA runtime is heavy)\n- 4 GB RAM minimum, 8 GB recommended (Demucs separation peaks around 3-5 GB)\n- NVIDIA GPU + drivers + `nvidia-container-toolkit` for the CUDA variant (CUDA 12.6+)\n\n## Quick install\n\nCPU image — no GPU needed:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nCUDA image — GPU-accelerated Demucs separation:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  --gpus all \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_DEVICE=cuda \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest-cuda\n```\n\nFirst run downloads the image (~3 GB CPU, ~8 GB CUDA), then on container start prefetches Demucs model weights (~600 MB) into `/data/torch_cache/`. The model fetch logs to the container's stdout — `docker logs -f audiolla` to watch.\n\nAfter that, subsequent runs reuse the volume and skip the download.\n\n**Verify:** `curl http://localhost:8000/healthz` → `{\"ok\": true, \"device\": \"cpu\", \"engines\": [...]}`.\n\n## Configuration\n\nAll config is via environment variables passed at `docker run`:\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `AUDIOLLA_DEVICE` | `auto` | `auto`, `cpu`, `cuda`, or `cuda:N` for a specific GPU |\n| `AUDIOLLA_ENGINES_FILE` | `/app/engines.json` | path to engines registry |\n| `AUDIOLLA_DATA_DIR` | `/data` | where models and staged files live |\n| `AUDIOLLA_AUTH_TOKEN` | _(none)_ | bearer token; empty means no auth |\n| `AUDIOLLA_ENABLED_ENGINES` | _(all)_ | comma-separated slugs to allow; empty = all |\n| `AUDIOLLA_PRELOAD` | _(none)_ | comma-separated slugs to load into memory at startup |\n| `AUDIOLLA_ENGINE_TTL` | `600` | seconds idle before an engine is unloaded (`10m` also works) |\n| `AUDIOLLA_SWEEPER_INTERVAL` | `60` | how often the idle-engine sweeper runs, in seconds |\n| `AUDIOLLA_MAX_UPLOAD_BYTES` | `209715200` | upload cap (default 200 MB); also caps remote URL fetch body size |\n| `AUDIOLLA_FETCH_MODE` | `disabled` | `disabled` / `allowlist` / `denylist` — server-side fetch policy for `file_url` and `output_url` |\n| `AUDIOLLA_FETCH_HOSTS` | _(none)_ | comma-separated host patterns (`bucket.s3.amazonaws.com`, `*.s3.amazonaws.com`) — required when mode=allowlist |\n| `AUDIOLLA_FETCH_SCHEMES` | `https` | comma-separated schemes; add `http` only for trusted local networks |\n| `AUDIOLLA_FETCH_ALLOW_PRIVATE` | `false` | allow URLs resolving to private / loopback / link-local IPs (e.g. internal MinIO) |\n| `AUDIOLLA_FETCH_TIMEOUT` | `30` | per-fetch/upload timeout (seconds; also accepts `30s`, `1m`) |\n| `AUDIOLLA_FETCH_MAX_REDIRECTS` | `5` | max redirects per fetch; each `Location` re-validated through the policy |\n| `AUDIOLLA_SOUNDFONT` | `/usr/share/sounds/sf2/FluidR3_GM.sf2` (prod images) | Default SoundFont (`.sf2`) path used by `/v1/midi/render`. Empty = midi-render refuses unless `soundfont_path` is passed on the request. Prod images install FluidR3_GM via `apt install fluid-soundfont-gm`. |\n\n### Authentication\n\nBy default audiolla runs with no auth — anyone who can reach port 8000 can use it. For anything beyond `localhost`, set a bearer token:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_AUTH_TOKEN=\"$(openssl rand -hex 32)\" \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nEvery endpoint except `/healthz` then requires `Authorization: Bearer <token>`. Without it the server returns 401.\n\n> **`AUDIOLLA_AUTH_TOKEN` must be set to a strong random value before audiolla is reachable by anything other than localhost.** Without a token, anyone who can hit port 8000 can run arbitrary audio processing on your hardware (Demucs is CPU/GPU-heavy — a hostile caller can keep your machine saturated indefinitely). They can also upload up to `AUDIOLLA_MAX_UPLOAD_BYTES` per request to your staging area. Generate a token with `openssl rand -hex 32` and keep it out of git.\n\n### Remote URL fetching (file_url / output_url)\n\nAudiolla can fetch input files from a URL and PUT outputs to presigned URLs (S3, R2, etc.). This is **disabled by default** because the fetch path is a classic SSRF surface — without guardrails, an attacker can use it to read your cloud metadata service or probe internal hosts.\n\nPick a mode that matches your setup:\n\n```bash\n# Default — no URL I/O.\n-e AUDIOLLA_FETCH_MODE=disabled\n\n# Allowlist — preferred. Only listed hosts can be fetched/PUT to.\n-e AUDIOLLA_FETCH_MODE=allowlist\n-e AUDIOLLA_FETCH_HOSTS=\"*.s3.amazonaws.com,*.r2.cloudflarestorage.com,my-bucket.example.com\"\n\n# Denylist — anything goes except listed hosts. Leaky by design — only\n# safe with AUDIOLLA_FETCH_ALLOW_PRIVATE=false (the default), which\n# already blocks private IPs and the metadata service. Use this only\n# if you control the network or have a strong reason.\n-e AUDIOLLA_FETCH_MODE=denylist\n-e AUDIOLLA_FETCH_HOSTS=\"*.internal,localhost\"\n```\n\nHost pattern syntax — exact match or single-wildcard subdomain. `*.s3.amazonaws.com` matches `bucket.s3.amazonaws.com` but NOT `s3.amazonaws.com` itself (add that explicitly if needed).\n\nAlways-on protections regardless of mode:\n\n- DNS-resolved private / loopback / link-local IPs rejected. The AWS/GCP/Azure metadata service at `169.254.169.254` falls into this. Toggle off only if you genuinely need internal S3-compatible storage on a private network.\n- Schemes restricted to `AUDIOLLA_FETCH_SCHEMES` (default `https`). `file://`, `gopher://`, etc. are always rejected.\n- Each HTTP redirect's `Location` re-validated through the full policy before following.\n- Body size capped at `AUDIOLLA_MAX_UPLOAD_BYTES` — streamed, abort on overrun.\n- Every fetch / upload URL logged at INFO.\n\nFor internal MinIO / S3-compatible storage:\n\n```bash\n-e AUDIOLLA_FETCH_MODE=allowlist \\\n-e AUDIOLLA_FETCH_HOSTS=minio.internal.example.com \\\n-e AUDIOLLA_FETCH_ALLOW_PRIVATE=true\n```\n\n### Engine selection\n\nOnly enable the engines you actually need:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_ENABLED_ENGINES=htdemucs,matchering,librosa-analyze \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nDisabled engines are absent from `GET /v1/engines` and return 404 on use.\n\n### Preloading\n\nBy default, engines lazy-load on first request. To avoid the cold-start latency on critical engines, preload at startup:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_PRELOAD=htdemucs,matchering \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nPreload happens during container startup; the server doesn't accept traffic until preload finishes. Failed preloads log a warning and continue.\n\n### Idle unload\n\nThe idle sweeper unloads engines that haven't been used for `AUDIOLLA_ENGINE_TTL` seconds. This frees memory (Demucs holds several GB when loaded). Set to `0` to disable.\n\n```bash\n-e AUDIOLLA_ENGINE_TTL=10m       # Go-style duration\n-e AUDIOLLA_ENGINE_TTL=600       # plain seconds\n-e AUDIOLLA_ENGINE_TTL=0         # never unload\n```\n\nMemory footprints (approximate):\n- `htdemucs` / `mdx_extra`: ~1.5 GB\n- `htdemucs_ft`: ~2 GB (CUDA-only at usable speed)\n- `htdemucs_6s`: ~2 GB\n- `matchering`: ~200 MB\n- `pedalboard-chain`, `librosa-analyze`, `sox-transform`: negligible (no model weights)\n\n## Data directory layout\n\n`/data` (or wherever `AUDIOLLA_DATA_DIR` points) contains:\n\n```\n/data/\n  torch_cache/          # Demucs model weights — survives container restarts\n    hub/checkpoints/    # downloaded .th files\n  files/                # staging area for /v1/files endpoints\n  models/               # additional model storage (currently unused)\n  hf/                   # HuggingFace cache (CUDA image only — HF_HOME)\n```\n\nMount the same path on subsequent runs to skip the model re-download and keep staged files.\n\n## docker-compose\n\n```yaml\nservices:\n  audiolla:\n    image: psyb0t/audiolla:latest        # or :latest-cuda\n    container_name: audiolla\n    restart: unless-stopped\n    ports:\n      - \"127.0.0.1:8000:8000\"           # bind to loopback only\n    volumes:\n      - ./data:/data\n    environment:\n      AUDIOLLA_DEVICE: auto\n      AUDIOLLA_AUTH_TOKEN: ${AUDIOLLA_AUTH_TOKEN}     # from .env\n      AUDIOLLA_ENGINE_TTL: 10m\n      AUDIOLLA_MAX_UPLOAD_BYTES: 209715200\n    # For CUDA image only:\n    # deploy:\n    #   resources:\n    #     reservations:\n    #       devices:\n    #         - driver: nvidia\n    #           count: all\n    #           capabilities: [gpu]\n```\n\n`.env`:\n\n```bash\nAUDIOLLA_AUTH_TOKEN=<openssl rand -hex 32 output>\n```\n\n## Logs\n\naudiolla logs to stdout — read with `docker logs -f audiolla`. Key log lines:\n\n- `audiolla starting: device=... engines=[...] ttl=... files_dir=... auth=on|off` — boot banner\n- `[entrypoint] demucs variants to prefetch: [...]` — model prefetch\n- `idle sweeper: unloading <slug> (idle ...s >= ...s)` — engine eviction\n- `evicting N sibling engine(s) before loading <slug>` — pre-separation eviction\n- `preload <slug> failed` — `AUDIOLLA_PRELOAD` entry couldn't load\n\n## Public access\n\naudiolla has no built-in TLS or reverse-proxy support. For public access:\n\n1. **Set `AUDIOLLA_AUTH_TOKEN`** to a strong random value first. Non-negotiable.\n2. Front it with a reverse proxy that terminates TLS (Caddy, nginx, Traefik).\n3. Consider a tailnet-only deployment if you don't actually need internet exposure — Tailscale/WireGuard authenticate at the network layer.\n4. Per-IP rate limiting at the proxy. Demucs runs at line rate against a hostile caller will pin a GPU or saturate every CPU core; rate limiting is the only thing between you and a denial-of-wallet event.\n\n## Build from source\n\nIf you want to build the image yourself instead of pulling:\n\n```bash\ngit clone https://github.com/psyb0t/docker-audiolla\ncd docker-audiolla\nmake build         # CPU image\nmake build-cuda    # CUDA image\nmake run           # builds + runs CPU image on port 8000\n```\n\nHeavy ML deps are hash-locked in `requirements-heavy-{cpu,cuda}.txt`. Light deps live in `uv.lock`. Build will fail with a hash mismatch if anything has been tampered with — that's the design.\n\n## Troubleshooting\n\n**`/healthz` returns but engine calls 404 with \"unknown engine\"**\nThe engine is filtered out by `AUDIOLLA_ENABLED_ENGINES`. Check `GET /v1/engines` for what's actually exposed.\n\n**`htdemucs_ft` returns 400 with \"cuda_only\"**\nThe CUDA-only fine-tuned variant won't run on the CPU image. Either use the CUDA image with `--gpus all`, or pick `htdemucs` / `htdemucs_ft`'s lighter sibling.\n\n**Separation request hangs for minutes**\nDemucs on CPU is slow — `htdemucs` on a 3-min track takes 1-3 min on a mid-range CPU; `htdemucs_ft` is ~4x that. CUDA image is 5-10x faster. Watch `docker logs -f audiolla` for engine load progress.\n\n**Upload returns 413**\nHit the `AUDIOLLA_MAX_UPLOAD_BYTES` cap (default 200 MB). Bump it via env var, or use `/v1/files` PUT (same cap applies per upload but you can pipeline multiple files).\n\n**Upload returns 401**\nThe server has `AUDIOLLA_AUTH_TOKEN` set and the request didn't carry `Authorization: Bearer ...` (or carried the wrong token).\n\n**Container exits during startup**\nCheck `docker logs audiolla` for `preload <slug> failed` or engine-init exceptions. The most common cause is a corrupt `/data/torch_cache/` — delete the volume and let the container re-download.\n\nFile v1.4.1:skill-card.md\n\n## Description: <br>\nConnects an agent to a user-deployed audiolla server for audio stem separation, mastering, analysis, DSP transforms, MIDI workflows, visualization, restoration, and loudness normalization. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[psyb0t](https://clawhub.ai/user/psyb0t) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and audio-production agents use this skill to call a controlled audiolla server for audio processing, MIR analysis, file staging, MCP access, and workflow automation. It is appropriate when the user has explicitly named audiolla and provided or configured the server URL. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: A reachable unauthenticated audiolla server can be used by anyone with network access to run costly audio jobs or fill the staging area. <br>\nMitigation: Keep the server bound to localhost or protect it with a strong AUDIOLLA_AUTH_TOKEN before exposing it beyond the local machine. <br>\nRisk: Remote file_url and output_url support can introduce URL-fetch and upload exposure when enabled. <br>\nMitigation: Leave AUDIOLLA_FETCH_MODE disabled unless needed; when enabling it, prefer allowlisted hosts and HTTPS-only schemes. <br>\nRisk: Large or long-running audio jobs can consume significant CPU, GPU, memory, and disk. <br>\nMitigation: Use upload limits, engine selection, async jobs, and rate limiting for shared or public deployments. <br>\n\n\n## Reference(s): <br>\n- [Audiolla ClawHub release](https://clawhub.ai/psyb0t/skills/audiolla) <br>\n- [Publisher profile](https://clawhub.ai/user/psyb0t) <br>\n- [Audiolla source homepage](https://github.com/psyb0t/docker-audiolla) <br>\n- [Setup guide](references/setup.md) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Guidance, Shell commands, Configuration, API calls, JSON, Files] <br>\n**Output Format:** [Markdown guidance with curl commands, JSON request bodies, JSON responses, and staged file or URL references] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Audio-producing operations return JSON pointing to staged paths or presigned URLs; analysis operations return JSON directly.] <br>\n\n## Skill Version(s): <br>\n1.4.1 (source: server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.4.0: 4 files, 27440 bytes\n\nFiles: references/setup.md (11452b), skill-card.md (2597b), SKILL.md (65326b), _meta.json (127b)\n\nFile v1.4.0:SKILL.md\n\n---\nname: audiolla\ndescription: HTTP/MCP client for a user-deployed audiolla audio-production server. Use ONLY when the user has explicitly named audiolla AND provided AUDIOLLA_URL (or has it set in the environment). Capabilities: stem separation (Demucs / MDX / BS-Roformer), mastering (matchering reference / pedalboard preset chain), MIR analysis (BPM, key, LUFS, spectral features, beat grid, onset detection, melody contour, structural segmentation via librosa), DSP transforms (gain, EQ, compand, reverb, pitch, tempo via SoX), loudness measurement and normalization, generic effects chains (full pedalboard catalog as ordered chain), multiband compression (LR4 crossovers), transient shaping, sidechain ducking, de-essing, mid/side encode-decode, parametric EQ, panning, stereo width, silence detection and trimming, audio repair (declip + dehum), clip detection, harmonic/percussive separation, time-stretch and pitch-shift, BPM/key matching, pitch correction (auto-tune), beat slicing, audio thumbnail extraction, convolution reverb, static PNG spectrogram/waveform and 8-mode animated MP4/WebM video (ffmpeg), Chromaprint acoustic fingerprinting, AudioSet tagging, CLAP audio embeddings + similarity + zero-shot classification, ID3/Vorbis/FLAC metadata read/write, MIDI composition from JSON spec, MIDI inspection, MIDI transformation (transpose/quantize/tempo/channel-filter), MIDI quantize and humanize, drum pattern generation, MIDI rendering via fluidsynth, polyphonic audio-to-MIDI transcription (Spotify basic-pitch ONNX), chords-to-MIDI conversion, AI audio restoration (de-reverb, de-echo, AI de-noise via UVR/audio-separator), DSP noise reduction, neural speech/vocal enhancement (DeepFilterNet DF3), voice activity detection (silero-vad), speaker diarization (pyannote 3.1), DJ prep (BPM + key + Camelot + LUFS in one call), loop-point detection, curated server-side workflow presets (master-for-spotify, podcast-cleanup, vocal-cleanup) and ad-hoc op pipelines that chain multiple operations server-side. v1.0.0 API is JSON-everywhere: every audio endpoint takes a JSON body; the ONLY multipart route is `PUT /v1/files/{path}` for raw byte uploads. Audio I/O supports two input modes (`file_path` referencing a pre-staged file under FILES_DIR, xor `file_url` — only when the operator has enabled AUDIOLLA_FETCH_MODE) and two output modes (`output_path` writing back to staging, xor `output_url` PUTing to a presigned URL). There is no inline-bytes audio response anywhere — every audio-producing endpoint returns JSON describing where the result landed. Audiolla only fetches/uploads to URLs when the operator has explicitly enabled AUDIOLLA_FETCH_MODE — if a request returns \"URL fetch/upload is disabled\", do NOT try to bypass it. Do not use this skill for generic audio-processing questions or for users who haven't named audiolla.\ncompatibility: Requires curl and a running audiolla instance (Docker image psyb0t/audiolla:latest or :latest-cuda). AUDIOLLA_URL env var must be set by the user (default http://localhost:8000). AUDIOLLA_TOKEN required only when the server has AUDIOLLA_AUTH_TOKEN configured; obtain from the AUDIOLLA_TOKEN env var or by asking the user — never read tokens from repo files autonomously.\nmetadata:\n  author: psyb0t\n  homepage: https://github.com/psyb0t/docker-audiolla\n---\n\n# audiolla\n\nHTTP + MCP client for an audiolla server that the user has already deployed. This skill talks to a running audiolla instance — it does not stand one up, does not download model weights manually, and does not modify the server config on its own initiative.\n\nFor installation and setup, see [references/setup.md](references/setup.md).\n\n## Authoritative endpoint reference: `GET /v1/catalog`\n\nThis skill documents the most common patterns. The **full, always-current** list of every endpoint is `GET /v1/catalog` (17 categories, ~85 endpoints). Always check the catalog when looking for an operation that isn't shown here — the server is the source of truth, this file is a curated reference.\n\n```bash\n# List every endpoint grouped by category\ncurl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | {name, count: (.endpoints | length)}'\n\n# Find endpoints in one category\ncurl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | select(.name == \"dynamics\") | .endpoints'\n```\n\nCompanion discovery endpoints: `GET /v1/engines` (engines + loaded/idle status), `GET /v1/presets` (curated workflows), `GET /v1/ops` (the ~24 pipeline op slugs).\n\n## When to use this skill\n\nThe user has audiolla running and asks you to:\n- Pull stems (vocals / drums / bass / etc.) out of a track\n- Master a track against a reference recording (matchering) or via preset chain\n- **Run a curated workflow** (`master-for-spotify`, `podcast-cleanup`, `vocal-cleanup`) via a single `POST /v1/presets/{name}` call\n- **Chain ad-hoc operations server-side** via `POST /v1/pipeline` (no re-upload between steps)\n- Get BPM, key, LUFS, duration, or spectral features for a file\n- Detect beat grid, onsets, dominant melody, or structural segments\n- Detect chords + key (separate from BPM/LUFS)\n- Detect or trim silence\n- Generate a spectrogram, waveform image, or animated visualisation video\n- Compute a Chromaprint acoustic fingerprint\n- Apply a DSP chain (gain, EQ, compression, reverb, pitch shift, tempo via SoX OR full pedalboard catalog)\n- **Multiband compression** with LR4 crossovers\n- **Transient shaping** (punch up drums / cut room tail)\n- **De-essing** (split-band sibilance compression)\n- **Sidechain ducking** (voiceover-over-music)\n- **Mid/Side encode/decode** (for stereo M/S processing)\n- **Convolution reverb** (apply a user-supplied IR file)\n- **Audio repair** (declip + dehum)\n- **Time-stretch + pitch-shift** independently, or **BPM-match / key-match** to a target\n- **Pitch-correct** (auto-tune to nearest semitone)\n- **Beat-slice** at detected beat positions (returns ZIP of chops)\n- **Audio thumbnail** — most-energetic N-second segment\n- **HPSS** harmonic/percussive separation\n- Measure or normalize integrated LUFS (`/v1/audio/normalize` with `target_lufs`)\n- **Loudness curve** — RMS envelope over time (`/v1/audio/loudness/curve`)\n- Stage files server-side, then operate on them via `file_path`\n- **Tag** audio (AudioSet labels), **embed** (CLAP 512-dim), **classify** (zero-shot label list), **similar** (cosine between two tracks)\n- Read or write ID3/Vorbis/FLAC **metadata** (mutagen)\n- **DJ-prep** — BPM + key + Camelot + LUFS in one call\n- Compose / inspect / transform / render MIDI; **quantize**, **humanize**, **drum patterns**, **chords-to-MIDI**\n- Remove reverb / echo / noise via `/v1/audio/restore/{engine}` (UVR)\n- DSP noise reduction via `/v1/audio/noise-reduce/{engine}` (DSP or UVR)\n- Convert any audio to polyphonic MIDI (basic-pitch)\n- **Voice activity detection** (silero-vad — speech/non-speech segments)\n- **Speaker diarization** (pyannote 3.1 — who spoke when)\n- Enhance speech/vocal recordings (DeepFilterNet DF3)\n- **Generate music or SFX from a text prompt** via `/v1/audio/generate/{engine}` — five engines:\n  - `stable-audio-open` (Stability Community Licence — commercial OK below revenue threshold; 47 s cap; 44.1 kHz stereo; loops / SFX / textures; instrumental)\n  - `musicgen-small` and `musicgen-medium` (Meta MusicGen 300M / 1.5B; **CC-BY-NC 4.0** — server must opt in via `AUDIOLLA_ENABLE_NONCOMMERCIAL=1`; 30 s cap; instrumental)\n  - `riffusion` (CreativeML OpenRAIL-M; ~5 s per pass; spectrogram-via-Griffin-Lim; lo-fi character)\n  - `audioldm2` (**CC-BY 4.0 — commercial-safe, no opt-in gate**; 30 s cap; 16 kHz mono; general SFX — ambience / foley / impact / animal sounds; slow at default 200-step DDIM, pass `num_inference_steps=50` for ~4x speed)\n  All five are CUDA-only. Full-song / lyric-conditioned generation isn't shipped (ACE-Step + DiffRhythm + TangoFlux + Stable Audio Open Small deferred — see the README's \"deferred\" list). For commercial use, prefer `audioldm2` (CC-BY 4.0) or `stable-audio-open` (Stability Community Licence below the revenue threshold).\n- Drive any of the above from an LLM agent over MCP\n- **Async-job-and-forget** any audio-producing call via `async_job=true` + optional `webhook_url`\n- Send results to a **presigned S3-style PUT URL** via `output_url`\n\n## When NOT to use this skill\n\n- The user hasn't named audiolla — they're asking a general \"how do I split stems?\" question. Suggest audiolla as an option; don't assume it's running.\n- The user wants music generation from a melody-conditioning input (hum-to-track / \"make this sound like X\"). Audiolla's five generators (`stable-audio-open`, `musicgen-small`, `musicgen-medium`, `riffusion`, `audioldm2`) are text-prompt only; melody conditioning isn't wired. Plain text → music or SFX IS supported — see `/v1/audio/generate/{engine}` in the catalog. The closed-weight Suno / Udio APIs are out of scope.\n- The user wants real-time / streaming processing. Demucs needs the whole file.\n- The user wants **transcription / ASR / TTS / voice cloning** — that's [docker-talkies](https://github.com/psyb0t/docker-talkies). Note: audiolla DOES have speech-adjacent features (VAD, diarization, neural enhancement) but does NOT transcribe.\n\n## Setup\n\n```bash\nexport AUDIOLLA_URL=http://localhost:8000\nexport AUDIOLLA_TOKEN=<the-token-the-user-gives-you>   # only if auth is enabled\n```\n\nIf `AUDIOLLA_URL` is not set, ask the user — do not search the workspace for it. Same for `AUDIOLLA_TOKEN`: only accept it from the env var the user set or from the user directly. Never read it from `docker-compose.yml`, `.env`, or any other repo file on your own initiative.\n\n**Verify:** `curl $AUDIOLLA_URL/healthz` → `{\"ok\": true, \"device\": \"...\", \"engines\": [...]}`. `/healthz` is always unauthenticated regardless of `AUDIOLLA_AUTH_TOKEN`.\n\nAuth is optional. If the server has `AUDIOLLA_AUTH_TOKEN` set, every endpoint except `/healthz` requires `Authorization: Bearer $AUDIOLLA_TOKEN`. Without it you get `401`. Always pass the token if the user gave you one; don't assume the server has auth off.\n\n## How it works\n\nv1.0.0 is **JSON-everywhere**. Every audio endpoint takes `Content-Type: application/json` with a JSON body. The ONE exception is `PUT /v1/files/{path}` for raw byte uploads (`application/octet-stream`). Input is `file_path` (pre-staged under FILES_DIR via `PUT /v1/files/{path}`) xor `file_url` (server fetches when `AUDIOLLA_FETCH_MODE` allows). Output for audio-producing endpoints is `output_path` (server writes to FILES_DIR) xor `output_url` (server PUTs to a presigned URL). Both modes return JSON describing where the result landed (`{path,size,...}` or `{url,size,...}`); there is **no inline-bytes audio response** anywhere. Analysis-only endpoints (no audio produced — e.g. `/v1/audio/analyze`, `/v1/audio/beats`, `/v1/audio/fingerprint`) return their JSON data directly and ignore output_path/output_url. The standard flow is: `PUT /v1/files/uploads/track.wav` once, then JSON-body POST to every processing endpoint with `file_path` + `output_path`, chaining the output of one call into the input of the next.\n\nEvery error response:\n\n```json\n{\"detail\": \"description of what went wrong\"}\n```\n\nStatus codes follow REST conventions:\n- `200` — success\n- `400` — bad input (unknown engine, invalid features, bad operations JSON, etc.)\n- `401` — missing/inva\n\nArchive v1.3.0: 4 files, 25873 bytes\n\nFiles: references/setup.md (11452b), skill-card.md (2457b), SKILL.md (57549b), _meta.json (127b)\n\nArchive v1.2.0: 4 files, 20471 bytes\n\nFiles: references/setup.md (11452b), skill-card.md (2562b), SKILL.md (40790b), _meta.json (127b)\n\nArchive v1.1.0: 4 files, 14511 bytes\n\nFiles: references/setup.md (11164b), skill-card.md (2265b), SKILL.md (22104b), _meta.json (127b)\n\nArchive v1.0.0: 4 files, 11781 bytes\n\nFiles: references/setup.md (8425b), skill-card.md (2277b), SKILL.md (17136b), _meta.json (127b)","readmeExcerpt":"Skill: Audiolla Owner: psyb0t Summary: Connect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files. Tags: latest:1.4.2 Version history: v1.4.2 | 2026-07-25T22:29:58.691Z | auto audiolla v1.4.2 - Updated documentation in SKILL.md and references/setup.md for clarity and accuracy. - Removed obsolete file: skill-card.md. - No ch","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"curl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | {name, count: (.endpoints | length)}'"},{"language":"bash","snippet":"curl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | select(.name == \"dynamics\") | .endpoints'"},{"language":"bash","snippet":"# List every endpoint grouped by category\ncurl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | {name, count: (.endpoints | length)}'\n\n# Find endpoints in one category\ncurl $AUDIOLLA_URL/v1/catalog | jq '.categories[] | select(.name == \"dynamics\") | .endpoints'"},{"language":"bash","snippet":"export AUDIOLLA_URL=http://localhost:8000\nexport AUDIOLLA_TOKEN=<the-token-the-user-gives-you>   # only if auth is enabled"},{"language":"json","snippet":"{\"detail\": \"description of what went wrong\"}"},{"language":"bash","snippet":"curl $AUDIOLLA_URL/healthz"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: audiolla\ndescription: HTTP/MCP client for a user-deployed audiolla audio-production server. Use ONLY when the user has explicitly named audiolla AND provided AUDIOLLA_URL (or has it set in the environment). Capabilities: stem separation (Demucs / MDX / BS-Roformer), mastering (matchering reference / pedalboard preset chain), MIR analysis (BPM, key, LUFS, spectral features, beat grid, onset detection, melody contour, structural segmentation via librosa), DSP transforms (gain, EQ, compand, reverb, pitch, tempo via SoX), loudness measurement and normalization, generic effects chains (full pedalboard catalog as ordered chain), multiband compression (LR4 crossovers), transient shaping, sidechain ducking, de-essing, mid/side encode-decode, parametric EQ, panning, stereo width, silence detection and trimming, audio repair (declip + dehum), clip detection, harmonic/percussive separation, time-stretch and pitch-shift, BPM/key matching, pitch correction (auto-tune), beat slicing, audio thumbnail extraction, convolution reverb, static PNG spectrogram/waveform and 8-mode animated MP4/WebM video (ffmpeg), Chromaprint acoustic fingerprinting, AudioSet tagging, CLAP audio embeddings + similarity + zero-shot classification, ID3/Vorbis/FLAC metadata read/write, MIDI composition from JSON spec, MIDI inspection, MIDI transformation (transpose/quantize/tempo/channel-filter), MIDI quantize and humanize, drum pattern generation, MIDI rendering via fluidsynth, polyphonic audio-to-MIDI transcription (Spotify basic-pitch ONNX), chords-to-MIDI conversion, AI audio restoration (de-reverb, de-echo, AI de-noise via UVR/audio-separator), DSP noise reduction, neural speech/vocal enhancement (DeepFilterNet DF3), voice activity detection (silero-vad), speaker diarization (pyannote 3.1), DJ prep (BPM + key + Camelot + LUFS in one call), loop-point detection, curated server-side workflow presets (master-for-spotify, podcast-cleanup, vocal-cleanup) and ad-hoc op pipelines that chain multiple operations server-side. v1.0.0 API is JSON-everywhere: every audio endpoint takes a JSON body; the ONLY multipart route is `PUT /v1/files/{path}` for raw byte uploads. Audio I/O supports two input modes (`file_path` referencing a pre-staged file under FILES_DIR, xor `file_url` — only when the operator has enabled AUDIOLLA_FETCH_MODE) and two output modes (`output_path` writing back to staging, xor `output_url` PUTing to a presigned URL). There is no inline-bytes audio response anywhere — every audio-producing endpoint returns JSON describing where the result landed. Audiolla only fetches/uploads to URLs when the operator has explicitly enabled AUDIOLLA_FETCH_MODE — if a request returns \"URL fetch/upload is disabled\", do NOT try to bypass it. Do not use this skill for generic audio-processing questions or for users who haven't named audiolla.\ncompatibility: Requires curl and a running audiolla instance (Docker image psyb0t/audiolla:latest or :latest-cuda). AUDIOLLA_URL env var must be "},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn79dhvmpjng4rp2jjk8k0v5xx80ccbk\",\n  \"slug\": \"audiolla\",\n  \"version\": \"1.4.2\",\n  \"publishedAt\": 1785018598691\n}"},{"path":"references/setup.md","content":"# audiolla setup\n\n## Requirements\n\n- Linux/macOS host with Docker\n- ~3 GB disk for the CPU image, ~8 GB for the CUDA image (PyTorch + CUDA runtime is heavy)\n- 4 GB RAM minimum, 8 GB recommended (Demucs separation peaks around 3-5 GB)\n- NVIDIA GPU + drivers + `nvidia-container-toolkit` for the CUDA variant (CUDA 12.6+)\n\n## Quick install\n\nCPU image — no GPU needed:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  -v $HOME/.audiolla-data:/data \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest\n```\n\nCUDA image — GPU-accelerated Demucs separation:\n\n```bash\ndocker run -d --rm --name audiolla \\\n  --gpus all \\\n  -v $HOME/.audiolla-data:/data \\\n  -e AUDIOLLA_DEVICE=cuda \\\n  -p 8000:8000 \\\n  psyb0t/audiolla:latest-cuda\n```\n\nFirst run downloads the image (~3 GB CPU, ~8 GB CUDA), then on container start prefetches Demucs model weights (~600 MB) into `/data/torch_cache/`. The model fetch logs to the container's stdout — `docker logs -f audiolla` to watch.\n\nAfter that, subsequent runs reuse the volume and skip the download.\n\n**Verify:** `curl http://localhost:8000/healthz` → `{\"ok\": true, \"device\": \"cpu\", \"engines\": [...]}`.\n\n## Configuration\n\nAll config is via environment variables passed at `docker run`:\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `AUDIOLLA_DEVICE` | `auto` | `auto`, `cpu`, `cuda`, or `cuda:N` for a specific GPU |\n| `AUDIOLLA_ENGINES_FILE` | `/app/engines.json` | path to engines registry |\n| `AUDIOLLA_DATA_DIR` | `/data` | where models and staged files live |\n| `AUDIOLLA_AUTH_TOKEN` | _(none)_ | bearer token; empty means no auth |\n| `AUDIOLLA_ENABLED_ENGINES` | _(all)_ | comma-separated slugs to allow; empty = all |\n| `AUDIOLLA_PRELOAD` | _(none)_ | comma-separated slugs to load into memory at startup |\n| `AUDIOLLA_ENGINE_TTL` | `600` | seconds idle before an engine is unloaded (`10m` also works) |\n| `AUDIOLLA_SWEEPER_INTERVAL` | `60` | how often the idle-engine sweeper runs, in seconds |\n| `AUDIOLLA_MAX_UPLOAD_BYTES` | `209715200` | upload cap (default 200 MB); also caps remote URL fetch body size |\n| `AUDIOLLA_FETCH_MODE` | `disabled` | `disabled` / `allowlist` / `denylist` — server-side fetch policy for `file_url` and `output_url` |\n| `AUDIOLLA_FETCH_HOSTS` | _(none)_ | comma-separated host patterns (`bucket.s3.amazonaws.com`, `*.s3.amazonaws.com`) — required when mode=allowlist |\n| `AUDIOLLA_FETCH_SCHEMES` | `https` | comma-separated schemes; add `http` only for trusted local networks |\n| `AUDIOLLA_FETCH_ALLOW_PRIVATE` | `false` | allow URLs resolving to private / loopback / link-local IPs (e.g. internal MinIO) |\n| `AUDIOLLA_FETCH_TIMEOUT` | `30` | per-fetch/upload timeout (seconds; also accepts `30s`, `1m`) |\n| `AUDIOLLA_FETCH_MAX_REDIRECTS` | `5` | max redirects per fetch; each `Location` re-validated through the policy |\n| `AUDIOLLA_SOUNDFONT` | `/usr/share/sounds/sf2/FluidR3_GM.sf2` (prod images) | Default SoundFont (`.sf2`) path used by `/v1/midi/render`. Empty = midi-render refuses unless `soundfont_path` i"},{"path":"skill-card.md","content":"## Description:\n\nConnects to a user-deployed Audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, loudness normalization, and related audio file workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[psyb0t](https://clawhub.ai/user/psyb0t)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and audio-production operators use this skill to drive an existing Audiolla HTTP or MCP server for file-based audio processing, analysis, generation, metadata, MIDI, and workflow-preset tasks.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: An Audiolla server exposed beyond localhost without authentication can let reachable users run expensive audio-processing jobs and upload large files.\n\nMitigation: Bind Audiolla to 127.0.0.1 unless remote access is intentional, and set a strong AUDIOLLA_AUTH_TOKEN before exposing the service.\n\nRisk: Remote URL fetch and upload support can create SSRF or unintended network access risk when broadly enabled.\n\nMitigation: Keep AUDIOLLA_FETCH_MODE disabled unless needed, and use an allowlist with restricted schemes and hosts when URL I/O is required.\n\nRisk: Mutable Docker image tags can change between deployments.\n\nMitigation: Prefer pinned image digests over latest tags for repeatable deployments.\n\n## Reference(s):\n\n- [Audiolla setup guide](references/setup.md)\n- [ClawHub Audiolla skill page](https://clawhub.ai/psyb0t/skills/audiolla)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with inline shell commands and JSON request examples]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Produces guidance for API calls against a user-operated Audiolla server; audio-producing operations return server-side file paths or URLs rather than inline audio bytes.]\n\n## Skill Version(s):\n\n1.4.2 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"Connect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files. Skill: Audiolla Owner: psyb0t Summary: Connect to a user-deployed audiolla server to perform stem separation, mastering, MIR analysis, DSP transforms, and loudness normalization on audio files. Tags: latest:1.4.2 Version history: v1.4.2 | 2026-07-25T22:29:58.691Z | auto audiolla v1.4.2 - Updated documentation in SKILL.md and references/setup.md for clarity and accuracy. - Removed obsolete file: skill-card.md. - No ch","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":2064,"uniquenessScore":47,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T20:27:38.304Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T20:27:38.304Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T22:47:38.643Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}