Agent skill · creative production · aiskillstore
ace-step
Generate, inpaint, and outpaint music with ACE Step on RunComfy via the `runcomfy` CLI. ACE Step is StepFun-AI's open-weights music foundation model — tag-driven composition (genre, mood, instruments), multilingual lyrics with section markers, 5 s to 4 min stereo output, $0.0002–0.0003 per second (≈ 27× cheaper than ElevenLabs Music). Four endpoints: ACE Step text-to-audio (the default), ACE Step 1.5 text-to-audio (50+ language lyrics, refined structured-lyric handling), ACE Step audio-inpaint (regenerate a time range inside an existing track), ACE Step audio-outpaint (extend an existing track before or after). Triggers on "ace step", "ace-step", "acestep", "ACE music", "open music model", "cheap AI music", "inpaint audio", "audio inpaint", "extend music", "audio outpaint", "lengthen track", "music with tags", or any explicit ask to generate or edit music with ACE Step.
Why this skill is useful
Adds executable commands for generating and editing music using the ACE Step model, which the AI cannot generate on its own.
What it needs
Requires @runcomfy/cli installed locally. Requires runcomfy account access. About 9k tokens when loaded. Last updated 2026-08-07. 405 stars on the source repository.
What this skill does
ACE Step — Pro Pack on RunComfy Tag-driven music generation, inpainting, and outpainting with StepFun-AI's ACE Step open-weights model. Four CLI-reachable endpoints, $0.0002–0.0003 per second of audio, up to 4 minutes per call. runcomfy.com · ACE Step base · ACE Step 1.5 · CLI docs Install this skill Powered by the RunComfy CLI Step 1 — install (one of, see the runcomfy-cli skill for details): Step 2 — sign in (or set RUNCOMFYTOKEN env var in CI / containers): Step 3 — generate: CLI deep dive: runcomfy-cli skill. --- Pick the right endpoint Listed newest first. ACE Step 1.5 (text-to-audio) — acestep-ai/ace-step-1.5/text-to-audio Latest ACE Step generation. 50+ language vocal support, refined structured-lyric handling, otherwise same shape as base. Slightly higher cost ($0.0003/s vs $0.0002/s). Pick for: multilingual lyrics, hero-quality vocal tracks, vocal songs that need clean section structure. Avoid for: cost-sensitive batches where the base model is good enough. ACE Step (text-to-audio) — acestep-ai/ace-step/text-to-audio (default — cheap & fast) Original ACE Step. Tag-driven composition, optional lyrics, 5–240 s stereo. $0.0002/s — ~27× cheaper than ElevenLabs Music. Pick for: high-volume drafts, background music, jingles, game loops, cost-sensitive iteration. Avoid for: maximally polished commercial vocal hooks — try ACE Step 1.5 or ElevenLabs Music for those. ACE Step (audio-inpaint) — acestep-ai/ace-step/audio-inpaint Regenerate a time range inside an existing track (not mask-based; uses starttime / endtime in seconds, each anchored to track start or end). Pick for: fix a bad chorus in the middle, swap the bridge, replace a 20 s section without re-rendering the whole song. Avoid for: edits that aren't time-bounded — those don't fit the schema. ACE Step (audio-outpaint) — acestep-ai/ace-step/audio-outpaint Extend an existing track bidirectionally — add intro before, outro after, or both. …
How to use it
Reference it in AdaL, Claude Code, Cursor or any coding agent — nothing to install:
@skills aiskillstore/ace-step