A fork of ACE-Step Studio, built around a local AI music generation workflow
Powered by ACE-Step 1.5 β The Open Source AI Music Generation Model
This is a fork of timoncool/ACE-Step-Studio, itself built on AmbsdOP's ACE-Step UI. It's been substantially reworked and extended β developed primarily on Linux Mint with an NVIDIA RTX GPU, with genuine Windows support built out in parallel rather than tacked on afterward.
Both platforms share the same underlying approach: a single uv-managed Python environment instead of the older embedded-Python-plus-manual-pip setup most Windows installers still rely on. Same fast, reliable dependency resolution, same GPU-aware flash-attn handling, same isolated environment for MIDI conversion β one installer per platform, one consistent approach underneath. macOS isn't currently tested or supported.
What's different from upstream:
- Playlists vs. Workspaces β a real separation between curated playlists and working sessions, with exclusive workspace membership (a song lives in one workspace at a time) and a virtual "default" view computed by exclusion.
- MIDI conversion, server-side β basic-pitch running in an isolated Python 3.11 environment (its TensorFlow dependency doesn't ship wheels for newer Python), converting stems to MIDI in seconds rather than tens of minutes in-browser.
- AudioMass, updated and wired in β upgraded to the multitrack build, with direct-load support: open a single stem or all four Demucs stems together as separate tracks, straight from the browser, no manual export/import round-trip.
- LoRA training from the UI β the full scan β label β preprocess β train pipeline, drivable from React without dropping into Gradio directly.
- One installer per platform, both
uv-based and GPU-aware β detects your actual compute capability and compiler version, not just a menu choice, and knows whenflash-attnwill and won't build correctly for your hardware (Blackwell/RTX 50-series needs CUDA 12.8+ to compile it at all on Linux, or a matching prebuilt wheel on Windows β both installers check this before attempting work that's doomed to fail).
Generation panel, workspace view, and the bottom player β all in one screen.
![]() |
![]() |
| Workspaces, organized visually | Stem extraction β WAV, MIDI, or straight to the editor |
All four Demucs stems, opened together as separate tracks in the AudioMass editor.
| Feature | Description |
|---|---|
| Full Song Generation | Create complete songs with vocals and lyrics up to 4+ minutes |
| Instrumental Mode | Generate instrumental tracks without vocals |
| Custom Mode | Fine-tune BPM, key, time signature, and duration |
| Style Tags | Define genre, mood, tempo, and instrumentation |
| Batch Generation | Generate multiple variations at once |
| AI Enhance | Enrich genre tags into detailed captions with proper BPM/key/time |
| Thinking Mode | Let AI reason about structure and generate audio codes |
| Feature | Description |
|---|---|
| Reference Audio | Use any audio file as a style reference |
| Audio Cover | Transform existing audio with new styles |
| Repainting | Regenerate specific sections of a track |
| Seed Control | Reproduce exact generations for consistency |
| Inference Steps | Control quality vs speed tradeoff |
| Feature | Description |
|---|---|
| Lyrics Editor | Write and format lyrics with structure tags |
| Format Assistant | AI-powered caption and lyrics formatting |
| Prompt Templates | Quick-start with genre presets |
| Reuse Prompts | Clone settings from any previous generation |
| Feature | Description |
|---|---|
| Playlists | Curated collections, a song can belong to several |
| Workspaces | Active working sessions β a song belongs to exactly one at a time |
| Default View | Everything not currently assigned to a workspace |
| Bottom Player | Full-featured player with waveform and progress |
| Real-time Progress | Live generation progress with queue position |
| LAN Access | Use from any device on your local network |
| Feature | Description |
|---|---|
| Multitrack Audio Editor | Trim, fade, and mix with AudioMass β open single stems or all four together as separate tracks |
| Stem Extraction | Separate vocals, drums, bass, and other with Demucs, in-browser |
| MIDI Conversion | Turn any stem into MIDI server-side, in seconds |
| LoRA Training | Full training pipeline, driven from the UI |
| Video Generator | Create music videos with Pexels backgrounds |
| Gradient Covers | Procedural album art, no internet needed |
| Layer | Technologies |
|---|---|
| Frontend | React 18, TypeScript, TailwindCSS, Vite |
| Backend | Express.js, SQLite, better-sqlite3 |
| AI Engine | ACE-Step 1.5 (Gradio API) |
| Audio Tools | AudioMass (multitrack), Demucs, basic-pitch, FFmpeg |
| Python tooling | uv β faster, more reliable dependency resolution than plain pip, same tool on both Linux and Windows |
| Requirement | Specification |
|---|---|
| OS | Linux (developed on Linux Mint / Ubuntu 24.04) or Windows 10/11 |
| Node.js | 22 LTS |
| Python | Managed automatically by uv β 3.12 on Linux, 3.11 on Windows for the main environment; a separate isolated 3.11 environment on both platforms for MIDI conversion |
| NVIDIA GPU | 4GB+ VRAM (works without LLM), 12GB+ recommended (with LLM) |
CUDA compiler (nvcc) |
Linux only, 12.8+ if you want flash-attn on Blackwell (RTX 50-series) β older cards work with older nvcc too, the installer checks and falls back to SDPA if not. Windows uses a prebuilt flash-attn wheel instead, no local compiler needed |
| FFmpeg, libsndfile | Installed automatically by the installer if missing |
| uv | Python package manager β installed automatically by both installers if missing |
# 1. Clone this repo and ACE-Step-1.5 side by side (see full install below)
git clone https://github.com/Sion971/ace-step-studio.git
cd ace-step-studio
# 2. Run the installer β handles GPU detection, PyTorch, dependencies,
# database migration, and the isolated MIDI conversion environment
./install.sh
# 3. Start everything (frontend + backend + AI engine) in one terminal
./run.sh# 1. Clone this repo and ACE-Step-1.5 side by side (see full install below)
git clone https://github.com/Sion971/ace-step-studio.git
cd ace-step-studio
# 2. Run the installer β same idea as Linux, uv-managed Python throughout
install.bat
# 3. Start everything (frontend + backend + AI engine) in one terminal
run.batThat's it β the UI opens automatically at http://localhost:3001.
git clone https://github.com/ace-step/ACE-Step-1.5.gitPlace it alongside this repo β the launcher expects ../ACE-Step-1.5 relative to this project by default (configurable).
Linux
git clone https://github.com/Sion971/ace-step-studio.git
cd ace-step-studio
./install.shThe installer walks through thirteen steps, all self-checking and safe to re-run:
- System dependencies (FFmpeg, libsndfile) via apt, only if missing
- Working directory structure
- GPU / CUDA selection (Pascal through Blackwell, or CPU-only) and Python virtual environment (via
uv) - Build tools
- PyTorch, matched to your selected CUDA version
5b. NVIDIA NPP (a
torchcodecruntime dependency that PyTorch doesn't pull in on its own) - ACE-Step dependencies, including a real compute-capability check before attempting
flash-attnβ skips it cleanly (falling back to SDPA) rather than burning hours on a build that can't succeed on your hardware pytorch_waveletspatch β works around apkg_resourcesremoval in modernsetuptoolsthat otherwise silently disables the optional DCW sampler correctiontorchcodecload verification- Node.js check
- npm install (frontend and server)
- Frontend build
- Database migration (playlist/workspace schema) β idempotent, safe on every reinstall
- Isolated
basic-pitchenvironment for MIDI conversion (Python 3.11 via deadsnakes PPA)
Windows
git clone https://github.com/Sion971/ace-step-studio.git
cd ace-step-studio
install.batThe installer walks through ten steps, all self-checking and safe to re-run:
uvinstall (if missing) and GPU / CUDA selection (Pascal through Blackwell, or CPU-only)- Python 3.11 virtual environment (via
uv) - PyTorch, matched to your selected CUDA version
- ACE-Step dependencies, including
flash-attnβ a prebuilt wheel on Blackwell (RTX 50-series), verified specifically for Python 3.11 + PyTorch 2.7 + CUDA 12.8, no local compiler needed pytorch_waveletspatch β samepkg_resourcesfix as Linux, same reasoning- Node.js
- npm install (frontend and server)
- Frontend build (FFmpeg is downloaded automatically around this point too, if missing)
- Database migration (playlist/workspace schema) β idempotent, safe on every reinstall
- Isolated
basic-pitchenvironment for MIDI conversion, its ownuv-managed venv to avoid atensorboard/tensorflowversion conflict with ACE-Step's own pin
If torchaudio fails to load with Could not find module ... (or one of its dependencies), install the Microsoft Visual C++ Redistributable β a very common missing piece for compiled Python extensions on a fresh Windows install, unrelated to this project specifically.
Models download automatically on first run (~5GB).
Linux
./run.shOptions:
| Flag | Effect |
|---|---|
--no-lm |
Skip the local 5Hz LM (0.6B) β frees ~1GB VRAM, good for LoRA training |
--gradio-only |
ACE-Step's own Gradio UI only (port 8001), no Express/React frontend β needed for dataset labeling |
--no-browser |
Don't auto-open a browser tab |
--port <n> |
Web server port (default 3001) |
Windows
run.batSame flags as Linux (--no-lm, --gradio-only, --no-browser, --port <n>).
Just describe your song in natural language β genre, mood, instruments β and let ACE-Step handle the rest.
Fine-grained control over BPM, key, time signature, duration, and structure tags in your lyrics.
AI Enhance enriches short genre tags into detailed captions with proper metadata. Thinking Mode lets the model reason about song structure before generating audio codes β better results, more VRAM.
Generate several variations of the same prompt in one pass to compare results quickly.
Audio Editor (AudioMass, multitrack) β trim, fade, apply effects. Open a single stem directly from your library, or send all four Demucs stems over together as separate tracks in one editor session.
Stem Extraction (Demucs) β runs in-browser via ONNX, separates vocals/drums/bass/other. Each stem can be downloaded, converted to MIDI, or sent straight to the editor.
MIDI Conversion (basic-pitch) β runs server-side in its own isolated environment, converts any stem to MIDI in seconds.
LoRA Training β scan your dataset, label, preprocess, and train, all from the UI.
Video Generator β turn a track into a music video with Pexels stock backgrounds.
| Issue | Solution |
|---|---|
| ACE-Step not reachable | Ensure Gradio server is running with --enable-api (handled automatically by the launcher) |
| CUDA out of memory | Set batch size to 1, reduce duration, or disable Thinking Mode |
| 4GB GPU β Out of memory | Batch size 1, Thinking Mode off. LLM features need 12GB+ |
flash-attn build fails or errors at runtime (Linux) |
Check your nvcc version supports your GPU's compute capability β see install.sh step 6, or fall back to --no-lm if you just need generation working now |
torchaudio fails to load with "Could not find module ... (or one of its dependencies)" (Windows) |
Install the Microsoft Visual C++ Redistributable |
DCW disabled with a pytorch_wavelets warning |
Both installers patch this automatically (step 7 on Linux, step 5 on Windows) β if it's still happening, run patch-pytorch-wavelets.py manually against the relevant environment |
| Songs show 0:00 duration | Linux: sudo apt install ffmpeg. Windows: delete the ffmpeg\ folder and re-run the installer |
| LAN access not working | Check firewall allows the port you're running on (default 3001) |
More detail in TROUBLESHOOTING.md.
- ACE-Step β the underlying open source AI music generation model
- timoncool/ACE-Step-Studio β the project this fork is built on
- AmbsdOP/ace-step-ui β the original UI this was in turn based on
- @bdsqlsz β Chinese localization, carried over from upstream
- AudioMass β web audio editor
- Demucs β audio source separation
- basic-pitch β audio-to-MIDI conversion
- Pexels β stock video backgrounds
Built with Claude (Anthropic) as a development pair β most of this fork's Linux port, features, and this very README were worked through together, session by session.
This started as a personal project to get ACE-Step Studio running well on Linux, and it's grown from there. Development happens in short, focused sessions rather than on a fixed schedule, so don't expect instant replies β but issues, questions, and pull requests are genuinely welcome. If something's broken, tell us. If you've got an idea, open a discussion. If you've fixed something yourself, a PR is very welcome.
MIT License β see LICENSE. Original copyright retained; this fork's changes are released under the same terms.
Built on the shoulders of ACE-Step, AudioMass, Demucs, and everyone who worked on this UI before it got here.




0 comments
log in to comment.