# Local Music GenAI

***Can an open music generation model running locally on one consumer GPU produce something I would actually want to listen to?***

*Everything runs locally.  
One Windows PC.  
One RTX 4080.  
One open-weight music model,* free for non-commercial use.

***Spoiler: yes.*** *Especially if you know a bit about computers and music.*

*This blog is about* ***YuE2***, an open-weight music generation model free for non-commercial use, and how I used it on my PC to create a complete track.

## **Why I did this experiment**

I wanted to see how far a small open music model could go on a normal high-end PC. The goal was simple: generate a complete song locally, understand how the model works, and see if the result could become something I would actually keep and finish in Ableton.

* * *

## **YuE2**

YuE2 is an open music generation model developed by M-A-P ([Multimodal Art Projection](https://huggingface.co/m-a-p)), an open AI research community working on foundation models and creative AI. It takes lyrics + a style prompt and generates a complete song with vocals and accompaniment. Under the hood, it uses one AR–NAR Transformer backbone to generate musical and semantic tokens, then acoustic latents, and finally a VAE turns them into 48 kHz stereo audio.

> YuE2 on [Hugging Face](https://huggingface.co/m-a-p/YuE2-3B)

YuE2 is open-weight, but that does not mean its complete training dataset is public. The team says YuE2 was trained on around **346,000 hours of primarily CC0 music (no copyright) and synthetic data**.

YuE2 works on a compressed musical representation first, then converts it into sound. This makes the model closer to a music-generation pipeline than a simple text-to-waveform model.

* * *

## The configuration

I used my Alienware desktop with an RTX 4080 and 16 GB of VRAM, running Windows 11.

For the model, I used YuE2-3B in Q8 format with the F16 VAE, running through audio.cpp + CUDA.

Q8 was a good compromise for this test. It keeps the model small enough to fit comfortably on the GPU while preserving much more precision than lower-bit quantization.

* * *

## Learning how to prompt YuE2

YuE2 works best when I separate lyrics from production instructions.

The lyrics define the structure with tags such as \[Verse\], \[Chorus\], \[Drum Solo\] or \[Guitar Solo\].

The style prompt carries the rest: scale, BPM, instruments, vocal tone, mix, energy and atmosphere.

I also tuned a few generation parameters. I kept `num_inference_steps=8` for fast iteration, fixed the `seed` when I wanted to compare two versions, and used the default sampling values: `temperature=1`, `top_p=0.95`, `top_k=100` and `repetition_penalty=1.2`.

One funny lesson: when I wrote detailed instructions inside the lyrics field, YuE2 sometimes sang them. So I kept the lyrics clean and moved all production details into style.

* * *

## The first surprise: it is fast (and good)

![](https://cdn.hashnode.com/uploads/covers/6802c3406275c65d6ecc73df/b76cb455-4083-4a3d-b04e-17a894d7fb6a.gif align="center")

The first real surprise was the speed.

One of my complete generations took about 59 seconds on the GPU. Most of that time was spent generating the semantic music tokens. The acoustic synthesis and VAE decoding were much faster.

And the second surprise was the quality.

The first tracks were basic, but already coherent enough to hear structure, groove, vocals and instrumentation. After a few prompt iterations, some generations became genuinely interesting rather than just technically impressive.

For a local 3B model running on one consumer GPU, that changed the experiment very quickly from “does it work?” to “how far can I push it?”

* * *

## The weird experiments were the best ones

The best results came when I stopped trying to write “proper” songs.

I tested French, English, Kryptonian-inspired lyrics, then a completely invented Nordic-sounding language called Vaersk. After that, I went one step further with “yaourt”: vocals that sound English without really being English.

That was surprisingly effective. Once meaning became less important, the voice could behave more like another instrument. Rhythm, phonetics and groove started to matter more than the words themselves.

That is where YuE2 became really fun.

* * *

## The final track: UpTown

The final version used pseudo-English lyrics, the *Yaourt language* ;-), over an early-2000s alternative hip-hop style: **A** major for a brighter mood, **94 BPM**, dark boom-bap drums, warm bass, dusty piano and a low male vocal.

YuE2 generated the full song locally. I then moved to Ableton Live for EQ, compression, balance and mastering.

At that point, the experiment had become a real track. Not a masterpiece, but that was never the purpose.

%[https://open.spotify.com/album/0mLOLjzKqw3E4rFbItNFku] 

* * *

## What YuE2 is good at, and bad at

YuE2 is good at structure, groove and experimentation. It understands sections, styles, tempo, vocal character and basic arrangement surprisingly well. It also reacts nicely to unusual phonetics, which made the fake-language tests much more interesting than expected.

Its limits are also clear. The songs tend to stay quite linear: little harmonic movement, almost no scale or key changes, few production FX, and transitions that a musician can often predict. The guitar and drum solos work, but they are usually simple rather than expressive or technically interesting (still better than me on guitar; on piano, I’m not so sure...).

High frequencies can also become metallic, as with other models. Vocals can repeat too much, and detailed structure instructions are followed with some freedom. The raw mix can be uneven.

For me, YuE2 is a very good creative generator. It gives you a strong first composition and useful material, then a DAW is still the place to add movement, surprise and final production.

* * *

## Conclusion XP

Yes, it is possible to run a capable music generation model for free, for non-commercial use, on a local PC.

For an artist, that is interesting as a source of material: ideas, textures, vocal directions, arrangements, strange experiments, or simply a way to get unstuck and try something quickly.

I would not use YuE2 to release a complete song directly. The output still needs musical judgment, editing and production.

And there is one practical condition: you need some computer skills. Installing the runtime, choosing the right model format, using CUDA, managing prompts and understanding what the model is doing still requires more than clicking a button.

For me, that is part of what makes the experiment interesting. It is fun. YuE2 feels less like a finished product and more like a creative laboratory for artists.

* * *
