Neural generative audio

2016-01-15 — 2026-08-09

quality 7.8

Wherein Neural Networks Are Taxed With the Generation of Sound, From Raw Waveform to Latent Diffusion, and the Peculiar Matter of a Raspberry Pi Rendering Timbre in Real Time Is Duly Noted.

generative art
machine learning
machine listening
making things
music
neural nets
signal processing
Figure 1

Neural networks generating audio: music, sound effects, the lot. This is a quick orientation page rather than a thorough survey; I have not been following closely. For the symbolic / MIDI side we have a separate notebook; for speech see voice fakes; for non‑NN signal models see analysis/resynthesis; and the underlying machinery sits in neural diffusion.

The unified audio‑text models below do not respect a speech/not‑speech division (how philosophically complex). They transcribe, synthesize speech and generate sound effects from one set of weights. I guess that means they fit here for want of a better idea?

1 Short clips

Here’s the current crop of open models that generate short bits of audio in ways I think are musically useful.

Model SFX Music Max length Sample rate Paper Code Weights
Stable Audio Open 1.0 47 s 44.1 kHz stereo (Evans et al. 2024) stability-ai/stable-audio-tools HF
AudioGen (AudioCraft) ~10 s 16 kHz mono (Kreuk, Synnaeve, et al. 2022) facebookresearch/audiocraft HF
MusicGen (AudioCraft) 30 s native, longer via sliding window 32 kHz (Copet et al. 2023) facebookresearch/audiocraft HF
AudioLDM / AudioLDM 2 ~10 s native 16 kHz (v1), 48 kHz checkpoint (v2) (H. Liu, Tian, et al. 2023; H. Liu, Chen, et al. 2023) haoheliu/AudioLDM2 HF

AudioCraft ships training scripts. stable-audio-tools also ships training scripts, but the released open weights are for non‑commercial use under the Stability AI Community License.

Stable Audio Open can also be served as an endpoint via vLLM‑Omni. Get ready for infinitely-long-form podcasts.

1.1 Open song generators

Models that write whole songs. Suno and Udio, etc. I am somewhat less excited about this category, but it is very hyped, so it needs mentioning.

Anyway, if there are to be models that write whole songs, I guess it’s best that there are open ones, from which we might learn.

YuE (乐, Chinese for “music” and “happiness”), out of HKUST and M‑A‑P, is a token‑LM model. A 7B model generates xcodec tokens from a genre‑tag‑plus‑lyrics prompt, a 1B stage refines, and out come several‑minute-long songs with separate vocal and accompaniment stems across English, Mandarin, Cantonese, Japanese, and Korean (Yuan et al. 2025). It is Apache‑2.0, which here means open in the sense that we can use the output commercially, not the “non‑commercial output usage” sense Stable Audio uses for “open” weights. It is slow — 30 s of audio runs about 150 s on an H800.

ACE‑Step (code, also Apache‑2.0) is a neural diffusion model. The authors assert that LLM song models get the lyrics aligned but run slowly and lose long‑range structure, so ACE‑Step pairs a deep‑compression autoencoder with a linear transformer and turns out roughly 4 minutes of music in about 20 s on an A100, some 15× faster. The stated aim is “the Stable Diffusion moment for music.” Since the model exposes readily accessible latent spaces, a number of cool features follow (as with image diffusions): voice cloning, repainting, lyric editing, lyric‑to‑vocal, accompaniment‑from‑vocal.

1.2 Unified audio‑text models

Everything above is audio‑out only: we prompt with text, a waveform comes back. But we can be more generally multi-modal.

Nemotron‑Labs‑Audex‑30B‑A3B is NVIDIA’s audio-text-pansexual design. [TODO clarify] Take a text LLM — Nemotron‑Cascade‑2‑30B‑A3B. Splice an audio encoder onto the input side which projects into the text embedding space. Extend the output vocabulary with the tokens of a neural codec. Then train the single decoder on all of it, so that text tokens and audio tokens are the same sort of object as far as the generation loop is concerned. This all happens, to be clear, in a unified architecture (almost — see below): what we are running is still a decoder emitting tokens. Out of that falls audio understanding, speech recognition, speech translation, text‑to‑speech, text‑to‑audio and speech‑to‑speech, plus the backbone’s thinking modes and 1M‑token context.

The paper is titled Unified Audio Intelligence Without Regressing on Text Intelligence (Kong et al. 2026). It’s surprising that they pulled this off without losing all that text facility.

Cons: As usual it has NVIDIA’s shitty Oneway Noncommercial license.

Also there was some sleight-of-hand in that “unified model” pitch up there. In fact the audio output path is not self‑contained, because the model emits codec tokens and something downstream has to turn those into sound: text‑to‑audio decodes through XCodec1 plus an optional 48 kHz enhancement VAE, text‑to‑speech through either a standalone causal speech decoder (streaming, and the default) or the original XCodec2 (better, not streaming).

1.3 In the DAW

OBSIDIAN Neural is a free AGPL‑3.0 VST3 / AU plugin (Windows / macOS / Linux) which wraps Stable Audio Open with MIDI triggering and tempo sync, plus several specialized fine‑tunes. The repository is still at innermost47/ai-dj under its old name; the homepage and README reflect the rebrand. It requires “30 seconds of patience per loop.”

A different approach is found in RAVE (code), Caillon and Esling’s real‑time variational autoencoder from IRCAM’s ACIDS group (Caillon and Esling 2021). Instead of prompting a big pretrained model and waiting, we train RAVE on our own corpus — a voice, a drum kit, a field recording, all of the above — and it learns a compact latent space to traverse live. Exported in streaming mode, it runs inside Max/MSP or PureData via the nn~ external at low enough latency that a Raspberry Pi 4 keeps up in real time (so I am told), and there is now a RAVE VST for grown-up DAWs. Because the latent is explicit, this makes it amenable to timbre transfer and haptic manipulation — the model as a playable instrument.

Figure 2

2 How we got here

A compressed timeline. NB: I have been checked out for the last 3 years and have missed much progress.

2.1 2016–18 — Raw waveform, sample by sample

WaveNet (DeepMind) and SampleRNN (Mehri et al. 2017) modelled audio one sample at a time with autoregressive networks. NSynth / WaveNet autoencoders (Engel et al. 2017) applied the idea to musical timbre. Dadabots’ SampleRNN metal albums are the entertaining proof‑of‑concept.

Sander Dieleman’s 2020 essay on waveform‑domain synthesis is still the best orientation piece for this era.

2.2 2018–20 — Alternatives to autoregression

Magenta’s DDSP (code) embedded classical synthesis modules — oscillators, filters, reverb — as differentiable layers, so we get parameter inference instead of waveform regression. The timbre transfer demo still demos well, even though Magenta itself has stagnated.

GANSynth used GANs over spectrogram representations; WaveGAN did the same for raw waveforms. MelNet went big on conditional spectrogram modelling. OpenAI’s Jukebox was the most ambitious of the era — a hierarchical VQ‑VAE plus autoregressive transformer trained on raw music — and is mostly historical now, superseded by latent diffusion.

2.3 2020–22 — Diffusion arrives

WaveGrad (N. Chen et al. 2020) and DiffWave (Kong et al. 2021) demonstrated denoising diffusion on raw audio. SaShiMi (Goel et al. 2022) showed that structured state‑space models could model raw audio as well as anything autoregressive — see the examples and code. NU‑Wave (Lee and Han 2021) and Pascual et al (Pascual et al. 2022) applied diffusion to upsampling and full‑band synthesis respectively.

2.4 2022–23 — Text conditioning

CLAP (Wu et al. 2023) gave us a contrastive joint embedding for text and audio, which is what most of the text‑to‑audio models use under the hood. Meta’s AudioGen (Kreuk, Synnaeve, et al. 2022) and Google’s MusicLM got the text‑to‑audio and text‑to‑music ball rolling. MusicGen (Copet et al. 2023) (the AudioCraft one) collapsed the cascade into a single transformer LM over EnCodec tokens.

2.5 2023–24 — Latent diffusion at scale

AudioLDM (H. Liu, Tian, et al. 2023) and AudioLDM 2 (H. Liu, Chen, et al. 2023) moved to latent diffusion, with AudioLDM 2’s “language of audio” abstraction unifying speech, music, and SFX. MusicLDM (K. Chen et al. 2023) added music‑specific tricks like tempo‑aware conditioning. Mousai (Schneider et al. 2023) is a similar latent‑diffusion approach. Stable Audio (Evans et al. 2024) went stereo and longer (47 s).

I’m not massively into spectral‑domain synthesis because I think the stationarity assumption is a stretch (heh). Or rather, my contrarian instinct says that working in the Fourier domain leaves audio quality on the table — the transient attack on percussion, the way a struck string rings out, that kind of thing — even though latent‑diffusion systems, for example, work on spectral‑adjacent latents and apparently get away with it.

2.6 2024–26 — Open weights and DAW integration

Stable Audio Open released the weights publicly. YuE and then ACE‑Step did the same for full songs with vocals, both under Apache‑2.0 — the first open weights to seriously rival Suno. OBSIDIAN Neural and friends started bringing model‑in‑plugin workflows into actual DAWs, with all the live‑performance implications that entails.

3 Methods

Seriously, this is not my field at the moment, so do not take my advice.

3.1 Latent diffusion

A VAE compresses waveforms into a much lower‑rate latent space; diffusion runs in the latent space; conditioning is by text embedding (CLAP, T5) cross‑attended into the diffusion transformer or U‑Net. Stable Audio, AudioLDM, MusicLDM, and Mousai are all variations on this theme. See neural diffusion for the underlying machinery.

3.2 Token‑based language models

Encode audio as discrete tokens with a neural codec (EnCodec, SoundStream, xcodec); train a transformer LM to predict the tokens autoregressively; decode the predicted tokens back to audio. MusicGen and AudioGen are archetypal examples.

Choice of codec turns out to be important here. Codecs tend to specialize on general waveforms, or on speech specifically. The X‑Codec team Ye et al. (2024) argues that the former case is leaving speech semantic information unexploited.

3.3 Differentiable DSP

DDSP threads a different needle: keep the classical synthesis topology (oscillator, filter, reverb), make the parameters differentiable, learn parameter trajectories from audio. The output is then by construction a real synthesis chain, not a regressed waveform.

3.4 State‑space models

SaShiMi (Goel et al. 2022) and successors, as seen in the state‑space models notebook, are a different way of modelling long sequences. Not currently the state of the art for music generation (which surprises me, TBH — so many connections to traditional synthesis methods) but a tidy alternative to attention for long sequences.

4 Tooling

5 Praxis and politics

Streaming platforms are now flooded with AI slop songs cranked out at low marginal cost, royalty pools are getting diluted, and working musicians’ incomes are getting squeezed. This is, I think we can all agree, what we might call “bad”.

At the same time, I am interested in composers and producers who treat the models as instruments rather than as cheap musician‑replacements. Examples: Jlin and Holly Herndon showed early on what an AI‑forward composition practice could look like — the model as collaborator and as instrument, the glitches as material. GAN.STYLE is in a similar vein.

6 References

Blaauw, and Bonada. 2017. A Neural Parametric Singing Synthesizer.” arXiv:1704.03809 [Cs].
Caillon, and Esling. 2021. RAVE: A Variational Autoencoder for Fast and High-Quality Neural Audio Synthesis.”
Carr, and Zukowski. 2018. Generating Albums with SampleRNN to Imitate Metal, Rock, and Punk Bands.” arXiv:1811.06633 [Cs, Eess].
Chen, Ke, Wu, Liu, et al. 2023. MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies.”
Chen, Nanxin, Zhang, Zen, et al. 2020. WaveGrad: Estimating Gradients for Waveform Generation.”
Copet, Kreuk, Gat, et al. 2023. Simple and Controllable Music Generation.”
Dieleman, Oord, and Simonyan. 2018. The Challenge of Realistic Music Generation: Modelling Raw Audio at Scale.” In Advances In Neural Information Processing Systems.
Du, Collins, Tenenbaum, et al. 2021. Learning Signal-Agnostic Manifolds of Neural Fields.” In Advances in Neural Information Processing Systems.
Dupont, Kim, Eslami, et al. 2022. From Data to Functa: Your Data Point Is a Function and You Can Treat It Like One.” In Proceedings of the 39th International Conference on Machine Learning.
Elbaz, and Zibulevsky. 2017. Perceptual Audio Loss Function for Deep Learning.” In Proceedings of the 18th International Society for Music Information Retrieval Conference (ISMIR’2017), Suzhou, China.
Engel, Resnick, Roberts, et al. 2017. Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders.” In PMLR.
Evans, Parker, Carr, et al. 2024. Stable Audio Open.”
Goel, Gu, Donahue, et al. 2022. It’s Raw! Audio Generation with State-Space Models.”
Grais, Ward, and Plumbley. 2018. Raw Multi-Channel Audio Source Separation Using Multi-Resolution Convolutional Auto-Encoders.” arXiv:1803.00702 [Cs].
Hernandez-Olivan, Hernandez-Olivan, and Beltran. 2022. A Survey on Artificial Intelligence for Music Generation: Agents, Domains and Perspectives.”
Kong, Lee, Kim, et al. 2026. Unified Audio Intelligence Without Regressing on Text Intelligence.”
Kong, Ping, Huang, et al. 2021. DiffWave: A Versatile Diffusion Model for Audio Synthesis.”
Kreuk, Synnaeve, Polyak, et al. 2022. AudioGen: Textually Guided Audio Generation.”
Kreuk, Taigman, Polyak, et al. 2022. Audio Language Modeling Using Perceptually-Guided Discrete Representations.”
Lee, and Han. 2021. NU-Wave: A Diffusion Probabilistic Model for Neural Audio Upsampling.” In Interspeech 2021.
Levy, Di Giorgi, Weers, et al. 2023. Controllable Music Production with Diffusion Models and Guidance Gradients.”
Liu, Haohe, Chen, Yuan, et al. 2023. AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.”
Liu, Yuzhou, Thoshkahna, Milani, et al. 2020. Voice and Accompaniment Separation in Music Using Self-Attention Convolutional Neural Network.”
Liu, Haohe, Tian, Yuan, et al. 2023. AudioLDM 2: Learning Holistic Audio Generation with Self-Supervised Pretraining.”
Liutkus, Badeau, and Richard. 2011. Gaussian Processes for Underdetermined Source Separation.” IEEE Transactions on Signal Processing.
Luo, Du, Tarr, et al. 2021. Learning Neural Acoustic Fields.” In.
Mehri, Kumar, Gulrajani, et al. 2017. SampleRNN: An Unconditional End-to-End Neural Audio Generation Model.” In Proceedings of International Conference on Learning Representations (ICLR) 2017.
Pascual, Bhattacharya, Yeh, et al. 2022. Full-Band General Audio Synthesis with Score-Based Diffusion.”
Sarroff, and Casey. 2014. Musical Audio Synthesis Using Autoencoding Neural Nets.” In.
Schlüter, and Böck. 2014. Improved Musical Onset Detection with Convolutional Neural Networks.” In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).
Schneider, Kamal, Jin, et al. 2023. Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion.”
Sprechmann, Bruna, and LeCun. 2014. Audio Source Separation with Discriminative Scattering Networks.” arXiv:1412.7022 [Cs].
Stöter, Uhlich, Liutkus, et al. 2019. Open-Unmix - A Reference Implementation for Music Source Separation.” Journal of Open Source Software.
Tenenbaum, and Freeman. 2000. Separating Style and Content with Bilinear Models.” Neural Computation.
Tzinis, Wang, and Smaragdis. 2020. “Sudo Rm -Rf: Efficient Networks for Universal Audio Source Separation.” In.
Venkataramani, and Smaragdis. 2017. End to End Source Separation with Adaptive Front-Ends.” arXiv:1705.02514 [Cs].
Venkataramani, Subakan, and Smaragdis. 2017. Neural Network Alternatives to Convolutive Audio Models for Source Separation.” arXiv:1709.07908 [Cs, Eess].
Verma, and Smith. 2018. Neural Style Transfer for Audio Spectograms.” In 31st Conference on Neural Information Processing Systems (NIPS 2017).
von Platen, Patil, Lozhkov, et al. 2022. Diffusers: State-of-the-Art Diffusion Models.”
Wu, Chen, Zhang, et al. 2023. Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation.”
Wyse. 2017. Audio Spectrogram Representations for Processing with Convolutional Neural Networks.” In Proceedings of the First International Conference on Deep Learning and Music, Anchorage, US, May, 2017 (arXiv:1706.08675v1 [Cs.NE]).
Xu, Wang, Jiang, et al. 2022. Signal Processing for Implicit Neural Representations.” In.
Ye, Sun, Lei, et al. 2024. Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model.”
Yuan, Lin, Guo, et al. 2025. YuE: Scaling Open Foundation Models for Long-Form Music Generation.”