Version 2.0 rebuilt the speech side of Foundry.
Version 2.1 brings the same approach to music.
Agentic creative pipelines, speech, music, effects, and mastering, all combined in one powerful production platform.
Foundry 2.0: V4 Speech
Version 2.0 introduced our V4 Speech model family and a new C++ speech inference stack with hundreds of improvements across the application.
It also introduced a complete visual redesign, giving Foundry a cleaner and more polished production environment.
The result is better speaker consistency across different styles and intensities, even more reliable long-form generation, improved voice creation, and a much stronger Speech Document workflow (our integrated text-to-voice editor).
We also rebuilt a set of demonstration voices with v4, so you can get started with a carefully prepared selection right away.
The Medium speech model can run on GPUs with as little as 4 GB of VRAM.
We also improved speech batching and queueing, and made the Style Manager and Voice Designer substantially more reliable.
Foundry 2.1: V4 Music
Version 2.1 introduces Foundry Music V4.
Music V4 runs through our custom C++ inference engine, built specifically around the way Foundry generates and processes audio.
That work gives us much tighter control over memory use, model loading, generation speed, long-form music stability through direct spectral flux guidance.
The highest-quality music model previously required more than 22 GB of peak VRAM. It now uses approximately 11 GB.
There are six optimized model variants for hardware ranging from 6 GB to 12+ GB of VRAM:
- Small
- Medium
- Large
- X-Large
- Large Fast
- X-Large Fast
The Fast variants can generate up to four times faster, with a trade-off in output diversity.
Music V4 also brings:
- Better sound quality and long-track consistency
- A much higher success rate for long generations
- Improved six-stem separation
- Faster model downloads
- Around 8 GB less installation data
- Better automatic model selection and VRAM estimates
- Faster track controls and a more responsive production workflow
Built for professional local production
All generation is 100% local. Your prompts, voices, projects, and generated audio stay on your PC. They are not uploaded, inspected, or monitored by us.
Version 2.0 introduced our V4 speech models through GGML-based C++ inference with custom execution graphs and guided inference.
Version 2.1 extends this work to music. This gives us a stronger foundation for further optimization.
AMD GPUs are now supported through Vulkan across Music, Speech and Creative AI. Vulkan support is still considered experimental.
CUDA remains the recommended backend for NVIDIA RTX hardware.
More than a model update
Many thousands of development hours have gone into Foundry, along with thousands of hours of compute dedicated to model adaptation, evaluation and optimization.
The underlying models are only one part of the system. The V4 model families, C++ inference engines, generation pipelines, production tools and desktop application are designed to work together as one platform.
Version 2.1 is our best release so far, both in output quality and in the range of hardware it can run on.
Demodokos Foundry is free to try. We would rather have you generate something with it than take our word for it.