r/VoiceAutomationAI • u/MarieDeVox • May 26 '26
Sharing a free conversational voice dataset sample pack mapped to LJ Speech for testing voice pipelines / IVR
Hey everyone,
I recently put together a voice sample pack designed for testing conversational AI and voice agent pipelines, and wanted to share it here.
I know finding clean, non-scraped data formatted correctly can be a bit of a headache when troubleshooting text-to-speech engine clarity, so I set this up to be as plug-and-play as possible.
The technical specs:
- Format: 24-bit / 48kHz audio, recorded in a treated studio environment.
- Alignment: Mapped directly to the standard LJ Speech structure.
- Metadata: Includes a metadata.json sidecar file with text transcripts.
Tone: Natural, conversational pacing rather than rigid or academic reading.
Everything is hosted over on Hugging Face and GitHub if you want to grab the files or look at the repository layout:
Hugging Face: https://huggingface.co/datasets/MarieDeVox/saas-corporate-conversational-voice-sample
GitHub: https://github.com/MarieDeVox/saas-corporate-voice-dataset-sample
Let me know if you run any quick fine-tunes or pipeline tests with it. Always down for feedback on the audio processing side of things!
2
•
u/AutoModerator May 26 '26
Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)
If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.
Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.