Home Projects AudioStory
AudioStory
Jupyter Notebook

AudioStory

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

by TencentARC · GitHub
Stars
Forks
Created
Last commit
audio-generationdiffusion-modelsmultimodal-large-language-modelsJupyter Notebook
View on GitHub
In plain words

Generate long audio stories from text, or dub videos with audio using a simple model.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
AudioStory — GitHub preview card
📈 Star history
302301
2026-07-202026-08-31
📈 Track AudioStory

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

AudioStory has 301 stars on GitHub. It has been forked 22 times. AudioStory is written mainly in Jupyter Notebook. It has been in active development since 2025. Its main topics are audio-generation, diffusion-models, multimodal-large-language-models, text-to-audio.

Frequently asked questions

What is AudioStory?

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

Is AudioStory open source?

AudioStory is an open-source project.

Is AudioStory free?

Yes. AudioStory is free and open source — you can use, modify and self-host it.

What language is AudioStory written in?

AudioStory is written mainly in Jupyter Notebook.

🏅 Maintainer of this project?
olud.ai badge — AudioStory

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=tencentarc-audiostory)](https://olud.ai/project/tencentarc-audiostory.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →
tango
A family of diffusion models for text-to-audio generation.
1.2k · audio-generation
sharestext-to-audioaudio-generation
EzAudio
High-quality Text-to-Audio Generation with Efficient Diffusion Transformer
332 · diffusion-models
sharestext-to-audiodiffusion-models
mustango
Mustango: Toward Controllable Text-to-Music Generation
395 · diffusion-models
sharestext-to-audiodiffusion-models
Make-An-Audio
PyTorch Implementation of Make-An-Audio (ICML'23) with a Text-to-Audio Generative Model
669 · diffusion-models
sharestext-to-audiodiffusion-models
MM-Diffusion
[CVPR'23] MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generat…
453 · audio-generation
sharesaudio-generationdiffusion-models
audio-ai-timeline
A timeline of the latest AI models for audio generation, starting in 2023!
1.9k · artificial-intelligence
sharesaudio-generation
nuwa-pytorch
Implementation of NÜWA, state of the art attention network for text to video synthesis, in Pyto…
548 · artificial-intelligence
sharestext-to-audio
YuE
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
6.3k · ai
sharesaudio-generation
soundstorm-pytorch
Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pyto…
1.5k · artificial-intelligence
sharesaudio-generation
Awesome-Medical-Large-Language-Models
Curated papers on Large Language Models in Healthcare and Medical domain
390 · large-language-models
sharesmultimodal-large-language-models

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.