Home Projects RPG-DiffusionMaster
RPG-DiffusionMaster
Jupyter Notebook

RPG-DiffusionMaster

[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)

by YangLing0818 · GitHub
Stars
Forks
Trending
License
Created
Last commit
image-edittinglarge-language-modelsmultimodal-large-language-modelsMITJupyter Notebook
View on GitHub
In plain words

Generate and edit images from text descriptions using advanced AI models in a Jupyter Notebook.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
RPG-DiffusionMaster — GitHub preview card
📈 Star history
1 8441 838
2026-07-202026-08-31
📈 Track RPG-DiffusionMaster

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)

RPG-DiffusionMaster has 1.8k stars on GitHub. It has been forked 100 times. RPG-DiffusionMaster is written mainly in Jupyter Notebook. It has been in active development since 2024. RPG-DiffusionMaster is available under the MIT license. Its main topics are image-editting, large-language-models, multimodal-large-language-models, text-to-image.

Frequently asked questions

What is RPG-DiffusionMaster?

[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)

Is RPG-DiffusionMaster open source?

RPG-DiffusionMaster is an open-source project. It is released under the MIT license.

Is RPG-DiffusionMaster free?

Yes. RPG-DiffusionMaster is free and open source — you can use, modify and self-host it.

What license does RPG-DiffusionMaster use?

RPG-DiffusionMaster is available under the MIT license.

What language is RPG-DiffusionMaster written in?

RPG-DiffusionMaster is written mainly in Jupyter Notebook.

🏅 Maintainer of this project?
olud.ai badge — RPG-DiffusionMaster

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=yangling0818-rpg-diffusionmaster)](https://olud.ai/project/yangling0818-rpg-diffusionmaster.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →
Ovis-U1
An unified model that seamlessly integrates multimodal understanding, text-to-image generation,…
450 · image-editing
sharesmultimodal-large-language-modelstext-to-image
Magic-Me
Codes for ID-Specific Video Customized Diffusion
460 · diffusion-models
sharesimage-editting
Ovis-Image
Ovis-Image is a 7B text-to-image model specifically optimized for high-quality text rendering,…
318 · image-generation
sharestext-to-image
DALLE2-pytorch
Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis neural network, in Pytorc…
11.3k · artificial-intelligence
sharestext-to-image
DF-GAN
[CVPR2022 oral] A Simple and Effective Baseline for Text-to-Image Synthesis
326 · generative-adversarial-network
sharestext-to-image
Awesome-Medical-Large-Language-Models
Curated papers on Large Language Models in Healthcare and Medical domain
390 · large-language-models
sharesmultimodal-large-language-models
Liquid
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
640 · autoregressive-models
sharesmultimodal-large-language-modelstext-to-image
Spatial-MLLM
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in…
477 · aigc
sharesmultimodal-large-language-models
VQGAN-CLIP
Just playing with getting VQGAN+CLIP running locally, rather than having to use colab.
2.7k · text-to-image
sharestext-to-image
Attend-and-Excite
Official Implementation for "Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-I…
772 · diffusion-models
sharestext-to-image

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.