multi-modality

15 projects share this GitHub topic

multi-modality — LLaVA ★25kmulti-modalityAwesome-Multimodal-Large-Language-Models — ★18kclip-as-service — ★12.8kdeep-daze — ★4.3kOtter — ★3.4kInternLM-XComposer — ★2.9k3DObjectTracking — ★1kVisRAG — ★975Long-RL — ★727Multi-Modality-Arena — ★565MM-Diffusion — ★453Collaborative-Diffusion — ★441Open-LLaVA-NeXT — ★439Olympus — ★428RLHF-V — ★310Awesome-Multimodal-Large…★ 18kclip-as-service★ 12.8kdeep-daze★ 4.3kOtter★ 3.4kInternLM-XComposer★ 2.9k3DObjectTracking★ 1kVisRAG★ 975Long-RL★ 727Multi-Modality-Arena★ 565MM-Diffusion★ 453Collaborative-Diffusion★ 441Open-LLaVA-NeXT★ 439Olympus★ 428RLHF-V★ 310

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
LLaVA
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25k
Awesome-Multimodal-Large-Language-Models
:sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18k
clip-as-service
🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
★ 12.8k
deep-daze
Simple command line tool for text to image generation using OpenAI's CLIP and Siren (Implicit neural…
★ 4.3k
Otter
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained…
★ 3.4k
InternLM-XComposer
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio…
★ 2.9k
3DObjectTracking
Algorithms and Publications on 3D Object Tracking
★ 1k
VisRAG
Parsing-free RAG supported by VLMs
★ 975
Long-RL
Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)
★ 727
Multi-Modality-Arena
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models…
★ 565
MM-Diffusion
[CVPR'23] MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation
★ 453
Collaborative-Diffusion
[CVPR 2023] Collaborative Diffusion
★ 441
Open-LLaVA-NeXT
An open-source implementation for training LLaVA-NeXT.
★ 439
Olympus
[CVPR 2025 Highlight] Official code for "Olympus: A Universal Task Router for Computer Vision Tasks"
★ 428
RLHF-V
[CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human…
★ 310
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.