multimodal-large-language-models

36 projets partagent ce topic GitHub

multimodal-large-language-models — Awesome-Multimodal-Large-Language-Models ★18kmultimodal-large-language-modelsstar-vector — ★4.6kBayLing-Speech — ★3.1kLLaMA-Omni — ★3.1kVideoPipe — ★2.9kcambrian — ★2kRPG-DiffusionMaster — ★1.8kOvis — ★1.5kawesome-multimodal-in-medical-imaging — ★974NEO — ★888Video-MME — ★791LLaVA-Plus-Codebase — ★770MovieChat — ★706unicom — ★702OmniVinci — ★677Woodpecker — ★649Liquid — ★640LLaVA-Mini — ★577cambrian-s — ★560Awesome-LLMs-meet-Multimodal-Generation — ★552EVF-SAM — ★505Spatial-MLLM — ★477awesome-vla-for-ad — ★468Ovis-U1 — ★450Awesome_Matching_Pretraining_Transfering — ★444Freeze-Omni — ★396Awesome-Medical-Large-Language-Models — ★391EVE — ★376Video-MME-v2 — ★367Awesome-Multimodal-LLM — ★355R1-VL — ★352VisionReasoner — ★348Awesome-LVLM-Hallucination — ★328Awesome-Multimodal-LLM-Autonomous-Driving — ★311AudioStory — ★301Awesome-AVI — ★86star-vector★ 4.6kBayLing-Speech★ 3.1kLLaMA-Omni★ 3.1kVideoPipe★ 2.9kcambrian★ 2kRPG-DiffusionMaster★ 1.8kOvis★ 1.5kawesome-multimodal-in-me…★ 974NEO★ 888Video-MME★ 791LLaVA-Plus-Codebase★ 770MovieChat★ 706unicom★ 702OmniVinci★ 677Woodpecker★ 649Liquid★ 640LLaVA-Mini★ 577cambrian-s★ 560Awesome-LLMs-meet-Multim…★ 552EVF-SAM★ 505Spatial-MLLM★ 477awesome-vla-for-ad★ 468Ovis-U1★ 450Awesome_Matching_Pretrai…★ 444Freeze-Omni★ 396Awesome-Medical-Large-La…★ 391EVE★ 376Video-MME-v2★ 367Awesome-Multimodal-LLM★ 355R1-VL★ 352VisionReasoner★ 348Awesome-LVLM-Hallucinati…★ 328Awesome-Multimodal-LLM-A…★ 311AudioStory★ 301Awesome-AVI★ 86

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
Awesome-Multimodal-Large-Language-Models
:sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18k
star-vector
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation…
★ 4.6k
BayLing-Speech
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon…
★ 3.1k
LLaMA-Omni
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon…
★ 3.1k
VideoPipe
A cross-platform video structuring (video analysis) framework based on CV models & mLLM.
★ 2.9k
cambrian
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2k
RPG-DiffusionMaster
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs…
★ 1.8k
Ovis
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and…
★ 1.5k
awesome-multimodal-in-medical-imaging
A collection of resources on applications of multi-modal learning in medical imaging.
★ 974
NEO
NEO Series: Native Vision-Language Models from First Principles
★ 888
Video-MME
✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video…
★ 791
LLaVA-Plus-Codebase
LLaVA-Plus: Large Language and Vision Assistants that Plug and Learn to Use Skills
★ 770
MovieChat
[CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
★ 706
unicom
Large-Scale Visual Representation Model
★ 702
OmniVinci
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 677
Woodpecker
✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
★ 649
Liquid
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
★ 640
LLaVA-Mini
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images,…
★ 577
cambrian-s
Cambrian-S: Towards Spatial Supersensing in Video
★ 560
Awesome-LLMs-meet-Multimodal-Generation
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
★ 552
EVF-SAM
Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"
★ 505
Spatial-MLLM
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based…
★ 477
awesome-vla-for-ad
🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
★ 468
Ovis-U1
An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image…
★ 450
Awesome_Matching_Pretraining_Transfering
The Paper List of Large Multi-Modality Model (Perception, Generation, Unification), Parameter-Efficient…
★ 444
Freeze-Omni
✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
★ 396
Awesome-Medical-Large-Language-Models
Curated papers on Large Language Models in Healthcare and Medical domain
★ 391
EVE
EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 376
Video-MME-v2
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
★ 367
Awesome-Multimodal-LLM
Research Trends in LLM-guided Multimodal Learning.
★ 355
R1-VL
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy…
★ 352
VisionReasoner
[ICLR 2026] VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
★ 348
Awesome-LVLM-Hallucination
up-to-date curated list of state-of-the-art Large vision language models hallucinations research work,…
★ 328
Awesome-Multimodal-LLM-Autonomous-Driving
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
★ 311
AudioStory
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
★ 301
Awesome-AVI
Awesome Audio-Visual Intelligence, Survey of Audio-Visual Intelligence
★ 86
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.