multimodal-large-language-models

44 projetos partilham este topic do GitHub

multimodal-large-language-models — Awesome-Multimodal-Large-Language-Models ★18kmultimodal-large-language-modelsMobileAgent — ★9kstar-vector — ★4.5kBayLing-Speech — ★3.1kVideoPipe — ★2.9kmPLUG-DocOwl — ★2.4kcambrian — ★2kRPG-DiffusionMaster — ★1.8kOvis — ★1.5kAwesome-MCoT — ★1kawesome-multimodal-in-medical-imaging — ★973NEO — ★878Video-MME — ★788LLaVA-Plus-Codebase — ★770MovieChat — ★706unicom — ★700MPP-LLaVA — ★685OmniVinci — ★675Woodpecker — ★649Liquid — ★642LLaVA-Mini — ★574cambrian-s — ★563Awesome-LLMs-meet-Multimodal-Generation — ★551EVF-SAM — ★505Spatial-MLLM — ★481awesome-vla-for-ad — ★453Ovis-U1 — ★450Awesome_Matching_Pretraining_Transfering — ★446Awesome-Medical-Large-Language-Models — ★393Freeze-Omni — ★388EVE — ★376lmms-finetune — ★373Video-MME-v2 — ★369Awesome-Multimodal-LLM — ★355R1-VL — ★352VisionReasoner — ★348Awesome-Multimodal-Papers — ★343Awesome-LVLM-Hallucination — ★325Awesome-Multimodal-LLM-Autonomous-Driving — ★313LLMVoX — ★308Youku-mPLUG — ★307MobileAgent★ 9kstar-vector★ 4.5kBayLing-Speech★ 3.1kVideoPipe★ 2.9kmPLUG-DocOwl★ 2.4kcambrian★ 2kRPG-DiffusionMaster★ 1.8kOvis★ 1.5kAwesome-MCoT★ 1kawesome-multimodal-in-me…★ 973NEO★ 878Video-MME★ 788LLaVA-Plus-Codebase★ 770MovieChat★ 706unicom★ 700MPP-LLaVA★ 685OmniVinci★ 675Woodpecker★ 649Liquid★ 642LLaVA-Mini★ 574cambrian-s★ 563Awesome-LLMs-meet-Multim…★ 551EVF-SAM★ 505Spatial-MLLM★ 481awesome-vla-for-ad★ 453Ovis-U1★ 450Awesome_Matching_Pretrai…★ 446Awesome-Medical-Large-La…★ 393Freeze-Omni★ 388EVE★ 376lmms-finetune★ 373Video-MME-v2★ 369Awesome-Multimodal-LLM★ 355R1-VL★ 352VisionReasoner★ 348Awesome-Multimodal-Paper…★ 343Awesome-LVLM-Hallucinati…★ 325Awesome-Multimodal-LLM-A…★ 313LLMVoX★ 308Youku-mPLUG★ 307

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
Awesome-Multimodal-Large-Language-Models
:sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18k
MobileAgent
Mobile-Agent: The Powerful GUI Agent Family
★ 9k
star-vector
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation…
★ 4.5k
BayLing-Speech
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon…
★ 3.1k
VideoPipe
A cross-platform video structuring (video analysis) framework. If you find it helpful, please give it a star:…
★ 2.9k
mPLUG-DocOwl
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
★ 2.4k
cambrian
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2k
RPG-DiffusionMaster
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs…
★ 1.8k
Ovis
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and…
★ 1.5k
Awesome-MCoT
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
★ 1k
awesome-multimodal-in-medical-imaging
A collection of resources on applications of multi-modal learning in medical imaging.
★ 973
NEO
NEO Series: Native Vision-Language Models from First Principles
★ 878
Video-MME
✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video…
★ 788
LLaVA-Plus-Codebase
LLaVA-Plus: Large Language and Vision Assistants that Plug and Learn to Use Skills
★ 770
MovieChat
[CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
★ 706
unicom
Large-Scale Visual Representation Model
★ 700
MPP-LLaVA
Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support…
★ 685
OmniVinci
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 675
Woodpecker
✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
★ 649
Liquid
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
★ 642
LLaVA-Mini
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images,…
★ 574
cambrian-s
Cambrian-S: Towards Spatial Supersensing in Video
★ 563
Awesome-LLMs-meet-Multimodal-Generation
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
★ 551
EVF-SAM
Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"
★ 505
Spatial-MLLM
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based…
★ 481
awesome-vla-for-ad
🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
★ 453
Ovis-U1
An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image…
★ 450
Awesome_Matching_Pretraining_Transfering
The Paper List of Large Multi-Modality Model (Perception, Generation, Unification), Parameter-Efficient…
★ 446
Awesome-Medical-Large-Language-Models
Curated papers on Large Language Models in Healthcare and Medical domain
★ 393
Freeze-Omni
✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
★ 388
EVE
EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 376
lmms-finetune
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave,…
★ 373
Video-MME-v2
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
★ 369
Awesome-Multimodal-LLM
Research Trends in LLM-guided Multimodal Learning.
★ 355
R1-VL
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy…
★ 352
VisionReasoner
[ICLR 2026] VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
★ 348
Awesome-Multimodal-Papers
A curated list of awesome Multimodal studies.
★ 343
Awesome-LVLM-Hallucination
up-to-date curated list of state-of-the-art Large vision language models hallucinations research work,…
★ 325
Awesome-Multimodal-LLM-Autonomous-Driving
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
★ 313
LLMVoX
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
★ 308
Youku-mPLUG
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
★ 307
colette
Multimodal RAG to search and interact locally with technical documents of any kind
★ 301
AudioStory
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
★ 301
Awesome-AVI
Awesome Audio-Visual Intelligence, Survey of Audio-Visual Intelligence
★ 84
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.