mllm

29 projects share this GitHub topic

mllm — unilm ★22.2kmllmAgent-S — ★12.2kSpatialLM — ★4.7kNExT-GPT — ★3.6kEagle — ★3.5kInternLM-XComposer — ★2.9kmPLUG-DocOwl — ★2.4kcambrian — ★2kawesome-yolo-object-detection — ★1.8kSa2VA — ★1.7kawesome-vlm-architectures — ★1.3kOpenEMMA — ★951NEO — ★888JarvisArt — ★8614KAgent — ★819FSDrive — ★801MPP-LLaVA — ★685Woodpecker — ★649Groma — ★586Awesome-LLMs-meet-Multimodal-Generation — ★552Awesome-Multimodal-Modeling — ★541Awesome-GUI-Agents — ★443JarvisEvo — ★415EVE — ★376mega-data-factory — ★372R1-VL — ★352Awesome-LVLM-Hallucination — ★328Youku-mPLUG — ★307OmniAgent — ★70Agent-S★ 12.2kSpatialLM★ 4.7kNExT-GPT★ 3.6kEagle★ 3.5kInternLM-XComposer★ 2.9kmPLUG-DocOwl★ 2.4kcambrian★ 2kawesome-yolo-object-dete…★ 1.8kSa2VA★ 1.7kawesome-vlm-architecture…★ 1.3kOpenEMMA★ 951NEO★ 888JarvisArt★ 8614KAgent★ 819FSDrive★ 801MPP-LLaVA★ 685Woodpecker★ 649Groma★ 586Awesome-LLMs-meet-Multim…★ 552Awesome-Multimodal-Model…★ 541Awesome-GUI-Agents★ 443JarvisEvo★ 415EVE★ 376mega-data-factory★ 372R1-VL★ 352Awesome-LVLM-Hallucinati…★ 328Youku-mPLUG★ 307OmniAgent★ 70

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22.2k
Agent-S
Agent S: an open agentic framework that uses computers like a human
★ 12.2k
SpatialLM
[NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
★ 4.7k
NExT-GPT
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3.6k
Eagle
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
★ 3.5k
InternLM-XComposer
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio…
★ 2.9k
mPLUG-DocOwl
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
★ 2.4k
cambrian
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2k
awesome-yolo-object-detection
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object…
★ 1.8k
Sa2VA
Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st…
★ 1.7k
awesome-vlm-architectures
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training…
★ 1.3k
OpenEMMA
OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.
★ 951
NEO
NEO Series: Native Vision-Language Models from First Principles
★ 888
JarvisArt
[NeurIPS' 2025] JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
★ 861
4KAgent
[NeurIPS 2025] 4KAgent: Agentic Any Image to 4K Super-Resolution. An intelligent computer vision agent that…
★ 819
FSDrive
[NeurIPS 2025 spotlight] Official implementation for "FutureSightDrive: Thinking Visually with…
★ 801
MPP-LLaVA
Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support…
★ 685
Woodpecker
✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
★ 649
Groma
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
★ 586
Awesome-LLMs-meet-Multimodal-Generation
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
★ 552
Awesome-Multimodal-Modeling
Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
★ 541
Awesome-GUI-Agents
A curated collection of resources, tools, and frameworks for developing GUI Agents.
★ 443
JarvisEvo
[CVPR' 2026] JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator…
★ 415
EVE
EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 376
mega-data-factory
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
★ 372
R1-VL
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy…
★ 352
Awesome-LVLM-Hallucination
up-to-date curated list of state-of-the-art Large vision language models hallucinations research work,…
★ 328
Youku-mPLUG
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
★ 307
OmniAgent
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that…
★ 70
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.