vision-language-model

60 Projekte teilen dieses GitHub-Topic

vision-language-model — LLaVA ★25kvision-language-modelInternVL — ★10.1kX-AnyLabeling — ★9.9kQwen-VL — ★6.7kMineContext — ★5.4kmlx-vlm — ★5.3kalign-anything — ★4.7klmms-eval — ★4.3kMiniMax-01 — ★3.4kMGM — ★3.3kVLM_survey — ★3.1kInternLM-XComposer — ★2.9kOGAM — ★2.8kcolpali — ★2.7kCradle — ★2.6kQwen-VL-Series-Finetune — ★1.9kAwesome-LLM4AD — ★1.9kAdvancedLiterateMachinery — ★1.8kvideo-search-and-summarization — ★1.8kthepipe — ★1.5kOvis — ★1.5kvllm-mlx — ★1.5kawesome-japanese-llm — ★1.4kmlx-tune — ★1.4kawesome-vlm-architectures — ★1.3kvlms-zero-to-hero — ★1.2kVisRAG — ★975Chat-UniVi — ★943top-cvpr-2025-papers — ★891VoxPoser — ★825Awesome-Robotics-3D — ★819X-VLA — ★696OmniVinci — ★675t2v_metrics — ★598Groma — ★585LLaVA-Mini — ★574Multi-Modality-Arena — ★565cambrian-s — ★563Flame-Code-VLM — ★561Senna — ★550awesome-knowledge-driven-AD — ★501InternVL★ 10.1kX-AnyLabeling★ 9.9kQwen-VL★ 6.7kMineContext★ 5.4kmlx-vlm★ 5.3kalign-anything★ 4.7klmms-eval★ 4.3kMiniMax-01★ 3.4kMGM★ 3.3kVLM_survey★ 3.1kInternLM-XComposer★ 2.9kOGAM★ 2.8kcolpali★ 2.7kCradle★ 2.6kQwen-VL-Series-Finetune★ 1.9kAwesome-LLM4AD★ 1.9kAdvancedLiterateMachiner…★ 1.8kvideo-search-and-summari…★ 1.8kthepipe★ 1.5kOvis★ 1.5kvllm-mlx★ 1.5kawesome-japanese-llm★ 1.4kmlx-tune★ 1.4kawesome-vlm-architecture…★ 1.3kvlms-zero-to-hero★ 1.2kVisRAG★ 975Chat-UniVi★ 943top-cvpr-2025-papers★ 891VoxPoser★ 825Awesome-Robotics-3D★ 819X-VLA★ 696OmniVinci★ 675t2v_metrics★ 598Groma★ 585LLaVA-Mini★ 574Multi-Modality-Arena★ 565cambrian-s★ 563Flame-Code-VLM★ 561Senna★ 550awesome-knowledge-driven…★ 501

Linien verbinden Mitglieder, die messbar miteinander verwandt sind. Die Punktgröße spiegelt die Sterne wider.

🧬 Mitglieder
LLaVA
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25k
InternVL
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. …
★ 10.1k
X-AnyLabeling
Open-source AI-assisted annotation platform for images, videos, text, and multimodal data.
★ 9.9k
Qwen-VL
The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by…
★ 6.7k
MineContext
MineContext is your proactive context-aware AI partner(Context-Engineering+ChatGPT Pulse)
★ 5.4k
mlx-vlm
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
★ 5.3k
align-anything
Align Anything: Training All-modality Model with Feedback
★ 4.7k
lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3k
MiniMax-01
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on…
★ 3.4k
MGM
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
★ 3.3k
VLM_survey
Collection of AWESOME vision-language models for vision tasks
★ 3.1k
InternLM-XComposer
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio…
★ 2.9k
OGAM
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs,…
★ 2.8k
colpali
The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.
★ 2.7k
Cradle
The Cradle framework is a first attempt at General Computer Control (GCC). Cradle supports agents to ace any…
★ 2.6k
Qwen-VL-Series-Finetune
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1.9k
Awesome-LLM4AD
A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually…
★ 1.9k
AdvancedLiterateMachinery
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project…
★ 1.8k
video-search-and-summarization
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for…
★ 1.8k
thepipe
Get clean data from tricky documents, powered by vision-language models ⚡
★ 1.5k
Ovis
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and…
★ 1.5k
vllm-mlx
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama,…
★ 1.5k
awesome-japanese-llm
日本語LLMまとめ - Overview of Japanese LLMs
★ 1.4k
mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR…
★ 1.4k
awesome-vlm-architectures
Famous Vision Language Models and Their Architectures
★ 1.3k
vlms-zero-to-hero
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge…
★ 1.2k
VisRAG
Parsing-free RAG supported by VLMs
★ 975
Chat-UniVi
[CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image…
★ 943
top-cvpr-2025-papers
About This repository is a curated collection of the most exciting and influential CVPR 2025 papers. 🔥…
★ 891
VoxPoser
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
★ 825
Awesome-Robotics-3D
A curated list of 3D Vision papers relating to Robotics domain in the era of large models i.e. LLMs/VLMs,…
★ 819
X-VLA
[ICLR 2026] The offical Implementation of "Soft-Prompted Transformer as Scalable Cross-Embodiment…
★ 696
OmniVinci
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 675
t2v_metrics
Evaluating text-to-image/video/3D models with VQAScore
★ 598
Groma
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
★ 585
LLaVA-Mini
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images,…
★ 574
Multi-Modality-Arena
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models…
★ 565
cambrian-s
Cambrian-S: Towards Spatial Supersensing in Video
★ 563
Flame-Code-VLM
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React…
★ 561
Senna
Bridging Large Vision-Language Models and End-to-End Autonomous Driving
★ 550
awesome-knowledge-driven-AD
A curated list of awesome knowledge-driven autonomous driving (continually updated)
★ 501
InstructCV
[ ICLR 2024 ] Official Codebase for "InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision…
★ 461
Open-LLaVA-NeXT
An open-source implementation for training LLaVA-NeXT.
★ 439
RoboFlamingo
Code for RoboFlamingo
★ 438
Olympus
[CVPR 2025 Highlight] Official code for "Olympus: A Universal Task Router for Computer Vision Tasks"
★ 428
InternVLA-M1
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
★ 419
OPERA
[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust…
★ 412
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 390
colpali-cookbooks
Recipes for learning, fine-tuning, and adapting ColPali to your multimodal RAG use cases. 👨🏻‍🍳
★ 357
R1-VL
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy…
★ 352
AlphaDrive
Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
★ 332
RISE
[RSS 2026] Code for RISE: Self-Improving Robot Policy with Compositional World Model
★ 329
Awesome-Multimodal-LLM-Autonomous-Driving
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
★ 313
colette
Multimodal RAG to search and interact locally with technical documents of any kind
★ 301
Thinking-with-Visual-Primitives
Archived snapshot of Thinking-with-Visual-Primitives
★ 273
ICLR2026-Guide-CN
不想啃 5000+ 全文?我已经替你和 LLM 啃完了 — ICLR 2026 全景中文导读
★ 155
Awesome-AVI
Awesome Audio-Visual Intelligence, Survey of Audio-Visual Intelligence
★ 84
OmniAgent
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that…
★ 60 · GitHub ↗
ThinkJEPA
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
★ 48 · GitHub ↗
neo-unify
Toy-scale unified multimodal model experiments — encoder-free understanding & generation with…
★ 47 · GitHub ↗
🔗 Verwandte Familien

Gemessen anhand der von beiden Projekten geteilten GitHub-Themen, gewichtet nach der Seltenheit jedes Themas.