vision-language-model

41 proyectos comparten este topic de GitHub

vision-language-model — InternVL ★10.1kvision-language-modelQwen-VL — ★6.7kmlx-vlm — ★5.4kalign-anything — ★4.7klmms-eval — ★4.4kMiniMax-01 — ★3.5kMGM — ★3.3kVLM_survey — ★3.1kOGAM — ★3kInternLM-XComposer — ★2.9kCradle — ★2.6kAwesome-LLM4AD — ★1.9kvideo-search-and-summarization — ★1.8kvllm-mlx — ★1.6kthepipe — ★1.5kawesome-japanese-llm — ★1.4kmlx-tune — ★1.4kdoc7 — ★1.2kagent-vision-toolkit — ★1.1kChat-UniVi — ★941VoxPoser — ★833OmniVinci — ★677t2v_metrics — ★600Groma — ★586LLaVA-Mini — ★577Multi-Modality-Arena — ★566Senna — ★551awesome-knowledge-driven-AD — ★499RoboFlamingo — ★437InternVLA-M1 — ★418Stream-Omni — ★392Awesome-Multimodal-LLM-Autonomous-Driving — ★311Thinking-with-Visual-Primitives — ★276GRACE — ★217ICLR2026-Guide-CN — ★168NEWTON — ★143VLMForge — ★89Awesome-AVI — ★86OmniAgent — ★70ThinkJEPA — ★53neo-unify — ★51Qwen-VL★ 6.7kmlx-vlm★ 5.4kalign-anything★ 4.7klmms-eval★ 4.4kMiniMax-01★ 3.5kMGM★ 3.3kVLM_survey★ 3.1kOGAM★ 3kInternLM-XComposer★ 2.9kCradle★ 2.6kAwesome-LLM4AD★ 1.9kvideo-search-and-summari…★ 1.8kvllm-mlx★ 1.6kthepipe★ 1.5kawesome-japanese-llm★ 1.4kmlx-tune★ 1.4kdoc7★ 1.2kagent-vision-toolkit★ 1.1kChat-UniVi★ 941VoxPoser★ 833OmniVinci★ 677t2v_metrics★ 600Groma★ 586LLaVA-Mini★ 577Multi-Modality-Arena★ 566Senna★ 551awesome-knowledge-driven…★ 499RoboFlamingo★ 437InternVLA-M1★ 418Stream-Omni★ 392Awesome-Multimodal-LLM-A…★ 311Thinking-with-Visual-Pri…★ 276GRACE★ 217ICLR2026-Guide-CN★ 168NEWTON★ 143VLMForge★ 89Awesome-AVI★ 86OmniAgent★ 70ThinkJEPA★ 53neo-unify★ 51

Las líneas conectan a los miembros que están mediblemente relacionados entre sí. El tamaño de los puntos refleja las estrellas.

🧬 Miembros
InternVL
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. …
★ 10.1k
Qwen-VL
The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by…
★ 6.7k
mlx-vlm
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
★ 5.4k
align-anything
Align Anything: Training All-modality Model with Feedback
★ 4.7k
lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.4k
MiniMax-01
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on…
★ 3.5k
MGM
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
★ 3.3k
VLM_survey
Collection of AWESOME vision-language models for vision tasks
★ 3.1k
OGAM
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs,…
★ 3k
InternLM-XComposer
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio…
★ 2.9k
Cradle
The Cradle framework is a first attempt at General Computer Control (GCC). Cradle supports agents to ace any…
★ 2.6k
Awesome-LLM4AD
A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually…
★ 1.9k
video-search-and-summarization
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for…
★ 1.8k
vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX,…
★ 1.6k
thepipe
Get clean data from tricky documents, powered by vision-language models ⚡
★ 1.5k
awesome-japanese-llm
日本語LLMまとめ - Overview of Japanese LLMs
★ 1.4k
mlx-tune
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR…
★ 1.4k
doc7
Turn documents into AI-ready Markdown with visual understanding
★ 1.2k
agent-vision-toolkit
★ 1.1k
Chat-UniVi
[CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image…
★ 941
VoxPoser
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
★ 833
OmniVinci
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 677
t2v_metrics
Evaluating text-to-image/video/3D models with VQAScore
★ 600
Groma
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
★ 586
LLaVA-Mini
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images,…
★ 577
Multi-Modality-Arena
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models…
★ 566
Senna
Bridging Large Vision-Language Models and End-to-End Autonomous Driving
★ 551
awesome-knowledge-driven-AD
A curated list of awesome knowledge-driven autonomous driving (continually updated)
★ 499
RoboFlamingo
Code for RoboFlamingo
★ 437
InternVLA-M1
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
★ 418
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 392
Awesome-Multimodal-LLM-Autonomous-Driving
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
★ 311
Thinking-with-Visual-Primitives
Archived snapshot of Thinking-with-Visual-Primitives
★ 276
GRACE
[ICML 2026] GRACE-VLM: deployable INT4 Qwen3-VL via quantization-aware distillation.
★ 217
ICLR2026-Guide-CN
不想啃 5000+ 全文?我已经替你和 LLM 啃完了 — ICLR 2026 全景中文导读
★ 168
NEWTON
NEWTON: Agentic Planning for Physically Grounded Video Generation
★ 143
VLMForge
Evaluate visual models on your own images, JSON Schema, and production constraints with LangGraph, Pareto…
★ 89
Awesome-AVI
Awesome Audio-Visual Intelligence, Survey of Audio-Visual Intelligence
★ 86
OmniAgent
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that…
★ 70
ThinkJEPA
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
★ 53
neo-unify
Toy-scale unified multimodal model experiments — encoder-free understanding & generation with…
★ 51
🔗 Familias relacionadas

Medido a partir de los temas de GitHub compartidos por ambos proyectos, ponderado por cuán raros son cada uno de los temas.