llava

26 Projekte teilen dieses GitHub-Topic

llava — LLaVA ★25kllavaSUPIR — ★5.6kmlx-vlm — ★5.3kVLMEvalKit — ★4.3kzero_nlp — ★3.8kLLamaSharp — ★3.8kEagle — ★3.3kOmAgent — ★2.7kVideo-ChatGPT — ★1.5ktaggui — ★1.3kawesome-vlm-architectures — ★1.3kUForm — ★1.2kMindPipe — ★1kTinyLLaVA_Factory — ★995PaddleMIX — ★724Vision-Language-Models-Overview — ★682awesome-foundation-and-multimodal-models — ★638LLaVA-Mini — ★574kolosal-cli — ★465RLAIF-V — ★458Kolosal — ★455Open-LLaVA-NeXT — ★439Olympus — ★428lmms-finetune — ★373HallusionBench — ★342ViP-LLaVA — ★339SUPIR★ 5.6kmlx-vlm★ 5.3kVLMEvalKit★ 4.3kzero_nlp★ 3.8kLLamaSharp★ 3.8kEagle★ 3.3kOmAgent★ 2.7kVideo-ChatGPT★ 1.5ktaggui★ 1.3kawesome-vlm-architecture…★ 1.3kUForm★ 1.2kMindPipe★ 1kTinyLLaVA_Factory★ 995PaddleMIX★ 724Vision-Language-Models-O…★ 682awesome-foundation-and-m…★ 638LLaVA-Mini★ 574kolosal-cli★ 465RLAIF-V★ 458Kolosal★ 455Open-LLaVA-NeXT★ 439Olympus★ 428lmms-finetune★ 373HallusionBench★ 342ViP-LLaVA★ 339

Linien verbinden Mitglieder, die messbar miteinander verwandt sind. Die Punktgröße spiegelt die Sterne wider.

🧬 Mitglieder
LLaVA
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25k
SUPIR
SUPIR aims at developing Practical Algorithms for Photo-Realistic Image Restoration In the Wild. Our new…
★ 5.6k
mlx-vlm
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
★ 5.3k
VLMEvalKit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3k
zero_nlp
中文nlp解决方案(大模型、数据、模型、训练、推理)
★ 3.8k
LLamaSharp
A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.
★ 3.8k
Eagle
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
★ 3.3k
OmAgent
[EMNLP-2024] Build multimodal language agents for fast prototype and production
★ 2.7k
Video-ChatGPT
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation…
★ 1.5k
taggui
Tag manager and captioner for image datasets
★ 1.3k
awesome-vlm-architectures
Famous Vision Language Models and Their Architectures
★ 1.3k
UForm
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and…
★ 1.2k
MindPipe
A powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
★ 1k
TinyLLaVA_Factory
A Framework of Small-scale Large Multimodal Models
★ 995
PaddleMIX
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end…
★ 724
Vision-Language-Models-Overview
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository.…
★ 682
awesome-foundation-and-multimodal-models
👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples…
★ 638
LLaVA-Mini
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images,…
★ 574
kolosal-cli
Super lightweight Ollama + Qwen Code alternative to run Llama 3.3, DeepSeek-R1, Phi-4, Gemma 3, Mistral Small…
★ 465
RLAIF-V
[CVPR'25 highlight] RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
★ 458
Kolosal
Kolosal AI is an OpenSource and Lightweight alternative to LM Studio to run LLMs 100% offline on your device.
★ 455
Open-LLaVA-NeXT
An open-source implementation for training LLaVA-NeXT.
★ 439
Olympus
[CVPR 2025 Highlight] Official code for "Olympus: A Universal Task Router for Computer Vision Tasks"
★ 428
lmms-finetune
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave,…
★ 373
HallusionBench
[CVPR'24] HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning…
★ 342
ViP-LLaVA
[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
★ 339
🔗 Verwandte Familien

Gemessen anhand der von beiden Projekten geteilten GitHub-Themen, gewichtet nach der Seltenheit jedes Themas.