vlm

47 Projekte teilen dieses GitHub-Topic

vlm — transformers ★164.7kvlmUI-TARS-desktop — ★38.8ksglang — ★33krunanywhere-sdks — ★10.3kPixelRAG — ★9.8kanomaly-detection-resources — ★9.4kGenieX — ★8.3kERNIE — ★7.7kVLM-R1 — ★6kUltraRAG — ★5.7kLLM-RL-Visualized — ★4.8kstar-vector — ★4.6klmms-eval — ★4.4kMiniMax-01 — ★3.5kevalscope — ★3.3kLocal-File-Organizer — ★3.3kSkywork-R1V — ★3.2kOSWorld — ★3.1kDeepCamera — ★3kOmAgent — ★2.7kCradle — ★2.6kpaperbanana — ★2.3kAwesome-LM-SSP — ★2.1ktokenspeed — ★2.1kvideo-search-and-summarization — ★1.8kawesome-yolo-object-detection — ★1.8kAwesome-Jailbreak-on-LLMs — ★1.6kAngelSlim — ★1.6kawesome-vlm-architectures — ★1.3kMindAct — ★920mindnlp — ★920NEO — ★888mobilegym — ★778OmniLottie — ★771SparkVSR — ★700awesome-ai-persona-skills — ★583rai — ★576Awesome-Multimodal-Modeling — ★541vla0 — ★488awesome-vla-for-ad — ★468simlingo — ★438UI-TARS-desktop★ 38.8ksglang★ 33krunanywhere-sdks★ 10.3kPixelRAG★ 9.8kanomaly-detection-resour…★ 9.4kGenieX★ 8.3kERNIE★ 7.7kVLM-R1★ 6kUltraRAG★ 5.7kLLM-RL-Visualized★ 4.8kstar-vector★ 4.6klmms-eval★ 4.4kMiniMax-01★ 3.5kevalscope★ 3.3kLocal-File-Organizer★ 3.3kSkywork-R1V★ 3.2kOSWorld★ 3.1kDeepCamera★ 3kOmAgent★ 2.7kCradle★ 2.6kpaperbanana★ 2.3kAwesome-LM-SSP★ 2.1ktokenspeed★ 2.1kvideo-search-and-summari…★ 1.8kawesome-yolo-object-dete…★ 1.8kAwesome-Jailbreak-on-LLM…★ 1.6kAngelSlim★ 1.6kawesome-vlm-architecture…★ 1.3kMindAct★ 920 · GitHub ↗mindnlp★ 920NEO★ 888mobilegym★ 778OmniLottie★ 771SparkVSR★ 700awesome-ai-persona-skill…★ 583rai★ 576Awesome-Multimodal-Model…★ 541vla0★ 488awesome-vla-for-ad★ 468simlingo★ 438

Linien verbinden Mitglieder, die messbar miteinander verwandt sind. Die Punktgröße spiegelt die Sterne wider.

🧬 Mitglieder
transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text,…
★ 164.7k
UI-TARS-desktop
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
★ 38.8k
sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 33k
runanywhere-sdks
Production ready toolkit to run AI locally
★ 10.3k
PixelRAG
https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search.…
★ 9.8k
anomaly-detection-resources
Anomaly detection related books, papers, videos, and toolboxes. Last update late 2025 for LLM and VLM works!
★ 9.4k
GenieX
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
★ 8.3k
ERNIE
The official repository for ERNIE 4.5 and ERNIEKit – its industrial-grade development toolkit based on…
★ 7.7k
VLM-R1
Solve Visual Understanding with Reinforced VLMs
★ 6k
UltraRAG
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
★ 5.7k
LLM-RL-Visualized
🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm…
★ 4.8k
star-vector
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation…
★ 4.6k
lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.4k
MiniMax-01
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on…
★ 3.5k
evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and…
★ 3.3k
Local-File-Organizer
An AI-powered file management tool that ensures privacy by organizing local texts, images. Using Llama3.2 3B…
★ 3.3k
Skywork-R1V
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in…
★ 3.2k
OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
★ 3.1k
DeepCamera
Open-Source AI Camera Skills Platform, AI NVR & CCTV Surveillance. Local VLM video analysis with Qwen,…
★ 3k
OmAgent
[EMNLP-2024] Build multimodal language agents for fast prototype and production
★ 2.7k
Cradle
The Cradle framework is a first attempt at General Computer Control (GCC). Cradle supports agents to ace any…
★ 2.6k
paperbanana
Open source implementation and extension of Google Research’s PaperBanana for automated academic figures,…
★ 2.3k
Awesome-LM-SSP
A reading list for large models safety, security, and privacy (including Awesome LLM Security, Safety, etc.).
★ 2.1k
tokenspeed
TokenSpeed is a speed-of-light LLM inference engine.
★ 2.1k
video-search-and-summarization
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for…
★ 1.8k
awesome-yolo-object-detection
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object…
★ 1.8k
Awesome-Jailbreak-on-LLMs
Awesome-Jailbreak-on-LLMs is a collection of state-of-the-art, novel, exciting jailbreak methods on LLMs. It…
★ 1.6k
AngelSlim
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
★ 1.6k
awesome-vlm-architectures
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training…
★ 1.3k
MindAct
MindSpore + 🤗Huggingface: Run any Transformers/Diffusers model on MindSpore with seamless compatibility…
★ 920 · GitHub ↗
mindnlp
MindSpore + 🤗Huggingface: Run any Transformers/Diffusers model on MindSpore with seamless compatibility…
★ 920
NEO
NEO Series: Native Vision-Language Models from First Principles
★ 888
mobilegym
[EMNLP 2026] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research ·…
★ 778
OmniLottie
[CVPR 2026🔥] 🧑‍🎨 OmniLottie, an open-sourced multi-modal instructed vector animation generator…
★ 771
SparkVSR
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
★ 700
awesome-ai-persona-skills
全网最全的、持续更新的、最火爆的 100+ 人格 蒸馏skills 合集| 多agent系统…
★ 583
rai
RAI is a vendor agnostic agentic framework for Physical AI robotics, utilizing ROS 2 tools to perform complex…
★ 576
Awesome-Multimodal-Modeling
Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
★ 541
vla0
VLA-0: Building State-of-the-Art VLAs with Zero Modification
★ 488
awesome-vla-for-ad
🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
★ 468
simlingo
[CVPR 2025, Spotlight] SimLingo (CarLLava): Vision-Only Closed-Loop Autonomous Driving with Language-Action…
★ 438
EVE
EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 376
mega-data-factory
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
★ 372
ICLR2026-Guide-CN
不想啃 5000+ 全文?我已经替你和 LLM 啃完了 — ICLR 2026 全景中文导读
★ 168
DataClaw0
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset &…
★ 117
MinerU-Skill
AI-Native document parser: PDF, Office & images → clean Markdown with LaTeX, tables & OCR. Zero-dependency…
★ 108
VLMForge
Evaluate visual models on your own images, JSON Schema, and production constraints with LangGraph, Pareto…
★ 89
🔗 Verwandte Familien

Gemessen anhand der von beiden Projekten geteilten GitHub-Themen, gewichtet nach der Seltenheit jedes Themas.