vlm

69 progetti condividono questo topic GitHub

vlm — transformers ★163.1kvlmUI-TARS-desktop — ★38.4ksglang — ★30.9krunanywhere-sdks — ★10.3knotebooks — ★9.6kanomaly-detection-resources — ★9.4kGenieX — ★8.3kPixelRAG — ★7.9kERNIE — ★7.7kVLM-R1 — ★6kUltraRAG — ★5.7kLLM-RL-Visualized — ★4.7kstar-vector — ★4.5klmms-eval — ★4.3kPromptEnhancer — ★3.7kMiniMax-01 — ★3.4kLocal-File-Organizer — ★3.3kevalscope — ★3.2kSkywork-R1V — ★3.2kOSWorld — ★3kDeepCamera — ★3kOmAgent — ★2.7kCradle — ★2.6kcomfyui_LLM_party — ★2.3kpaperbanana — ★2.2kAwesome-LM-SSP — ★2kQwen-VL-Series-Finetune — ★1.9kawesome-yolo-object-detection — ★1.8kvideo-search-and-summarization — ★1.8ktokenspeed — ★1.8kreact-native-executorch — ★1.7kAwesome-Jailbreak-on-LLMs — ★1.5kAngelSlim — ★1.5kunblink — ★1.5kawesome-vlm-architectures — ★1.3kAeroSandbox — ★1.3kInternNav — ★1kmindnlp — ★920gpt-assistant-android — ★888NEO — ★878UniPic — ★871UI-TARS-desktop★ 38.4ksglang★ 30.9krunanywhere-sdks★ 10.3knotebooks★ 9.6kanomaly-detection-resour…★ 9.4kGenieX★ 8.3kPixelRAG★ 7.9kERNIE★ 7.7kVLM-R1★ 6kUltraRAG★ 5.7kLLM-RL-Visualized★ 4.7kstar-vector★ 4.5klmms-eval★ 4.3kPromptEnhancer★ 3.7kMiniMax-01★ 3.4kLocal-File-Organizer★ 3.3kevalscope★ 3.2kSkywork-R1V★ 3.2kOSWorld★ 3kDeepCamera★ 3kOmAgent★ 2.7kCradle★ 2.6kcomfyui_LLM_party★ 2.3kpaperbanana★ 2.2kAwesome-LM-SSP★ 2kQwen-VL-Series-Finetune★ 1.9kawesome-yolo-object-dete…★ 1.8kvideo-search-and-summari…★ 1.8ktokenspeed★ 1.8kreact-native-executorch★ 1.7kAwesome-Jailbreak-on-LLM…★ 1.5kAngelSlim★ 1.5kunblink★ 1.5kawesome-vlm-architecture…★ 1.3kAeroSandbox★ 1.3kInternNav★ 1kmindnlp★ 920gpt-assistant-android★ 888NEO★ 878UniPic★ 871

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text,…
★ 163.1k
UI-TARS-desktop
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
★ 38.4k
sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 30.9k
runanywhere-sdks
Production ready toolkit to run AI locally
★ 10.3k
notebooks
A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from…
★ 9.6k
anomaly-detection-resources
Anomaly detection related books, papers, videos, and toolboxes. Last update late 2025 for LLM and VLM works!
★ 9.4k
GenieX
Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
★ 8.3k
PixelRAG
The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
★ 7.9k
ERNIE
The official repository for ERNIE 4.5 and ERNIEKit – its industrial-grade development toolkit based on…
★ 7.7k
VLM-R1
Solve Visual Understanding with Reinforced VLMs
★ 6k
UltraRAG
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
★ 5.7k
LLM-RL-Visualized
🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm…
★ 4.7k
star-vector
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation…
★ 4.5k
lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3k
PromptEnhancer
[CVPR 2026] PromptEnhancer is a prompt-rewriting tool, refining prompts into clearer, structured versions for…
★ 3.7k
MiniMax-01
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on…
★ 3.4k
Local-File-Organizer
An AI-powered file management tool that ensures privacy by organizing local texts, images. Using Llama3.2 3B…
★ 3.3k
evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and…
★ 3.2k
Skywork-R1V
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in…
★ 3.2k
OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
★ 3k
DeepCamera
Open-Source AI Camera Skills Platform, AI NVR & CCTV Surveillance. Local VLM video analysis with Qwen,…
★ 3k
OmAgent
[EMNLP-2024] Build multimodal language agents for fast prototype and production
★ 2.7k
Cradle
The Cradle framework is a first attempt at General Computer Control (GCC). Cradle supports agents to ace any…
★ 2.6k
comfyui_LLM_party
LLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt…
★ 2.3k
paperbanana
Open source implementation and extension of Google Research’s PaperBanana for automated academic figures,…
★ 2.2k
Awesome-LM-SSP
A reading list for large models safety, security, and privacy (including Awesome LLM Security, Safety, etc.).
★ 2k
Qwen-VL-Series-Finetune
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1.9k
awesome-yolo-object-detection
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object…
★ 1.8k
video-search-and-summarization
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for…
★ 1.8k
tokenspeed
TokenSpeed is a speed-of-light LLM inference engine.
★ 1.8k
react-native-executorch
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
★ 1.7k
Awesome-Jailbreak-on-LLMs
Awesome-Jailbreak-on-LLMs is a collection of state-of-the-art, novel, exciting jailbreak methods on LLMs. It…
★ 1.5k
AngelSlim
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
★ 1.5k
unblink
Camera monitoring with VLM
★ 1.5k
awesome-vlm-architectures
Famous Vision Language Models and Their Architectures
★ 1.3k
AeroSandbox
Aircraft design optimization made fast through computational graph transformations (e.g., automatic…
★ 1.3k
InternNav
InternRobotics' open platform for building generalized navigation foundation models.
★ 1k
mindnlp
MindSpore + 🤗Huggingface: Run any Transformers/Diffusers model on MindSpore with seamless compatibility…
★ 920
gpt-assistant-android
★ 888
NEO
NEO Series: Native Vision-Language Models from First Principles
★ 878
UniPic
Open-source SOTA multi-image editing model
★ 871
OpenWorldLib
Unified Codebase for Advanced World Models.
★ 847
Awesome-Robotics-3D
A curated list of 3D Vision papers relating to Robotics domain in the era of large models i.e. LLMs/VLMs,…
★ 819
awesome-llm-and-aigc
🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language…
★ 811
Automodel
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
★ 771
mobilegym
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research ·…
★ 741
OmniLottie
[CVPR 2026🔥] 🧑‍🎨 OmniLottie, an open-sourced multi-modal instructed vector animation generator…
★ 730
dingo
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
★ 730
SparkVSR
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
★ 692
VLM2Vec
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and…
★ 669
optimum-intel
🤗 Optimum Intel: Accelerate inference with Intel optimization tools
★ 609
FluxVLA
An all-in-one VLA engineering platform for embodied AI — from data to real-robot deployment.
★ 572
Flame-Code-VLM
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React…
★ 561
rai
RAI is a vendor agnostic agentic framework for Physical AI robotics, utilizing ROS 2 tools to perform complex…
★ 559
vlmrun-hub
A hub for various industry-specific schemas to be used with VLMs.
★ 554
Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 550
Fabric
Node Creative Coding / 3D / Image Processing tool inspired by Quartz Composer
★ 534
Awesome-Embodied-AI
A curated list of awesome papers on Embodied AI and related research/industry-driven resources.
★ 527
Awesome-Multimodal-Modeling
Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
★ 508
vla0
VLA-0: Building State-of-the-Art VLAs with Zero Modification
★ 489
awesome-ai-persona-skills
全网最全的、持续更新的、最火爆的 100+ 人格 蒸馏skills 合集| 多agent系统…
★ 475
awesome-vla-for-ad
🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
★ 453
simlingo
[CVPR 2025, Spotlight] SimLingo (CarLLava): Vision-Only Closed-Loop Autonomous Driving with Language-Action…
★ 442
JarvisEvo
[CVPR' 2026] JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator…
★ 420
LightRFT
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
★ 404
EVE
EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 376
mega-data-factory
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
★ 370
ICLR2026-Guide-CN
不想啃 5000+ 全文?我已经替你和 LLM 啃完了 — ICLR 2026 全景中文导读
★ 155
DataClaw0
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset &…
★ 117
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.