multi-modal

43 projets partagent ce topic GitHub

multi-modal — agentscope ★28.4kmulti-modalten-framework — ★11kInternVL — ★10.1kmodelscope — ★9.1kbig-AGI — ★7.1kdata-juicer — ★6.8kCogVLM — ★6.7kChinese-CLIP — ★6kDALLE-pytorch — ★5.6kmarqo — ★5kDeepKE — ★4.5kOmniGen — ★4.3kVLMEvalKit — ★4.3kVisualGLM-6B — ★4.2kLLamaSharp — ★3.8kpy-xiaozhi — ★3.4kdocarray — ★3.1kLISA — ★2.7kRecSysPapers — ★2.2kMotionGPT — ★1.9kGPTDiscord — ★1.9kSALMONN — ★1.5ktransfusion-pytorch — ★1.4kTransformer-in-Vision — ★1.3kAwesome-Knowledge-Distillation-of-LLMs — ★1.3kVisRAG — ★975byaldi — ★851OmniLottie — ★730Versatile-OCR-Program — ★677forge-film — ★645Aether — ★604RT-2 — ★581Robust-R1 — ★530ACTalker — ★461awesome-vla-for-ad — ★453chat2graph — ★426unisondb — ★419LightRFT — ★404LLMGA — ★395MeshXL — ★339ViP-LLaVA — ★339ten-framework★ 11kInternVL★ 10.1kmodelscope★ 9.1kbig-AGI★ 7.1kdata-juicer★ 6.8kCogVLM★ 6.7kChinese-CLIP★ 6kDALLE-pytorch★ 5.6kmarqo★ 5kDeepKE★ 4.5kOmniGen★ 4.3kVLMEvalKit★ 4.3kVisualGLM-6B★ 4.2kLLamaSharp★ 3.8kpy-xiaozhi★ 3.4kdocarray★ 3.1kLISA★ 2.7kRecSysPapers★ 2.2kMotionGPT★ 1.9kGPTDiscord★ 1.9kSALMONN★ 1.5ktransfusion-pytorch★ 1.4kTransformer-in-Vision★ 1.3kAwesome-Knowledge-Distil…★ 1.3kVisRAG★ 975byaldi★ 851OmniLottie★ 730Versatile-OCR-Program★ 677forge-film★ 645Aether★ 604RT-2★ 581Robust-R1★ 530ACTalker★ 461awesome-vla-for-ad★ 453chat2graph★ 426unisondb★ 419LightRFT★ 404LLMGA★ 395MeshXL★ 339ViP-LLaVA★ 339

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
agentscope
Build and run agents you can see, understand and trust.
★ 28.4k
ten-framework
Open-source framework for conversational voice AI agents
★ 11k
InternVL
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. …
★ 10.1k
modelscope
ModelScope: bring the notion of Model-as-a-Service to life.
★ 9.1k
big-AGI
AI suite powered by state-of-the-art models and providing advanced AI/AGI functions. Includes AI personas,…
★ 7.1k
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6.8k
CogVLM
a state-of-the-art-level open visual language model | 多模态预训练模型
★ 6.7k
Chinese-CLIP
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
★ 6k
DALLE-pytorch
Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
★ 5.6k
marqo
Ecommerce Search and Discovery - marqo.ai
★ 5k
DeepKE
[EMNLP 2022] An Open Toolkit for Knowledge Graph Extraction and Construction
★ 4.5k
OmniGen
OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
★ 4.3k
VLMEvalKit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3k
VisualGLM-6B
Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型
★ 4.2k
LLamaSharp
A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.
★ 3.8k
py-xiaozhi
Open-source AI assistant ecosystem with MCP integrations, multimodal workflows, IoT support, and…
★ 3.4k
docarray
Represent, send, store and search multimodal data
★ 3.1k
LISA
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
★ 2.7k
RecSysPapers
推荐/广告/搜索领域工业界经典以及最前沿论文集合。A collection of industry classics and…
★ 2.2k
MotionGPT
[NeurIPS 2023] MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model…
★ 1.9k
GPTDiscord
A robust, all-in-one GPT interface for Discord. ChatGPT-style conversations, image generation, AI-moderation,…
★ 1.9k
SALMONN
SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5k
transfusion-pytorch
Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal…
★ 1.4k
Transformer-in-Vision
Recent Transformer-based CV and related works.
★ 1.3k
Awesome-Knowledge-Distillation-of-LLMs
This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break…
★ 1.3k
VisRAG
Parsing-free RAG supported by VLMs
★ 975
byaldi
Use late-interaction multi-modal models such as ColPali in just a few lines of code.
★ 851
OmniLottie
[CVPR 2026🔥] 🧑‍🎨 OmniLottie, an open-sourced multi-modal instructed vector animation generator…
★ 730
Versatile-OCR-Program
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
★ 677
forge-film
Multi-model DAG-driven parallel AI film generation — parallel speedup scales with scene independence;…
★ 645
Aether
[ICCV 2025 & ICCV 2025 RIWM Outstanding Paper] Aether: Geometric-Aware Unified World Modeling
★ 604
RT-2
Democratization of RT-2 "RT-2: New model translates vision and language into action"
★ 581
Robust-R1
🔥🔥🔥[AAAI 2026 Oral] Official Implementation of Robust-R1: Degradation-Aware Reasoning for Robust…
★ 530
ACTalker
ICCV 2025 ACTalker: an end-to-end video diffusion framework for talking head synthesis that supports both…
★ 461
awesome-vla-for-ad
🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
★ 453
chat2graph
Chat2Graph: Graph Native Agentic System.
★ 426
unisondb
A streaming multimodal database for Edge AI, and Edge Computing.
★ 419
LightRFT
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
★ 404
LLMGA
This project is the official implementation of 'LLMGA: Multimodal Large Language Model based Generation…
★ 395
MeshXL
[NeurIPS 2024] MeshXL: Neural Coordinate Field for Generative 3D Foundation Models, a 3D fundamental model…
★ 339
ViP-LLaVA
[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
★ 339
LL3DA
[CVPR 2024] "LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and…
★ 319
GenDB
GenDB, an LLM-Powered Generative Query Engine Built for the Future
★ 71 · GitHub ↗
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.