video-understanding

18 projets partagent ce topic GitHub

video-understanding — mmaction2 ★5.1kvideo-understandinglmms-eval — ★4.3kAsk-Anything — ★3.3kInternVideo — ★2.3kvideo-search-and-summarization — ★1.8kVideoAgent — ★1.6kSALMONN — ★1.5kvllm-mlx — ★1.5kLance — ★1.3kawesome-grounding — ★1.1kAwesome-CV-MasterHub — ★963Chat-UniVi — ★943telemem — ★475video-understanding-dataset — ★469Video-RAG-master — ★448watch-skill — ★241OmniAgent — ★60ThinkJEPA — ★48lmms-eval★ 4.3kAsk-Anything★ 3.3kInternVideo★ 2.3kvideo-search-and-summari…★ 1.8kVideoAgent★ 1.6kSALMONN★ 1.5kvllm-mlx★ 1.5kLance★ 1.3kawesome-grounding★ 1.1kAwesome-CV-MasterHub★ 963Chat-UniVi★ 943telemem★ 475video-understanding-data…★ 469Video-RAG-master★ 448watch-skill★ 241OmniAgent★ 60 · GitHub ↗ThinkJEPA★ 48 · GitHub ↗

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
mmaction2
OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark
★ 5.1k
lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3k
Ask-Anything
[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as…
★ 3.3k
InternVideo
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2.3k
video-search-and-summarization
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for…
★ 1.8k
VideoAgent
"VideoAgent: All-in-One Agentic Framework for Video Understanding, Editing, and Remaking"
★ 1.6k
SALMONN
SALMONN family: A suite of advanced multi-modal LLMs
★ 1.5k
vllm-mlx
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama,…
★ 1.5k
Lance
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and…
★ 1.3k
awesome-grounding
awesome grounding: A curated list of research papers in visual grounding
★ 1.1k
Awesome-CV-MasterHub
:fire: :fire: :fire: A paper list of some recent Computer Vision(CV) works
★ 963
Chat-UniVi
[CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image…
★ 943
telemem
TeleMem is a high-performance drop-in replacement for Mem0, featuring semantic deduplication, long-term…
★ 475
video-understanding-dataset
A collection of recent video understanding datasets, under construction!
★ 469
Video-RAG-master
✨✨[NeurIPS 2025] This is the official implementation of our paper "Video-RAG: Visually-aligned…
★ 448
watch-skill
Video understanding and self-verification for AI agents. Turn videos, streams, and agent screen recordings…
★ 241
OmniAgent
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that…
★ 60 · GitHub ↗
ThinkJEPA
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
★ 48 · GitHub ↗
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.