Home Projects OmniVinci
OmniVinci
Python

OmniVinci

OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.

by NVlabs · GitHub
Stars
Forks
Trending
License
Created
Last commit
Language
audio-language-modeldeep-learninglarge-language-modelsApache-2.0Python
View on GitHub
In plain words

Understand and analyze information from images, audio, and text together using a multi-modal AI model.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
OmniVinci — GitHub preview card
📈 Star history
677674
2026-07-202026-08-31
📈 Track OmniVinci

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.

OmniVinci has 677 stars on GitHub. It has been forked 56 times. OmniVinci is written mainly in Python. It has been in active development since 2025. OmniVinci is available under the Apache-2.0 license. Its main topics are audio-language-model, deep-learning, large-language-models, multimodal-large-language-models.

Frequently asked questions

What is OmniVinci?

OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.

Is OmniVinci open source?

OmniVinci is an open-source project. It is released under the Apache-2.0 license.

Is OmniVinci free?

Yes. OmniVinci is free and open source — you can use, modify and self-host it.

What license does OmniVinci use?

OmniVinci is available under the Apache-2.0 license.

What language is OmniVinci written in?

OmniVinci is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — OmniVinci

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=nvlabs-omnivinci)](https://olud.ai/project/nvlabs-omnivinci.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →
Qwen-VL
The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed b…
6.7k · large-language-models
sharesvision-language-model
t2v_metrics
Evaluating text-to-image/video/3D models with VQAScore
600 · generative-ai
sharesvision-language-model
SoundMind
We introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotate…
1.1k · audio-language-model
sharesaudio-language-model
Fun-ASR
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, w…
1.4k · 31-languages
sharesaudio-language-model
RoboFlamingo
Code for RoboFlamingo
437 · artificial-intelligence
sharesvision-language-model
Awesome-Medical-Large-Language-Models
Curated papers on Large Language Models in Healthcare and Medical domain
390 · large-language-models
sharesmultimodal-large-language-models
Awesome-Multimodal-LLM-Autonomous-Driving
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
311 · autonomous-car
sharesmultimodal-large-language-modelsvision-language-model
Spatial-MLLM
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in…
477 · aigc
sharesmultimodal-large-language-models
MGM
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
3.3k · generation
sharesvision-language-model
Ovis-U1
An unified model that seamlessly integrates multimodal understanding, text-to-image generation,…
450 · image-editing
sharesmultimodal-large-language-models

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.