[EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Use a language model to understand and analyze video and audio content.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytime[EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Video-LLaMA has 3.1k stars on GitHub. It has been forked 287 times. Video-LLaMA is written mainly in Python. It has been in active development since 2023. Video-LLaMA is available under the BSD-3-Clause license. Its main topics are blip2, cross-modal-pretraining, large-language-models, llama.
[EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Video-LLaMA is an open-source project. It is released under the BSD-3-Clause license.
Yes. Video-LLaMA is free and open source — you can use, modify and self-host it.
Video-LLaMA is available under the BSD-3-Clause license.
Video-LLaMA is written mainly in Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/damo-nlp-sg-video-llama.html)