clip

41 proyectos comparten este topic de GitHub

clip — X-AnyLabeling ★9.9kclipChinese-CLIP — ★6kVLMEvalKit — ★4.3kmmpretrain — ★3.8kzero_nlp — ★3.8kVLM_survey — ★3.1kclip-retrieval — ★2.8khcaptcha-challenger — ★2.4kcambrian — ★2kawesome-openai-vision-api-experiments — ★1.7kVideo-ChatGPT — ★1.5kawesome-vlm-architectures — ★1.3kUForm — ★1.2kvlms-zero-to-hero — ★1.2kStable-Diffusion-NCNN — ★1.1knatural-language-image-search — ★1kCLIP4Clip — ★1knatural-language-youtube-search — ★935Transformer-MM-Explainability — ★911aphantasia — ★789PaddleMIX — ★724Vision-Language-Models-Overview — ★682SkyPaint-AI-Diffusion — ★648awesome-foundation-and-multimodal-models — ★638keras_cv_attention_models — ★627tidy — ★586clip.cpp — ★563cliport — ★547PicQuery — ★506Transformers-for-NLP-and-Computer-Vision-3rd-Edition — ★499diffusion-explainer — ★486OCRAutoScore — ★483CLIP_Surgery — ★482EVE — ★376Instruct2Act — ★374ViralCutter — ★358GenSim — ★350ViP-LLaVA — ★339FiT3D — ★329Disco_Diffusion_Local — ★316LodeDB — ★89Chinese-CLIP★ 6kVLMEvalKit★ 4.3kmmpretrain★ 3.8kzero_nlp★ 3.8kVLM_survey★ 3.1kclip-retrieval★ 2.8khcaptcha-challenger★ 2.4kcambrian★ 2kawesome-openai-vision-ap…★ 1.7kVideo-ChatGPT★ 1.5kawesome-vlm-architecture…★ 1.3kUForm★ 1.2kvlms-zero-to-hero★ 1.2kStable-Diffusion-NCNN★ 1.1knatural-language-image-s…★ 1kCLIP4Clip★ 1knatural-language-youtube…★ 935Transformer-MM-Explainab…★ 911aphantasia★ 789PaddleMIX★ 724Vision-Language-Models-O…★ 682SkyPaint-AI-Diffusion★ 648awesome-foundation-and-m…★ 638keras_cv_attention_model…★ 627tidy★ 586clip.cpp★ 563cliport★ 547PicQuery★ 506Transformers-for-NLP-and…★ 499diffusion-explainer★ 486OCRAutoScore★ 483CLIP_Surgery★ 482EVE★ 376Instruct2Act★ 374ViralCutter★ 358GenSim★ 350ViP-LLaVA★ 339FiT3D★ 329Disco_Diffusion_Local★ 316LodeDB★ 89

Las líneas conectan a los miembros que están mediblemente relacionados entre sí. El tamaño de los puntos refleja las estrellas.

🧬 Miembros
X-AnyLabeling
Open-source AI-assisted annotation platform for images, videos, text, and multimodal data.
★ 9.9k
Chinese-CLIP
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
★ 6k
VLMEvalKit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3k
mmpretrain
OpenMMLab Pre-training Toolbox and Benchmark
★ 3.8k
zero_nlp
中文nlp解决方案(大模型、数据、模型、训练、推理)
★ 3.8k
VLM_survey
Collection of AWESOME vision-language models for vision tasks
★ 3.1k
clip-retrieval
Easily compute clip embeddings and build a clip retrieval system with them
★ 2.8k
hcaptcha-challenger
🥂 Gracefully face hCaptcha challenge with multimodal large language model.
★ 2.4k
cambrian
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2k
awesome-openai-vision-api-experiments
Must-have resource for anyone who wants to experiment with and build on the OpenAI vision API 🔥
★ 1.7k
Video-ChatGPT
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation…
★ 1.5k
awesome-vlm-architectures
Famous Vision Language Models and Their Architectures
★ 1.3k
UForm
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and…
★ 1.2k
vlms-zero-to-hero
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge…
★ 1.2k
Stable-Diffusion-NCNN
Stable Diffusion in NCNN with c++, supported txt2img and img2img
★ 1.1k
natural-language-image-search
Search photos on Unsplash using natural language
★ 1k
CLIP4Clip
An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
★ 1k
natural-language-youtube-search
Search inside YouTube videos using natural language
★ 935
Transformer-MM-Explainability
[ICCV 2021- Oral] Official PyTorch implementation for Generic Attention-model Explainability for Interpreting…
★ 911
aphantasia
CLIP + FFT/DWT/RGB = text to image/video
★ 789
PaddleMIX
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end…
★ 724
Vision-Language-Models-Overview
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository.…
★ 682
SkyPaint-AI-Diffusion
★ 648
awesome-foundation-and-multimodal-models
👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples…
★ 638
keras_cv_attention_models
Keras beit,caformer,CMT,CoAtNet,convnext,davit,dino,efficientdet,edgenext,efficientformer,efficientnet,eva,fas…
★ 627
tidy
Offline semantic Text-to-Image and Image-to-Image search on Android powered by quantized state-of-the-art…
★ 586
clip.cpp
CLIP inference in plain C/C++ with no extra dependencies
★ 563
cliport
CLIPort: What and Where Pathways for Robotic Manipulation
★ 547
PicQuery
🔍 Search local images with natural language on Android, powered by OpenAI's CLIP model. / 在 Android…
★ 506
Transformers-for-NLP-and-Computer-Vision-3rd-Edition
Transformers 3rd Edition
★ 499
diffusion-explainer
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
★ 486
OCRAutoScore
OCR自动化阅卷项目
★ 483
CLIP_Surgery
[Pattern Recognition 25] CLIP Surgery for Better Explainability with Enhancement in Open-Vocabulary Tasks
★ 482
EVE
EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 376
Instruct2Act
Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model
★ 374
ViralCutter
Free tool to create viral videos from YouTube, generating clips optimized for TikTok and Instagram with…
★ 358
GenSim
Generating Robotic Simulation Tasks via Large Language Models
★ 350
ViP-LLaVA
[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
★ 339
FiT3D
[ECCV 2024] Improving 2D Feature Representations by 3D-Aware Fine-Tuning
★ 329
Disco_Diffusion_Local
Getting the latest versions of Disco Diffusion to work locally, instead of colab. Including how I run this on…
★ 316
LodeDB
World's fastest and most compact embedded vector database: exact by default, multimodal, local-first, and…
★ 89
🔗 Familias relacionadas

Medido a partir de los temas de GitHub compartidos por ambos proyectos, ponderado por cuán raros son cada uno de los temas.