vision-transformer

39 progetti condividono questo topic GitHub

vision-transformer — LaTeX-OCR ★16.5kvision-transformerTransformers-Tutorials — ★11.7kVAR — ★8.7kmlx-vlm — ★5.3kAwesome-Transformer-Attention — ★5.1kmmpretrain — ★3.8kscenic — ★3.8ktowhee — ★3.5kInternLM-XComposer — ★2.9kInternVideo — ★2.3kMambaVision — ★2.2kViTPose — ★2.1kTransformer-Explainability — ★2kEasyCV — ★1.9klightly-train — ★1.6kthepipe — ★1.5kawesome-attention-mechanism-in-cv — ★1.3kVoxFormer — ★1.2kAwesome-Foundation-Models — ★1.2kONE-PEACE — ★1.1kAwesome-CV-MasterHub — ★963YOLOS — ★903parseq — ★727unicom — ★700swin2sr — ★690RAG-Driven-Generative-AI — ★617eomt — ★615actionformer_release — ★572GFNet — ★511Transformers-for-NLP-and-Computer-Vision-3rd-Edition — ★499CLIP_Surgery — ★482logit-standardization-KD — ★385NanoLLM — ★381LFM — ★357Awesome-MIM — ★354v2x-vit — ★347Awesome-Multimodal-LLM-Autonomous-Driving — ★313pytorch-vit — ★309ImageFolder — ★307Transformers-Tutorials★ 11.7kVAR★ 8.7kmlx-vlm★ 5.3kAwesome-Transformer-Atte…★ 5.1kmmpretrain★ 3.8kscenic★ 3.8ktowhee★ 3.5kInternLM-XComposer★ 2.9kInternVideo★ 2.3kMambaVision★ 2.2kViTPose★ 2.1kTransformer-Explainabili…★ 2kEasyCV★ 1.9klightly-train★ 1.6kthepipe★ 1.5kawesome-attention-mechan…★ 1.3kVoxFormer★ 1.2kAwesome-Foundation-Model…★ 1.2kONE-PEACE★ 1.1kAwesome-CV-MasterHub★ 963YOLOS★ 903parseq★ 727unicom★ 700swin2sr★ 690RAG-Driven-Generative-AI★ 617eomt★ 615actionformer_release★ 572GFNet★ 511Transformers-for-NLP-and…★ 499CLIP_Surgery★ 482logit-standardization-KD★ 385NanoLLM★ 381LFM★ 357Awesome-MIM★ 354v2x-vit★ 347Awesome-Multimodal-LLM-A…★ 313pytorch-vit★ 309ImageFolder★ 307

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
LaTeX-OCR
pix2tex: Using a ViT to convert images of equations into LaTeX code.
★ 16.5k
Transformers-Tutorials
This repository contains demos I made with the Transformers library by HuggingFace.
★ 11.7k
VAR
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official…
★ 8.7k
mlx-vlm
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
★ 5.3k
Awesome-Transformer-Attention
An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related…
★ 5.1k
mmpretrain
OpenMMLab Pre-training Toolbox and Benchmark
★ 3.8k
scenic
Scenic: A Jax Library for Computer Vision Research and Beyond
★ 3.8k
towhee
Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.
★ 3.5k
InternLM-XComposer
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio…
★ 2.9k
InternVideo
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2.3k
MambaVision
[CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
★ 2.2k
ViTPose
The official repo for [NeurIPS'22] "ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation"…
★ 2.1k
Transformer-Explainability
[CVPR 2021] Official PyTorch implementation for Transformer Interpretability Beyond Attention Visualization,…
★ 2k
EasyCV
An all-in-one toolkit for computer vision
★ 1.9k
lightly-train
All-in-one training for vision models (YOLO, ViTs, RT-DETR, DINOv3): pretraining, fine-tuning, distillation.
★ 1.6k
thepipe
Get clean data from tricky documents, powered by vision-language models ⚡
★ 1.5k
awesome-attention-mechanism-in-cv
Awesome List of Attention Modules and Plug&Play Modules in Computer Vision
★ 1.3k
VoxFormer
Official PyTorch implementation of VoxFormer [CVPR 2023 Highlight]
★ 1.2k
Awesome-Foundation-Models
A curated list of foundation models for vision and language tasks
★ 1.2k
ONE-PEACE
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One…
★ 1.1k
Awesome-CV-MasterHub
:fire: :fire: :fire: A paper list of some recent Computer Vision(CV) works
★ 963
YOLOS
[NeurIPS 2021] You Only Look at One Sequence
★ 903
parseq
Scene Text Recognition with Permuted Autoregressive Sequence Models (ECCV 2022)
★ 727
unicom
Large-Scale Visual Representation Model
★ 700
swin2sr
[ECCV] Swin2SR: SwinV2 Transformer for Compressed Image Super-Resolution and Restoration. Advances in Image…
★ 690
RAG-Driven-Generative-AI
This repository provides programs to build Retrieval Augmented Generation (RAG) code for Generative AI with…
★ 617
eomt
[CVPR 2025 Highlight] Official code and models for Encoder-only Mask Transformer (EoMT).
★ 615
actionformer_release
Code release for ActionFormer (ECCV 2022)
★ 572
GFNet
[NeurIPS 2021] [T-PAMI] Global Filter Networks for Image Classification
★ 511
Transformers-for-NLP-and-Computer-Vision-3rd-Edition
Transformers 3rd Edition
★ 499
CLIP_Surgery
[Pattern Recognition 25] CLIP Surgery for Better Explainability with Enhancement in Open-Vocabulary Tasks
★ 482
logit-standardization-KD
[CVPR 2024 Highlight] Logit Standardization in Knowledge Distillation
★ 385
NanoLLM
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models,…
★ 381
LFM
Official PyTorch implementation of the paper: Flow Matching in Latent Space
★ 357
Awesome-MIM
[Survey] Masked Modeling for Self-supervised Representation Learning on Vision and Beyond…
★ 354
v2x-vit
[ECCV2022] Official Implementation of paper "V2X-ViT: Vehicle-to-Everything Cooperative Perception with…
★ 347
Awesome-Multimodal-LLM-Autonomous-Driving
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
★ 313
pytorch-vit
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
★ 309
ImageFolder
High-performance Image Tokenizers for VAR and AR
★ 307
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.