vision

36 projetos partilham este topic do GitHub

vision — LibreChat ★41.4kvisionUI-TARS-desktop — ★38.4kcaffe — ★34.6kskyvern — ★22.6kPixelRAG — ★7.9kSimpleMem — ★3.7kTorch-Pruning — ★3.3kSimpleCV — ★2.7kha-llmvision — ★1.4kvisual-pushing-grasping — ★1.1kcalvin — ★964deepdrive — ★926iowncode — ★9113dmatch-toolbox — ★906caer — ★812bottleneck-transformer-pytorch — ★677design2code — ★677myvision — ★610vector-python-sdk — ★603pytorch-dense-correspondence — ★576LLaVA-Mini — ★574cliport — ★547PythonFromSpace — ★471apriltag_ros — ★459GLM-skills — ★452photonvision — ★419potato — ★399Stream-Omni — ★390GRIP — ★387World-Simulator — ★383VectorDB-Plugin — ★369rowfill — ★368ChatGPT-OpenAI-Smart-Speaker — ★318apc-vision-toolbox — ★309cc-VisionRouter — ★94Awesome-AVI — ★84UI-TARS-desktop★ 38.4kcaffe★ 34.6kskyvern★ 22.6kPixelRAG★ 7.9kSimpleMem★ 3.7kTorch-Pruning★ 3.3kSimpleCV★ 2.7kha-llmvision★ 1.4kvisual-pushing-grasping★ 1.1kcalvin★ 964deepdrive★ 926iowncode★ 9113dmatch-toolbox★ 906caer★ 812bottleneck-transformer-p…★ 677design2code★ 677myvision★ 610vector-python-sdk★ 603pytorch-dense-correspond…★ 576LLaVA-Mini★ 574cliport★ 547PythonFromSpace★ 471apriltag_ros★ 459GLM-skills★ 452photonvision★ 419potato★ 399Stream-Omni★ 390GRIP★ 387World-Simulator★ 383VectorDB-Plugin★ 369rowfill★ 368ChatGPT-OpenAI-Smart-Spe…★ 318apc-vision-toolbox★ 309cc-VisionRouter★ 94Awesome-AVI★ 84

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
LibreChat
Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure,…
★ 41.4k
UI-TARS-desktop
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
★ 38.4k
caffe
Caffe: a fast open framework for deep learning.
★ 34.6k
skyvern
Automate browser based workflows with AI
★ 22.6k
PixelRAG
The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
★ 7.9k
SimpleMem
SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
★ 3.7k
Torch-Pruning
[CVPR 2023] DepGraph: Towards Any Structural Pruning; LLMs, Vision Foundation Models, etc.
★ 3.3k
SimpleCV
The Open Source Framework for Machine Vision
★ 2.7k
ha-llmvision
Visual intelligence for your home.
★ 1.4k
visual-pushing-grasping
Train robotic agents to learn to plan pushing and grasping actions for manipulation with deep reinforcement…
★ 1.1k
calvin
CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
★ 964
deepdrive
Deepdrive is a simulator that allows anyone with a PC to push the state-of-the-art in self-driving
★ 926
iowncode
A curated collection of iOS, ML, AR resources sprinkled with some UI additions
★ 911
3dmatch-toolbox
3DMatch - a 3D ConvNet-based local geometric descriptor for aligning 3D meshes and point clouds.
★ 906
caer
High-performance Vision library in Python. Scale your research, not boilerplate.
★ 812
bottleneck-transformer-pytorch
Implementation of Bottleneck Transformer in Pytorch
★ 677
design2code
Convert any web design screenshot to clean HTML/CSS code
★ 677
myvision
Computer vision based ML training data generation tool :rocket:
★ 610
vector-python-sdk
Anki Vector Python SDK
★ 603
pytorch-dense-correspondence
Code for "Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation"
★ 576
LLaVA-Mini
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images,…
★ 574
cliport
CLIPort: What and Where Pathways for Robotic Manipulation
★ 547
PythonFromSpace
Python Examples for Remote Sensing
★ 471
apriltag_ros
A ROS wrapper of the AprilTag 3 visual fiducial detector
★ 459
GLM-skills
Official skills for the GLM family of models.
★ 452
photonvision
PhotonVision is the free, fast, and easy-to-use computer vision solution for the FIRST Robotics Competition.
★ 419
potato
potato: the portable annotation tool
★ 399
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 390
GRIP
Program for rapidly developing computer vision applications
★ 387
World-Simulator
[IEEE TPAMI 2026] Simulating the Real World: Survey & Resources, which contains our survey "Simulating the…
★ 383
VectorDB-Plugin
Program that lets you ask questions about your documents, audio, and video files.
★ 369
rowfill
Open-source spreadsheets platform for deep research and document processing
★ 368
ChatGPT-OpenAI-Smart-Speaker
This AI Smart Speaker uses speech recognition, TTS (text-to-speech), and STT (speech-to-text) to enable voice…
★ 318
apc-vision-toolbox
MIT-Princeton Vision Toolbox for the Amazon Picking Challenge 2016 - RGB-D ConvNet-based object segmentation…
★ 309
cc-VisionRouter
Transparent proxy for Claude Code that auto-routes image-bearing requests to a multimodal model — so a…
★ 94
Awesome-AVI
Awesome Audio-Visual Intelligence, Survey of Audio-Visual Intelligence
★ 84
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.