vision

22 proyectos comparten este topic de GitHub

vision — LibreChat ★42.7kvisionUI-TARS-desktop — ★38.8kcaffe — ★34.6kskyvern — ★22.9kPixelRAG — ★9.8kmodlens — ★3.8kSimpleMem — ★3.7kTorch-Pruning — ★3.3kha-llmvision — ★1.4kagent-vision-toolkit — ★1.1kdsh-vision-router — ★1kLLaVA-Mini — ★577GLM-skills — ★469apriltag_ros — ★457Stream-Omni — ★392claude-code-vision-skill — ★171dsh-vision — ★88Awesome-AVI — ★86opencode-senses — ★82cc-VisionRouter — ★78mere-run — ★69dsh-vision-complete — ★42UI-TARS-desktop★ 38.8kcaffe★ 34.6kskyvern★ 22.9kPixelRAG★ 9.8kmodlens★ 3.8kSimpleMem★ 3.7kTorch-Pruning★ 3.3kha-llmvision★ 1.4kagent-vision-toolkit★ 1.1kdsh-vision-router★ 1kLLaVA-Mini★ 577GLM-skills★ 469apriltag_ros★ 457Stream-Omni★ 392claude-code-vision-skill★ 171dsh-vision★ 88Awesome-AVI★ 86opencode-senses★ 82cc-VisionRouter★ 78mere-run★ 69dsh-vision-complete★ 42

Las líneas conectan a los miembros que están mediblemente relacionados entre sí. El tamaño de los puntos refleja las estrellas.

🧬 Miembros
LibreChat
Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure,…
★ 42.7k
UI-TARS-desktop
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
★ 38.8k
caffe
Caffe: a fast open framework for deep learning.
★ 34.6k
skyvern
Automate browser based workflows with AI
★ 22.9k
PixelRAG
https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search.…
★ 9.8k
modlens
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste…
★ 3.8k
SimpleMem
SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
★ 3.7k
Torch-Pruning
[CVPR 2023] DepGraph: Towards Any Structural Pruning; LLMs, Vision Foundation Models, etc.
★ 3.3k
ha-llmvision
Visual intelligence for your home.
★ 1.4k
agent-vision-toolkit
★ 1.1k
dsh-vision-router
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools…
★ 1k
LLaVA-Mini
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images,…
★ 577
GLM-skills
Official skills for the GLM family of models.
★ 469
apriltag_ros
A ROS wrapper of the AprilTag 3 visual fiducial detector
★ 457
Stream-Omni
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across…
★ 392
claude-code-vision-skill
为 Claude Code 赋能多模态视觉能力,适配 纯文本 LLM 底座,用于截图 / UI /…
★ 171
dsh-vision
Near-native image understanding for DeepSeek Harness
★ 88
Awesome-AVI
Awesome Audio-Visual Intelligence, Survey of Audio-Visual Intelligence
★ 86
opencode-senses
The vision plugin for OpenCode that truly understands images. Inspect, read, and reason about any screenshot…
★ 82
cc-VisionRouter
Transparent proxy for Claude Code that auto-routes image-bearing requests to a multimodal model — so a…
★ 78
mere-run
Run local image, text, speech, vision, music, and video workflows, plus model management and a loopback…
★ 69
dsh-vision-complete
给 DeepSeek 补上「眼睛和耳朵」的多模态视觉插件:看图 / OCR / 物体检测 / 视频理解…
★ 42
🔗 Familias relacionadas

Medido a partir de los temas de GitHub compartidos por ambos proyectos, ponderado por cuán raros son cada uno de los temas.