Home Projects EVF-SAM
EVF-SAM
Python

EVF-SAM

Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"

by hustvl · GitHub
Stars
Forks
License
Created
Last commit
Category
Language
multimodalmultimodal-large-language-modelsreferring-image-segmentationApache-2.0Python
View on GitHub
In plain words

Use this code to enhance image segmentation based on text prompts.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
EVF-SAM — GitHub preview card
📈 Star history
506505
2026-07-202026-08-31
📈 Track EVF-SAM

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"

EVF-SAM has 505 stars on GitHub. It has been forked 26 times. EVF-SAM is written mainly in Python. It has been in active development since 2024. EVF-SAM is available under the Apache-2.0 license. Its main topics are multimodal, multimodal-large-language-models, referring-image-segmentation, segment-anything.

Frequently asked questions

What is EVF-SAM?

Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"

Is EVF-SAM open source?

EVF-SAM is an open-source project. It is released under the Apache-2.0 license.

Is EVF-SAM free?

Yes. EVF-SAM is free and open source — you can use, modify and self-host it.

What license does EVF-SAM use?

EVF-SAM is available under the Apache-2.0 license.

What language is EVF-SAM written in?

EVF-SAM is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — EVF-SAM

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=hustvl-evf-sam)](https://olud.ai/project/hustvl-evf-sam.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →
Spatial-MLLM
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in…
477 · aigc
sharesmultimodal-large-language-modelsmultimodal
Ovis
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align vis…
1.5k · chatbot
sharesmultimodal-large-language-modelsmultimodal
VisionReasoner
[ICLR 2026] VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Le…
348 · counting-objects
sharesmultimodal-large-language-modelsmultimodal
3D-Box-Segment-Anything
We extend Segment Anything to 3D perception by combining it with VoxelNeXt.
564 · 3d
sharessegment-anything
finetune-anything
Fine-tune SAM (Segment Anything Model) for computer vision tasks such as semantic segmentation,…
866 · computer-vision
sharessegment-anything
SegAnyGAussians
The official implementation of Segment Any 3D GAussians (AAAI-25)
986 · 3d-segmentation
sharessegment-anything
sd-webui-segment-anything
Segment Anything for Stable Diffusion WebUI
3.5k · segment-anything
sharessegment-anything
Awesome-Medical-Large-Language-Models
Curated papers on Large Language Models in Healthcare and Medical domain
390 · large-language-models
sharesmultimodal-large-language-models
Ovis-U1
An unified model that seamlessly integrates multimodal understanding, text-to-image generation,…
450 · image-editing
sharesmultimodal-large-language-models
NEO
NEO Series: Native Vision-Language Models from First Principles
888 · agi
sharesmultimodal-large-language-modelsmultimodal

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.