Unified Open-World Segmentation with Multi-Modal Prompts

Editor
1 Min Read


Recent years have witnessed the rapid development of open-world image segmentation, including open-vocabulary segmentation and in-context segmentation. Nonetheless, existing methods are limited to a single modality prompt, which lacks the flexibility and accuracy needed for complex object-aware prompting. In this work, we present COSINE, a unified open-world segmentation model that Consolidates Open-vocabulary Segmentation and IN-context sEgmentation. By framing open-vocabulary task and in-context segmentation task as promptable segmentation tasks, COSINE supports diverse modalities of input, such as images and text. Containing a model pool and a segdecoder, COSINE makes full use of the representation capability of foundations models and is able to accurately segment specific concept based on diverse modalities of input, such as images and text, offering powerful open-world perception capabilities. Experiments on various segmentation tasks show the effectiveness of the proposed method.

Share this Article
Please enter CoinGecko Free Api Key to get this plugin works.