Overview of the All-Seeing Project V2
mainThe All-Seeing Project V2 is a framework designed for panoptic visual recognition and general relation comprehension. It introduces three core components:
- All-Seeing Dataset V2 (AS-V2): A dataset of 127K high-quality Relation Conversation (ReC) samples. ReC unifies text generation, object localization, and relation comprehension.
- All-Seeing Model v2 (ASMv2): A Multi-modal Large Language Model (MLLM) that integrates grounding and referring capabilities, making it suitable for region-level tasks and open-ended Scene Graph Generation.
- Circular-based Relation Probing Evaluation (CRPE) benchmark: A benchmark for quantitatively evaluating object recognition and relation comprehension by probing the elements of relation triplets
(subject, predicate, object).