What is DINO-X and its core capabilities
mainDINO-X is a unified vision model designed for open-world object detection and understanding. It serves as a general object-centric vision model that supports multiple input types and outputs various semantic representations.
Input Modalities
- Text prompts: Describe objects via natural language.
- Visual prompts: Use visual cues for grounding.
- Customized prompts: Specialized input formats.
- Prompt-Free: Supports "Prompt-Free Anything Detection and Recognition" using a universal object prompt.
Output Representations
- Bounding boxes for object detection.
- Segmentation masks for instance segmentation.
- Pose keypoints for pose estimation.
- Object captions for region captioning.
Supported Tasks
- Open-Set Object Detection and Segmentation
- Phrase Grounding
- Visual-Prompt Counting
- Pose Estimation
- Region Captioning