Overview of the Eagle VLM Family
mainEagle is a family of frontier vision-language models (VLMs) from NVIDIA designed for multimodal understanding, long-context reasoning, and embodied applications. The family includes several specialized models:
- LocateAnything: A generalist grounding model focused on detection and pointing using Parallel Box Decoding.
- Eagle 2.5: A frontier VLM optimized for state-of-the-art image and video understanding, utilizing specific frameworks and data strategies for long-context multimodal reasoning.
- Eagle 2: A frontier VLM focused on state-of-the-art image understanding and exploring post-training data strategies.
- Eagle: A VLM architecture utilizing a mixture-of-encoders to explore the design space for vision-centric models.