groundingLMM
Observation pool · not formally ranked
This project does not currently meet the main index requirement of two current dimensions across two source types.
About
Grounding Large Multimodal Model (GLaMM) is an end-to-end trained LMM which provides visual grounding capabilities with the flexibility to process both image and region inputs. This enables the new unified task of Grounded Conversation Generation that combines phrase grounding, referring expression segmentation, and vision-language conversations. Equipped with the capability for detailed region understanding, pixel-level groundings, and conversational abilities, GLaMM offers a versatile capability to interact with visual inputs provided by the user at multiple granularity levels.
Delve into the core of GLaMM with our detailed guides on the model's Training and Evaluation methodologies.…
Across sources
- Stars 968
- Forks 56
- Commits 48
- Releases 0