aiagent.club
English
GitHub

groundingLMM

观察池 · 暂无正式排名

该项目当前未满足“两个有效维度 + 两种数据源”的主榜门槛。

未排名
可观测采用度缺失
动量当前有效 · 2026-09-12
0
关注度当前有效 · 2026-09-12
0
信号可信度 依据当前有数据的独立评分维度数量计算。
2/3 · 1 种数据源

方法论 v2.0 · 快照 2026-09-12 · 超过 2 天视为过期

项目介绍

Grounding Large Multimodal Model (GLaMM) is an end-to-end trained LMM which provides visual grounding capabilities with the flexibility to process both image and region inputs. This enables the new unified task of Grounded Conversation Generation that combines phrase grounding, referring expression segmentation, and vision-language conversations. Equipped with the capability for detailed region understanding, pixel-level groundings, and conversational abilities, GLaMM offers a versatile capability to interact with visual inputs provided by the user at multiple granularity levels.

Delve into the core of GLaMM with our detailed guides on the model's Training and Evaluation methodologies.…

各数据源

968 Star
  • Star 968
  • Fork 56
  • 提交 48
  • 发布 0