Evidence
同一模型与写作位置 · 提供检索文献片段 · 首次输出
Despite this progress, VLN agents remain largely evaluated in isolation, and their extension to collaborative settings where instructions are jointly negotiated with a human partner is comparatively underexplored. Prior work on human-object interaction detection has begun to model multivariate relationships involving auxiliary entities such as tools, explicitly capturing the functional role of these objects through triplet structures.
提供给 Evidence 版本的文献片段
Contextualized Representation Learning for Effective Human-Object Interaction Detection ↗
Human-Object Interaction (HOI) detection aims to simultaneously localize human-object pairs and recognize their interactions. While recent two-stage approaches have made significant progress, they still face challenges due to incomplete context modeling. In this work, we introduce a Contextualized R…
展开完整文献摘录
Human-Object Interaction (HOI) detection aims to simultaneously localize human-object pairs and recognize their interactions. While recent two-stage approaches have made significant progress, they still face challenges due to incomplete context modeling. In this work, we introduce a Contextualized Representation Learning Network that integrates both affordance-guided reasoning and contextual prompts with visual cues to better capture complex interactions. We enhance the conventional HOI detection framework by expanding it beyond simple human-object pairs to include multivariate relationships involving auxiliary entities like tools. Specifically, we explicitly model the functional role (affordance) of these auxiliary objects through triplet structures < < human, tool, object > > .