TY - GEN
T1 - IntRec
T2 - 25th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2026
AU - Shamsolmoali, Pourya
AU - Zareapoor, Masoumeh
AU - Granger, Eric
AU - Lu, Yue
N1 - Publisher Copyright:
© 2026 International Foundation for Autonomous Agents and Multiagent Systems.
PY - 2026/5/24
Y1 - 2026/5/24
N2 - Retrieving user-specified objects from complex scenes remains a challenging task, especially when queries are ambiguous or involve multiple similar objects. Existing open-vocabulary detectors operate in a one-shot manner, lacking the ability to refine predictions based on user feedback. To address this, we propose IntRec, an interactive object retrieval framework that refines predictions based on user feedback. At its core is an Intent State (IS) that maintains dual memory sets for positive anchors (confirmed cues) and negative constraints (rejected hypotheses). A contrastive alignment function ranks candidate objects by maximizing similarity to positive cues while penalizing rejected ones, enabling fine-grained disambiguation in cluttered scenes. Our interactive framework provides substantial improvements in retrieval accuracy without additional supervision. On LVIS, IntRec achieves 35.4 AP, outperforming OVMR, CoDet, and CAKE by +2.3, +3.7, and +0.5, respectively. On the challenging LVIS-Ambiguous benchmark, it improves performance by +7.9 AP over its one-shot baseline after a single corrective feedback, with less than 30 ms of added latency per interaction.
AB - Retrieving user-specified objects from complex scenes remains a challenging task, especially when queries are ambiguous or involve multiple similar objects. Existing open-vocabulary detectors operate in a one-shot manner, lacking the ability to refine predictions based on user feedback. To address this, we propose IntRec, an interactive object retrieval framework that refines predictions based on user feedback. At its core is an Intent State (IS) that maintains dual memory sets for positive anchors (confirmed cues) and negative constraints (rejected hypotheses). A contrastive alignment function ranks candidate objects by maximizing similarity to positive cues while penalizing rejected ones, enabling fine-grained disambiguation in cluttered scenes. Our interactive framework provides substantial improvements in retrieval accuracy without additional supervision. On LVIS, IntRec achieves 35.4 AP, outperforming OVMR, CoDet, and CAKE by +2.3, +3.7, and +0.5, respectively. On the challenging LVIS-Ambiguous benchmark, it improves performance by +7.9 AP over its one-shot baseline after a single corrective feedback, with less than 30 ms of added latency per interaction.
KW - Contrastive alignment
KW - Interactive learning
KW - Object detection
UR - https://www.scopus.com/pages/publications/105041455818
U2 - 10.65109/AGOG8131
DO - 10.65109/AGOG8131
M3 - Contribution to conference proceedings
AN - SCOPUS:105041455818
T3 - AAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems
SP - 1628
EP - 1636
BT - AAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems
PB - Association for Computing Machinery, Inc
Y2 - 25 May 2026 through 29 May 2026
ER -