Object-Oriented Architecture
HandMeThat: Human-RobotCommunication inPhysicalandSocialEnvironments
While previousbenchmarks insimilar domains havebeenprimarily focusing onthelanguage grounding of object properties (e.g., "table"), relations (e.g., "on"), and planning (e.g., object search and manipulation) [6,7],inthispaper,wehighlights theadditional challenge forunderstanding human instructions withambiguities (i.e., recognizing the subgoal) based on physical states and human actionsandgoals. Each episode in HandMeThat contains twostages.
3341f6f048384ec73a7ba2e77d2db48b-Paper.pdf
Instance segmentation, which seeks to obtain both class and instance labels for each pixelinthe input image, isachallenging task incomputer vision. State-ofthe-art algorithms often employ a search-based strategy, which first divides the output image with a regular grid and generate proposals at each grid cell, then the proposals are classified and boundaries refined.