EDITH achieves 48.6% average success rate and 75.9% task progress by translating nonverbal human signals into keyframe-grounded subtasks.
Muffin-Serving
Tumbler-Sorting
Tool-Passing
Average
The language-only baselines barely get off the ground: \(\pi_l^{\mathrm{lang}}\) and \(\pi_h^{\mathrm{lang}} + \pi_l^{\mathrm{lang}}\) average 4.2% and 6.9% success. Without the egocentric stream there is nothing in the instruction that identifies the target.
Conditioning an end-to-end policy directly on the current egocentric context is not enough either — \(\pi_l^{\mathrm{ego+lang}}\) reaches 6.2%. It helps only while gaze stays on the target and degrades as soon as attention becomes intermittent.
\(\pi_{\mathrm{FAM\text{-}HRI}} + \pi_l\), which infers intent from gaze and speech before acting, is the strongest baseline at 25.0% SR and 57.1% TP. EDITH roughly doubles that success rate by monitoring intent separately in the high-level policy and handing the low-level policy a keyframe-grounded subtask.





