Gaze-to-Task Inference in Chart Reading: Best Practices for Integrating Human Attention with Multimodal LLMs

概要

Understanding how humans interact with charts is crucial for designing effective visualization systems. While identifying a user’s task from gaze is fundamental, traditional methods rely on labor-intensive feature engineering, showing limited performance and high task-specificity. In this work, we introduce a framework for MLLM-based gaze-to-task inference, including an automatic few-shot sample generation process that creates structured demonstrations for in-context learning. We present the first systematic investigation into gaze-based task inference on charts, benchmarking five MLLMs and rigorously exploring gaze encoding and prompting strategies to establish optimal design principles. Our findings identify the heatmap representation as the optimal visual gaze encoding, and we demonstrate the necessity of Chain-of-Thought prompting. Notably, MLLMs can autonomously decode cognitive intent without manual AOI definitions, exceeding traditional baseline performance. This study offers actionable insights for integrating human gaze into MLLMs, guiding the design of future systems that adapt to a user’s analytical focus.

収録
Proceedings of the ACM in Computer Graphics and Interactive Techniques (PACMCGIT), Vol. 9, Issue 2 (ETRA 2026), pp. 1-16