Sato Lab./Sugano Lab.
Sato Lab./Sugano Lab.
Y. Sato Lab.
Sugano Lab.
News
Publications
Contact
Resources
Internal Wiki
English
日本語
Paper-Conference
Multi-speaker Attention Alignment for Multimodal Social Interaction
Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, …
Liangyang Ouyang
,
Yifei Huang
,
Mingfang Zhang
,
Caixin Kang
,
Ryosuke Furuta
,
Yoichi Sato
PDF
Cite
Physically Plausible Human-Object Interaction Generation via Attribute Classifier Guidance
Human motion during daily interactions is inherently shaped by physical attributes such as mass, friction, and fragility. While the …
Kengo Ikeuchi
,
Takehiko Ohkawa
,
Risa Shinoda
,
Yoichi Sato
PDF
Cite
Exploring a Collaborative Gamified Approach to Vision-Language Model Evaluation
Vision-language models demonstrate impressive capabilities, yet constructing evaluation datasets that capture their failure cases still …
Fan Gao
,
Jenna Ren Mei Wang
,
Tomomi Sayuda
,
Miles Pennington
,
Yusuke Sugano
PDF
Cite
DOI
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has …
Nie Lin
,
Yansen Wang
,
Dongqi Han
,
Weibang Jiang
,
Jingyuan Li
,
Ryosuke Furuta
,
Yoichi Sato
,
Dongsheng Li
PDF
Cite
UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training
Despite decades of research on data collection and model architectures, current gaze estimation models encounter significant challenges …
Jiawei Qin
,
Xucong Zhang
,
Yusuke Sugano
PDF
Cite
Code
Robust Long-term Test-Time Adaptation for 3D Human Pose Estimation through Motion Discretization
Online test-time adaptation addresses the train-test domain gap by adapting the model on unlabeled streaming test inputs before making …
Yilin Wen
,
Kechuan Dong
,
Yusuke Sugano
PDF
Cite
DOI
Generative Modeling of Shape-Dependent Self-Contact Human Poses
One can hardly model self-contact of human poses without considering underlying body shapes. For example, the pose of rubbing a belly …
Takehiko Ohkawa
,
Jihyun Lee
,
Shunsuke Saito
,
Jason Saragih
,
Fabian Prada
,
Yichen Xu
,
Shoou-I Yu
,
Ryosuke Furuta
,
Yoichi Sato
,
Takaaki Shiratori
PDF
Cite
Code
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
Bimanual human activities inherently involve coordinated movements of both hands and body. However, the impact of this coordination in …
Tatsuro Banno
,
Takehiko Ohkawa
,
Ruicong Liu
,
Ryosuke Furuta
,
Yoichi Sato
PDF
Cite
DOI
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical …
Yuki Sakai
,
Ryosuke Furuta
,
Juichun Yen
,
Yoichi Sato
PDF
Cite
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
As AI systems become increasingly integrated into human lives, endowing them with robust social intelligence has emerged as a critical …
Caixin Kang
,
Yifei Huang
,
Liangyang Ouyang
,
Mingfang Zhang
,
Yoichi Sato
PDF
Cite
DOI
«
»
Cite
×