Sato Lab./Sugano Lab.
Sato Lab./Sugano Lab.
Y. Sato Lab.
Sugano Lab.
News
Publications
Contact
Resources
Internal Wiki
English
日本語
Paper-Conference
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
As AI systems become increasingly integrated into human lives, endowing them with robust social intelligence has emerged as a critical …
Caixin Kang
,
Yifei Huang
,
Liangyang Ouyang
,
Mingfang Zhang
,
Yoichi Sato
PDF
Cite
DOI
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
How can we reconstruct 3D hand poses when large portions of the hand are heavily occluded by itself or by objects? Humans often resolve …
Naru Suzuki
,
Takehiko Ohkawa
,
Tatsuro Banno
,
Jihyun Lee
,
Ryosuke Furuta
,
Yoichi Sato
PDF
Cite
DOI
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages …
Mingfang Zhang
,
Ryo Yonetani
,
Yifei Huang
,
Liangyang Ouyang
,
Ruicong Liu
,
Yoichi Sato
PDF
Cite
Code
Data-driven Head Motion Generation through Natural Gaze-Head Coordination
We present the first data-driven approach to model temporal gaze-head coordination from large-scale in-the-wild facial videos. To …
Xiaohan Liu
,
Yilin Wen
,
Yusuke Sugano
PDF
Cite
Project
DOI
ChartQC: Question Classification from Human Attention Data on Charts
Understanding how humans interact with information visualizations is crucial for improving user experience and designing effective …
Takumi Nishiyasu*
,
Tobias Kostorz*
,
Yao Wang
,
Yoichi Sato
,
Andreas Bulling
PDF
Cite
DOI
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
We address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze …
Michele Mazzamuto
,
Antonino Furnari
,
Yoichi Sato
,
Giovanni Maria Farinella
PDF
Cite
SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training
We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand …
Nie Lin
,
Takehiko Ohkawa
,
Yifei Huang
,
Mingfang Zhang
,
Minjie Cai
,
Ming Li
,
Ryosuke Furuta
,
Yoichi Sato
PDF
Cite
Code
A Multimodal LLM-based Assistant for User-Centric Interactive Machine Learning
This paper proposes a system based on a multimodal large language model (MLLM) to assist non-expert users without prior experience in …
Wataru Kawabe
,
Yusuke Sugano
PDF
Cite
DOI
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
We propose a novel benchmark for cross-view knowledge transfer of dense video captioning, adapting models from web instructional videos …
Takehiko Ohkawa
,
Takuma Yagi
,
Taichi Nishimura
,
Ryosuke Furuta
,
Atsushi Hashimoto
,
Yoshitaka Ushiku
,
Yoichi Sato
PDF
Cite
Learning Multiple Object States from Actions via Large Language Models
Recognizing the states of objects in a video is crucial in understanding the scene beyond actions and objects. For instance, an egg can …
Masatoshi Tateno
,
Takuma Yagi
,
Ryosuke Furuta
,
Yoichi Sato
PDF
Cite
Code
DOI
«
»
Cite
×