佐藤研究室/菅野研究室
佐藤研究室/菅野研究室
佐藤 (洋) 研究室
菅野研究室
ニュース
発表文献
連絡先
リソース
内部ページ
日本語
English
Paper-Conference
Multi-speaker Attention Alignment for Multimodal Social Interaction
Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, …
Liangyang Ouyang
,
Yifei Huang
,
Mingfang Zhang
,
Caixin Kang
,
Ryosuke Furuta
,
Yoichi Sato
PDF
引用
Physically Plausible Human-Object Interaction Generation via Attribute Classifier Guidance
Human motion during daily interactions is inherently shaped by physical attributes such as mass, friction, and fragility. While the …
Kengo Ikeuchi
,
Takehiko Ohkawa
,
Risa Shinoda
,
Yoichi Sato
PDF
引用
Exploring a Collaborative Gamified Approach to Vision-Language Model Evaluation
Vision-language models demonstrate impressive capabilities, yet constructing evaluation datasets that capture their failure cases still …
Fan Gao
,
Jenna Ren Mei Wang
,
Tomomi Sayuda
,
Miles Pennington
,
Yusuke Sugano
PDF
引用
DOI
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has …
Nie Lin
,
Yansen Wang
,
Dongqi Han
,
Weibang Jiang
,
Jingyuan Li
,
Ryosuke Furuta
,
Yoichi Sato
,
Dongsheng Li
PDF
引用
UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training
Despite decades of research on data collection and model architectures, current gaze estimation models encounter significant challenges …
Jiawei Qin
,
Xucong Zhang
,
Yusuke Sugano
PDF
引用
ソースコード
Robust Long-term Test-Time Adaptation for 3D Human Pose Estimation through Motion Discretization
Online test-time adaptation addresses the train-test domain gap by adapting the model on unlabeled streaming test inputs before making …
Yilin Wen
,
Kechuan Dong
,
Yusuke Sugano
PDF
引用
DOI
Data-driven Head Motion Generation through Natural Gaze-Head Coordination
We present the first data-driven approach to model temporal gaze-head coordination from large-scale in-the-wild facial videos. To …
Xiaohan Liu
,
Yilin Wen
,
Yusuke Sugano
PDF
引用
プロジェクト
DOI
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
We address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze …
Michele Mazzamuto
,
Antonino Furnari
,
Yoichi Sato
,
Giovanni Maria Farinella
PDF
引用
SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training
We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand …
Nie Lin
,
Takehiko Ohkawa
,
Yifei Huang
,
Mingfang Zhang
,
Minjie Cai
,
Ming Li
,
Ryosuke Furuta
,
Yoichi Sato
PDF
引用
ソースコード
A Multimodal LLM-based Assistant for User-Centric Interactive Machine Learning
This paper proposes a system based on a multimodal large language model (MLLM) to assist non-expert users without prior experience in …
Wataru Kawabe
,
Yusuke Sugano
PDF
引用
DOI
«
»
引用
×