NeurIPS 2026

GRASP: Learning to Ground Social Reasoning in Multi‑Person Non‑Verbal Interactions

Junho Kim1 Xu Cao1 Houze Yang1 Bikram Boote1 Ana Jojic1 Fiona Ryan2 Bolin Lai3 Sangmin Lee4 James M. Rehg1
1University of Illinois Urbana-Champaign 2Georgia Institute of Technology 3Amazon AGI 4Korea University
📄 arXiv 💻 Code 🤗 HF Paper 📦 Dataset (coming soon)
GRASP teaser: grounding social reasoning in the correct participants over time
Multi-person social reasoning requires grounding subtle non-verbal cues in the correct participants over time. Existing MLLMs often take spurious scene-level shortcuts, whereas ours leverages evidence-aware supervision to reason from the relevant social event.

Abstract

Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with whom in multi-person videos. We introduce GRASP, a large-scale social reasoning dataset that connects high-level social QA with fine-grained gaze and deictic gesture events. GRASP contains 290K question–answer pairs over 46K videos totaling 749 hours, organized by a 16-category taxonomy spanning gaze, gesture, and joint gaze–gesture reasoning, together with GRASP-Bench for evaluation. Unlike prior resources that focus on either isolated cues or high-level social QA, GRASP builds questions from identity-consistent gaze trajectories, deictic gestures, and their joint compositions into social events. Moreover, we propose Social Grounding Reward (SGR), a learning signal that uses these social events to encourage models to reason about the participants involved in each interaction. Experiments show that SGR improves performance on GRASP-Bench while maintaining zero-shot performance on related social video QA benchmarks.

The GRASP Dataset

GRASP converts multi-person videos into person-consistent gaze and gesture events, composes them into structured social QA pairs, and applies subset validation with human feedback for quality control. Questions are built from identity-consistent gaze trajectories, deictic gestures, and their joint compositions into social events — not from scene-level captions.

290K
QA pairs
46K
videos
749h
total footage
215K
gaze events
88K
deictic gestures
16
QA categories
GRASP construction pipeline
GRASP construction pipeline. Multi-person videos are converted into person-consistent gaze and gesture events, composed into structured social QA pairs, with subset validation and human feedback for quality control.
GRASP taxonomy and statistics
Taxonomy & statistics. 16 categories across three reasoning types: Gaze (T1–T6), Gesture (G1–G6), and Joint gaze–gesture (J1–J4), spanning easy/medium/hard difficulty.
Per-category comparison on GRASP-Bench
Per-category comparison on GRASP-Bench across SFT, reasoning, GRPO-tuned baselines, and ours.

Explore GRASP-Bench

One representative example per category, drawn from the GRASP-Bench test split. Play the clip, then reveal the answer. Overlaid boxes show person-consistent identities used to ground each question; clips are shown at the 2 fps rate models receive.

Social Grounding Reward (SGR)

Standard outcome-only RL rewards let models guess the right option while attending to the wrong people. SGR is a participant-level learning signal built directly on GRASP's structured social events: during GRPO post-training, the model is rewarded for reasoning about the correct participants involved in each gaze or gesture interaction, not just for the final answer.

Because every GRASP question is generated from an underlying social event graph, the ground-truth participants are known for free — the dataset and the reward are coupled by design. SGR improves performance on GRASP-Bench while maintaining zero-shot performance on related social video QA benchmarks (MMSI, Online-MMSI, TVQA+).

Results on GRASP-Bench

Accuracy (%) grouped by reasoning type. With SGR, open 8–9B models surpass the strongest proprietary model evaluated (Gemini 3.1 Pro). Full 16-category results are in the paper.

MethodGazeGestureJointOverall
Proprietary
Claude Sonnet 4.630.246.537.337.0
GPT-5.434.960.439.043.9
Gemini 3.1 Pro41.867.644.350.5
Open-source (SFT / instruct)
Qwen3-VL-8B-Instruct28.352.738.038.3
Qwen3.5-9B (Instruct)32.859.040.843.0
Open-source (RL post-trained)
Video-R1-7B37.134.836.636.3
Qwen3-VL-8B-Thinking28.148.438.036.9
Socially-grounded RL (ours)
Qwen3-VL-8B + SGR46.358.048.150.4
Qwen3.5-9B + SGR48.064.445.652.6

BibTeX

@article{kim2026grasp, title={GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions}, author={Kim, Junho and Cao, Xu and Yang, Houze and Boote, Bikram and Jojic, Ana and Ryan, Fiona and Lai, Bolin and Lee, Sangmin and Rehg, James M}, journal={arXiv preprint arXiv:2605.15764}, year={2026} }