Literature
Display
← All papers

CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views

Alexey Gavryushin* † Dingxi Zhang* Zhao Huang* · Alexandros Delitzas Jiaqi Chen Ben Ellis Cedric Zöllner · Manthan Patel Manuel Kaufmann Marc Pollefeys Xi Wang

6 min · CondensedOriginal paper ↗
Contents

Overview

The overview figure aligns collaborative recordings, scene geometry, and the three prediction tasks.

Refer to caption
Figure 1: The CoMind dataset captures 41 hours of human-human collaboration across 80 sessions, combining egocentric/exocentric videos, audio, gaze, and high-quality 3D scene and object scans, as well as rich annotations for three novel social reasoning tasks.

At a glance

CoMind records unscripted collaborative cooking with synchronized first- and third-person views, audio, gaze, hand tracking, and 3D scans. Three benchmarks test joint attention, socially triggered action anticipation, and handovers before reaching begins. Current vision-language models struggle especially with localization; fine-tuning on CoMind improves several social and action predictions.

1 Introduction

Theory of Mind in physical collaboration requires inferring another person’s attention, intent, and need for assistance from evolving social cues. CoMind supplies aligned observations of both participants and their environment, moving beyond isolated actions or text-only belief questions.

Related references: [83], [5], [2], [77], [65], [84], [42], [96], [15], [17], [74], [18], [76], [86], [32], [95], [39], [9], [38], [33], [59], [40], [44], [88], [41], [34], [12], [62], [94].

2 Related Work

Existing datasets emphasize individual activity, short interactions, conversation, or handover mechanics. CoMind combines long collaborative sessions with gaze, dialogue, object grounding, and social annotations. Its tasks ask about shared attention and impending assistance before the physical interaction is already underway.

Related references: [19], [59], [10], [76], [32], [33], [38], [86], [40], [89], [13], [18], [95], [39], [9], [64], [41], [47], [94], [12], [62], [79], [61], [44], [88], [60], [75], [77], [65], [84], [42], [96], [15], [67], [16], [26], [25], [20], [45], [48], [92], [37], [68], [85], [1], [30], [14], [69], [70], [63], [31], [29], [58], [28], [57], [78], [46], [93], [50], [97], [36], [87], [81], [82], [27], [56], [11], [43].

Table 1: Overview of egocentric datasets. For Modality, [Uncaptioned image] denotes video, [Uncaptioned image] denotes gaze, [Uncaptioned image] denotes IMU, [Uncaptioned image] denotes semi-dense point cloud, [Uncaptioned image] denotes dense 3D scene scans, [Uncaptioned image] denotes 3D object scans, [Uncaptioned image] denotes hand tracking. Hours (Ego) refers to the cumulative duration of distinct egocentric video recordings. Our dataset provides unprecedented depth by offering rich social cues and collaborative annotations alongside multi-view 3D grounding.
  Benchmark     Domain     Modality     Hours (Ego)     
  Average
  Duration
  
  Exo     
  Collaboration
  Dynamics
  
  
  Social Cue
  Annotation
  
  
  #Humans
  per session
  
  
  Total
  participants
  
  EPIC-KITCHENS-100 [19]     Kitchen     [Uncaptioned image]     100     8.5 min                    1     32  
  HD-EPIC [59]     Kitchen     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     41     15.9 min                    1     9  
  EPFL-Smart-Kitchen-30 [10]     Kitchen     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     29.7     35.9 min                    1     16  
  Assembly101 [76]     Task Execution     [Uncaptioned image]     167     7.1 min                    1     53  
  Ego4D [32]     Daily Activities     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     3,670     22.8 min                    1     923  
  EgoExo4D [33]     Skilled Activities     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     221     2.6 min                    1-2     740  
  EgoExoLearn [38]     Task Execution     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     120     13.4 min                    1     -  
  HoloAssist [86]     Assistive Task     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     166     4.8 min                    1-2     222  
  LEMMA [40]     Task Execution     [Uncaptioned image] [Uncaptioned image]     10     2 min                    1-2     8  
  EgoLife [89]     Daily Life     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     266     44.3 h                    6     6  
  IndEgo [13]     Industrial Tasks     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     197     3.4 min                    1-2     20  
  Ours     Kitchen     [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image] [Uncaptioned image]     81     30.5 min                    2     125  

3 The CoMind Tasks and Dataset

The dataset formalizes three complementary tasks and provides recordings, annotations, and shared spatial calibration.

3.1 Proposed Tasks

Joint attention predicts an attended object, its boxes in both views, and eliciting cue type from a target frame plus ten seconds of speech context. Social action anticipation predicts the helper’s next verb, object, box, and cue from ten seconds of history. Handover prediction estimates time, sender/receiver, initiator, object, box, and cue before reaching starts. These operational tasks probe social inference; they do not directly measure a complete theory of mind.

Refer to caption
Figure 2: Joint Attention Estimation Task. From the audio or transcribed speech in a context window, together with a target frame (green), the model predicts the jointly attended object, its bounding box in the left and right views, and the social cue type (blue).
Refer to caption
Figure 3: Example annotations. Frames show joint attention cues, gaze points, and object bounding boxes with labels.
Refer to caption
Figure 4: Socially Conditioned Object Interaction Anticipation Task. From context frames (or video) and audio or transcribed speech together with a prediction frame (green), the model predicts the noun and verb of the action to be performed, the interacted object’s bounding box, and the social cue type (blue).
Refer to caption
Figure 5: Example annotations. The prediction frame shows the ground-truth action verb, noun, object bounding box, and cue type(s).
Refer to caption
Figure 6: Collaborative Handover Prediction Task. From context frames (or video) and audio or transcribed speech, together with a prediction frame (green), the model predicts the delivery flow (who hands to whom), the object category and bounding box in the view of the handing participant, the initiator, and the cue type (blue).
Refer to caption
Figure 7: Example annotations. Frames show handover initiator, cue type, delivery flow, and object bounding boxes with labels.

3.2 Data Recording Protocol and Hardware Setup

Both participants wear calibrated Aria glasses. Two GoPro cameras capture front and rear views. A Leica scanner records each kitchen, while an Artec scanner captures shared utensils. Sessions are unscripted, monitored remotely, and recorded with participant consent.

Related references: [24], [6].

Refer to caption
Figure 8: Sensors used to create the CoMind dataset. Two Aria glasses record the scene from each participant’s view. Two GoPro Hero cameras capture exocentric views. The Leica BLK2Go and Artec 3D Leo are used for 3D scene and object scanning respectively.
Refer to caption
Figure 9: 3D scans of objects used in the recordings. Participants work with objects that have been 3D-scanned.

3.3 Data processing

Aria streams are synchronized to sub-millisecond precision; external views use audio cross-correlation. Machine Perception Services produce camera poses, point clouds, gaze, and hand tracking in a shared frame. WhisperX generates transcripts. Kitchen scans align through image localization followed by point-to-plane registration, with manual verification.

Related references: [49], [80], [66], [8], [72], [73], [71].

3.4 Data Annotation

In-house and professional annotators label the three tasks using dedicated interfaces and training instructions. Quality teams review external work, and a main contributor reviews all annotations.

Related references: [23].

Refer to caption
Figure 10: Modalities in the CoMind dataset. The dataset provides synchronized multimodal recordings of collaborative sessions, including egocentric video with eye gaze and hand tracking, audio and IMU signals, exocentric camera views, camera trajectories and spatial mapping, dense 3D scans of the scene, and a set of high-quality 3D scans of commonly used objects.

3.5 Data Statistics

The dataset contains eighty paired recordings from 55 kitchens and 125 participants, with about 40 hours 43 minutes of sessions or 81 hours 26 minutes summed over two egocentric streams. Sessions average roughly thirty minutes. The hierarchy contains 281 fine, 76 intermediate, and fourteen coarse object categories. Appendix statistics report 8,513 joint-attention, 1,747 social-action, and 586 handover events; some main-text captions use different category or verb totals.

Related references: [32], [33].

Original figure
Figure 11: Histogram of recording lengths. Our lengthy recordings allow for the study of human-human collaborative work in long-context settings. In total, we record 41 hours of two-person resp. 81 hours of single-view data.
Original figure
Figure 12: Participant gender and ethnicity distribution. We balance participants from Eastern and Western backgrounds, and maintain an equal gender distribution of our participants.
Original figure
Figure 13: Action verb and action noun distribution of the Socially Conditioned Object Interaction Anticipation task. Our annotations feature a wide range of objects from 13 coarse (L3) and over 180 fine (L1) object categories, as well as 22 unique verbs.
Original figure
Figure 14: Annotation statistics for the Joint Attention Estimation task. The CoMind dataset allows investigating multiple indicators of joint attention between humans on a broad variety of objects.
Original figure
Figure 15: Annotation statistics for the Collaborative Handover Prediction task. In addition to the category and bounding box of the handed-over object, we provide annotations about the Time-To-Handover (TTH), delivery flow direction, initiator and initiation cue for each handover event.

4 Baselines and Experiments

Training and test participants are disjoint; 25 videos provide test annotations. Models receive task-specific prompts and structured output requirements, usually five frames and transcripts from the preceding ten seconds. Baselines include pretrained proprietary and open models, fine-tuned Qwen3-VL, random sampling, and most-frequent predictions.

Related references: [3], [4], [22], [53], [54], [55], [21], [7].

4.1 Joint Attention Estimation

Localization is measured as the fraction of boxes exceeding 0.5 intersection over union, separately for each view. Gemini 3 Flash leads at 0.2164 and 0.2071. Fine-tuned Qwen3-VL-32B ranks second spatially; the 8B model has the best cue accuracy (0.4775). Object identification is easier than grounding the shared target in both views.

Table 2: Quantitative results on Joint Attention Estimation. Best, second-best, and third-best results are highlighted. L-IoU@0.5\mathrm{L\text{-}IoU@0.5} and R-IoU@0.5\mathrm{R\text{-}IoU@0.5} denote the fraction of predicted bounding boxes with IoU>0.5\mathrm{IoU}>0.5 with the ground-truth jointly attended object in the left and right egocentric streams, respectively. Cue Type reports cue classification accuracy, and Cat. L1 measures the fraction of correct L1 object category predictions. Models marked FT are fine-tuned on CoMind.
Model L-IoU@0.5 \uparrow R-IoU@0.5 \uparrow Cue Type \uparrow Cat. (L1\uparrow
Random Sampling 0.1772 0.1337 0.3163 0.0533
Most Frequent 0.0004 0.0004 0.4647 0.2067
Claude Opus 4.5 0.0015 0.0025 0.2419 0.3433
Claude Opus 4.6 0.0051 0.0015 0.3439 0.3487
Gemini 3 Flash 0.2164\mathbf{{0.2164}} 0.2071\mathbf{{0.2071}} 0.3685 0.42060.4206
GPT 4o 0.0026 0.0003 0.0870 0.3451
GPT 5.1 0.0042 0.0017 0.1133 0.42810.4281
GPT 5.2 0.0044 0.0017 0.1474 0.4285\mathbf{{0.4285}}
Gemma 4 31B IT 0.1295 0.1174 0.46280.4628 0.3590
Qwen3-VL 8B Instruct 0.0000 0.0012 0.0074 0.2701
Qwen3-VL 8B Instruct FT 0.14830.1483 0.13450.1345 0.4775\mathbf{{0.4775}} 0.3007
Qwen3-VL 32B Instruct 0.0001 0.0012 0.0740 0.3494
Qwen3-VL 32B Instruct FT 0.15620.1562 0.14960.1496 0.45720.4572 0.3323

4.2 Socially Conditioned Object Interaction Anticipation

The fine-tuned 32B model leads in localization (0.2513), cue classification (0.5503), and action-verb accuracy (0.6508). Fine-tuning raises verb accuracy from 0.0873 to 0.6429 for 8B and from 0.1380 to 0.6508 for 32B. This indicates substantial benefit from training on temporal and social context.

Table 3: Quantitative results on Socially Conditioned Object Interaction Anticipation. Best, second-best, and third-best results are highlighted. IoU@0.5\mathrm{IoU@0.5} denotes the fraction of predicted bounding boxes with IoU>0.5\mathrm{IoU}>0.5 with the ground-truth object to be interacted with. Cue Type reports cue classification accuracy, and Act. Verb resp. Act. Noun (L1) measures the fraction of correct verb resp. L1 object category predictions. Models marked FT are fine-tuned on CoMind.
Model IoU@0.5 \uparrow Cue Type \uparrow Act. Verb \uparrow Act. Noun (L1\uparrow
Random Sampling 0.0152 0.3801 0.1854 0.0411
Most Frequent 0.0005 0.4040 0.3430 0.0662
Claude Opus 4.5 0.0034 0.53020.5302 0.0819 0.3490
Claude Opus 4.6 0.0091 0.53470.5347 0.0987 0.3507
Gemini 3 Flash 0.18590.1859 0.5298 0.0781 0.3775
GPT 4o 0.0009 0.4358 0.1046 0.3351
GPT 5.1 0.0202 0.4556 0.1325 0.41190.4119
GPT 5.2 0.0038 0.5113 0.1060 0.4371\mathbf{{0.4371}}
Gemma 4 31B IT 0.1468 0.4173 0.0679 0.3225
Qwen3-VL 8B Instruct 0.0909 0.1151 0.0873 0.2288
Qwen3-VL 8B Instruct FT 0.24250.2425 0.4299 0.64290.6429 0.4021
Qwen3-VL 32B Instruct 0.0035 0.4901 0.13800.1380 0.3509
Qwen3-VL 32B Instruct FT 0.2513\mathbf{{0.2513}} 0.5503\mathbf{{0.5503}} 0.6508\mathbf{{0.6508}} 0.43120.4312

4.3 Collaborative Handover Prediction

Handover localization remains very weak: the best box-success fraction is 0.0313. Timing counts as correct within 0.25 seconds; the best timing score is 0.1317. Fine-tuned Qwen models lead on initiator (0.7286) and cue type (0.3714), but all methods leave large gaps in proactive assistance.

Table 4: Quantitative results on Collaborative Handover Prediction. Best, second-best, and third-best results are highlighted. IoU@0.5\mathrm{IoU@0.5} denotes the fraction of predicted bounding boxes with IoU>0.5\mathrm{IoU}>0.5 with the ground-truth object. Del. Flow, Initiator, and Init. Type report the fraction of correct predictions for the delivery flow, the initiating participant, and the initiation cue type, respectively. Cat. (L1) measures the fraction of correct L1 object category predictions. TTH reports the fraction of time-to-handover predictions within 0.25s0.25\,\mathrm{s} of the ground-truth. Models marked FT are fine-tuned on CoMind.
Model IoU@0.5 \uparrow Del. Flow \uparrow Initiator \uparrow Init. Type \uparrow Cat. (L1\uparrow TTH \uparrow
Random Sampling 0.0000 0.4476 0.6810 0.2905 0.0429 0.0714
Most Frequent 0.0000 0.4857 0.7857 0.4667 0.1048 0.0524
Claude Opus 4.5 0.0000 0.5659 0.6195 0.30730.3073 0.3610 0.09760.0976
Claude Opus 4.6 0.00250.0025 0.61460.6146 0.64880.6488 0.2976 0.3610 0.1317\mathbf{{0.1317}}
Gemini 3 Flash 0.0313\mathbf{{0.0313}} 0.6524\mathbf{{0.6524}} 0.5810 0.34760.3476 0.3524 0.0667
GPT 4o 0.0008 0.5433 0.5240 0.2644 0.3317 0.0962
GPT 5.1 0.0000 0.5857 0.5952 0.2333 0.4000\mathbf{{0.4000}} 0.0810
GPT 5.2 0.0000 0.5952 0.5667 0.2905 0.37140.3714 0.0810
Gemma 4 31B IT 0.01300.0130 0.5050 0.4100 0.2800 0.3400 0.0750
Qwen3-VL 8B Instruct 0.0000 0.5381 0.4667 0.1286 0.1905 0.0524
Qwen3-VL 8B Instruct FT 0.00250.0025 0.5714 0.66670.6667 0.1429 0.1952 0.11430.1143
Qwen3-VL 32B Instruct 0.0000 0.5381 0.6143 0.2952 0.3095 0.0381
Qwen3-VL 32B Instruct FT 0.01300.0130 0.61430.6143 0.7286\mathbf{{0.7286}} 0.3714\mathbf{{0.3714}} 0.36190.3619 0.0905

5 Conclusion

CoMind provides complementary sensory evidence and benchmark labels for social collaboration. Fine-tuning improves prediction, while weak spatial grounding and handover anticipation leave substantial room for progress.

Acknowledgements

The work acknowledges Swiss and European research funding and institutional support.

Supplementary Material

The supplement retains qualitative comparisons, training settings, task prompts, annotation interfaces, and consent materials.

Appendix S1 Qualitative Examples

Figures S1–S5 compare pretrained and fine-tuned predictions with ground truth. They illustrate the benefits and remaining errors of task-specific training.

Refer to caption
Refer to caption
Refer to caption
Refer to caption
Figure S1: Baseline comparison for the Joint Attention Estimation task. We visualize and compare the predictions of Qwen3-VL 8B with those of Qwen3-VL 8B finetuned on the CoMind training set and Gemini 3 Flash, and provide the ground-truth solution given by our dataset. Finetuning on our training set improves the test set predictions.
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Figure S2: Baseline comparison for the Socially Conditioned Object Interaction Anticipation task. We visualize and compare the predictions of Qwen3-VL 8B with those of Qwen3-VL 8B finetuned on the CoMind training set and Gemini 3 Flash, and provide the ground-truth solution given by our dataset. Finetuning on our training set improves the test set predictions.
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Figure S3: Baseline comparison for the Collaborative Handover Prediction task. We visualize and compare the predictions of Qwen3-VL 8B with those of Qwen3-VL 8B finetuned on the CoMind training set and Gemini 3 Flash, and provide the ground-truth solution given by our dataset. Finetuning on our training set improves the test set predictions.
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Figure S4: Further predictions on the Joint Attention Estimation task. We visualize the predictions of Qwen3-VL 8B finetuned on our training set in red, together with the ground-truth annotations in blue.
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Refer to caption
Figure S5: Further predictions on the Socially Conditioned Object Interaction Anticipation task. We visualize the predictions of Qwen3-VL 8B finetuned on our training set in red, together with the ground-truth annotations in blue.

Appendix S2 Downstream Applications and Future Tasks

The synchronized geometry supports future body-pose, object-pose, gaze-following, and 3D grounding tasks. The illustrated reconstructions are examples of potential derived annotations, rather than additional validated benchmark ground truth.

Related references: [90], [51], [67], [49], [52], [91].

Refer to caption
(a) 6D object pose estimation
Refer to caption
(b) Human body pose estimation
Figure S6: Pseudo ground-truth extraction with foundation models. Our dataset provides modalities well-suited for pseudo ground-truth extraction using foundation models.
Refer to caption
Figure S7: Gaze Following as a precursor to Joint Attention. In collaborative tasks, gaze following frequently serves as a vital geometric and social prior before joint attention is fully established. CoMind’s synchronous dual views capture the partner’s head orientation and line of sight (blue vectors). The specific gaze points of the leader and the helper are indicated by red and blue dots, respectively. This setup effectively grounds the target objects (bounding boxes) in the shared workspace prior to the interaction.
Refer to caption
Figure S8: Potential for 3D Spatial Reasoning. By aligning dense depth estimates with the sparse 3D correspondences provided by the Aria MPS, our 2D bounding box annotations can be geometrically lifted into the 3D space. This figure shows reconstructed 3D bounding boxes within the scene point cloud, demonstrating CoMind’s capability to support future 3D prediction tasks. Only representative static object bounding boxes shown for clarity.

Appendix S3 Finetuning Hyperparameters

Qwen3-VL fine-tuning uses LoRA rank 64, scale 128, learning rate 0.0002, and batch size four. Joint attention trains for three epochs; the other tasks train for five.

Related references: [35].

Appendix S4 Communication Analysis

Table S5 checks whether the preceding ten-second transcript directly names the target object at each hierarchy level. Exact string matching excludes synonyms, paraphrases, and implicit references, so it measures explicit mentions only.

Table S5: Last-10s transcript matches to ground-truth object labels (rates). For each annotation, we check whether the transcript in the prior 10 seconds contains a direct string match to the ground-truth label at each level. Any level counts a single match if any of the Level 1–3 ground-truth labels match.
Task Level 1 Level 2 Level 3 Any level
Socially Conditioned Object Interaction 13.57% 7.96% 0.06% 13.97%
Collaborative Handover Prediction 14.16% 9.04% 0.00% 14.85%
Joint Attention Detection  5.39% 2.81% 0.02%  5.61%

Appendix S5 Noun Hierarchy & Consolidation Details

Verb consolidation produces 21 canonical action classes. Nouns retain object identity at the finest level, group related functions at the middle level, and use fourteen broad categories at the top. Figure S9 preserves the complete hierarchy and parent–child relationships.

Refer to caption
Figure S9: Object hierarchy tree for the CoMind Dataset. Boxes show L3 (left), L2 (middle), and L1 (right) categories; lines indicate parent–child relations. Colors are keyed to L3 groups.

Appendix S6 Input Ablations

Adding transcripts improves cue and object prediction; context frames improve verb and object prediction. The target frame remains available in every ablation, and prompts adjust to the supplied modalities.

Table S6: Ablation study on context frames and transcripts on the Socially Conditioned Object Interaction Anticipation task. IoU@0.5\mathrm{IoU@0.5} denotes the fraction of predicted bounding boxes with IoU>0.5\mathrm{IoU}>0.5 with the ground-truth object to be interacted with. Cue Type reports cue classification accuracy, and Act. Verb resp. Act. Noun (L1) measures the fraction of correct verb resp. L1 object category predictions. Providing transcripts greatly improves cue type and act. noun accuracy. Context frames improve act. verb and act. noun accuracy.
Model Context Transcript IoU Cue Act. Act.
Frames @0.50 Type Verb Noun (L1)
Gemini 3 Flash 0.1867 0.2954 0.0437 0.2384
0.1954 0.5404 0.0623 0.3748
0.1449 0.3272 0.0517 0.2874
0.1859 0.5298 0.0781 0.3775
Qwen3-VL 8B I. 0.0909 0.0887 0.0596 0.2199
0.0024 0.4159 0.0848 0.3258
0.0020 0.1192 0.0768 0.2450
0.0114 0.4172 0.0993 0.3298

Appendix S7 Prompting Details

Every model receives the same prompt for a given task. Figures S10–S12 retain the three complete prompt templates.

You are provided with a dual-view (side-by-side) image frame showing two egocentric views from two persons (left person and right person). Both halves show the EXACT SAME physical scene from two different spatial perspectives.
Additionally, here is the transcribed speech of the two persons during the time period leading up to the moment in the provided image frame: "{transcript}"
TASK:
Carefully observe the visual interactions and social cues in the frame and analyze the moment of joint attention shown in the frame. Answer the following questions sequentially:
1. Cue type: Analyze the frame and classify the underlying cue(s) that reveal the joint attention. Choose from one or more of these four categories (1 to 4 choices):
- <gaze>: The eye gaze or head pose of both individuals clearly converges onto the same target object or spatial area.
- <action>: The joint attention is established through a physical, functional manipulation (e.g. grasping, handing over, or manipulating a tool).
- <gestural>: The attention is directed via a communicative body movement (e.g. pointing a finger, holding an object up for display).
- <verbal>: The shared focus is established via verbal means (e.g. following a spoken instruction or comment).
2. Object grounding: Look at the provided image frame. Based on your answers to the previous questions, locate and identify the precise object of joint attention in BOTH views SEPARATELY:
- Object category: Provide a short, precise category name of the object being mutually attended to by both persons.
- Left object bounding box: Provide the bounding box coordinates of the object of joint attention specifically within the LEFT persons egocentric view.
- Right object bounding box: Provide the bounding box coordinates of the EXACT SAME object specifically within the RIGHT persons egocentric view.
OUTPUT FORMAT:
Provide your final answer as a JSON list containing a single dictionary. You may think step-by-step beforehand, but do not provide any additional text in the output.
[
{
"cue type": correct subset of ["<gaze>", "<action>", "<gestural>", "<verbal>"],
"object category": "<string>",
"left object bounding box": [<ymin>, <xmin>, <ymax>, <xmax>],
"right object bounding box": [<ymin>, <xmin>, <ymax>, <xmax>]
}
]
CRITICAL RULES:
- All bounding box coordinates MUST be normalized integers between [0, 1000] relative to the overall CONCATENATED image dimensions.
- For left object bounding box‘, its ‘<xmin>‘ and ‘<xmax>‘ values MUST STRICTLY fall between 0 and 500.
- For right object bounding box‘, its ‘<xmin>‘ and ‘<xmax>‘ values MUST STRICTLY fall between 500 and 1000.
Figure S10: Prompt for the Joint Attention Estimation task.
You are provided with a series of sequential dual-view (side-by-side) images showing two egocentric views from two persons (left person and right person) over the span of 10 seconds, and a final single reference image frame.
Additionally, here is the transcribed speech of the two persons during the considered time period: "{transcript}"
TASK:
Carefully observe the human behaviors and social cues and answer the following questions sequentially:
1. Handover Timing: Approximately how many seconds after the reference image frame will the object handover between two persons occur? (Provide a decimal time offset).
2. Delivering Flow: Whether the object will be handed over from left person to right person or from right person to left person? (Choose STRICTLY: <left to right> or <right to left>).
3. Initiator: Who initiates the handover? Initiation refers to the first cue that prompts the handover to occur. (Choose STRICTLY: <left> or <right>).
4. Initiation Type: Classify the initiators cue STRICTLY into one of these four categories:
- <verbal>: The cue is PURELY spoken (e.g., a verbal request/command).
- <gestural>: The cue is PURELY a physical gesture (e.g., reaching out, pointing) with NO speech.
- <verbal and gestural>: BOTH speech and a physical gesture are used simultaneously to demand the object.
- <implicit>: ALL OTHER cases. The handover occurs naturally due to situational context, mutual routine, or proactive offering, with NO explicit prior speaking or reaching out from the receiver.
5. Object Grounding: Look at the PROVIDED SINGLE REFERENCE IMAGE FRAME. Based on the context of the video clip and your answers to the previous questions, locate the precise object that is most likely going to be handed over in the future:
- Provide a short, precise object category name.
- Then, provide the bounding box of this object within the PROVIDED SINGLE REFERENCE IMAGE FRAME, specifically in the GIVERS view (determine who the giver is, left view or right view, based on your answer to the Delivering Flow’).
OUTPUT FORMAT:
Provide your final answer as a JSON list containing a single dictionary. You may think step-by-step beforehand, but do not provide any additional text in the output.
[
{
"handover happen offset": <int>,
"delivering flow": "<left to right> or <right to left>",
"initiator": "<left> or <right>",
"initiation type": "<verbal>, <gestural>, <verbal and gestural> or <implicit>",
"object category": "<string>",
"object bounding box": [<ymin>, <xmin>, <ymax>, <xmax>]
}
]
CRITICAL RULES:
- The bounding box coordinates MUST be normalized integers between [0, 1000] relative to the overall image dimensions.
Figure S11: Prompt for the Socially Conditioned Object Interaction Anticipation task.
You are provided with a series of sequential dual-view (side-by-side) images showing two egocentric views from two persons (left person and right person) over the span of 10 seconds, and a final single reference image frame.
Additionally, here is the transcribed speech of the two persons during the considered time period: "{transcript}"
TASK:
Carefully observe the human behaviors and social cues and answer the following questions sequentially:
1. Handover Timing: Approximately how many seconds after the reference image frame will the object handover between two persons occur? (Provide a decimal time offset).
2. Delivering Flow: Whether the object will be handed over from left person to right person or from right person to left person? (Choose STRICTLY: <left to right> or <right to left>).
3. Initiator: Who initiates the handover? Initiation refers to the first cue that prompts the handover to occur. (Choose STRICTLY: <left> or <right>).
4. Initiation Type: Classify the initiators cue STRICTLY into one of these four categories:
- <verbal>: The cue is PURELY spoken (e.g., a verbal request/command).
- <gestural>: The cue is PURELY a physical gesture (e.g., reaching out, pointing) with NO speech.
- <verbal and gestural>: BOTH speech and a physical gesture are used simultaneously to demand the object.
- <implicit>: ALL OTHER cases. The handover occurs naturally due to situational context, mutual routine, or proactive offering, with NO explicit prior speaking or reaching out from the receiver.
5. Object Grounding: Look at the PROVIDED SINGLE REFERENCE IMAGE FRAME. Based on the context of the video clip and your answers to the previous questions, locate the precise object that is most likely going to be handed over in the future:
- Provide a short, precise object category name.
- Then, provide the bounding box of this object within the PROVIDED SINGLE REFERENCE IMAGE FRAME, specifically in the GIVERS view (determine who the giver is, left view or right view, based on your answer to the Delivering Flow’).
OUTPUT FORMAT:
Provide your final answer as a JSON list containing a single dictionary. You may think step-by-step beforehand, but do not provide any additional text in the output.
[
{
"handover happen offset": <int>,
"delivering flow": "<left to right> or <right to left>",
"initiator": "<left> or <right>",
"initiation type": "<verbal>, <gestural>, <verbal and gestural> or <implicit>",
"object category": "<string>",
"object bounding box": [<ymin>, <xmin>, <ymax>, <xmax>]
}
]
CRITICAL RULES:
- The bounding box coordinates MUST be normalized integers between [0, 1000] relative to the overall image dimensions.
Figure S12: Prompt for the Collaborative Handover Prediction task.

Appendix S8 Annotation Details

Annotations cover 8,513 joint-attention events, 1,747 social interactions, and 586 handovers, at respective densities of 209.1, 42.9, and 14.4 events per hour. Task-specific browser interfaces, tutorials, review, and shared edge-case instructions support consistent labels.

Related references: [23].

Refer to caption
Figure S13: Annotation interface for the Joint Attention Estimation Task. The task-specific annotation interface is intuitive, accessible over the internet and synchronizes outputs to a central database, making annotation and supervision easily scalable. Transcripts are provided as subtitles to facilitate and speed up annotation.
Refer to caption
Figure S14: Annotation interface for the Socially Conditioned Object Interaction Anticipation task. The task-specific annotation interface is intuitive, accessible over the internet and synchronizes outputs to a central database, making annotation and supervision easily scalable. Transcripts are provided as subtitles to facilitate and speed up annotation.
Refer to caption
Figure S15: Annotation interface for the Collaborative Handover Prediction task. The task-specific annotation interface is intuitive, accessible over the internet and synchronizes outputs to a central database, making annotation and supervision easily scalable.
Table S7: Annotation statistics for the benchmark tasks. For each task we report the number of recordings containing annotations, the total duration of the corresponding streams, the number of annotated interaction events, and the resulting annotation density measured as events per hour of video.
Task # Events Events / hour
Socially Conditioned Object Interaction 1,747   42.9
Collaborative Handover Prediction   586   14.4
Joint Attention Detection 8,513 209.1

Appendix S9 Ethical Considerations

Participants consented to release of identifiable video; faces remain visible because gaze and expression are central signals. The project received institutional ethics approval. Figure S16 reproduces the consent form. The source supplement also contains the full annotation instructions.

Original figure
Figure S16: Consent form.

Annotation guide: joint attention

Annotators identify sustained shared focus using action, dialogue, gestures, and gaze. When cues conflict, physical interaction takes priority over explicit verbal or gestural intent, then gaze overlays. Mark the interval, draw one object box per view, and assign object, cue, and confidence labels. The original confidence table and illustrative output schema are retained below; the schema is source pseudocode rather than executable JSON.

{
"start_time": 12.345,        // seconds 
  "end_time": 12.400,        // seconds  
  "bbox_left": [x, y, w, h],    // pixel coordinates for helper view on the left 
  "bbox_right": [x, y, w, h],    // pixel coordinates for leader view on the right 
  "object_label": "knife",     // gazed object category 
  “cue type”:  [1, 3]. // used cues, multi-choice, [1: gaze, 2: action, 3: gestures, 4: 
communication] 
  "confidence": 0.0-1.0,       // annotator confidence (optional) 
}
Original annotation guide table, source page 43, panel 1.
Original annotation guide table, source page 43, panel 1.

Annotation guide: socially cued actions

Label only the helper’s first goal-directed action caused by the leader’s cue, beginning within ten seconds of cue completion. Ignore unrelated preparatory actions. The start frame follows cue completion and must show the object continuously until action begins; the prediction frame precedes obvious action. Draw the box on that prediction frame, label verb and noun, and select the cue tag. The tables preserve role definitions, timing rules, examples, and edge cases.

Original annotation guide table, source page 46, panel 1.
Original annotation guide table, source page 46, panel 1.
Original annotation guide table, source page 46, panel 2.
Original annotation guide table, source page 46, panel 2.
Original annotation guide table, source page 47, panel 1.
Original annotation guide table, source page 47, panel 1.
Original annotation guide table, source page 47, panel 2.
Original annotation guide table, source page 47, panel 2.
Original annotation guide table, source page 48, panel 1.
Original annotation guide table, source page 48, panel 1.
Original annotation guide table, source page 48, panel 2.
Original annotation guide table, source page 48, panel 2.
Original annotation guide table, source page 49, panel 1.
Original annotation guide table, source page 49, panel 1.
Original annotation guide table, source page 49, panel 2.
Original annotation guide table, source page 49, panel 2.
Original annotation guide table, source page 49, panel 3.
Original annotation guide table, source page 49, panel 3.
Original annotation guide table, source page 50, panel 1.
Original annotation guide table, source page 50, panel 1.
Original annotation guide table, source page 51, panel 1.
Original annotation guide table, source page 51, panel 1.

Annotation guide: handovers

Mark the transfer instant when both people touch the object, transfer direction, prediction timestamp, object box and category, initiation cue, and initiator. Select a prediction frame before reaching starts, with the object fully visible. The initiator is whoever first signals the need, regardless of who gives the object. The screenshots preserve the interface’s timestamp, add, and delete operations.

Handover annotation interface, source page 52, screenshot 1.
Handover annotation interface, source page 52, screenshot 1.
Handover annotation interface, source page 52, screenshot 2.
Handover annotation interface, source page 52, screenshot 2.
Handover annotation interface, source page 53, screenshot 1.
Handover annotation interface, source page 53, screenshot 1.
Handover annotation interface, source page 53, screenshot 2.
Handover annotation interface, source page 53, screenshot 2.