CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
Alexey Gavryushin* † Dingxi Zhang* Zhao Huang* · Alexandros Delitzas Jiaqi Chen Ben Ellis Cedric Zöllner · Manthan Patel Manuel Kaufmann Marc Pollefeys Xi Wang
The overview figure aligns collaborative recordings, scene geometry, and the three prediction tasks.
Figure 1: The CoMind dataset captures 41 hours of human-human collaboration across 80 sessions, combining egocentric/exocentric videos, audio, gaze, and high-quality 3D scene and object scans, as well as rich annotations for three novel social reasoning tasks.
At a glance
CoMind records unscripted collaborative cooking with synchronized first- and third-person views, audio, gaze, hand tracking, and 3D scans. Three benchmarks test joint attention, socially triggered action anticipation, and handovers before reaching begins. Current vision-language models struggle especially with localization; fine-tuning on CoMind improves several social and action predictions.
1 Introduction
Theory of Mind in physical collaboration requires inferring another person’s attention, intent, and need for assistance from evolving social cues. CoMind supplies aligned observations of both participants and their environment, moving beyond isolated actions or text-only belief questions.
Existing datasets emphasize individual activity, short interactions, conversation, or handover mechanics. CoMind combines long collaborative sessions with gaze, dialogue, object grounding, and social annotations. Its tasks ask about shared attention and impending assistance before the physical interaction is already underway.
Table 1: Overview of egocentric datasets.
For Modality,
denotes video,
denotes gaze,
denotes IMU,
denotes semi-dense point cloud,
denotes dense 3D scene scans,
denotes 3D object scans,
denotes hand tracking. Hours (Ego) refers to the cumulative duration of distinct egocentric video recordings.
Our dataset provides unprecedented depth by offering rich social cues and collaborative annotations alongside multi-view 3D grounding.
The dataset formalizes three complementary tasks and provides recordings, annotations, and shared spatial calibration.
3.1 Proposed Tasks
Joint attention predicts an attended object, its boxes in both views, and eliciting cue type from a target frame plus ten seconds of speech context. Social action anticipation predicts the helper’s next verb, object, box, and cue from ten seconds of history. Handover prediction estimates time, sender/receiver, initiator, object, box, and cue before reaching starts. These operational tasks probe social inference; they do not directly measure a complete theory of mind.
Figure 2: Joint Attention Estimation Task. From the audio or transcribed speech in a context window, together with a target frame (green), the model predicts the jointly attended object, its bounding box in the left and right views, and the social cue type (blue).
Figure 3: Example annotations. Frames show joint attention cues, gaze points, and object bounding boxes with labels.
Figure 4: Socially Conditioned Object Interaction Anticipation Task. From context frames (or video) and audio or transcribed speech together with a prediction frame (green), the model predicts the noun and verb of the action to be performed, the interacted object’s bounding box, and the social cue type (blue).
Figure 5: Example annotations. The prediction frame shows the ground-truth action verb, noun, object bounding box, and cue type(s).
Figure 6: Collaborative Handover Prediction Task. From context frames (or video) and audio or transcribed speech, together with a prediction frame (green), the model predicts the delivery flow (who hands to whom), the object category and bounding box in the view of the handing participant, the initiator, and the cue type (blue).
Figure 7: Example annotations. Frames show handover initiator, cue type, delivery flow, and object bounding boxes with labels.
3.2 Data Recording Protocol and Hardware Setup
Both participants wear calibrated Aria glasses. Two GoPro cameras capture front and rear views. A Leica scanner records each kitchen, while an Artec scanner captures shared utensils. Sessions are unscripted, monitored remotely, and recorded with participant consent.
Figure 8: Sensors used to create the CoMind dataset. Two Aria glasses record the scene from each participant’s view. Two GoPro Hero cameras capture exocentric views. The Leica BLK2Go and Artec 3D Leo are used for 3D scene and object scanning respectively.
Figure 9: 3D scans of objects used in the recordings. Participants work with objects that have been 3D-scanned.
3.3 Data processing
Aria streams are synchronized to sub-millisecond precision; external views use audio cross-correlation. Machine Perception Services produce camera poses, point clouds, gaze, and hand tracking in a shared frame. WhisperX generates transcripts. Kitchen scans align through image localization followed by point-to-plane registration, with manual verification.
In-house and professional annotators label the three tasks using dedicated interfaces and training instructions. Quality teams review external work, and a main contributor reviews all annotations.
Figure 10: Modalities in the CoMind dataset. The dataset provides synchronized multimodal recordings of collaborative sessions, including egocentric video with eye gaze and hand tracking, audio and IMU signals, exocentric camera views, camera trajectories and spatial mapping, dense 3D scans of the scene, and a set of high-quality 3D scans of commonly used objects.
3.5 Data Statistics
The dataset contains eighty paired recordings from 55 kitchens and 125 participants, with about 40 hours 43 minutes of sessions or 81 hours 26 minutes summed over two egocentric streams. Sessions average roughly thirty minutes. The hierarchy contains 281 fine, 76 intermediate, and fourteen coarse object categories. Appendix statistics report 8,513 joint-attention, 1,747 social-action, and 586 handover events; some main-text captions use different category or verb totals.
Figure 11: Histogram of recording lengths. Our lengthy recordings allow for the study of human-human collaborative work in long-context settings. In total, we record 41 hours of two-person resp. 81 hours of single-view data.
Figure 12: Participant gender and ethnicity distribution. We balance participants from Eastern and Western backgrounds, and maintain an equal gender distribution of our participants.
Figure 13: Action verb and action noun distribution of the Socially Conditioned Object Interaction Anticipation task. Our annotations feature a wide range of objects from 13 coarse (L3) and over 180 fine (L1) object categories, as well as 22 unique verbs.Figure 14: Annotation statistics for the Joint Attention Estimation task. The CoMind dataset allows investigating multiple indicators of joint attention between humans on a broad variety of objects.Figure 15: Annotation statistics for the Collaborative Handover Prediction task. In addition to the category and bounding box of the handed-over object, we provide annotations about the Time-To-Handover (TTH), delivery flow direction, initiator and initiation cue for each handover event.
4 Baselines and Experiments
Training and test participants are disjoint; 25 videos provide test annotations. Models receive task-specific prompts and structured output requirements, usually five frames and transcripts from the preceding ten seconds. Baselines include pretrained proprietary and open models, fine-tuned Qwen3-VL, random sampling, and most-frequent predictions.
Localization is measured as the fraction of boxes exceeding 0.5 intersection over union, separately for each view. Gemini 3 Flash leads at 0.2164 and 0.2071. Fine-tuned Qwen3-VL-32B ranks second spatially; the 8B model has the best cue accuracy (0.4775). Object identification is easier than grounding the shared target in both views.
Table 2: Quantitative results on Joint Attention Estimation.Best, second-best, and third-best results are highlighted. and denote the fraction of predicted bounding boxes with with the ground-truth jointly attended object in the left and right egocentric streams, respectively.
Cue Type reports cue classification accuracy, and Cat. L1 measures the fraction of correct L1 object category predictions.
Models marked FT are fine-tuned on CoMind.
The fine-tuned 32B model leads in localization (0.2513), cue classification (0.5503), and action-verb accuracy (0.6508). Fine-tuning raises verb accuracy from 0.0873 to 0.6429 for 8B and from 0.1380 to 0.6508 for 32B. This indicates substantial benefit from training on temporal and social context.
Table 3: Quantitative results on Socially Conditioned Object Interaction Anticipation.Best, second-best, and third-best results are highlighted. denotes the fraction of predicted bounding boxes with with the ground-truth object to be interacted with.
Cue Type reports cue classification accuracy, and Act. Verb resp. Act. Noun (L1) measures the fraction of correct verb resp. L1 object category predictions.
Models marked FT are fine-tuned on CoMind.
Handover localization remains very weak: the best box-success fraction is 0.0313. Timing counts as correct within 0.25 seconds; the best timing score is 0.1317. Fine-tuned Qwen models lead on initiator (0.7286) and cue type (0.3714), but all methods leave large gaps in proactive assistance.
Table 4: Quantitative results on Collaborative Handover Prediction.Best, second-best, and third-best results are highlighted. denotes the fraction of predicted bounding boxes with with the ground-truth object.
Del. Flow, Initiator, and Init. Type report the fraction of correct predictions for the delivery flow, the initiating participant, and the initiation cue type, respectively.
Cat. (L1) measures the fraction of correct L1 object category predictions.
TTH reports the fraction of time-to-handover predictions within of the ground-truth.
Models marked FT are fine-tuned on CoMind.
CoMind provides complementary sensory evidence and benchmark labels for social collaboration. Fine-tuning improves prediction, while weak spatial grounding and handover anticipation leave substantial room for progress.
Acknowledgements
The work acknowledges Swiss and European research funding and institutional support.
Supplementary Material
The supplement retains qualitative comparisons, training settings, task prompts, annotation interfaces, and consent materials.
Appendix S1 Qualitative Examples
Figures S1–S5 compare pretrained and fine-tuned predictions with ground truth. They illustrate the benefits and remaining errors of task-specific training.
Figure S1: Baseline comparison for the Joint Attention Estimation task. We visualize and compare the predictions of Qwen3-VL 8B with those of Qwen3-VL 8B finetuned on the CoMind training set and Gemini 3 Flash, and provide the ground-truth solution given by our dataset. Finetuning on our training set improves the test set predictions.
Figure S2: Baseline comparison for the Socially Conditioned Object Interaction Anticipation task. We visualize and compare the predictions of Qwen3-VL 8B with those of Qwen3-VL 8B finetuned on the CoMind training set and Gemini 3 Flash, and provide the ground-truth solution given by our dataset. Finetuning on our training set improves the test set predictions.
Figure S3: Baseline comparison for the Collaborative Handover Prediction task. We visualize and compare the predictions of Qwen3-VL 8B with those of Qwen3-VL 8B finetuned on the CoMind training set and Gemini 3 Flash, and provide the ground-truth solution given by our dataset. Finetuning on our training set improves the test set predictions.
Figure S4: Further predictions on the Joint Attention Estimation task. We visualize the predictions of Qwen3-VL 8B finetuned on our training set in red, together with the ground-truth annotations in blue.
Figure S5: Further predictions on the Socially Conditioned Object Interaction Anticipation task. We visualize the predictions of Qwen3-VL 8B finetuned on our training set in red, together with the ground-truth annotations in blue.
Appendix S2 Downstream Applications and Future Tasks
The synchronized geometry supports future body-pose, object-pose, gaze-following, and 3D grounding tasks. The illustrated reconstructions are examples of potential derived annotations, rather than additional validated benchmark ground truth.
Figure S6: Pseudo ground-truth extraction with foundation models. Our dataset provides modalities well-suited for pseudo ground-truth extraction using foundation models.Figure S7: Gaze Following as a precursor to Joint Attention. In collaborative tasks, gaze following frequently serves as a vital geometric and social prior before joint attention is fully established. CoMind’s synchronous dual views capture the partner’s head orientation and line of sight (blue vectors). The specific gaze points of the leader and the helper are indicated by red and blue dots, respectively. This setup effectively grounds the target objects (bounding boxes) in the shared workspace prior to the interaction.Figure S8: Potential for 3D Spatial Reasoning. By aligning dense depth estimates with the sparse 3D correspondences provided by the Aria MPS, our 2D bounding box annotations can be geometrically lifted into the 3D space. This figure shows reconstructed 3D bounding boxes within the scene point cloud, demonstrating CoMind’s capability to support future 3D prediction tasks. Only representative static object bounding boxes shown for clarity.
Appendix S3 Finetuning Hyperparameters
Qwen3-VL fine-tuning uses LoRA rank 64, scale 128, learning rate 0.0002, and batch size four. Joint attention trains for three epochs; the other tasks train for five.
Table S5 checks whether the preceding ten-second transcript directly names the target object at each hierarchy level. Exact string matching excludes synonyms, paraphrases, and implicit references, so it measures explicit mentions only.
Table S5: Last-10s transcript matches to ground-truth object labels (rates).
For each annotation, we check whether the transcript in the prior 10 seconds
contains a direct string match to the ground-truth label at each level.
Any level counts a single match if any of the Level 1–3 ground-truth labels match.
Verb consolidation produces 21 canonical action classes. Nouns retain object identity at the finest level, group related functions at the middle level, and use fourteen broad categories at the top. Figure S9 preserves the complete hierarchy and parent–child relationships.
Figure S9: Object hierarchy tree for the CoMind Dataset. Boxes show L3 (left), L2 (middle), and L1 (right) categories; lines indicate parent–child relations. Colors are keyed to L3 groups.
Appendix S6 Input Ablations
Adding transcripts improves cue and object prediction; context frames improve verb and object prediction. The target frame remains available in every ablation, and prompts adjust to the supplied modalities.
Table S6: Ablation study on context frames and transcripts on the Socially Conditioned Object Interaction Anticipation task. denotes the fraction of predicted bounding boxes with with the ground-truth object to be interacted with. Cue Type reports cue classification accuracy, and Act. Verb resp. Act. Noun (L1) measures the fraction of correct verb resp. L1 object category predictions. Providing transcripts greatly improves cue type and act. noun accuracy. Context frames improve act. verb and act. noun accuracy.
Figure S12: Prompt for the Collaborative Handover Prediction task.
Appendix S8 Annotation Details
Annotations cover 8,513 joint-attention events, 1,747 social interactions, and 586 handovers, at respective densities of 209.1, 42.9, and 14.4 events per hour. Task-specific browser interfaces, tutorials, review, and shared edge-case instructions support consistent labels.
Figure S13: Annotation interface for the Joint Attention Estimation Task. The task-specific annotation interface is intuitive, accessible over the internet and synchronizes outputs to a central database, making annotation and supervision easily scalable. Transcripts are provided as subtitles to facilitate and speed up annotation.Figure S14: Annotation interface for the Socially Conditioned Object Interaction Anticipation task. The task-specific annotation interface is intuitive, accessible over the internet and synchronizes outputs to a central database, making annotation and supervision easily scalable. Transcripts are provided as subtitles to facilitate and speed up annotation.Figure S15: Annotation interface for the Collaborative Handover Prediction task. The task-specific annotation interface is intuitive, accessible over the internet and synchronizes outputs to a central database, making annotation and supervision easily scalable.
Table S7: Annotation statistics for the benchmark tasks. For each task we report the number of recordings containing annotations, the total duration of the corresponding streams, the number of annotated interaction events, and the resulting annotation density measured as events per hour of video.
Task
# Events
Events / hour
Socially Conditioned Object Interaction
1,747
42.9
Collaborative Handover Prediction
586
14.4
Joint Attention Detection
8,513
209.1
Appendix S9 Ethical Considerations
Participants consented to release of identifiable video; faces remain visible because gaze and expression are central signals. The project received institutional ethics approval. Figure S16 reproduces the consent form. The source supplement also contains the full annotation instructions.
Figure S16: Consent form.
Annotation guide: joint attention
Annotators identify sustained shared focus using action, dialogue, gestures, and gaze. When cues conflict, physical interaction takes priority over explicit verbal or gestural intent, then gaze overlays. Mark the interval, draw one object box per view, and assign object, cue, and confidence labels. The original confidence table and illustrative output schema are retained below; the schema is source pseudocode rather than executable JSON.
{
"start_time": 12.345, // seconds
"end_time": 12.400, // seconds
"bbox_left": [x, y, w, h], // pixel coordinates for helper view on the left
"bbox_right": [x, y, w, h], // pixel coordinates for leader view on the right
"object_label": "knife", // gazed object category
“cue type”: [1, 3]. // used cues, multi-choice, [1: gaze, 2: action, 3: gestures, 4:
communication]
"confidence": 0.0-1.0, // annotator confidence (optional)
}
Original annotation guide table, source page 43, panel 1.
Annotation guide: socially cued actions
Label only the helper’s first goal-directed action caused by the leader’s cue, beginning within ten seconds of cue completion. Ignore unrelated preparatory actions. The start frame follows cue completion and must show the object continuously until action begins; the prediction frame precedes obvious action. Draw the box on that prediction frame, label verb and noun, and select the cue tag. The tables preserve role definitions, timing rules, examples, and edge cases.
Mark the transfer instant when both people touch the object, transfer direction, prediction timestamp, object box and category, initiation cue, and initiator. Select a prediction frame before reaching starts, with the object fully visible. The initiator is whoever first signals the need, regardless of who gives the object. The screenshots preserve the interface’s timestamp, add, and delete operations.
[1]Y. Abu Farha, A. Richard, and J. Gall (2018)When will you do what?-anticipating temporal occurrences of activities.
In Proceedings of the IEEE conference on computer vision and pattern recognition,
pp. 5343–5352.
[2]R. B. Adams Jr, D. N. Albohn, and K. Kveraga (2017)Social vision: applying a social-functional approach to face and expression perception.
Current directions in psychological science26 (3), pp. 243–248.
[7]S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. (2025)Qwen3-vl technical report.
arXiv preprint arXiv:2511.21631.
[9]P. Banerjee, S. Shkodrani, P. Moulon, S. Hampali, S. Han, F. Zhang, L. Zhang, J. Fountain, E. Miller, S. Basol, et al. (2024)HOT3D: hand and object tracking in 3d from egocentric multi-view videos.
arXiv preprint arXiv:2411.19167.
[10]A. Bonnetto, H. Qi, F. Leong, M. Tashkovska, M. Rad, S. Shokur, F. Hummel, S. Micera, M. Pollefeys, and A. Mathis (2025)EPFL-smart-kitchen-30: densely annotated cooking dataset with 3d kinematics to challenge video and language models.
arXiv preprint arXiv:2506.01608.
[11]M. Cakmak, S. S. Srinivasa, M. K. Lee, J. Forlizzi, and S. Kiesler (2011)Human preferences for robot-human hand-over configurations.
In 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems,
pp. 1986–1993.
[12]M. Chang, G. Chhablani, A. Clegg, M. D. Cote, R. Desai, M. Hlavac, V. Karashchuk, J. Krantz, R. Mottaghi, P. Parashar, et al. (2024)PARTNR: a benchmark for planning and reasoning in embodied multi-agent tasks.
arXiv preprint arXiv:2411.00081.
[13]V. Chavan, Y. Imgrund, T. Dao, S. Bai, B. Wang, Z. Lu, O. Heimann, and J. Krüger (2025)IndEgo: a dataset of industrial scenarios and collaborative work for egocentric assistants.
In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track,
[14]J. Chen, G. Mittal, Y. Yu, Y. Kong, and M. Chen (2022)Gatehub: gated history unit with background suppression for online action detection.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,
pp. 19925–19934.
[15]Z. Chen, J. Wu, J. Zhou, B. Wen, G. Bi, G. Jiang, Y. Cao, M. Hu, Y. Lai, Z. Xiong, et al. (2024)ToMBench: benchmarking theory of mind in large language models.
In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),
pp. 15959–15983.
[16]E. Chong, Y. Wang, N. Ruiz, and J. M. Rehg (2020)Detecting attended visual targets in video.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),
pp. 5396–5406.
[18]D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al. (2018)Scaling egocentric vision: the epic-kitchens dataset.
In Proceedings of the European conference on computer vision (ECCV),
pp. 720–736.
[19]D. Damen, H. Doughty, G. M. Farinella, A. Furnari, J. Ma, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray (2022)Rescaling egocentric vision: collection, pipeline and challenges for epic-kitchens-100.
International Journal of Computer Vision (IJCV)130, pp. 33–55.
External Links: Link
[20]R. A. de Belen, G. Mohammadi, and A. Sowmya (2023)Temporal understanding of gaze communication with gazetransformer.
In NeuRIPS Workshop on Gaze Meets ML,
[23]A. Dutta and A. Zisserman (2019)The VIA annotation software for images, audio and video.
In Proceedings of the 27th ACM International Conference on Multimedia,
MM ’19, New York, NY, USA.
External Links: ISBN 978-1-4503-6889-6/19/10
[24]J. Engel, K. Somasundaram, M. Goesele, A. Sun, A. Gamino, A. Turner, A. Talattof, A. Yuan, B. Souti, B. Meredith, et al. (2023)Project aria: a new tool for egocentric multi-modal ai research.
arXiv preprint arXiv:2308.13561.
[25]L. Fan, W. Wang, S. Huang, X. Tang, and S. Zhu (2019)Understanding human gaze communication by spatio-temporal graph reasoning.
In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV),
pp. 5724–5733.
[26]A. Fathi, Y. Li, and J. M. Rehg (2012)Learning to recognize daily actions using gaze.
In Proceedings of the European Conference on Computer Vision (ECCV) Workshops,
pp. 314–327.
[27]R. Fu, D. Zhang, A. Jiang, W. Fu, A. Funk, D. Ritchie, and S. Sridhar (2025)Gigahands: a massive annotated dataset of bimanual hand activities.
In Proceedings of the Computer Vision and Pattern Recognition Conference,
pp. 17461–17474.
[28]A. Furnari and G. M. Farinella (2020)Rolling-unrolling lstms for action anticipation from first-person video.
IEEE transactions on pattern analysis and machine intelligence43 (11), pp. 4021–4036.
[29]A. Furnari and G. M. Farinella (2022)Towards streaming egocentric action anticipation.
In 2022 26th International Conference on Pattern Recognition (ICPR),
pp. 1250–1257.
[31]R. Girdhar and K. Grauman (2021)Anticipative video transformer.
In Proceedings of the IEEE/CVF international conference on computer vision,
pp. 13505–13515.
[32]K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al. (2022)Ego4d: around the world in 3,000 hours of egocentric video.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,
pp. 18995–19012.
[33]K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, K. Ashutosh, V. Baiyya, S. Bansal, B. Boote, et al. (2024)Ego-exo4d: understanding skilled human activity from first-and third-person perspectives.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,
pp. 19383–19400.
[34]Z. Guo, Y. Hou, P. Wang, Z. Gao, M. Xu, and W. Li (2023)FT-hid: a large-scale rgb-d dataset for first-and third-person human interaction analysis.
Neural Computing and Applications35 (2), pp. 2007–2024.
[35]E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022)LoRA: low-rank adaptation of large language models.
In ICLR,
[36]D. Huang, O. Hilliges, L. Van Gool, and X. Wang (2023)Palm: predicting actions through language models@ ego4d long-term action anticipation challenge 2023.
arXiv preprint arXiv:2306.16545.
[37]Y. Huang, M. Cai, and Y. Sato (2020)An ego-vision system for discovering human joint attention.
IEEE Transactions on Human-Machine Systems50 (4), pp. 306–316.
[38]Y. Huang, G. Chen, J. Xu, M. Zhang, L. Yang, B. Pei, H. Zhang, D. Lu, Y. Wang, L. Wang, and Y. Qiao (2024)EgoExoLearn: a dataset for bridging asynchronous ego- and exo-centric view of procedural activities in real world.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,
[39]Y. Jang, B. Sullivan, C. Ludwig, I. Gilchrist, D. Damen, and W. Mayol-Cuevas (2019)Epic-tent: an egocentric video dataset for camping tent assembly.
In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops,
pp. 0–0.
[40]B. Jia, Y. Chen, S. Huang, Y. Zhu, and S. Zhu (2020)Lemma: a multi-view dataset for le arning m ulti-agent m ulti-task a ctivities.
In European Conference on Computer Vision,
pp. 767–786.
[41]R. Khirodkar, A. Bansal, L. Ma, R. Newcombe, M. Vo, and K. Kitani (2023)EgoHumans: an egocentric 3d multi-human benchmark.
arXiv preprint arXiv:2305.16487.
[42]Y. Li, V. Veerabadran, M. L. Iuzzolino, B. D. Roads, A. Celikyilmaz, and K. Ridgeway (2025)Egotom: benchmarking theory of mind reasoning from egocentric videos.
arXiv preprint arXiv:2503.22152.
[43]S. Liu, S. Tripathi, S. Majumdar, and X. Wang (2022)Joint hand motion and interaction hotspots prediction from egocentric videos.
In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition,
pp. 3282–3292.
[44]Y. Liu, C. Zhang, R. Xing, B. Tang, B. Yang, and L. Yi (2024)Core4d: a 4d human-object-human interaction dataset for collaborative object rearrangement.
arXiv preprint arXiv:2406.19353.
[45]M. J. Marin-Jimenez, A. Zisserman, M. Eichner, and V. Ferrari (2014)Detecting people looking at each other in videos.
International Journal of Computer Vision (IJCV)106, pp. 282–296.
[46]E. V. Mascaró, H. Ahn, and D. Lee (2023)Intention-conditioned long-term human egocentric action anticipation.
In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,
pp. 6048–6057.
[47]C. McLean, M. Meendering, T. Swartz, O. Gabbay, A. Olsen, R. Jacobs, N. Rosen, P. de Bree, T. Garcia, G. Merrill, J. Sandakly, J. Buffalini, N. Jain, S. Krenn, M. Kumar, D. Markovic, E. Ng, F. Prada, A. Saba, S. Zhang, V. Agrawal, T. Godisart, A. Richard, and M. Zollhoefer (2025)Embody 3d: a large-scale multimodal motion and behavior dataset.
Technical ReportarXiv.
Note: arXiv preprintExternal Links: Link
[48]M. M. K. V. Medina (2021)Suárez p zisserman a laeo-net++: revisiting people looking at each other in videos.
IEEE transactions on pattern analysis and machine intelligence (TPAMI)10.
[50]H. Mittal, P. Morgado, U. Jain, and A. Gupta (2022)Learning state-aware visual representations from audible interactions.
Advances in Neural Information Processing Systems35, pp. 23765–23779.
[52]T. D. Ngo, B. Hua, and K. Nguyen (2023)Isbnet: a 3d point cloud instance segmentation network with instance-aware sampling and box-aware dynamic convolution.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,
pp. 13550–13559.
[56]V. Ortenzi, A. Cosgun, T. Pardi, W. P. Chan, E. Croft, and D. Kulić (2021)Object handovers: a review for robotics.
IEEE Transactions on Robotics37 (6), pp. 1855–1873.
[57]N. Osman, G. Camporese, P. Coscia, and L. Ballan (2021)Slowfast rolling-unrolling lstms for action anticipation in egocentric videos.
In Proceedings of the IEEE/CVF International Conference on Computer Vision,
pp. 3437–3445.
[58]R. Pasca, A. Gavryushin, M. Hamza, Y. Kuo, K. Mo, L. Van Gool, O. Hilliges, and X. Wang (2024)Summarize the past to predict the future: natural language descriptions of context boost multimodal object interaction anticipation.
In Conference on Computer Vision and Pattern Recognition 2024,
External Links: Link
[59]T. Perrett, A. Darkhalil, S. Sinha, O. Emara, S. Pollard, K. K. Parida, K. Liu, P. Gatti, S. Bansal, K. Flanagan, et al. (2025)Hd-epic: a highly-detailed egocentric video dataset.
In Proceedings of the Computer Vision and Pattern Recognition Conference,
pp. 23901–23913.
[61]X. Puig, K. Ra, M. Boben, J. Li, T. Wang, S. Fidler, and A. Torralba (2018)VirtualHome: simulating household activities via programs.
External Links: 1806.07011
[62]X. Puig, T. Shu, S. Li, Z. Wang, Y. Liao, J. B. Tenenbaum, S. Fidler, and A. Torralba (2020)Watch-and-help: a challenge for social perception and human-ai collaboration.
arXiv preprint arXiv:2010.09890.
[63]Z. Qi, S. Wang, C. Su, L. Su, Q. Huang, and Q. Tian (2021)Self-regulated learning for egocentric video activity anticipation.
IEEE Transactions on Pattern Analysis and Machine Intelligence.
[64]H. Qiu, Z. Shi, L. Wang, H. Xiong, X. Li, and H. Li (2025)EgoMe: a new dataset and challenge for following me via egocentric view in real world.
arXiv preprint arXiv:2501.19061.
[65]N. Rabinowitz, F. Perbet, F. Song, C. Zhang, S. A. Eslami, and M. Botvinick (2018)Machine theory of mind.
In International conference on machine learning,
pp. 4218–4227.
[66]A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever (2023)Robust speech recognition via large-scale weak supervision.
In International conference on machine learning,
pp. 28492–28518.
[69]M. N. Rizve, G. Mittal, Y. Yu, M. Hall, S. Sajeev, M. Shah, and M. Chen (2023)Pivotal: prior-driven supervision for weakly-supervised temporal action localization.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,
pp. 22992–23002.
[70]I. Rodin, A. Furnari, D. Mavroeidis, and G. M. Farinella (2022)Untrimmed action anticipation.
In International Conference on Image Analysis and Processing,
pp. 337–348.
[71]S. Rusinkiewicz and M. Levoy (2001)Efficient variants of the icp algorithm.
In Proceedings third international conference on 3-D digital imaging and modeling,
pp. 145–152.
[74]L. Schilbach, B. Timmermans, V. Reddy, A. Costall, G. Bente, T. Schlicht, and K. Vogeley (2013)Toward a second-person neuroscience.
Behavioral and Brain Sciences36 (4), pp. 393–414.
[75]N. Sebanz, H. Bekkering, and G. Knoblich (2006)Joint action: bodies and minds moving together.
Trends in Cognitive Sciences10 (2), pp. 70–76.
External Links: ISSN 1364-6613
[76]F. Sener, D. Chatterjee, D. Shelepov, K. He, D. Singhania, R. Wang, and A. Yao (2022)Assembly101: a large-scale multi-view video dataset for understanding procedural activities.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,
pp. 21096–21106.
[77]H. Shi, S. Ye, X. Fang, C. Jin, L. Isik, Y. Kuo, and T. Shu (2025)Muma-tom: multi-modal multi-agent theory of mind.
In Proceedings of the AAAI Conference on Artificial Intelligence,
[78]Y. Shi, B. Fernando, and R. Hartley (2018)Action anticipation with rbf kernelized feature mapping rnn.
In Proceedings of the European Conference on Computer Vision (ECCV),
pp. 301–317.
[79]M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox (2020)Alfred: a benchmark for interpreting grounded instructions for everyday tasks.
In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition,
pp. 10740–10749.
[80]J. Six and M. Leman (2015)Synchronizing Multimodal Recordings Using Audio-To-Audio Alignment.
Journal of Multimodal User Interfaces9 (3), pp. 223–229.
External Links: ISSN 1783-7677,
Document
[81]K. W. Strabala, M. K. Lee, A. D. Dragan, J. L. Forlizzi, S. Srinivasa, M. Cakmak, and V. Micelli (2013)Towards seamless human-robot handovers.
Journal of Human-Robot Interaction2 (1), pp. 112–132.
[82]O. Taheri, N. Ghorbani, M. J. Black, and D. Tzionas (2020)GRAB: a dataset of whole-body human grasping of objects.
In European conference on computer vision,
pp. 581–600.
[83]M. Tomasello, M. Carpenter, J. Call, T. Behne, and H. Moll (2005)Understanding and sharing intentions: the origins of cultural cognition.
Behavioral and Brain Sciences28, pp. 675–735.
[84]E. Villa-Cueva, S. Ahmed, R. Chevi, J. C. B. Cruz, K. Elzeky, F. Cristobal, A. F. Aji, S. Wang, R. Mihalcea, and T. Solorio (2025)MOMENTS: a comprehensive multimodal benchmark for theory of mind.
arXiv preprint arXiv:2507.04415.
[85]C. Vondrick, H. Pirsiavash, and A. Torralba (2016)Anticipating visual representations from unlabeled video.
In Proceedings of the IEEE conference on computer vision and pattern recognition,
pp. 98–106.
[86]X. Wang, T. Kwon, M. Rad, B. Pan, I. Chakraborty, S. Andrist, D. Bohus, A. Feniello, B. Tekin, F. V. Frujeri, et al. (2023)Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world.
In Proceedings of the IEEE/CVF International Conference on Computer Vision,
pp. 20270–20281.
[87]Z. Wang, X. Zhao, S. Stepputtis, W. Kim, T. Wu, K. Sycara, and Y. Xie (2024)HiMemFormer: hierarchical memory-aware transformer for multi-agent action anticipation.
arXiv preprint arXiv:2411.01455.
[88]N. Wiederhold, A. Megyeri, D. Paris, S. Banerjee, and N. Banerjee (2023)Hoh: markerless multimodal human-object-human handover dataset with large object count.
Advances in Neural Information Processing Systems36, pp. 68736–68748.
[89]J. Yang, S. Liu, H. Guo, Y. Dong, X. Zhang, S. Zhang, P. Wang, Z. Zhou, B. Xie, Z. Wang, et al. (2025)Egolife: towards egocentric life assistant.
In Proceedings of the Computer Vision and Pattern Recognition Conference,
pp. 28885–28900.
[90]X. Yang, D. Kukreja, D. Pinkus, A. Sagar, T. Fan, J. Park, S. Shin, J. Cao, J. Liu, N. Ugrinovic, M. Feiszli, J. Malik, P. Dollar, and K. Kitani (2026)SAM 3d body: robust full-body human mesh recovery.
arXiv preprint arXiv:2602.15989.
[91]Y. Yang, L. Piccinelli, M. Segu, S. Li, R. Huang, Y. Fu, M. Pollefeys, H. Blum, and Z. Bauer (2025)3D-mood: lifting 2d to 3d for monocular open-set object detection.
In ICCV,
[92]Z. Ye, Y. Li, Y. Liu, C. Bridges, A. Rozga, and J. M. Rehg (2015)Detecting bids for eye contact using a wearable camera.
In IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG),
Vol. 1, pp. 1–8.
[93]C. Zhang, C. Fu, S. Wang, N. Agarwal, K. Lee, C. Choi, and C. Sun (2024)Object-centric video representation for long-term action anticipation.
In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV),
pp. 6751–6761.
[94]H. Zhang, W. Du, J. Shan, Q. Zhou, Y. Du, J. B. Tenenbaum, T. Shu, and C. Gan (2023)Building cooperative embodied agents modularly with large language models.
arXiv preprint arXiv:2307.02485.
[95]S. Zhang, Q. Ma, Y. Zhang, Z. Qian, T. Kwon, M. Pollefeys, F. Bogo, and S. Tang (2022)Egobody: human body shape and motion of interacting people from head-mounted devices.
In European conference on computer vision,
pp. 180–200.
[96]X. Zhang, Y. Chen, S. Yeh, and S. Li (2025)Metamind: modeling human social thoughts with metacognitive multi-agent systems.
arXiv preprint arXiv:2505.18943.
[97]Q. Zhao, S. Wang, C. Zhang, C. Fu, M. Q. Do, N. Agarwal, K. Lee, and C. Sun (2024)AntGPT: can large language models help long-term action anticipation from videos?.
In The Twelfth International Conference on Learning Representations,