Naturalistic Social Dyads Assessment In Free Play: A Multi-view Framework For Recognizing Co-located Eye Contact Between Children and Their Assessors
Abstract
Mutual eye contact is a clinically relevant behavioral marker in social engagement assessments. Existing clinical tools often rely on manual observation, which may introduce subjective biases. No validated video-derived algorithm currently exists for objectively recognizing clinically defined eye contact behaviors in multi-individual free play settings. This study introduced a novel multi-view deep learning framework for automated recognition of clinically defined eye contact behaviors between children and assessors during free play. We investigated whether multi-view video recordings and spatial-temporal behavioral patterns improved recognition performance. The study included video recordings from 103 children diagnosed with autism (mean age = 6.96 years) and 33 typically developing children (mean age = 7.63 years). Mutual eye contact behaviors between children and assessors were analyzed during five-minute free play sessions, with caregivers present in the assessment room. The framework was trained using sparsely sampled spatial- and temporal-domain images from multiple fixed cameras capturing complementary views of the assessment environment. The fused single-view model achieved F1 scores of 0.85 [0.83, 0.88] for eye contact behavior and 0.90 [0.89, 0.91] for non-eye-contact behavior. Multi-view fusion improved performance to 0.92 [0.89, 0.95] and 0.94 [0.93, 0.96], respectively. These findings suggest that multi-view video analysis can objectively identify clinically defined eye contact behaviors among multiple individuals during clinical free play assessments. The framework may support scalable behavioral coding and quantitative assessment of clinically relevant social behaviors consistent with established criteria.
Similar content being viewed by others
Introduction
Mutual eye contact is broadly acknowledged as an important indicator of social intention, emotion and intimacy in human social interactions [1]. It is an essential non-verbal communication behavior that can convey information about othersâ social attention [2], change in emotion states [3], intimacy [4, 5], and trust [5, 6]. While mutual eye contact has been employed as a broad marker of social behavior for a range of experimental studies, its absence early in child development has also been linked clinically to atypical social development [6], neurocognitive divergence [7, 8], and social withdrawal [9,10,11]. For example, reduced eye contact patterns are commonly observed in individuals with autism spectrum disorders (ASD), beginning early in life and continuing across development.
There are, however, significant limitations to current methods for collecting and quantifying clinically defined eye contact. Limitations are particularly evident in multi-individual social interactions, where eye contact must be interpreted within dynamic exchanges between dyads (e.g., among children, assessors, and caregivers), as well as across diverse clinical assessment environments. Current observational approaches to evaluating mutual eye contact in individuals with ASD typically rely on subjective assessment or manual coding of recorded interactions. Such approaches are time-consuming and subject to human error in coding [12, 13], making these resource-intensive research measures impractical for routine clinical use. Alternatively, studies have used questionnaire-based measures such as the Social Responsiveness Scale-2 [14] to assess childrenâs social behavior. However, these measures can be subject to reporter biases when completed by individuals themselves, parents, teachers, or caregivers. The development of objective, video-derived behavioral markers of clinically defined eye contact is therefore critical for reducing the burden and variability of manual observation and supporting quantitative assessment of social engagement in ASD.
For the study of child social interaction, free play tasks involving children, their caregivers or an independent assessor is typically used. In clinical assessment contexts, free play allows observation of spontaneous interactions with other people and objects in the surrounding environment [15]. In clinically administered semi-structured free play tasks, measurement of eye contact has traditionally relied on observing video-recorded sessions to identify eye contact episodes and manual coding to determine the duration of each observed interaction [5, 16].
Several studies have used controlled or simulated interactive environments to support more efficient assessment [17, 18]. In these settings, eye-tracking systems are commonly used to record eye movements [19] and quantify facial responses and gaze fixation on predefined video stimuli [20]. Other studies have introduced robot-assisted treatment mechanisms [21], implementing programmable or remotely controlled robots to engage participants in social interactions and respond to their reactions. However, these methods may not fully capture spontaneous multi-person behaviors during clinical free play because they often rely on constrained frontal views or predefined interaction structures.
In recent years, wearable technology has gained popularity in clinical research for recording gaze-related behaviors during human-to-human social interactions [22, 23]. One example is the use of an egocentric camera embedded in wearable glasses to capture the gaze direction of child participants [24]. This approach can effectively capture frontal facial information during face-to-face interaction [25], addressing the need for close-view facial information. However, because wearable glasses move with the wearer, hardware calibration and stable gaze estimation can be challenging. In addition, wearable devices introduce equipment to the face, potentially obscuring facial information [26]. Such devices may also be challenging to use with some people, including young children or individuals with sensory sensitivities to glasses [27, 28]. Recent non-wearable, deep-learning-based eye-tracking approaches have addressed some limitations of wearable systems by using small table-mounted dual cameras to estimate gaze towards facial regions and mutual eye contact during face-to-face interactions [29]. These developments demonstrate the value of camera-based approaches for quantifying gaze-related behaviors in structured face-to-face interactions.
Building on this broader movement toward non-intrusive video-based assessment, this study examines naturalistic social interactions among children, assessors, and caregivers during a hands-free, semi-structured free play task using multiple fixed cameras. In this study, ânaturalisticâ refers to spontaneous, semi-structured free play interactions within a clinical assessment setting, rather than in an uncontrolled everyday home environment. We introduce a synergistic multi-modal, multi-view deep learning framework for recognizing clinically defined mutual eye contact behaviors among multiple individuals in a free play clinical setting. The framework is designed to provide an objective, video-derived behavioral marker based on spatial and temporal information from synchronized multi-view recordings, rather than to measure gaze vectors, ocular fixation, or gaze angle directly. In many video-based assessment settings, participant movement is often restricted to maintain facial visibility and reduce occlusion, which can limit the observation of spontaneous social behaviors. Digitized behavior patterns extracted through deep learning algorithms may support objective analysis of behavioral markers and developmental monitoring in children.
In this study, we first evaluate the effectiveness of the proposed multi-view and multi-modal fusion framework for extracting spatio-temporal features of clinically defined mutual eye contact behaviors. We then evaluate model performance on participant-independent held-out test data and use Grad-CAM visualization to examine whether the trained model focuses on visually meaningful regions during classification, thereby enhancing transparency and interpretability. This research demonstrates the potential of objective, video-based behavioral recognition to support scalable coding and quantitative assessment of clinically relevant social behaviors, informing clinical assessments and future clinical decision-support tools.
Methods
Participants
ASD diagnoses were confirmed using the Diagnostic and Statistical Manual of Mental Disorders (DSM-5) criteria [30], with the Autism Diagnostic Observation Schedule 2nd edition [31] administered by research-reliable assessors alongside independent interviews. As a control, age-matched typically developing (TD) children were recruited from the broader community through newspapers and flyers and were assessed by telephone interview, including a structured clinical interview for DSM-5. No participants were receiving psychotropic medication.
Procedures
Each child participant completed a 60-minute semi-structured social interaction session in a clinical assessment room. Conversations and behavioral actions involving the child participant, assessor, and caregivers were recorded using four synchronized wall-mounted high-definition AXIS IP cameras (AXIS Communications; 1920 Ă 1080 pixels; 25 Hz), positioned to capture complementary views of the assessment room. Video synchronization was performed using ObserverXT (Noldus). The overall data collection flow is shown in Fig. 1 (Step A), and an illustration of the room setup is provided in Fig. 2. The room was illuminated by standard indoor clinical lighting, with no additional task-specific lighting introduced during the assessment. All cameras remained fixed throughout the session at a height suitable for capturing both child and adult movements, and were integrated into the room setup to minimize interference with the child-assessor interaction.
The Free Play Task
This study focused on the free play task derived from the baseline social interaction session. The task was designed as a semi-structured clinical interaction to elicit spontaneous child-assessor engagement while allowing the child to move freely within the assessment room. Within this setting, the interaction was considered naturalistic because it was minimally scripted, children could select and manipulate toys, and their movement was unrestricted. Here, ânaturalisticâ refers to spontaneous free play embedded within a clinical assessment setting rather than an uncontrolled everyday home environment. The free play task began when the assessor removed the relevant toys from the storage box and placed them on the table or, if the child preferred to play on the floor, on a nearby blanket (see examples in Fig. 1 (Step D)). Children were instructed to play on their own for at least 3 min. The assessor could join the play spontaneously by interacting with the child or directing the childâs attention to various toys. If the child had difficulty engaging with the toys, parents could support the childâs interest and participation.
During the five-minute free play task, an additional observer in the adjacent room monitored the video recording and signaled the end of the session to the assessor by knocking on the door. The four-camera setup provided complementary views of the assessment room, maximizing visual coverage of the child throughout the interaction. After each session, a researcher identified and recorded the start time of the free play task in each camera recording. Using these recorded start times, the FFmpeg (version 4.2.4) library [32] was used to extract the standardized free play segment from each 60-minute social interaction session.
Eye Contact Definitions and Data Annotations
Based on the clinical coding criteria used in this study, an eye contact event was defined as the child directing visual attention toward the eyes of the co-located assessor while the assessor maintained eye contact with the child. In this study, eye contact refers to a clinically defined, video-observed behavioral event rather than an eye-tracking-derived measurement of gaze vectors, ocular fixation, or gaze angle. Accordingly, the model was designed to recognize clinically defined eye contact behaviors from multi-view video recordings rather than directly measure physiological gaze. In ASD assessments, clinicians code the duration and frequency of gaze-related behaviors during social engagement, distinguishing brief or unintentional glances that do not involve social engagement from sustained mutual eye contact. Two independent coders with domain knowledge were trained by the lead clinical researchers. They used the MATLAB Video Labeler (version R2021a) to annotate the start and end times of each mutual eye contact event within the segmented videos (Fig. 1 (STEP B)). The free play task ended at the onset of the still-face paradigm. If the still-face paradigm occurred within a standardized video segment, coders marked its start time and excluded the corresponding samples from model training. When coders could not reliably distinguish mutual eye contact from general face-looking or ambiguous gaze behavior, the event was reviewed with the lead clinical researchers and resolved in accordance with the clinical coding definition.
Inter-rater reliability between the two trained coders was assessed using Cohenâs Kappa. Mutual eye contact events were independently annotated across 33 free play sessions in the TD group and 16 sessions in the ASD group. Inter-rater reliability was moderate to high, with a mean Cohenâs Kappa of 0.72 (SD = 0.25) for the TD group and 0.84 (SD = 0.15) for the ASD group. All video annotations were subsequently cross-checked by the lead clinical researchers to resolve any remaining discrepancies before model development.
Behavior Pattern Extraction for Model Development
The segmented free play videos were further trimmed into two behavior categories: âEye Contactâ and âOther.â âEye Contactâ clips corresponded to clinically defined eye contact events identified during free play, whereas âOtherâ clips corresponded to video segments without annotated social interaction events. The Python library DenseFlow [33] was used to decode each trimmed clip into RGB frames representing the true color of the image, and optical flow images [34, 35] presenting motion patterns between consecutive frames. The dataset consisted of 1,726 clips of mutual eye contact events (42,110 RGB frames with 40,384 optical flow images in each axis direction) and 2,435 clips of âOtherâ events (3,849,732 RGB frames with 3,847,297 optical flow images in each axis direction). This distribution reflects the natural imbalance in eye contact behaviors during free play, in which most recorded time did not involve mutual eye contact.
Model development used 5-fold stratified cross-validation with participant-level partitioning. In each fold, a predetermined 60% of unique participants were allocated to model training, 20% to validation, and 20% to held-out testing. To prevent data leakage, all clips, RGB frames, and optical flow images derived from the same participant were assigned exclusively to a single partition within each fold. Thus, no participant contributed data to more than one of the training, validation, or test sets within the same fold. Participants from both the ASD and TD groups were stratified across the partitions to maintain group representation and minimize potential sampling bias during model development. A detailed summary of the sample distribution is provided in Supplementary Table 1. Samples within each partition were randomly shuffled to prevent the model from learning consecutive behaviors from the same participant. The held-out test set was reserved for evaluating the final model trained in each fold, providing a robust assessment of performance on distinct stratified samples.
Proposed Deep Learning Framework for Eye Contact Recognition
The Multi-Modal Social Behavior Recognition (MSBR) framework, outlined in Fig. 2, was developed to recognize clinically defined mutual eye contact behaviors between assessors and children aged 3â12 years, with or without an ASD diagnosis, during the free play task, using synchronized multi-view videos. The MSBR framework is divided into two components: a spatial-domain branch (highlighted in blue) that employs multiple single-frame RGB images, and a temporal-domain branch (highlighted in grey) that uses segmented clips of consecutive optical flow images. The framework was based on the Temporal Segment Network architecture [36], with ResNet-50 [37] used as the backbone. Human activity recognition research uses information acquired from video- or sensor-based systems, and deep learning approaches have been used to reduce manual feature engineering by learning higher-level representations from sequential data on human behavior [38]. This supports the use of spatial and temporal feature-learning approaches for automated behavior recognition tasks. To reduce the influence of sample imbalance and variable clip lengths, the model branches were trained on randomly sampled, uniformly length sequences from long video clips. This sampling strategy increased exposure to positive eye contact events while preventing the model from being dominated by long âOtherâ segments. ResNet-50 was chosen because it outperforms many other backbones in action recognition tasks [39].
The MSBR framework integrates predictions from multi-view input videos by processing sequences of segments, each of which is decoded into two-dimensional images. During input data pre-processing (Fig. 2), each video clip was decoded into RGB frames and optical flow images, assigned its corresponding action class and camera-view labels, and then shuffled. Random shuffling was used to improve model generalizability and reduce the risk of overfitting to irrelevant sequential features. Data augmentation was introduced to increase the diversity of the training samples and reduce the risk of severe overfitting associated with the limited dataset.
Data augmentation techniques, including MultiScaleCrop (with scales 1, 0.875, 0.75, 0.66) for the RGB branch (spatial domain CNN branch) and RandomResizedCrop for the optical flow branch (temporal domain CNN branch), were applied only to the training data. All images were then resized to 224 Ă 224 pixels to match the input requirements of the ResNet-50 backbone. During training, images were further horizontally flipped with a probability of 50% to enhance data diversity, followed by batch normalization to scale the image features uniformly, supporting faster model convergence. For the validation and test sets, only resizing and normalization were applied, ensuring that model performance was evaluated on unseen participant data without augmentation.
Implementation Details
The proposed method was implemented in PyTorch and trained on two 11GB NVIDIA GeForce RTX 2080Ti GPUs. Hyperparameters were selected based on validation set performance within the 5-fold participant-level cross-validation framework. Held-out test sets were not used for hyperparameter tuning or model selection. For the spatial-domain CNN branch, the final configuration used a learning rate of 3 Ă 10â 5, a batch size of 10, a dropout ratio of 0.1, and 10 RGB images extracted per segment at one-frame intervals. For the temporal-domain CNN branch, the best-performing configuration used a learning rate of 8 Ă 10â 6, a batch size of 32, a dropout rate of 0.6, and three input clips per video segment, each containing four consecutive optical flow frames sampled at one-frame intervals. Adam optimization was used with a CosineAnnealing learning rate scheduler to support stable convergence and reduce manual learning rate tuning. Training was conducted for up to 300 epochs, using categorical cross-entropy loss for model optimization.
Single-View Feature Learning
During the single-modal prediction phase (Fig. 2), the framework randomly selected equally spaced RGB images and consecutive optical flow images from each single-view video segment for input to the spatial- and temporal-domain CNN branches, respectively. The spatial domain CNN branch generated predictions for each of the selected n RGB frames (\(\:{\varvec{M}}_{\varvec{i}}\)). Frame-level prediction scores were then averaged to produce the final prediction for the input video (FRGB[each-view]). The temporal domain CNN branch processed equally spaced sequences of consecutive optical flow images (\(\:{\varvec{P}}_{\varvec{i}}\)) of length \(\:\varvec{l}\). Similarly, segment-level prediction scores were averaged to produce the prediction for the same input video (FFlow[each-view]). Strict temporal alignment between the spatial- and temporal-domain features was not required because the consensus predictions were derived from the same input video segment.
Multi-View Fusion
Behavioral event predictions from multiple camera views were obtained by aggregating segment-level predictions from individual views. To ensure temporal alignment, all video streams were synchronized during data collection using the commercial ObserverXT software platform. The multi-view fusion process analyzed video segments aligned to the same timestamps. The input video segments are denoted by the function \(\:{\varvec{H}}_{\varvec{c}}\), where âcâ represents each camera view. Spatial domain (\(\:{\varvec{M}\varvec{S}\varvec{B}\varvec{R}}_{\varvec{R}\varvec{G}\varvec{B}}\left({\varvec{V}}_{\varvec{s}\varvec{e}\varvec{g}\varvec{m}\varvec{e}\varvec{n}\varvec{t}}\right)\)) (1) and temporal domain (\(\:{\varvec{M}\varvec{S}\varvec{B}\varvec{R}}_{\varvec{F}\varvec{l}\varvec{o}\varvec{w}}\left({\varvec{V}}_{\varvec{s}\varvec{e}\varvec{g}\varvec{m}\varvec{e}\varvec{n}\varvec{t}}\right)\)) (2) predictions from each view were averaged to perform late fusion across all camera views (\(\:{\varvec{M}\varvec{S}\varvec{B}\varvec{R}}_{\varvec{F}\varvec{u}\varvec{s}\varvec{i}\varvec{o}\varvec{n}}\left({\varvec{V}}_{\varvec{s}\varvec{e}\varvec{g}\varvec{m}\varvec{e}\varvec{n}\varvec{t}}\right)\)) (3), producing final prediction scores for each behavioral class. The purpose of multi-view fusion was to improve behavioral coverage and reduce the impact of occlusion across camera angles, rather than to alter the clinical structure of the free play task.
Evaluation of Eye Contact Behavior Recognition
Classification Performance
This study evaluated the performance of the proposed MSBR framework in identifying clinically defined mutual eye contact events during free play. Within each fold, models were optimized using the training set and selected using the validation set. Evaluation metrics included Top-1 accuracy, precision (positive predictive value), recall (sensitivity), and F1 score. The best-performing social behavior recognition model was then selected for each branch within each fold. Because mutual eye contact events occurred less frequently than âOtherâ behaviors, precision, recall, and F1 score were emphasized as the primary evaluation metrics. Accuracy and receiver operating characteristic (ROC) curves may be less informative in the presence of class imbalance, as they can yield an overly optimistic impression of performance when models are evaluated on imbalanced samples. The trained models were tested on participant-independent held-out test sets to evaluate model performance on unseen participants within the collected dataset, as illustrated in Fig. 1 (STEP D). Predictions from single-view and multi-view models for the same time-stamped segments were analyzed across three configurations: spatial domain, temporal domain, and fused spatial-temporal models. The F1 score represents the harmonic mean of precision and recall. It accounts for the âfalse positivesâ (the segment did not show eye contact, but the model predicted it as eye contact) and âfalse negativesâ errors (the model predicted the input segment as âOtherâ, but the segment did show eye contact) through these two measures. F1 scores were used to evaluate statistical differences between model settings.
Statistical Analysis
F1 scores for single-modal and multi-modal predictions were compared using two-sided paired-samples t-tests across the five cross-validation folds within each diagnostic group (ASD and TD). 95% confidence intervals (95% CIs) were calculated for the main performance metrics. For paired model performance comparisons across the five held-out test sets, effect sizes were calculated using Cohenâs d and interpreted as fold-level model effects rather than participant- or group-level clinical effects.
Grad-CAM Visualization
To aid model interpretability, we applied Gradient-weighted Class Activation Mapping (Grad-CAM [40]), a gradient-based localization technique, to generate visual heatmaps highlighting image regions that contributed most strongly to the MSBR frameworkâs predictions. Grad-CAM was used to visualize regions associated with the classification of frames containing mutual eye contact. These visualizations were interpreted qualitatively to assess whether behaviorally meaningful regions, such as faces, bodies, and interaction-relevant areas, were highlighted during prediction.
Results
Social interaction assessments of 136 separate videos that totaled 733.14 min of video recording were annotated and analyzed. These recordings included 103 children diagnosed with ASD and 33 TD children aged 3â12 years. Mutual eye contact events between the child and assessor were identified in 77 children in the ASD group and 25 children in the TD group. The total duration of mutual eye contact events was 567.11 s overall (424.75 s in the ASD group and 142.36 s in the TD group). Across all included participants, the average duration of mutual eye contact events was 4.17 s (SD = 6.80), and the average frequency was 3.94 events (SD = 5.27). The average duration of the free play task was 5.39 min (SD = 2.20).
Demographics
Descriptive statistics were calculated for age, gender, and Nonverbal Intelligence Quotient (NVIQ) for the sample (see Table 1). Independent-samples t-tests showed no statistically significant differences in age (t(134) = 1.245, p = 0.215 > 0.05). However, statistically significant differences were found in NVIQ (t(70.895) = 7.088, p < 0.001, two-sided), with 25 missing data points. There was a statistically higher proportion of males than females, which is expected given that it was drawn from the ASD population, which has a higher prevalence of males (Ď²(1, N = 136) = 4.976, p = 0.026 < 0.05). The samples were ethnically diverse; however, not all individual ethnicity data were available.
MSBR Performance in Eye Contact Recognition
Table 2 summarizes the average participant-independent held-out test performance of the MSBR framework for recognizing clinically defined eye contact and non-eye-contact (âOtherâ) behaviors in single-view and multi-view settings. Overall model performance was evaluated across the five stratified cross-validation folds. Top-1 accuracy, precision, recall, and F1 scores were calculated separately for the spatial, temporal, and fused spatial-temporal configurations (Table 2). Detailed performance results for each held-out test set are provided in Supplementary Tables 2â7, along with the AUC curves in Supplementary Fig. 1.
Across model configurations, multi-view prediction improved performance compared with single-view prediction. The strongest performance was achieved using fused spatial-temporal information across multiple views, with a Top-1 accuracy of 0.94 [0.91,0.96], and an F1 score of 0.92 [0.89, 0.95] for clinically defined eye contact behavior, and an F1 score of 0.94 [0.93,0.96] for non-eye-contact (âOtherâ) behavior.
To evaluate the magnitude and consistency of the multi-view improvement, paired fold-level comparisons were conducted using F1 scores across the five held-out test sets. As shown in Supplementary Table 8, multi-view prediction improved F1 scores compared with single-view prediction across the spatial, temporal, and fused model configurations. For the fused spatial-temporal model, multi-view prediction increased the F1 score by a mean paired difference of 0.065 [0.053, 0.078] for eye contact behavior and 0.044 [0.034, 0.055] for non-eye-contact (âOtherâ) behavior. These differences corresponded to fold-level Cohenâs d values of 6.70 and 5.45, respectively. These large effect sizes reflect the very small variability in paired fold-level performance differences across the five cross-validation folds. They should not be interpreted as participant-level clinical effect sizes.
In addition, single-view and multi-view predictions were compared across feature modalities, as illustrated in Fig. 3. The spatial domain model using RGB frames performed better on average than the temporal domain model using optical images. Fusion of the spatial and temporal predictions, maintaining a spatial-to-temporal ratio of 1, achieved competitive or improved performance for both eye contact and non-eye-contact (âOtherâ) behaviors under the single-view and multi-view settings. Multi-view prediction performance was statistically significantly better than single-view performance across all modalities and diagnostic groups. No statistically significant differences in model performance were observed between the ASD and TD groups across modalities.
Discussion
Principal Results
In this study, we demonstrated the capability of deep learning algorithms to objectively recognize clinically defined mutual eye contact behaviors between multiple individuals in a semi-structured free play clinical assessment environment. Our proposed multi-view, multi-modal feature-learning framework provided a feasible video-derived approach for quantitatively parsing social engagement behaviors from long duration panoramic video recordings captured by fixed wall-mounted cameras. Rather than directly measuring gaze vectors or ocular fixation, the framework recognizes clinically defined eye contact behaviors using spatial and temporal information extracted from synchronized multi-view video recordings. While the labels do track continuous gaze movement, it does accurately capture the mutual gaze between dyads that is critical for observational assessments in experimental studies of social interaction and in clinical assessment of ASD. This approach offers a scalable complement to subjective manual coding by enabling automated recognition of heterogeneous movement patterns of mutual gaze during multi-person interactions.
Our findings show that the multi-view approach significantly outperformed single-view models across spatial and temporal feature modalities, demonstrating robust and reliable performance for both ASD and TD groups. This suggests that more reliable recognition of brief, spatially variable mutual eye contact behaviors can be achieved by integrating complementary camera views in clinical environments. The results further indicate that the framework can recognize clinically relevant eye contact behaviors across both ASD and TD groups, supporting its potential use as a scalable tool for behavior coding in clinical research and assessment settings. In particular, this framework has the potential to help clinicians quantify social engagement patterns associated with ASD, thereby supporting more detailed behavioral characterization and informing future clinical decision-support tools. In addition, the framework may contribute to a more comprehensive understanding of social interactions among multiple individuals in dyadic interaction and reduce the extensive labor costs associated with manual video annotation, making clinical assessments more efficient and scalable while maintaining the complexity of multi-person interactions.
Comparison with Prior Work
Previous literature has examined face-looking, mutual gaze, and social attention during free play and other structured social paradigms as early behavioral markers for neurodevelopmental conditions [41]. Other video-based approaches have proposed multi-cue features to assess social stress responses and attention to caregivers [42]. Recent deep learning studies have also demonstrated automated recognition of gaze-related behaviors using constrained camera views or wearable/egocentric recordings [24]. These approaches have provided important evidence that computer vision methods can support quantitative analysis of social behavior. However, these methods may limit direct analysis of social behaviors between interacting individuals.
Our study differs from prior single-view gaze estimation and eye-tracking approaches in several ways. First, the framework was designed to recognize clinically defined eye contact behaviors during multi-person free play interactions, rather than to estimate gaze direction from a constrained frontal view or an eye-tracking device. Second, the use of synchronized fixed cameras allowed the child, assessor, and caregiver(s) to remain in the assessment room without requiring wearable devices, minimizing potential participant distraction associated with introducing additional equipment, such as eye-tracking devices or robots. Third, the multi-view design improved behavior coverage and reduced the impact of occlusion across camera angles. Therefore, the benefit of the four-camera setup should be interpreted as improved visual coverage and robustness for behavioral recognition, rather than as evidence that adding cameras alone increases ecological validity. Within this clinical context, the free play task preserves important elements of spontaneous child-adult interaction, including toy play, variable body orientation, movement around the room, and unscripted social engagement. These features make the task more representative of clinical social assessment than highly constrained gaze-fixation paradigms.
Model Interpretability and Clinical Relevance
The model demonstrated robust performance despite the visual complexity of uncertain room occupancy during social interactions. This suggests that the learned video-derived digital features captured behaviorally relevant information across different camera angles and interaction contexts. The multi-view and multi-modal fusion results across diagnostic groups, with F1 scores exceeding 0.90, suggest competitive performance for automated behavioral coding of clinically defined eye contact behaviors. As shown in Fig. 4, Grad-CAM visualizations suggest that behaviorally meaningful visual regions were associated with the classification of eye contact and no-eye-contact behaviors. In the input convolutional layers, the highlighted regions captured broader textural and contextual information from the video recordings, including the shapes of individual bodies and the structure of the trial room. In the output convolutional layers, the highlighted regions became more concentrated around areas where relevant social behaviors occurred, particularly the faces, upper bodies, and interaction-relevant regions (e.g., objects or toys being manipulated) of the child and assessors. This pattern is consistent with the types of visual information that clinicians and trained coders consider when perceptually distinguishing mutual eye contact from its absence during clinical behavior coding.
However, these visualizations should be interpreted qualitatively. Grad-CAM does not demonstrate that the model directly estimated gaze vectors, ocular fixation, or eye angle. Grad-CAM saliency maps indicate image regions associated with model predictions but should not be interpreted as providing causal explanations of the modelâs decision-making process. Rather, these visualizations indicate that the modelâs classification decisions were associated with visually and behaviorally meaningful regions within the multi-view recordings. Future work could integrate fixed-camera behavioral recognition with eye-tracking or gaze-estimation methods to further examine the correspondence between clinically defined social behaviors and objective gaze patterns.
Furthermore, the paired samples t-test results provide promising evidence that the proposed framework can recognize clinically defined eye contact behaviors in complex multi-person clinical environments. These findings raise critical questions for future automated clinical assessment and treatment, including the generalizability of video-based behavioral assessment tools across populations and settings (remote areas, clinical environments, home settings); ethical considerations related to privacy and data security; and broader applications to other social behavior assessments. The development of such tools may support objective and comprehensive analysis of social behaviors while reducing the burden of fully manual video coding.
Limitations
Several limitations should be considered. First, although the proposed MSBR framework provides an objective video-derived behavioral marker for clinically defined mutual eye contact behaviors, it does not directly measure gaze vectors, ocular fixation, or gaze angle. The framework should therefore be interpreted as an automated behavioral recognition tool rather than an eye-tracking system. In some cases, video-based observation may not fully distinguish eye-to-eye contact from gaze toward the face, particularly when fine-grained ocular orientation is not visible from a given camera view. Future work should validate these clinically defined video-observed behaviors against eye-tracking or gaze-estimation methods.
Second, it is important to note that eye contact alone does not provide a full assessment or understanding of ASD. Integrating additional objective markers into the assessment process has advantages to enhance precision and accuracy, but such tools only represent one part of a comprehensive developmental evaluation. The dataset also contained an inherent imbalance between eye contact and non-eye-contact behaviors because mutual eye contact occurred relatively infrequently during free play interactions. The raw frame counts were highly imbalanced due to the longer duration of non-eye-contact periods. However, model training used sampled, fixed-length segments rather than all decoded frames, reducing the effective class imbalance during optimization. At the segment level, the effective class distribution was substantially less imbalanced (1,726 eye contact clips versus 2,435 non-eye-contact clips; approximately 1:1.4). In addition, precision, recall, and F1 score were emphasized because these metrics are more informative than accuracy under class imbalance. The observation that spatial domain models outperformed temporal domain models contrasts with findings reported in previous action recognition articles [36, 43]. This may be due to the transient nature of eye contact behaviors (small movement amplitudes and trajectories), which makes it challenging to extract representative temporal features. Future studies should investigate alternative temporal feature learning strategies and extend the framework to additional behavioral modalities, including gesture dynamics, body movement, speech and vocal features, audio-visual interaction patterns, and broader multimodal social interaction cues. Integrating these complementary signals may provide a more comprehensive characterization of social behavior and support the development of additional quantifiable digital behavioral markers.
Third, the framework was trained and evaluated within a specific four-camera clinical assessment setup and a child sample aged 3â12 years. Although participant-level held-out testing was used to reduce data leakage, external validation across independent sites was not available. Future studies should test generalizability across different room layouts, camera positions, lighting conditions, occlusions, age ranges, and clinical or home settings. Performance may also vary across clinical sites due to differences in participant populations, camera calibration procedures, recording configurations, and assessor behavior, which may introduce variation in both the visual inputs and the social interaction patterns observed by the model. The high demand for manual annotation also remains a challenge. Future studies should develop action localization algorithms or weakly supervised learning approaches to identify clinically relevant behaviors in longer unlabeled video recordings. These approaches may help locate key frames, reduce annotation requirements, and support quantitative analysis of social engagement behaviors using gold standard clinical assessments.
Finally, video-based behavioral recognition requires careful consideration of privacy, secure data governance, and deployment feasibility, including inference latency, memory usage, and real-time scalability. The RGB and optical flow-based models we proposed in this study each contained approximately 23.5 million learnable parameters, corresponding to 89.7 MiB of FP32 parameter storage per modality. This relatively modest model storage requirement suggests that deployment on GPU-enabled workstations or higher-performance edge-computing platforms may be feasible. For clinical assessments such as the free play tasks, the intended workflow does not necessarily require continuous real-time inference. Hardware-specific benchmarking will therefore be required in future clinical implementation studies to establish end-to-end inference latency in hospital environments, including data transfer, feature pre-processing, and model inference. In addition, because raw video recordings contain identifiable participant information, future deployment should consider privacy-preserving computational frameworks, such as the use of de-identified secondary features. Recent studies on secure and decentralized learning [44] and electronic medical record security [45] highlight potential technical directions for protecting sensitive healthcare and video-derived data in AI systems. It is also important to evaluate the systemâs adaptability to resource-limited settings, considering factors such as camera-view selection, image-resolution requirements, and ease of use.
Conclusion
In this paper, we developed a multi-view, multi-modal feature learning framework for recognizing clinically defined mutual eye contact behaviors between children and assessors during a semi-structured clinical free play assessment. The framework achieved strong participant-independent held-out test performance, with multi-view fusion improving recognition compared with single-view models. These findings suggest that synchronized fixed-camera recordings can provide objective video-derived behavioral markers for scalable coding of clinically relevant social behaviors. Grad-CAM visualizations indicated that behaviorally meaningful regions, particularly facial and interaction-relevant areas, were associated with the modelâs classification of mutual eye contact behaviors. This pattern is consistent with the visual information considered by clinicians and trained coders during behavioral assessments. However, the framework was designed for multi-view social behavior recognition rather than to evaluate video-derived facial information as a biomarker for detecting neurodevelopmental conditions. Future work should validate the approach across independent clinical sites, different recording configurations, and complementary gaze-measurement technologies. Future studies may also examine whether video-derived facial and social behavioral representations extracted during free play interactions can contribute to diagnostic decision support.
Data Availability
No datasets were generated or analysed during the current study.
References
Del Prette ZAP, Del Prette A. Social Competence and Social Skills: A Theoretical and Practical Guide. Cham: Springer; 2021.
Tomasello M, Carpenter M, Call J, Behne T, Moll H. Understanding and sharing intentions: The origins of cultural cognition. Behav Brain Sci. 2005;28(5):675â91.
Spence SH. Social skills training with children and young people: Theory, evidence and practice. Child Adolesc Mental Health. 2003;8(2):84â96.
Croes EAJ, Antheunis ML, Schouten AP, Krahmer EJ. The role of eye-contact in the development of romantic attraction: Studying interactive uncertainty reduction strategies during speed-dating. Comput Hum Behav. 2020;105:106218.
Kleinke CL. Gaze and eye contact: a research review. Psychol Bull. 1986;100(1):78.
Senju A, Vernetti A, Ganea N, Hudry K, Tucker L, Charman T, et al. Early social experience affects the development of eye gaze processing. Curr Biol. 2015;25(23):3086â91.
Senju A, Johnson MH. The eye contact effect: mechanisms and development. Trends Cogn Sci. 2009;13(3):127â34.
Kuboshita R, Fujisawa TX, Makita K, Kasaba R, Okazawa H, Tomoda A. Intrinsic brain activity associated with eye gaze during mother-child interaction. Sci Rep. 2020;10(1):18903.
Milne L, Greenway P, Guedeney A, Larroque B. Long term developmental impact of social withdrawal in infants. Infant Behav Dev. 2009;32(2):159â66.
Buitelaar JK. Attachment and Social Withdrawal in Autism: Hypotheses and Findings. Behaviour. 1995;132(5â6):319â50.
Guedeney A, Marchand-Martin L, Cote SJ, Larroque B. Perinatal risk factors and social withdrawal behaviour. Eur Child Adolesc Psychiatry. 2012;21(4):185â91.
White J, Hegarty J, Beasley N, EYE CONTACT AND OBSERVER. BIAS: A RESEARCH NOTE. Br J Psychol. 1970;61(2):271.
Knight DJ, Langmeyer D, Lundgren DC. Eye-Contact, Distance, and Affiliation: The Role of Observer Bias. Sociometry. 1973;36(3):390â401.
Constantino JN, Gruber CP. Social responsiveness scale: SRS-2: Western psychological services Torrance. CA. 2012.
Anderson A, Moore DW, Godfrey R, Fletcher-Flinn CM. Social skills assessment of children with autism in free-play situations. Autism. 2004;8(4):369â85.
Jongerius C, Hessels RS, Romijn JA, Smets EMA, Hillen MA. The Measurement of Eye Contact in Human Interactions: A Scoping Review. J Nonverbal Behav. 2020;44(3):363â89.
Caruana N, McArthur G, Woolgar A, Brock J. Simulating social interactions for the experimental investigation of joint attention. Neurosci Biobehavioral Reviews. 2017;74:115â25.
Pfeiffer UJ, Vogeley K, Schilbach L. From gaze cueing to dual eye-tracking: Novel approaches to investigate the neural correlates of gaze in social interaction. Neurosci Biobehavioral Reviews. 2013;37(10):2516â28.
Wieser MJ, Pauli P, Alpers GW, MĂźhlberger A. Is eye to eye contact really threatening and avoided in social anxiety?âAn eye-tracking and psychophysiology study. J Anxiety Disord. 2009;23(1):93â103.
Wang Q, Lu L, Zhang Q, Fang F, Zou XB, Yi L. Eye avoidance in young children with autism spectrum disorder is modulated by emotional facial expressions. J Abnorm Psychol. 2018;127(7):722â32.
Ricks DJ, Colton MB. Trends and considerations in robot-assisted autism therapy. IEEE Int Conf Robot Autom. 2010;4354-9.
Ye Z, Li Y, Fathi A, Han Y, Rozga A, Abowd GD, et al. editors. Detecting eye contact using wearable eye-tracking glasses. Ubicomp â12: The 2012 ACM Conference on Ubiquitous Computing: Association for Computing Machinery. 2012.
Noris B, Keller J-B, Billard A. A wearable gaze tracking system for children in unconstrained environments. Comput Vis Image Underst. 2011;115(4):476â86.
Chong E, Clark-Whitney E, Southerland A, Stubbs E, Miller C, Ajodan EL, et al. Detection of eye contact with deep neural networks is as accurate as human experts. Nat Commun. 2020;11(1):6386.
MĂźller P, Huang MX, Zhang X, Bulling A, editors. Robust eye contact detection in natural multi-person interactions using gaze and speaking behaviour. Proceedings of the 2018 ACM Symposium on Eye Tracking Research & Applications. Association for Computing Machinery. 2018.
Niehorster DC, Santini T, Hessels RS, Hooge IT, Kasneci E, NystrĂśm M. The impact of slippage on the data quality of head-worn eye trackers. Behav Res Methods. 2020;52(3):1140â60.
Perkovich E, Laakman A, Mire S, Yoshida H. Conducting head-mounted eye-tracking research with young children with autism and children with increased likelihood of later autism diagnosis. J Neurodevelopmental Disorders. 2024;16(1):7.
Ben-Sasson A, Hen L, Fluss R, Cermak SA, Engel-Yeger B, Gal E. A meta-analysis of sensory modulation symptoms in individuals with autism spectrum disorders. J Autism Dev Disord. 2009;39(1):1â11.
Thorsson M, Galazka MA, Ă
sberg Johnels J, Hadjikhani N. Influence of autistic traits and communication role on eye contact behavior during face-to-face interaction. Sci Rep. 2024;14(1):8162.
American Psychiatric Association. Diagnostic and Statistical Manual of Mental Disorders (5th ed.); 2013.
Lord C, Rutter M, DiLavore P, Risi S, Gotham K, Bishop S. Autism diagnostic observation scheduleâ2nd edition (ADOS-2). Volume 284. Western Psychological Corporation; 2012.
Tomar S. Converting video formats with FFmpeg. Linux J. 2006;2006(146):10.
Zhu Y, Newsam S, editors. DenseNet for dense flow. IEEE International Conference on Image Processing (ICIP); 2017 Sept. 2017. IEEE: IEEE; 2017.
Barron JL, Fleet DJ, Beauchemin SS. Performance of optical flow techniques. Int J Comput Vision. 1994;12(1):43â77.
Sun S, Kuang Z, Sheng L, Ouyang W, Zhang W, editors. Optical flow guided feature: A fast and robust motion representation for video action recognition. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2018.
Wang L, Xiong Y, Wang Z, Qiao Y, Lin D, Tang X, et al. editors. Temporal segment networks: Towards good practices for deep action recognition. Computer Vision â ECCV 2016; 2016 October: Springer.
He K, Zhang X, Ren S, Sun J, editors. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 27â30 June: IEEE.
Jaber TA, editor. Sensor based human action recognition and known public datasets a comprehensive survey. AIP Conference Proceedings: AIP Publishing LLC; 2023.
Chen C-FR, Panda R, Ramakrishnan K, Feris R, Cohn J, Oliva A, et al. editors. Deep analysis of cnn-based spatio-temporal representations for action recognition. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021.
Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D, editors. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. 2017 IEEE International Conference on Computer Vision (ICCV); 2017.
Franchak JM, Kretch KS, Adolph KE. See and be seen: Infant-caregiver social looking during locomotor free play. Dev Sci. 2018;21(4):e12626.
Tang C, Zheng W, Zong Y, Qiu N, Lu C, Zhang X, et al. Automatic identification of high-risk autism spectrum disorder: A feasibility study using video and audio data under the still-face paradigm. IEEE Trans Neural Syst Rehabil Eng. 2020;28(11):2401â10.
Kong Y, Fu Y. Human Action Recognition and Prediction: A Survey. Int J Comput Vision. 2022;130(5):1366â401.
Alzubaidi L, Jebur SA, Jaber TA, Mohammed MA, Alwzwazy HA, Saihood A, et al. ATD Learning: A secure, smart, and decentralised learning method for big data environments. Inform Fusion. 2025;118:102953.
Mohammed MA, Abdul Wahab HB. A novel approach for electronic medical records based on NFT-EMR. Int J Online Biomedical Eng. 2023;19(5):93â104.
Acknowledgements
We thank Ayesha Sadozai, Zahava Ambarchi, Qi Wu, and Taylor Lowry for their contributions to data annotation.
Funding
Open Access funding enabled and organized by CAUL and its Member Institutions. This study was supported by the Brain and Mind Centre Child Neurodevelopment and Mental Health team, the National Health and Medical Research Council (Project grants 1043664 and 1125449) and the Bupa Health Foundation independent research grants program.
Author information
Authors and Affiliations
Contributions
C.S., A.J.G., A.M., W.O., and L.Z. participated in the design of the study; R.T., and E.E.T. collected the assessment data, C.S. contributed to all aspects of coding and data annotation, and H.Z. helped with data visualisation. All authors read and approved the manuscript.
Corresponding authors
Ethics declarations
Competing interests
This paper has data that is the subject of an ongoing patent filing.
Additional information
Publisherâs Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary Information
Below is the link to the electronic supplementary material.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the articleâs Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the articleâs Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
About this article
Cite this article
Sun, C., Guastella, A.J., Ouyang, W. et al. Naturalistic Social Dyads Assessment In Free Play: A Multi-view Framework For Recognizing Co-located Eye Contact Between Children and Their Assessors. Cogn Comput 18, 110 (2026). https://doi.org/10.1007/s12559-026-10659-7
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s12559-026-10659-7
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.