Multimodal behaviour analysis

Combining several observable channels — speech, face, gaze, gesture, body movement, object use — into a single temporally structured description of what a person did during an interaction. It is NIPG’s central technical competence and the method layer beneath most clinical applications in this ecosystem.

The ADOS-2 study

Multimodal Framework for Automatic Behavior Analysis of Children with Autism During ADOS-2 (DOI) is the flagship reference. It combines verbal and non-verbal behavioural analysis — speech, body and hand behaviour, object use, gaze-related signals, and other observable events — during a structured clinical interaction, and documents collaboration involving ELTE-related researchers, Rush, and Sorbonne (including Bruno Melício).

Safeguard. Event-detection performance must not be presented as the diagnostic accuracy of an automated autism test. Detecting that a gesture occurred is not the same as diagnosing a condition.

Composite AI

The approach layers rule-based reasoning over learned features to segment and interpret requests, gesture, object manipulation and eye contact. This matters for clinical acceptability: the learned components do perception, while the interpretable layer does the reasoning that a clinician can inspect and contest. It is also why the Department of Artificial Intelligence’s stated focus on composite AI is directly relevant to this work.

The channel inventory

ChannelNIPG capabilityEvidence
Speech and acousticsMFCC, eGeMAPS, acoustic emotion features, transcription, forced alignmentExordium (public code)
Facial behaviourFace and iris tracking, landmarks, action unitsPublic components
Gaze and head poseGaze estimation, head posePublic components, ADOS-2 study
Blink and eye stateTransformer-based blink detectionBlinkLinMulT (public code)
Body and 3D poseMulti-view skeleton reconstruction, edge pose estimationMVMB-NRSFM, DeepRehab
FusionLinear-complexity attention over fused audiovisual and text sequencesLinMulT (public code)

Full table with links at Capability portfolio.

Why fusion needs efficiency

Multimodal transformers over long interaction recordings are expensive. LinMulT — linear-complexity attention for fused audiovisual and text sequences — is NIPG’s answer, and it is what makes minute- or session-scale multimodal modelling tractable rather than theoretical.

Route into the clinical ecosystem

Multimodal analysis is currently an opportunity, not an agreed project, for the BPD setting. A camera-based or broader multimodal extension might later improve the precision of therapy monitoring, but there is no agreed study and no project plan. → Semmelweis / VIKOTE

NIPG · Capability portfolio · Argus Cognitive · Bruno Melício · Autism heterogeneity · Perceived personality