Multimodal behaviour analysis
Combining several observable channels — speech, face, gaze, gesture, body movement, object use — into a single temporally structured description of what a person did during an interaction. It is NIPG’s central technical competence and the method layer beneath most clinical applications in this ecosystem.
The ADOS-2 study
Multimodal Framework for Automatic Behavior Analysis of Children with Autism During ADOS-2 (DOI) is the flagship reference. It combines verbal and non-verbal behavioural analysis — speech, body and hand behaviour, object use, gaze-related signals, and other observable events — during a structured clinical interaction, and documents collaboration involving ELTE-related researchers, Rush, and Sorbonne (including Bruno Melício).
Safeguard. Event-detection performance must not be presented as the diagnostic accuracy of an automated autism test. Detecting that a gesture occurred is not the same as diagnosing a condition.
Composite AI
The approach layers rule-based reasoning over learned features to segment and interpret requests, gesture, object manipulation and eye contact. This matters for clinical acceptability: the learned components do perception, while the interpretable layer does the reasoning that a clinician can inspect and contest. It is also why the Department of Artificial Intelligence’s stated focus on composite AI is directly relevant to this work.
The channel inventory
| Channel | NIPG capability | Evidence |
|---|---|---|
| Speech and acoustics | MFCC, eGeMAPS, acoustic emotion features, transcription, forced alignment | Exordium (public code) |
| Facial behaviour | Face and iris tracking, landmarks, action units | Public components |
| Gaze and head pose | Gaze estimation, head pose | Public components, ADOS-2 study |
| Blink and eye state | Transformer-based blink detection | BlinkLinMulT (public code) |
| Body and 3D pose | Multi-view skeleton reconstruction, edge pose estimation | MVMB-NRSFM, DeepRehab |
| Fusion | Linear-complexity attention over fused audiovisual and text sequences | LinMulT (public code) |
Full table with links at Capability portfolio.
Why fusion needs efficiency
Multimodal transformers over long interaction recordings are expensive. LinMulT — linear-complexity attention for fused audiovisual and text sequences — is NIPG’s answer, and it is what makes minute- or session-scale multimodal modelling tractable rather than theoretical.
Route into the clinical ecosystem
Multimodal analysis is currently an opportunity, not an agreed project, for the BPD setting. A camera-based or broader multimodal extension might later improve the precision of therapy monitoring, but there is no agreed study and no project plan. → Semmelweis / VIKOTE
Related pages
NIPG · Capability portfolio · Argus Cognitive · Bruno Melício · Autism heterogeneity · Perceived personality