NIPG capability portfolio
The full technical inventory behind NIPG, with a per-row evidence pointer for each capability.
Read this first. Capability maturity varies a lot across these rows, from public repository to research prototype with code not verified. This is not a single validated product suite. The evidence column is what establishes readiness for any given row. See Capability readiness and Editorial safeguards.
What the portfolio does
NIPG develops multimodal perception and composite-AI systems that convert speech, video, human movement and environmental observation into temporally structured and interpretable information for healthcare, rehabilitation, social interaction and robotics.
The inventory
| # | Capability | Demonstrated functionality | Evidence / readiness |
|---|---|---|---|
| 1 | Multimodal AI pipelines | Audio, video and text preprocessing, feature extraction, GPU execution, caching and integration | Public code: Exordium |
| 2 | Speech and acoustic analysis | MFCC, eGeMAPS, acoustic emotion features, transcription, forced alignment, transcript assessment | Public code / third-party model integration: Exordium |
| 3 | Facial behaviour and gaze | Face and iris tracking, facial landmarks, gaze, head pose, action units | Public components; clinical research (ADOS-2) |
| 4 | Blink and eye-state recognition | Transformer-based blink detection, efficient inference, live demonstration | Public code: BlinkLinMulT |
| 5 | Efficient multimodal transformers | Linear-complexity attention for fused audiovisual and text sequences | Public code: LinMulT |
| 6 | Affect and personality modelling | Multimodal sentiment and Big Five trait estimation; experimental rather than clinical assessment | Public code: PersonalityLinMulT |
| 7 | Composite AI for social interaction | Rule-based reasoning over learned features to segment and interpret requests, gesture, object manipulation and eye contact | Published ADOS-2 study; full clinical pipeline not verified public |
| 8 | Label-efficient video segmentation | Object-instance tracking and segmentation with sparse scribble supervision | Public research code: Cluster2Former |
| 9 | Multi-view, multi-person 3D pose | 3D skeleton reconstruction from uncalibrated cameras without 3D ground-truth supervision | Publication and public code: MVMB-NRSFM |
| 10 | Skeleton-guided view synthesis | Novel-view generation guided by skeleton information from single imagery | Skel3D: repository exists; models/instructions incomplete |
| 11 | Semantic 3D reconstruction | RGB-D segmentation, visual SLAM, cross-view semantic transfer, human–robot viewpoints | Public research pipeline: Semantic Matching |
| 12 | Temporally coherent semantic SLAM | Memory-efficient, temporally consistent 3D semantic mapping from video | 2025 VGGT-based SLAM preprint |
| 13 | Ambient-intelligence rehabilitation | Home scanning, neural scene representations, navigation, exercise/camera placement, avatar feedback | AIRS demonstration |
| 14 | Edge human pose estimation | 23-keypoint body/face/foot estimation with Edge TPU deployment pathway | Public legacy research code: DeepRehab |
| 15 | Animal tracking and behaviour | Video tracking of similar animals, occlusion-tolerant behavioural extraction | Public research code: rat_tracking |
| 16 | Embodied human–AI interaction | AR virtual characters, gesture-based nonverbal interaction with LLMs | Keep Gesturing demonstration |
| 17 | Interpretable signal processing | Parameterized wavelets, variable projection, interpretable neural signal features for ECG | 2025 IEEE publication |
| 18 | Hungarian-language NLP | Hungarian transformer fine-tuning, document and reflective-writing classification | 2026 publication |
| 19 | Research software integration | GPU pipelines, ROS, pretrained-model integration, Apptainer containers, configuration and checks | Visible across rows 1, 10, 11; maturity varies |
Further repositories: Fodor GitHub profile.
How to read this table
Rows 1–6 are the measurement layer most relevant to the clinical work: everything needed to turn a recording into structured features.
Rows 7 and 16 are the interpretation layer — composite AI reasoning over learned features, and embodied interaction.
Rows 8–15 are perception and 3D, relevant to camera-based therapy monitoring should that opportunity mature, and to the rehabilitation line.
Rows 17–19 are adjacent competences that establish breadth rather than feeding the psychiatry narrative directly.
Mapping to clinical use
| Clinical need | Rows that serve it |
|---|---|
| Speech Gap measurement and the respiratory extension | 1, 2, 17 |
| Multimodal behaviour analysis in structured interaction | 3, 4, 5, 6, 7 |
| Camera-based therapy monitoring (opportunity only) | 8, 9, 11, 12, 14 |
| Avatar therapy and interaction | 13, 16 |
Related pages
NIPG · Capability readiness · Multimodal behaviour analysis · Bibliography · Editorial safeguards