NIPG capability portfolio

The full technical inventory behind NIPG, with a per-row evidence pointer for each capability.

Read this first. Capability maturity varies a lot across these rows, from public repository to research prototype with code not verified. This is not a single validated product suite. The evidence column is what establishes readiness for any given row. See Capability readiness and Editorial safeguards.

What the portfolio does

NIPG develops multimodal perception and composite-AI systems that convert speech, video, human movement and environmental observation into temporally structured and interpretable information for healthcare, rehabilitation, social interaction and robotics.

The inventory

#CapabilityDemonstrated functionalityEvidence / readiness
1Multimodal AI pipelinesAudio, video and text preprocessing, feature extraction, GPU execution, caching and integrationPublic code: Exordium
2Speech and acoustic analysisMFCC, eGeMAPS, acoustic emotion features, transcription, forced alignment, transcript assessmentPublic code / third-party model integration: Exordium
3Facial behaviour and gazeFace and iris tracking, facial landmarks, gaze, head pose, action unitsPublic components; clinical research (ADOS-2)
4Blink and eye-state recognitionTransformer-based blink detection, efficient inference, live demonstrationPublic code: BlinkLinMulT
5Efficient multimodal transformersLinear-complexity attention for fused audiovisual and text sequencesPublic code: LinMulT
6Affect and personality modellingMultimodal sentiment and Big Five trait estimation; experimental rather than clinical assessmentPublic code: PersonalityLinMulT
7Composite AI for social interactionRule-based reasoning over learned features to segment and interpret requests, gesture, object manipulation and eye contactPublished ADOS-2 study; full clinical pipeline not verified public
8Label-efficient video segmentationObject-instance tracking and segmentation with sparse scribble supervisionPublic research code: Cluster2Former
9Multi-view, multi-person 3D pose3D skeleton reconstruction from uncalibrated cameras without 3D ground-truth supervisionPublication and public code: MVMB-NRSFM
10Skeleton-guided view synthesisNovel-view generation guided by skeleton information from single imagerySkel3D: repository exists; models/instructions incomplete
11Semantic 3D reconstructionRGB-D segmentation, visual SLAM, cross-view semantic transfer, human–robot viewpointsPublic research pipeline: Semantic Matching
12Temporally coherent semantic SLAMMemory-efficient, temporally consistent 3D semantic mapping from video2025 VGGT-based SLAM preprint
13Ambient-intelligence rehabilitationHome scanning, neural scene representations, navigation, exercise/camera placement, avatar feedbackAIRS demonstration
14Edge human pose estimation23-keypoint body/face/foot estimation with Edge TPU deployment pathwayPublic legacy research code: DeepRehab
15Animal tracking and behaviourVideo tracking of similar animals, occlusion-tolerant behavioural extractionPublic research code: rat_tracking
16Embodied human–AI interactionAR virtual characters, gesture-based nonverbal interaction with LLMsKeep Gesturing demonstration
17Interpretable signal processingParameterized wavelets, variable projection, interpretable neural signal features for ECG2025 IEEE publication
18Hungarian-language NLPHungarian transformer fine-tuning, document and reflective-writing classification2026 publication
19Research software integrationGPU pipelines, ROS, pretrained-model integration, Apptainer containers, configuration and checksVisible across rows 1, 10, 11; maturity varies

Further repositories: Fodor GitHub profile.

How to read this table

Rows 1–6 are the measurement layer most relevant to the clinical work: everything needed to turn a recording into structured features.

Rows 7 and 16 are the interpretation layer — composite AI reasoning over learned features, and embodied interaction.

Rows 8–15 are perception and 3D, relevant to camera-based therapy monitoring should that opportunity mature, and to the rehabilitation line.

Rows 17–19 are adjacent competences that establish breadth rather than feeding the psychiatry narrative directly.

Mapping to clinical use

Clinical needRows that serve it
Speech Gap measurement and the respiratory extension1, 2, 17
Multimodal behaviour analysis in structured interaction3, 4, 5, 6, 7
Camera-based therapy monitoring (opportunity only)8, 9, 11, 12, 14
Avatar therapy and interaction13, 16

NIPG · Capability readiness · Multimodal behaviour analysis · Bibliography · Editorial safeguards