Datasets previously analysed or used in collaborative research
External corpora NIPG has worked with, and what each one actually establishes. This is an experience record, not a claim of ownership or continued access.
External benchmark and interaction corpora
| Dataset | Experience demonstrated | Appropriate wiki claim |
|---|---|---|
| ChaLearn First Impressions | First-impression and perceived-personality modelling | Previously used external benchmark; identify the exact training or evaluation role from the relevant methods paper |
| MULTISIMO | Evaluation against external-observer labels in a different data domain | Out-of-distribution evaluation experience |
| UDIVA | Dyadic interaction across different tasks | Analysis of task-dependent perceived personality |
| ELEA | Small-group collaboration and group performance | Group-level behavioural analysis |
| AMI Meeting Corpus | A sequence of task-oriented meetings | Analysis of changes across interactions and tasks |
| MeMo | Memory and repeated social interaction | Joint dataset publication and Delft collaboration |
The original ChaLearn dataset is described in ChaLearn LAP 2016 First Round Challenge on First Impressions.
These corpora are the evidence base behind perceived personality, including the ELEA result on group performance.
MeMo in detail
Introducing MeMo: A Multimodal Dataset for Memory Modelling in Multiparty Conversations (Tsfasman, Dudzik, Fenech, Lőrincz, Jonker, Oertel — arXiv). 31 hours of small-group discussion, repeated three times over two weeks, with participant memory reports and multimodal behavioural and perceptual measures.
Safeguard. Not a clinical dataset.
The repeated-measurement structure is methodologically the closest external analogue to what the Speech Gap longitudinal study would produce.
Internal archived material
The 2013 educational-games recordings — 180 primary-school participants in the study, 57 analysed on video. Potentially reusable with current methods, pending a feasibility check on file and permission quality.
Safeguard. No raw video of children in the public wiki. Aggregate findings and dataset descriptions only.
→ Adaptive assessment and serious games
Prospective data
| Source | Status |
|---|---|
| Speech Gap / BPD longitudinal recordings | Planned; protocol not yet registered |
| Danish registers and biobanks | Opportunity only; no access established, and registers do not necessarily contain repeated speech recordings |
| Camera-based therapy monitoring | Opportunity only; no agreed study, no project plan |
Related pages
Perceived personality · Supporting collaborations · Data governance and annotation · Bibliography