HUMADEX/dementia_education_chatbot_sources
Dementia Education Chatbot Sources Dataset Summary This dataset contains the reproducibility and analysis artifacts used in the AI4HOPE Dementia Companion project. It includes the merged multilingual source metadata table, the normalized study export tables, and the accompanying codebook used to document the columns in each release file. The dataset was created from language-specific crawl and preprocessing outputs and then organized into publication-ready tabular… See the full description on the dataset page: https://huggingface.co/datasets/HUMADEX/dementia_education_chatbot_sources.
Dementia Education Chatbot Sources
Dataset Summary
This dataset contains the reproducibility and analysis artifacts used in the AI4HOPE Dementia Companion project. It includes the merged multilingual source metadata table, the normalized study export tables, and the accompanying codebook used to document the columns in each release file.
The dataset was created from language-specific crawl and preprocessing outputs and then organized into publication-ready tabular files for source analysis, provenance tracking, retrieval indexing, and downstream statistical analysis.
Dataset Structure
Source metadata
- File name:
sources.parquet - Rows: 4,470
- Columns: 45
Normalized study exports
participants_experts.csvexpert_prompts.csvsource_ratings.csvexpert_sus.csvparticipants_users.csvsession_logs.csvresponse_ratings.csvcodebook.csv
Language Distribution
For sources.parquet:
- de: 312
- en: 3,018
- es: 451
- pt: 206
- sl: 483
Dataset Creation
The source metadata was produced by processing language-specific crawl outputs from the project pipeline. The resulting metadata parquet files were transformed into a unified source table for publication and reproducibility.
The normalized CSV exports were created from the study master tables by separating participants, prompts, source ratings, usability scores, response ratings, and session-level event logs into tidy, analysis-ready files. The codebook.csv file provides a schema description for all published columns and supports re-use by external researchers.
Intended Use
This dataset is intended for:
- source analysis
- multilingual retrieval and indexing
- reproducible research workflows
- publication support material
- quality control and provenance inspection
- study analysis and re-use of normalized export tables
Notes
- This is a research-oriented dataset, not a raw text corpus.
session_logs.csvis a derived event log created from available timestamps in the study exports.- It is suitable for research and reproducibility purposes.
- Please cite the associated AI4HOPE project and publication when using this dataset.
Citation
If you use this dataset, please cite the AI4HOPE project and the associated publication.
Repos
License
Use a license consistent with the publication and the underlying source material. If the dataset includes only derived metadata and normalized exports, a permissive research-friendly license is usually appropriate, but the final license should match your project policy and source constraints.
Funding
Funded by the European Union (AI4HOPE, 101136769). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the Health and Digital Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.
This work was funded by UK Research and Innovation (UKRI) under the UK government's Horizon Europe funding guarantee [Grant No. 101136769].
Contact
- izidor.mlakar@um.si
- rigon.sallauka@um.si
