datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CommonVoice_elderlychinese-american-elder-fraud-qa
chinese-american-elder-fraud-qa
A hand-authored, trilingual (Mandarin / Cantonese / English) fraud-recognition dataset for first-generation Chinese-American elders and the adult children who help them. 235 rows authored, 207 adapted through the Adaption Labs platform with reasoning traces. Grounded in FBI, IC3, and SFPD reports on Chinese-community elder fraud.
Adaption Labs Uncharted Data Challenge submission.
Metric
Value
Rows authored
235
Rows adapted (training… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/chinese-american-elder-fraud-qa.multilingual-elder-safety-msgs
multilingual-elder-safety-msgs
A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation.
Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.the_elder_scrolls_v_skyrim_special_edition_recordings_01
上古卷轴5 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_7216b537a4b31f0d12deb0d9d0837226
Collection: general (泛数据)
Recordings: 57
Layout: recordings/<recording_id>/<raw component>
the_elder_scrolls_v_skyrim_special_edition_bc_01
上古卷轴5 BC parquet archives
Collection mode: general
Subset: default
Archives: 21
Encrypted bytes: 692373802346
Generated by the game data platform BC repository consolidator.
so-vits-svc-4.0-ru-The_Elder_Scrolls_V_Skyrimaugmented_above_70yo_elderly_people_dataset
Dataset Card for "augmented_above_70yo_elderly_people_dataset"
More Information needed
persian-elderly-asr
Final gathered Persian elderly speech
Final corpus: 1,980 train / 294 validation / 329 test chunks. Another 956 uncertain chunks are quarantined under portable/review/ and excluded from these splits. This revision replaces the earlier gathered corpus; earlier data remains available through repository commit history.
80 paired recordings from four speaker folders. Reference transcripts were aligned with a historical Persian Wav2Vec2-base checkpoint, then cut at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-elderly-asr.prepared_above_70yo_elderly_people_datasetV2
Dataset Card for "prepared_above_70yo_elderly_people_datasetV2"
More Information needed
Masri-Elders
Masri Elders Speech-to-Text Dataset
Dataset Description
This dataset contains Egyptian Arabic (Masri) speech recordings from elderly speakers, paired with their transcriptions. It is designed to help improve Automatic Speech Recognition (ASR) models for this specific dialect and demographic, which is often underrepresented in standard datasets.
Language: Egyptian Arabic (Masri)
Demographic: Elders
Sampling Rate: 16kHz (Mono)
Format: WAV audio + Text transcripts… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/Masri-Elders.asia-owid-population-young-working-elderly-with-projections
Population Young Working Elderly With Projections | Asia (Our World in Data)
🌏 7,399 observations · 49 Asia countries · 1950–2100 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 7,399 observations of Population Young Working Elderly With Projections data across 49 Asia countries, spanning 1950–2100.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Population Young Working Elderly… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-population-young-working-elderly-with-projections.aide_elderly
Aide Elderly Dataset
Description
Small dataset for experiments related to project Aide.
Authors
Fabien Allemand (fabien.allemand@cea.fr)
thai_elderly_speechThe Thai Elderly Speech dataset by Data Wow and VISAI Version 1 dataset aims at
advancing Automatic Speech Recognition (ASR) technology specifically for the
elderly population. Researchers can use this dataset to advance ASR technology
for healthcare and smart home applications. The dataset consists of 19,200 audio
files, totaling 17 hours and 11 minutes of recorded speech. The files are
divided into 2 categories: Healthcare (relating to medical issues and services
in 30 medical categories) and Smart Home (relating to smart home devices in 7
household contexts). The dataset contains 5,156 unique sentences spoken by 32
seniors (10 males and 22 females), aged 57-60 years old (average age of 63
years).above_70yo_elderly_people_dataset
Dataset Card for "above_70yo_elderly_people_dataset"
More Information needed
Hydrus-The-Elder-Scrolls-Instruct
TES instruct
First attempt at mass-data-gen, wanted to get some instruct data related to games and such so i used a scrape for the elder scrolls wiki's including oblivion, morrowind, skyrim and then i deduped for the most unique possible entries + So i don't get killed by the guy who uses my deepseek API key
1. Initial Data gen
The base file containing source texts (e.g., lore articles) was used as input. For each entry, deepseek V3.1 was prompted to generate a… See the full description on the dataset page: https://huggingface.co/datasets/Delta-Vector/Hydrus-The-Elder-Scrolls-Instruct.record-testing2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 7,
"total_frames": 5850,
"total_tasks": 1,
"total_videos": 7,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:7"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ElderCareAI/record-testing2.stress_ball_new_room_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 33,
"total_frames": 32829,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:33"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ElderCareAI/stress_ball_new_room_v3.thai_Elderly_Speech_datasetThai speech datasets Thai_Elderly_Speech_dataset: https://github.com/VISAI-DATAWOW/Thai-Elderly-Speech-dataset/releases/tag/v1.0.0
iranian-elderly-psychospiritual-interviews
Iranian Elderly Psycho-Spiritual Interviews
A culturally grounded, fully synthetic conversational interview dataset for assessing the mental and spiritual health of Iranian older adults, generated using Large Language Models.
Dataset Summary
This dataset introduces a culturally grounded, fully synthetic conversational interview corpus designed for the assessment and analysis of mental and spiritual health among Iranian older adults. All interviews are conducted in… See the full description on the dataset page: https://huggingface.co/datasets/liamirali/iranian-elderly-psychospiritual-interviews.18-12-25-v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 100,
"total_frames": 53057,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:100"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ElderCareAI/18-12-25-v3.Elder_Scrolls_Wiki_Dataset
The Elder Scrolls Wiki Dataset
This dataset is a snapshot of the Unofficial Elder Scrolls Pages (UESP) wiki, curated and organized by game: Skyrim, Morrowind, and Oblivion. It is designed for use in retrieval-augmented generation (RAG) pipelines, knowledge base construction, semantic search, and other natural language processing (NLP) tasks relevant to The Elder Scrolls universe.
📚 Dataset Summary
Source: UESP Wiki
Snapshot Date: April 2025
License: CC BY-SA 2.5… See the full description on the dataset page: https://huggingface.co/datasets/RoyalCities/Elder_Scrolls_Wiki_Dataset.mascir_elderly_voice
Dataset Card for "mascir_elderly_voice"
More Information needed
15-12-25-v5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 20153,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ElderCareAI/15-12-25-v5.diasynth-vi-elderly-orpo
DiaSynth Vietnamese Elderly Care — ORPO
Preference pairs for ORPO finetuning, focused on Vietnamese elderly-care
safety hazards. Each record is a single-turn (prompt, chosen, rejected)
tuple where chosen is a safe and persona-consistent response and rejected
is an unsafe / off-persona variant for the same hazard prompt.
Splits
Split
Pairs
Bytes
train
15,923
30,460,387
val
331
633,493
test
332
635,798
Schema
{
"pair_id":… See the full description on the dataset page: https://huggingface.co/datasets/quannguyen204/diasynth-vi-elderly-orpo.hinglish-elderly-empathy---
license: cc-by-nc-4.0
task_categories:
- text-generation
- conversational
language:
- hi
- en
tags:
- hinglish- elderly- empathy- mental-health- dialogue
size_categories:
- 10K<n<100K
pretty_name: Hinglish Elderly Empathy
---
# Hinglish Elderly Empathy
## Dataset Description
A Hinglish (Hindi-English) conversational dataset designed for an **Empathetic AI Companion for the Elderly**. The model mimics a respectful, patient, and warm persona ('Buzurg Saathi').
### Dataset Structure
-… See the full description on the dataset page: https://huggingface.co/datasets/hbpkillerX/hinglish-elderly-empathy.15-12-25-v4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 14,
"total_frames": 6768,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:14"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ElderCareAI/15-12-25-v4.2-2-26-test2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 995,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ElderCareAI/2-2-26-test2.diasynth-vi-elderly-sft
DiaSynth Vietnamese Elderly Care — SFT
Synthetic Vietnamese multi-turn dialogues for finetuning a voice-assistant LLM
(Qwen3-4B base) targeting elderly care / companionship. Generated with the
DiaSynth pipeline (Persona × Subtopic × CoT) and rewritten for a consistent
persona where the assistant addresses itself as "con" and the user as
ông / bà / bác.
Splits
Split
Records
Bytes
train
47,532
263,324,285
val
991
5,496,336
test
990
5,466,399… See the full description on the dataset page: https://huggingface.co/datasets/quannguyen204/diasynth-vi-elderly-sft.farahan-yoruba-elder-corpus
Apakose Ezekiel Imoleayo — Farahàn
Appear. Be found.
Lagos, Nigeria · UNILAG Yoruba Studies ·
Graduating 2030
What I Build
Yoruba Oral Knowledge Corpus
I document primary-source Yoruba knowledge
directly from elder speakers in Lagos and
southwest Nigeria. Structured interviews
covering proverbs (Owe), oral history (Itan),
praise poetry (Oriki), and cultural knowledge
systems that no web scrape produces.
What makes this corpus different:… See the full description on the dataset page: https://huggingface.co/datasets/Apakose-Ezekiel/farahan-yoruba-elder-corpus.2-2-26-100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 100,
"total_frames": 55257,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:100"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ElderCareAI/2-2-26-100.
