datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.medical_asr_recording_datasetData Source
Kaggle Medical Speech, Transcription, and Intent
Context
8.5 hours of audio utterances paired with text for common medical symptoms.
Content
This data contains thousands of audio utterances for common medical symptoms like “knee pain” or “headache,” totaling more than 8 hours in aggregate. Each utterance was created by individual human contributors based on a given symptom. These audio snippets can be used to train conversational agents in the medical field.
This Figure Eight… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/medical_asr_recording_dataset.HowFarAreYou_3DSpeakerTrain_fullhan-instruct-dataset-v4.0
Dataset Card for Han Instruct Dataset v4.0 🪿🪿🪿🪿
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others.
Data sources:
Reference desk at Thai wikipedia.
Law from… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v4.0.han-instruct-dataset-v1.0
Dataset Card for "han-instruct-dataset-v1.0"
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
Dataset Summary
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. It collect the instruction following in Thai from many source.
Many question are collect from Reference desk at Thai wikipedia.
Data sources:
Reference desk at Thai wikipedia.
Law from justicechannel.org… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v1.0.han-instruction-dataset
Han Instruction Dataset
Han instruction dataset: Thai instruction dataset
🪿 Han (ห่าน or goose) Instruction Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others.
The final dataset of han instruction dataset was released!
GitHub: https://github.com/wannaphong/han-instruction-dataset
Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruction-dataset.LaMini-Instruction-Indonesian-Google-Translated
Dataset Card for "LaMini-Instruction-Indonesian-Google-Translated"
This dataset is on development: the are miss translation in some question answering case like please add whitespaces to this text: iwanttoplayfootball. It will be translated to harap tambahkan spasi pada teks ini: iwanttoplayfootball or translated but the whitespaces exist harap tambahkan spasi pada teks ini: saya ingin bermain sepak bola
han-instruct-dataset-v2.0
Dataset Card for Han Instruct Dataset v2.0
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collect all Thai instruct dataset that made by human and our old model. The dataset can use to train Instruction Following model like ChatGPT or other.
Many question are collect from Reference desk at Thai wikipedia.
Data sources:
Reference desk… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v2.0.han-instruct-dataset-v3.0
Dataset Card for Han Instruct Dataset v3.0
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others.
Many questions are collect from Reference desk at Thai wikipedia.
Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v3.0.quran_dataset_hani_clean
Quranic Dataset by Tanzil Project (Qari: Hani)
Overview
The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision.
Text Verification Process
To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification process:… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_hani_clean.rohingya-hanifi-rohingyalish-english
Rohingya Hanifi–Rohingyalish–English Lexicon
A multilingual lexical dataset from RohingyaLanguage.org connecting English dictionary headwords with Rohingyalish (Latin-script Rohingya) and Hanifi Rohingya script.
Dataset summary
15,926 validated rows
Based on 6,510 English dictionary entries
Languages: English and Rohingya (rhg)
Scripts: Rohingyalish/Latin and Hanifi Rohingya
Hanifi forms are generated using the same rule-based converter used by… See the full description on the dataset page: https://huggingface.co/datasets/rohingyalanguage/rohingya-hanifi-rohingyalish-english.SynthaticPipelines
For mor info follow the below link at Github
(SyntheticData@Github)[]
26295 Row
~5.5 GB
~34H:22M
EnvironmentalSoundClassification_ESC50-HumanAndNonSpeechSounds_TTSlq-decide-data
LQ-Decide training data
141,038 rows for training models that answer typed decisions: given a state and a question with a fixed option set,
return a probability over the options rather than generated text. Built for
LQ-Decide 0.6B by Hanish Keloth.
Every source is licence-checked and named. Non-commercial and unclear-licence sources were excluded by a flag rather
than being quietly included; the excluded list is below so you can decide for yourself.
Files… See the full description on the dataset page: https://huggingface.co/datasets/Hanish/lq-decide-data.nil-hover-cmn-hani
aarontseng/nil-hover-cmn-hani
Hover word-sense choice labels for EN↔ZH dictionary hover UI.
Each row: one FineTranslations sentence, one Intl-segmented hovered token, noisy lexicon candidates, DeepSeek Flash integer label (0=none, 1..N=candidate index).
split
rows
train
44999
validation
5000
Fields
side: en or zh (50/50)
sentence, query, char_start, char_end
candidates, candidate_counts
label, label_text (label_text null when label==0)
Built by… See the full description on the dataset page: https://huggingface.co/datasets/aarontseng/nil-hover-cmn-hani.PronounciationEvaluationFluency_Speechocean762StressDetection_MIRSD_TTSsnips_slu_v1.0nil-hover-reject-cmn-hani
aarontseng/nil-hover-reject-cmn-hani
Hover reject-unsuitable labels for EN↔ZH dictionary hover UI.
Each row: FineTranslations sentence + hovered token + noisy lexicon candidates;
DeepSeek Flash marks unsuitable candidate indices (0=none unsuitable, else 1,3,5).
split
rows
train
44999
validation
5000
Fields
side: en or zh (50/50)
sentence, query, char_start, char_end
candidates, candidate_counts
reject, keep (1-based indices)
reject_texts… See the full description on the dataset page: https://huggingface.co/datasets/aarontseng/nil-hover-reject-cmn-hani.han-instruct-dataset-v4.0-chatmlDirect Folk From: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v4.0
I added "text" column here to reformating to chatml format
Med-REFL-DPO
News
[2025/06/10] We are releasing the Med-REFL dataset, which is split into two subsets: Reasoning Enhancement Data and Reflection Enhancement Data.
Introduction
This is the Direct Preference Optimization (DPO) dataset created by the Med-REFL framework, designed to improve the reasoning and reflection capabilities of Large Language Models in the medical field.
The dataset is constructed using a low-cost, scalable pipeline that leverages a Tree-of-Thought (ToT) approach… See the full description on the dataset page: https://huggingface.co/datasets/HANI-LAB/Med-REFL-DPO.university-1652
University-1652: Drone-based Geo-localization Benchmark 🚁
University-1652 is a multi-view dataset for drone-based geo-localization, annotating 1652 buildings across 72 universities (ACM Multimedia 2020, paper). Cited in 50+ papers, it supports Drone → Satellite localization and Satellite → Drone navigation.
Dataset Structure
Splits:
Train: 50,218 images (drone, satellite, street, google; 33 universities)
Test:
query_drone: 37,855 images
gallery_drone: 51,355… See the full description on the dataset page: https://huggingface.co/datasets/Hani6999/university-1652.DialogueEmotionClassification_DailyTalkDialogueActClassification_DailyTalkSpeakerVerification_LibriSpeech-TestClean_TTSAccentClassification_AccentdbExtended_TTSag_news_annotated
Dataset Card for ag_news_annotated
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Haniehedi/ag_news_annotated.DialogueEmotionClassification_DailyTalk_testparalinguistic_datasetgithub-issues
