CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hani89 /Synthetic-Medical-Speech-Dataset Synthetic Medical Speech Dataset Overview Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.audioautomatic-speech-recognition10K<n<100K4 likes226 downloads1y agoHugging Face02Hani89 /medical_asr_recording_datasetData Source Kaggle Medical Speech, Transcription, and Intent Context 8.5 hours of audio utterances paired with text for common medical symptoms. Content This data contains thousands of audio utterances for common medical symptoms like “knee pain” or “headache,” totaling more than 8 hours in aggregate. Each utterance was created by individual human contributors based on a given symptom. These audio snippets can be used to train conversational agents in the medical field. This Figure Eight… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/medical_asr_recording_dataset.textautomatic-speech-recognition1K<n<10K11 likes211 downloads3y agoHugging Face03HaninZ /HowFarAreYou_3DSpeakerTrain_fullaudio10K<n<100K0 likes199 downloads2y agoHugging Face04pythainlp /han-instruct-dataset-v4.0 Dataset Card for Han Instruct Dataset v4.0 🪿🪿🪿🪿 The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset. 🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others. Data sources: Reference desk at Thai wikipedia. Law from… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v4.0.texttext-generation1K<n<10K5 likes88 downloads6d agoHugging Face05pythainlp /han-instruct-dataset-v1.0 Dataset Card for "han-instruct-dataset-v1.0" The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset. Dataset Summary 🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. It collect the instruction following in Thai from many source. Many question are collect from Reference desk at Thai wikipedia. Data sources: Reference desk at Thai wikipedia. Law from justicechannel.org… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v1.0.texttext-generation1K<n<10K4 likes86 downloads6d agoHugging Face06pythainlp /han-instruction-dataset Han Instruction Dataset Han instruction dataset: Thai instruction dataset 🪿 Han (ห่าน or goose) Instruction Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others. The final dataset of han instruction dataset was released! GitHub: https://github.com/wannaphong/han-instruction-dataset Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruction-dataset.texttext-generation10K<n<100K1 likes67 downloads6d agoHugging Face07hanifabdlh /LaMini-Instruction-Indonesian-Google-Translated Dataset Card for "LaMini-Instruction-Indonesian-Google-Translated" This dataset is on development: the are miss translation in some question answering case like please add whitespaces to this text: iwanttoplayfootball. It will be translated to harap tambahkan spasi pada teks ini: iwanttoplayfootball or translated but the whitespaces exist harap tambahkan spasi pada teks ini: saya ingin bermain sepak bola text1M<n<10M2 likes57 downloads3y agoHugging Face08pythainlp /han-instruct-dataset-v2.0 Dataset Card for Han Instruct Dataset v2.0 The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset. 🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collect all Thai instruct dataset that made by human and our old model. The dataset can use to train Instruction Following model like ChatGPT or other. Many question are collect from Reference desk at Thai wikipedia. Data sources: Reference desk… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v2.0.texttext-generation1K<n<10K3 likes57 downloads6d agoHugging Face09pythainlp /han-instruct-dataset-v3.0 Dataset Card for Han Instruct Dataset v3.0 The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset. 🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others. Many questions are collect from Reference desk at Thai wikipedia. Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v3.0.texttext-generation1K<n<10K3 likes54 downloads6d agoHugging Face10Nash-pAnDiTa /quran_dataset_hani_clean Quranic Dataset by Tanzil Project (Qari: Hani) Overview The Tanzil Project is an international initiative aimed at providing a highly accurate and verified Quranic text in Unicode. The Tanzil text is refined through rigorous verification processes to ensure adherence to the Medina Mushaf and to achieve exceptional precision. Text Verification Process To achieve a high level of accuracy, the Tanzil Project has implemented a three-phase verification process:… See the full description on the dataset page: https://huggingface.co/datasets/Nash-pAnDiTa/quran_dataset_hani_clean.audio1K<n<10K0 likes52 downloads2y agoHugging Face11rohingyalanguage /rohingya-hanifi-rohingyalish-english Rohingya Hanifi–Rohingyalish–English Lexicon A multilingual lexical dataset from RohingyaLanguage.org connecting English dictionary headwords with Rohingyalish (Latin-script Rohingya) and Hanifi Rohingya script. Dataset summary 15,926 validated rows Based on 6,510 English dictionary entries Languages: English and Rohingya (rhg) Scripts: Rohingyalish/Latin and Hanifi Rohingya Hanifi forms are generated using the same rule-based converter used by… See the full description on the dataset page: https://huggingface.co/datasets/rohingyalanguage/rohingya-hanifi-rohingyalish-english.text10K<n<100K2 likes43 downloads5d agoHugging Face12Hani89 /SynthaticPipelines For mor info follow the below link at Github (SyntheticData@Github)[] 26295 Row ~5.5 GB ~34H:22M audiotext-to-speech10K<n<100K0 likes41 downloads11mo agoHugging Face13HaninZ /EnvironmentalSoundClassification_ESC50-HumanAndNonSpeechSounds_TTSaudion<1K1 likes39 downloads2y agoHugging Face14Hanish /lq-decide-data LQ-Decide training data 141,038 rows for training models that answer typed decisions: given a state and a question with a fixed option set, return a probability over the options rather than generated text. Built for LQ-Decide 0.6B by Hanish Keloth. Every source is licence-checked and named. Non-commercial and unclear-licence sources were excluded by a flag rather than being quietly included; the excluded list is below so you can decide for yourself. Files… See the full description on the dataset page: https://huggingface.co/datasets/Hanish/lq-decide-data.texttext-classification100K<n<1M0 likes39 downloads5d agoHugging Face15aarontseng /nil-hover-cmn-hani aarontseng/nil-hover-cmn-hani Hover word-sense choice labels for EN↔ZH dictionary hover UI. Each row: one FineTranslations sentence, one Intl-segmented hovered token, noisy lexicon candidates, DeepSeek Flash integer label (0=none, 1..N=candidate index). split rows train 44999 validation 5000 Fields side: en or zh (50/50) sentence, query, char_start, char_end candidates, candidate_counts label, label_text (label_text null when label==0) Built by… See the full description on the dataset page: https://huggingface.co/datasets/aarontseng/nil-hover-cmn-hani.tabulartext-classification10K<n<100K0 likes38 downloads5d agoHugging Face16HaninZ /PronounciationEvaluationFluency_Speechocean762audio1K<n<10K1 likes34 downloads2y agoHugging Face17HaninZ /StressDetection_MIRSD_TTSaudion<1K0 likes30 downloads2y agoHugging Face18HaninZ /snips_slu_v1.0audio1K<n<10K0 likes30 downloads2y agoHugging Face19aarontseng /nil-hover-reject-cmn-hani aarontseng/nil-hover-reject-cmn-hani Hover reject-unsuitable labels for EN↔ZH dictionary hover UI. Each row: FineTranslations sentence + hovered token + noisy lexicon candidates; DeepSeek Flash marks unsuitable candidate indices (0=none unsuitable, else 1,3,5). split rows train 44999 validation 5000 Fields side: en or zh (50/50) sentence, query, char_start, char_end candidates, candidate_counts reject, keep (1-based indices) reject_texts… See the full description on the dataset page: https://huggingface.co/datasets/aarontseng/nil-hover-reject-cmn-hani.tabulartext-classification10K<n<100K0 likes30 downloads4d agoHugging Face20pacozaa /han-instruct-dataset-v4.0-chatmlDirect Folk From: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v4.0 I added "text" column here to reformating to chatml format text1K<n<10K0 likes29 downloads2y agoHugging Face21HANI-LAB /Med-REFL-DPO News [2025/06/10] We are releasing the Med-REFL dataset, which is split into two subsets: Reasoning Enhancement Data and Reflection Enhancement Data. Introduction This is the Direct Preference Optimization (DPO) dataset created by the Med-REFL framework, designed to improve the reasoning and reflection capabilities of Large Language Models in the medical field. The dataset is constructed using a low-cost, scalable pipeline that leverages a Tree-of-Thought (ToT) approach… See the full description on the dataset page: https://huggingface.co/datasets/HANI-LAB/Med-REFL-DPO.textquestion-answering10K<n<100K0 likes29 downloads1y agoHugging Face22Hani6999 /university-1652 University-1652: Drone-based Geo-localization Benchmark 🚁 University-1652 is a multi-view dataset for drone-based geo-localization, annotating 1652 buildings across 72 universities (ACM Multimedia 2020, paper). Cited in 50+ papers, it supports Drone → Satellite localization and Satellite → Drone navigation. Dataset Structure Splits: Train: 50,218 images (drone, satellite, street, google; 33 universities) Test: query_drone: 37,855 images gallery_drone: 51,355… See the full description on the dataset page: https://huggingface.co/datasets/Hani6999/university-1652.text100K<n<1M0 likes29 downloads5mo agoHugging Face23HaninZ /DialogueEmotionClassification_DailyTalkaudio10K<n<100K0 likes28 downloads2y agoHugging Face24HaninZ /DialogueActClassification_DailyTalkaudio10K<n<100K0 likes26 downloads2y agoHugging Face25HaninZ /SpeakerVerification_LibriSpeech-TestClean_TTSaudion<1K0 likes25 downloads2y agoHugging Face26HaninZ /AccentClassification_AccentdbExtended_TTSaudion<1K0 likes23 downloads2y agoHugging Face27Haniehedi /ag_news_annotated Dataset Card for ag_news_annotated This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Using this dataset with Argilla To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code: import argilla as rg ds =… See the full description on the dataset page: https://huggingface.co/datasets/Haniehedi/ag_news_annotated.text100K<n<1M0 likes22 downloads4mo agoHugging Face28HaninZ /DialogueEmotionClassification_DailyTalk_testaudion<1K0 likes20 downloads2y agoHugging Face29HaninZ /paralinguistic_datasetaudio1K<n<10K0 likes20 downloads2y agoHugging Face30Haniehedi /github-issuestextn<1K0 likes20 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.