CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cis-lmu /udhr-lid UDHR-LID Why UDHR-LID? You can access UDHR (Universal Declaration of Human Rights) here, but when a verse is missing, they have texts such as "missing" or "?". Also, about 1/3 of the sentences consist only of "articles 1-30" in different languages. We cleaned the entire dataset from XML files and selected only the paragraphs. We cleared any unrelated language texts from the data and also removed the cases that were incorrect. Incorrect? Look at the ckb and kmr files in the UDHR.… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/udhr-lid.text10K<n<100K8 likes238 downloads2y agoHugging Face02ud-synthetic /indian-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction The Synthetic India Passports Dataset assembles more than 1,000 AI-generated passport images intended for training OCR and computer vision models on identity documents. Because every record is fully synthetic — with no… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/indian-passports.textimage-to-textn<1K1 likes77 downloads2mo agoHugging Face03ud-nlp /human-robot-conversation-korean Human-Robot Conversation Dataset (Korean) - 660+ Hours Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes51 downloads6mo agoHugging Face04uditjain /veneer-bench VENEER — a format-bias benchmark for LLM judges Same facts. Different clothes. Watch the judge change its mind. VENEER measures how much of an LLM judge's verdict is bought by presentation rather than substance. Every answer in this dataset is rendered from one fixed list of atomic claims. Nine renderers turn that same claim list into plain prose, bullets, a numbered list, generic markdown headings, a table, bolded key terms, maximal markdown, emoji bullets, and a padded… See the full description on the dataset page: https://huggingface.co/datasets/uditjain/veneer-bench.tabular1K<n<10K0 likes44 downloads2mo agoHugging Face05ud-synthetic /philippine-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Philippines The Synthetic Philippines Passports Dataset assembles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Every record is fully synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/philippine-passports.textimage-to-textn<1K1 likes38 downloads2mo agoHugging Face06ud-synthetic /ukrainian-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Ukraine The Synthetic Ukraine Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/ukrainian-passports.textimage-to-textn<1K1 likes37 downloads2mo agoHugging Face07udayl /donors_choose_datatabular100K<n<1M0 likes34 downloads4y agoHugging Face08ud-nlp /russian-speech-recognition-dataset Russian Telephone Dialogues Dataset - 338 Hours The Russian speech dataset includes 338 hours of telephone dialogues in Russian from 460 native speakers, offering high-quality audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for building accurate Russian dialogue and audio datasets. - Get the data Dataset characteristics: Characteristic Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/russian-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes34 downloads8mo agoHugging Face09elliotnorrevik /swedish-ud-pos-csv Swedish POS Dataset (CSV) Swedish Part-of-Speech tagging dataset in CSV format, converted from CoNLL-U. Format sentence_id: Sentence ID token_id: Token position form: Word form lemma: Lemma upos: Universal POS tag xpos: Language-specific POS tag feats: Morphological features head: Dependency head deprel: Dependency relation deps: Enhanced dependencies misc: Miscellaneous annotations license: mit texttoken-classification1K<n<10K0 likes32 downloads1y agoHugging Face10fdemelo /ud-conll2017-aligned UD CoNLL-U 2017 aligned This dataset is a processed version of the dataset made available for the UD CoNLL Shared Task 2017, entitled "Multilingual Parsing from Raw Text to Universal Dependencies." This task is based on the Universal Dependencies 2.0 dataset. The processing aligns words and tokens along with their morphological annotations using the universal part-of-speech set for sequence-to-sequence task training. The dataset fields are described in the table below: Field… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ud-conll2017-aligned.text100K<n<1M0 likes32 downloads11mo agoHugging Face11irlab-udc /metahategated MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection This is MetaHate: a meta-collection of 36 hate speech datasets from social media comments. What's New in Version 2.0 Data Refinement and Size Adjustment: The original publication reported 1,226,202 total instances and 1,101,165 public instances. Due to refined cross-dataset deduplication and overlapping instance resolution, the current dataset sizes are 1,226,203 (full) and 1,084,236 (public).… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/metahate.texttext-classification1M<n<10M21 likes30 downloads2mo agoHugging Face12ud-nlp /hindi-speech-recognition-dataset Hindi Telephone Dialogues Dataset - 760 Hours Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes30 downloads8mo agoHugging Face13ud-nlp /human-robot-conversation-english Human-Robot Conversation Dataset (English) - 660+ Hours Dataset (English) contains 660+ hours of audio featuring dialogues between AI and a human in English across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-english.audioautomatic-speech-recognitionn<1K1 likes30 downloads6mo agoHugging Face14ud-nlp /british-english-speech-recognition-dataset British English Telephone Dialogues Dataset - 200 Hours The dataset consists of 200 hours of high-quality telephone dialogues from 310 native speakers in the UK, with detailed annotations (transcriptions, timestamps, speaker ID, gender, and background noise) to support speech recognition systems, NLP tasks, and machine learning models requiring diverse British English audio datasets. - Get the data Dataset characteristics: Characteristic Data Description Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/british-english-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes29 downloads8mo agoHugging Face15UDACA /AF Data Splits The dataset is split into two sets: Train: The training set consists of 4330 samples. Test: The testing set consists of 1082 samples. This split ensures a comprehensive evaluation, allowing models trained on this data to be thoroughly tested on unseen examples. Data Fields text: The text entry contains the sample with the structure : The Formal Statement is: ______ The Informal Statement is: _______ text1K<n<10K0 likes28 downloads3y agoHugging Face16ud-nlp /LLM-Text-Generation-Dataset Generated Text Dataset - 4 Millions+ Logs Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data Dataset characteristics: Characteristic Data Description Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.texttext-generation1K<n<10K0 likes24 downloads1y agoHugging Face17ud-nlp /korean-speech-recognition Korean Speech Recognition Dataset - 10+ hours Dataset comprises 10 hours of high-quality telephone audio recordings in Korean, featuring 20 native speakers. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/korean-speech-recognition.audioautomatic-speech-recognitionn<1K0 likes21 downloads10mo agoHugging Face18UDACA /Split-Isa-AF-MMA About This dataset is an amalgamation of every entry of the DQ Round Trip Problem Selection spreadsheet (https://docs.google.com/spreadsheets/d/1dEWWzjuEXwf9s4II0CixH4sqopc1flIMFx19UjHiyNU/edit?usp=sharing&resourcekey=0-_G7oxmbh7szV5jx-HxhepQ) with formal and informal columns combined and concatenated with the formal and informal statements from the Isabelle train and val sets from the Multilingual Mathematical Autoformalization dataset, to which a text column has been added in… See the full description on the dataset page: https://huggingface.co/datasets/UDACA/Split-Isa-AF-MMA.text100K<n<1M0 likes19 downloads3y agoHugging Face19ud-nlp /german-speech-recognition-dataset German Telephone Dialogues Dataset - 431 Hours Dataset comprises 431 hours of high-quality audio recordings from 590+ native German speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating German speech recognition systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in German for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/german-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes19 downloads8mo agoHugging Face20ud-nlp /french-speech-recognition-dataset French Telephone Dialogues Dataset - 547 Hours his speech recognition dataset comprises 547 hours of telephone dialogues in French from 964 native speakers, providing audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for training and evaluating automatic speech recognition technology. - Get the data Dataset characteristics: Characteristic Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/french-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes19 downloads8mo agoHugging Face21ud-synthetic /vietnamese-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Vietnam The Synthetic Vietnam Passports Dataset gathers more than 1,000 AI-generated passport images tailored for training OCR and computer vision pipelines on identity documents. Every entry is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/vietnamese-passports.textimage-to-textn<1K1 likes19 downloads2mo agoHugging Face22ud-synthetic /indonesian-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Indonesia The Synthetic Indonesia Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/indonesian-passports.textimage-to-textn<1K3 likes19 downloads2mo agoHugging Face23UDACA /AF-split About This dataset is an amalgamation of every entry of the DQ Round Trip Problem Selection spreadsheet (https://docs.google.com/spreadsheets/d/1dEWWzjuEXwf9s4II0CixH4sqopc1flIMFx19UjHiyNU/edit?usp=sharing&resourcekey=0-_G7oxmbh7szV5jx-HxhepQ) with the informal statement column concatenated to the bottom of the formal statement column and all other metadata deleted. `statement': either a formal statement in Isabelle or an informal statement language: - en text10K<n<100K0 likes18 downloads3y agoHugging Face24WOOJYE /udiatimagen<1K0 likes18 downloads2mo agoHugging Face25ud-synthetic /turkish-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction The Synthetic Turkey Passports Dataset offers more than 1,000 AI-generated passport images, designed to support the training of OCR and computer vision models on identity documents. Because every record is fully… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/turkish-passports.textimage-to-textn<1K1 likes17 downloads2mo agoHugging Face26ud-synthetic /chinese-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - China The Synthetic China Passports Dataset brings together more than 1,000 AI-generated passport images, purpose-built for training OCR and computer vision systems on identity documents. Because every record is… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/chinese-passports.textimage-to-textn<1K1 likes17 downloads2mo agoHugging Face27ud-nlp /vietnamese-speech-recognition Vietnamese Speech Dataset - 10+ hours Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/vietnamese-speech-recognition.audioautomatic-speech-recognitionn<1K0 likes16 downloads10mo agoHugging Face28udit-k /Amazon-ML-Challenge-2024image100K<n<1M0 likes15 downloads2y agoHugging Face29UDACA /Isa-MMA Isa-MMA: Isabelle Components of the Multilingual Mathematical Autoformalization (MMA) Dataset Dataset Description Jiang et. al. published a paper entitled Multilingual Mathematical Autoformalization that included datasets in Isabelle and Lean. This dataset is the combination of the Isabelle test and Isabelle val datasets merged into one. Data Fields input: The words "Statement in natural language:" followed by a statement in natural language. output: The… See the full description on the dataset page: https://huggingface.co/datasets/UDACA/Isa-MMA.text100K<n<1M0 likes14 downloads3y agoHugging Face30ud-nlp /japanese-speech-recognition-dataset Japanese Telephone Dialogues Dataset - 10 Hours Dataset comprises 10 hours of high-quality telephone audio recordings in Japanese, featuring 20+ native speakers and achieving a 95% sentence accuracy rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/japanese-speech-recognition-dataset.audioautomatic-speech-recognitionn<1K0 likes14 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.