datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
udhr-lid
UDHR-LID
Why UDHR-LID?
You can access UDHR (Universal Declaration of Human Rights) here, but when a verse is missing, they have texts such as "missing" or "?". Also, about 1/3 of the sentences consist only of "articles 1-30" in different languages. We cleaned the entire dataset from XML files and selected only the paragraphs. We cleared any unrelated language texts from the data and also removed the cases that were incorrect.
Incorrect? Look at the ckb and kmr files in the UDHR.… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/udhr-lid.indian-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction
The Synthetic India Passports Dataset assembles more than 1,000 AI-generated passport images intended for training OCR and computer vision models on identity documents. Because every record is fully synthetic — with no… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/indian-passports.human-robot-conversation-korean
Human-Robot Conversation Dataset (Korean) - 660+ Hours
Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.veneer-bench
VENEER — a format-bias benchmark for LLM judges
Same facts. Different clothes. Watch the judge change its mind.
VENEER measures how much of an LLM judge's verdict is bought by presentation
rather than substance.
Every answer in this dataset is rendered from one fixed list of atomic claims.
Nine renderers turn that same claim list into plain prose, bullets, a numbered
list, generic markdown headings, a table, bolded key terms, maximal markdown,
emoji bullets, and a padded… See the full description on the dataset page: https://huggingface.co/datasets/uditjain/veneer-bench.philippine-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Philippines
The Synthetic Philippines Passports Dataset assembles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Every record is fully synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/philippine-passports.ukrainian-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Ukraine
The Synthetic Ukraine Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/ukrainian-passports.donors_choose_datarussian-speech-recognition-dataset
Russian Telephone Dialogues Dataset - 338 Hours
The Russian speech dataset includes 338 hours of telephone dialogues in Russian from 460 native speakers, offering high-quality audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for building accurate Russian dialogue and audio datasets. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/russian-speech-recognition-dataset.swedish-ud-pos-csv
Swedish POS Dataset (CSV)
Swedish Part-of-Speech tagging dataset in CSV format, converted from CoNLL-U.
Format
sentence_id: Sentence ID
token_id: Token position
form: Word form
lemma: Lemma
upos: Universal POS tag
xpos: Language-specific POS tag
feats: Morphological features
head: Dependency head
deprel: Dependency relation
deps: Enhanced dependencies
misc: Miscellaneous annotations
license: mit
ud-conll2017-aligned
UD CoNLL-U 2017 aligned
This dataset is a processed version of the dataset made available for the UD CoNLL Shared Task 2017, entitled "Multilingual Parsing from Raw Text to Universal Dependencies." This task is based on the Universal Dependencies 2.0 dataset.
The processing aligns words and tokens along with their morphological annotations using the universal part-of-speech set for sequence-to-sequence task training.
The dataset fields are described in the table below:
Field… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ud-conll2017-aligned.metahate
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
This is MetaHate: a meta-collection of 36 hate speech datasets from social media comments.
What's New in Version 2.0
Data Refinement and Size Adjustment: The original publication reported 1,226,202 total instances and 1,101,165 public instances. Due to refined cross-dataset deduplication and overlapping instance resolution, the current dataset sizes are 1,226,203 (full) and 1,084,236 (public).… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/metahate.hindi-speech-recognition-dataset
Hindi Telephone Dialogues Dataset - 760 Hours
Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.human-robot-conversation-english
Human-Robot Conversation Dataset (English) - 660+ Hours
Dataset (English) contains 660+ hours of audio featuring dialogues between AI and a human in English across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-english.british-english-speech-recognition-dataset
British English Telephone Dialogues Dataset - 200 Hours
The dataset consists of 200 hours of high-quality telephone dialogues from 310 native speakers in the UK, with detailed annotations (transcriptions, timestamps, speaker ID, gender, and background noise) to support speech recognition systems, NLP tasks, and machine learning models requiring diverse British English audio datasets. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/british-english-speech-recognition-dataset.AF
Data Splits
The dataset is split into two sets:
Train: The training set consists of 4330 samples.
Test: The testing set consists of 1082 samples.
This split ensures a comprehensive evaluation, allowing models trained on this data to be thoroughly tested on unseen examples.
Data Fields
text: The text entry contains the sample with the structure : The Formal Statement is: ______ The Informal Statement is: _______
LLM-Text-Generation-Dataset
Generated Text Dataset - 4 Millions+ Logs
Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data
Dataset characteristics:
Characteristic
Data
Description
Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.korean-speech-recognition
Korean Speech Recognition Dataset - 10+ hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Korean, featuring 20 native speakers. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/korean-speech-recognition.Split-Isa-AF-MMA
About
This dataset is an amalgamation of every entry of the DQ Round Trip Problem Selection spreadsheet (https://docs.google.com/spreadsheets/d/1dEWWzjuEXwf9s4II0CixH4sqopc1flIMFx19UjHiyNU/edit?usp=sharing&resourcekey=0-_G7oxmbh7szV5jx-HxhepQ)
with formal and informal columns combined and concatenated with the formal and informal statements from the Isabelle train and val sets from the Multilingual Mathematical Autoformalization dataset, to which a text column has been added in… See the full description on the dataset page: https://huggingface.co/datasets/UDACA/Split-Isa-AF-MMA.german-speech-recognition-dataset
German Telephone Dialogues Dataset - 431 Hours
Dataset comprises 431 hours of high-quality audio recordings from 590+ native German speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating German speech recognition systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in German for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/german-speech-recognition-dataset.french-speech-recognition-dataset
French Telephone Dialogues Dataset - 547 Hours
his speech recognition dataset comprises 547 hours of telephone dialogues in French from 964 native speakers, providing audio recordings with detailed annotations (text, speaker ID, gender, age) to support speech recognition systems, natural language processing, and deep learning models for training and evaluating automatic speech recognition technology. - Get the data
Dataset characteristics:
Characteristic
Data… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/french-speech-recognition-dataset.vietnamese-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Vietnam
The Synthetic Vietnam Passports Dataset gathers more than 1,000 AI-generated passport images tailored for training OCR and computer vision pipelines on identity documents. Every entry is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/vietnamese-passports.indonesian-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Indonesia
The Synthetic Indonesia Passports Dataset compiles more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Each record is fully synthetic, so… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/indonesian-passports.AF-split
About
This dataset is an amalgamation of every entry of the DQ Round Trip Problem Selection spreadsheet (https://docs.google.com/spreadsheets/d/1dEWWzjuEXwf9s4II0CixH4sqopc1flIMFx19UjHiyNU/edit?usp=sharing&resourcekey=0-_G7oxmbh7szV5jx-HxhepQ)
with the informal statement column concatenated to the bottom of the formal statement column and all other metadata deleted.
`statement': either a formal statement in Isabelle or an informal statement
language:
- en
udiatturkish-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction
The Synthetic Turkey Passports Dataset offers more than 1,000 AI-generated passport images, designed to support the training of OCR and computer vision models on identity documents. Because every record is fully… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/turkish-passports.chinese-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - China
The Synthetic China Passports Dataset brings together more than 1,000 AI-generated passport images, purpose-built for training OCR and computer vision systems on identity documents. Because every record is… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/chinese-passports.vietnamese-speech-recognition
Vietnamese Speech Dataset - 10+ hours
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/vietnamese-speech-recognition.Amazon-ML-Challenge-2024Isa-MMA
Isa-MMA: Isabelle Components of the Multilingual Mathematical Autoformalization (MMA) Dataset
Dataset Description
Jiang et. al. published a paper entitled Multilingual Mathematical Autoformalization that included datasets in Isabelle and Lean. This dataset is the combination of the Isabelle test and Isabelle val datasets merged into one.
Data Fields
input: The words "Statement in natural language:" followed by a statement in natural language.
output: The… See the full description on the dataset page: https://huggingface.co/datasets/UDACA/Isa-MMA.japanese-speech-recognition-dataset
Japanese Telephone Dialogues Dataset - 10 Hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Japanese, featuring 20+ native speakers and achieving a 95% sentence accuracy rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/japanese-speech-recognition-dataset.
