datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polysemous-words
Polysemous Words
Polysemous Words is a large-scale collection of contextual examples for 200 common polysemous English words. Each word in this dataset has multiple, distinct senses and appears across a wide variety of natural web text contexts, making the dataset ideal for research in word sense induction (WSI), word sense disambiguation (WSD), and probing the contextual understanding capabilities of large language models (LLMs).
Dataset Overview
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/tsivakar/polysemous-words.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.SPEEED_s3_words_french_100k-300kSPEEED_s3_words_russian_0k-300kSPEEED_s3_words_spanish_1200k-1400k
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (spanish).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_spanish_1200k-1400k.SPEEED_s3_words_spanish-600k-900SPEEED_s3_words_russian_300k-600kSPEEED_s3_words_thai_0k-250kSPEEED_s3_words_italian_0k_200kSPEEED_s3_words_italian_200k_400kSPEEED_s3_words_spanish_1000k-1200k
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (spanish).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_spanish_1000k-1200k.probability_words_nliWords of estimative probability NLI (3 configs). Parquet conversion; CSVs retained.
SPEEED_s3_words_multi_0-200000
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (english).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_multi_0-200000.heb-words27-spaced
heb-words27-diffpen
Word-level Hebrew handwriting for DiffusionPen finetuning -- the step up from
cyttic/heb-connected-bigrams17-diffpen, which covered 2-character units. Here the unit
is a whole word.
Content
20,000 most frequent words of LLMGen2 (sentences_llm2.txt, 968,204 LLM-generated
Hebrew sentences, 167,879 unique words)
643 numbers -- LLMGen2 contains no digits at all, because the generation
prompt forbade them, so numerals had to be added separately or… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/heb-words27-spaced.SPEEED_s3_words_spanish_200k-400k
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (spanish).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_spanish_200k-400k.SPEEED_s3_words_italian_400k_650kWordSim353
📖 WordSim353
The WordSim dataset contains word pairs with human annotated similarity scores.
The columns are Word 1, Word 2, Human (Mean).
353 rows of word similarieties.
Although small, this dataset is very popular and have been used in many researches, including fastText (that called this WS353).
Lisense
This dataset licensed under a Creative Commons Attribution 4.0 International License.
The English WordSim-353 word pairs and instructions are credited to:… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/WordSim353.frequency-words-2018
Frequency Words 2018
This dataset is a clone of the data provided by hermitdave's FrequencyWords.
The original dataset can be found on https://opus.nlpl.eu/OpenSubtitles2018.php.
Supported languages
The table below shows the ISO codes for the languages that are included in this dataset
Code
Language
sq
Albanian
af
Afrikaans
am
Amharic
ar
Arabic
hy
Armenian
az
Azerbaijani
bn
Bengali
bs
Bosnian
br
Breton
bg
Bulgarian
ca
Catalan
zh_cn
Chinese… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/frequency-words-2018.SPEEED_s3_words_mandarin-750k-850k
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (mandarin).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_mandarin-750k-850k.SPEEED_s3_words_mandarin_0k-130kSPEEED_s3_words_mandarin_540k-750k
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (mandarin).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_mandarin_540k-750k.voxpopuli-test-words-temptask044_essential_terms_identifying_essential_words
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task044_essential_terms_identifying_essential_words
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task044_essential_terms_identifying_essential_words.persian-abusive-words
Persian Abusive Words Dataset
This is a labeled dataset of Persian Abusive Words, originally sourced from Persian Abusive Words GitHub repository. The dataset has been split by the contributor into two subsets: train and test.
This dataset can be used for developing systems to detect and filter offensive or abusive language in various contexts. It is particularly useful for identifying inappropriate words and managing content moderation in applications where Persian language… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaFzp/persian-abusive-words.malayalam-ocr-words
Malayalam OCR Words
A word-level Malayalam OCR dataset: cropped word images paired with their
transcribed text label, split into train/validation/test sets.
Dataset structure
train.csv / val.csv / test.csv # tab-separated: <relative image path>\t<Malayalam word>
train/ val/ test/ # image files referenced by the corresponding CSV
Each CSV row maps one image file (path relative to its split folder) to its
ground-truth Malayalam word transcription.… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-ocr-words.SPEEED_s3_words_japanese_400k-600kSPEEED_s3_words_french_0k-100k
Multilingual Audio Alignments - Processed (Mixed Text/Phonemes)
This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (french).
Curriculum Learning
This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule:
p_start: 0.0 (starting probability of using phonemes)
p_end: 0.0 (ending probability of using phonemes)
curriculum_rows: 400000 (rows over which probability increases)
Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_french_0k-100k.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/muhammadrizo5721/101_billion_arabic_words_dataset.gregg-anniversary-words
Gregg Anniversary Words
This dataset is derived from the 1930 Gregg Shorthand Dictionary1.
Structure
The dataset contains three columns:
image: The image of a shorthand form in the dictionary
grascii_normalized: The normalized grascii of the shorthand form
longhand: The English longhand represented by the shorthand form
Issues
If you notice any problems in the dataset, open an issue in the
datasets repository.
1Gregg, John Robert. Gregg Shorthand Dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/grascii/gregg-anniversary-words.noisy-gt-missing-words
Noisy Ground Truth - Missing Words
Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator.
In Noisy Ground Truth - Missing Words, each variation column is affected by the noise, without considering the split between train, validation and test.
Data structure
The dataset is composed of the… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-missing-words.
