CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tsivakar /polysemous-words Polysemous Words Polysemous Words is a large-scale collection of contextual examples for 200 common polysemous English words. Each word in this dataset has multiple, distinct senses and appears across a wide variety of natural web text contexts, making the dataset ideal for research in word sense induction (WSI), word sense disambiguation (WSD), and probing the contextual understanding capabilities of large language models (LLMs). Dataset Overview This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/tsivakar/polysemous-words.text100M<n<1B1 likes31k downloads1y agoHugging Face02ClusterlabAi /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.texttext-generation10M<n<100M73 likes1.9k downloads2y agoHugging Face03AdoCleanCode /SPEEED_s3_words_french_100k-300ktext100K<n<1M0 likes941 downloads8mo agoHugging Face04AdoCleanCode /SPEEED_s3_words_russian_0k-300ktext100K<n<1M0 likes893 downloads7mo agoHugging Face05AdoCleanCode /SPEEED_s3_words_spanish_1200k-1400k Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (spanish). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_spanish_1200k-1400k.text100K<n<1M0 likes805 downloads7mo agoHugging Face06AdoCleanCode /SPEEED_s3_words_spanish-600k-900text100K<n<1M0 likes792 downloads8mo agoHugging Face07AdoCleanCode /SPEEED_s3_words_russian_300k-600ktext100K<n<1M0 likes698 downloads7mo agoHugging Face08AdoCleanCode /SPEEED_s3_words_thai_0k-250ktext100K<n<1M0 likes698 downloads7mo agoHugging Face09AdoCleanCode /SPEEED_s3_words_italian_0k_200ktext100K<n<1M0 likes697 downloads7mo agoHugging Face10AdoCleanCode /SPEEED_s3_words_italian_200k_400ktext100K<n<1M0 likes659 downloads7mo agoHugging Face11AdoCleanCode /SPEEED_s3_words_spanish_1000k-1200k Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (spanish). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_spanish_1000k-1200k.text100K<n<1M0 likes641 downloads7mo agoHugging Face12sileod /probability_words_nliWords of estimative probability NLI (3 configs). Parquet conversion; CSVs retained. tabular10K<n<100K6 likes592 downloads3d agoHugging Face13AdoCleanCode /SPEEED_s3_words_multi_0-200000 Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (english). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_multi_0-200000.text100K<n<1M0 likes589 downloads8mo agoHugging Face14cyttic /heb-words27-spaced heb-words27-diffpen Word-level Hebrew handwriting for DiffusionPen finetuning -- the step up from cyttic/heb-connected-bigrams17-diffpen, which covered 2-character units. Here the unit is a whole word. Content 20,000 most frequent words of LLMGen2 (sentences_llm2.txt, 968,204 LLM-generated Hebrew sentences, 167,879 unique words) 643 numbers -- LLMGen2 contains no digits at all, because the generation prompt forbade them, so numerals had to be added separately or… See the full description on the dataset page: https://huggingface.co/datasets/cyttic/heb-words27-spaced.imageimage-to-text100K<n<1M0 likes571 downloads3d agoHugging Face15AdoCleanCode /SPEEED_s3_words_spanish_200k-400k Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (spanish). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_spanish_200k-400k.text100K<n<1M0 likes570 downloads8mo agoHugging Face16AdoCleanCode /SPEEED_s3_words_italian_400k_650ktext100K<n<1M0 likes534 downloads7mo agoHugging Face17almogtavor /WordSim353 📖 WordSim353 The WordSim dataset contains word pairs with human annotated similarity scores. The columns are Word 1, Word 2, Human (Mean). 353 rows of word similarieties. Although small, this dataset is very popular and have been used in many researches, including fastText (that called this WS353). Lisense This dataset licensed under a Creative Commons Attribution 4.0 International License. The English WordSim-353 word pairs and instructions are credited to:… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/WordSim353.textn<1K1 likes480 downloads1y agoHugging Face18StephanAkkerman /frequency-words-2018 Frequency Words 2018 This dataset is a clone of the data provided by hermitdave's FrequencyWords. The original dataset can be found on https://opus.nlpl.eu/OpenSubtitles2018.php. Supported languages The table below shows the ISO codes for the languages that are included in this dataset Code Language sq Albanian af Afrikaans am Amharic ar Arabic hy Armenian az Azerbaijani bn Bengali bs Bosnian br Breton bg Bulgarian ca Catalan zh_cn Chinese… See the full description on the dataset page: https://huggingface.co/datasets/StephanAkkerman/frequency-words-2018.text10M<n<100M0 likes456 downloads2y agoHugging Face19AdoCleanCode /SPEEED_s3_words_mandarin-750k-850k Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (mandarin). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_mandarin-750k-850k.text100K<n<1M0 likes431 downloads7mo agoHugging Face20AdoCleanCode /SPEEED_s3_words_mandarin_0k-130ktext100K<n<1M0 likes396 downloads8mo agoHugging Face21AdoCleanCode /SPEEED_s3_words_mandarin_540k-750k Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (mandarin). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_mandarin_540k-750k.text100K<n<1M0 likes395 downloads7mo agoHugging Face22AdrienB134 /voxpopuli-test-words-temptext1M<n<10M0 likes392 downloads2y agoHugging Face23Lots-of-LoRAs /task044_essential_terms_identifying_essential_words Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task044_essential_terms_identifying_essential_words Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task044_essential_terms_identifying_essential_words.texttext-generation1K<n<10K0 likes379 downloads2y agoHugging Face24AlirezaFzp /persian-abusive-words Persian Abusive Words Dataset This is a labeled dataset of Persian Abusive Words, originally sourced from Persian Abusive Words GitHub repository. The dataset has been split by the contributor into two subsets: train and test. This dataset can be used for developing systems to detect and filter offensive or abusive language in various contexts. It is particularly useful for identifying inappropriate words and managing content moderation in applications where Persian language… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaFzp/persian-abusive-words.text10K<n<100K2 likes377 downloads2y agoHugging Face25icfoss /malayalam-ocr-words Malayalam OCR Words A word-level Malayalam OCR dataset: cropped word images paired with their transcribed text label, split into train/validation/test sets. Dataset structure train.csv / val.csv / test.csv # tab-separated: <relative image path>\t<Malayalam word> train/ val/ test/ # image files referenced by the corresponding CSV Each CSV row maps one image file (path relative to its split folder) to its ground-truth Malayalam word transcription.… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/malayalam-ocr-words.imageimage-to-text10K<n<100K0 likes340 downloads8d agoHugging Face26AdoCleanCode /SPEEED_s3_words_japanese_400k-600ktext100K<n<1M0 likes334 downloads8mo agoHugging Face27AdoCleanCode /SPEEED_s3_words_french_0k-100k Multilingual Audio Alignments - Processed (Mixed Text/Phonemes) This dataset contains processed audio alignments from AAdonis/multilingual_audio_alignments (french). Curriculum Learning This dataset uses mixed text/phoneme conditioning with a curriculum learning schedule: p_start: 0.0 (starting probability of using phonemes) p_end: 0.0 (ending probability of using phonemes) curriculum_rows: 400000 (rows over which probability increases) Early in the dataset, more words… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/SPEEED_s3_words_french_0k-100k.text100K<n<1M0 likes305 downloads8mo agoHugging Face28muhammadrizo5721 /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/muhammadrizo5721/101_billion_arabic_words_dataset.texttext-generation10M<n<100M0 likes278 downloads7mo agoHugging Face29grascii /gregg-anniversary-words Gregg Anniversary Words This dataset is derived from the 1930 Gregg Shorthand Dictionary1. Structure The dataset contains three columns: image: The image of a shorthand form in the dictionary grascii_normalized: The normalized grascii of the shorthand form longhand: The English longhand represented by the shorthand form Issues If you notice any problems in the dataset, open an issue in the datasets repository. 1Gregg, John Robert. Gregg Shorthand Dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/grascii/gregg-anniversary-words.imageimage-to-text10K<n<100K0 likes276 downloads5mo agoHugging Face30alix-tz /noisy-gt-missing-words Noisy Ground Truth - Missing Words Dataset of synthetic data for experimentation with noisy ground truth. The text in the dataset is based on Colette's Sido and Les Vignes, also the data was processed prior to generating images with the TextRecognitionDataGenerator. In Noisy Ground Truth - Missing Words, each variation column is affected by the noise, without considering the split between train, validation and test. Data structure The dataset is composed of the… See the full description on the dataset page: https://huggingface.co/datasets/alix-tz/noisy-gt-missing-words.imageimage-to-text10K<n<100K0 likes263 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.