CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01moumeneb1 /large_vocabulary_datasettext100M<n<1B0 likes732 downloads5y agoHugging Face02egcortes /asr-jargon-specialized-vocabulary A Dataset for Evaluating ASR on Specialized Vocabulary Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026). Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code Configs Config Language Description synthetic_terms_en English Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms synthetic_terms_pt Portuguese Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.audioautomatic-speech-recognition10K<n<100K0 likes632 downloads3mo agoHugging Face03Synthyra /SwissProt-Annotation-Vocabulary Swiss-Prot Annotation Vocabulary 2026_02 This release converts a pinned Swiss-Prot snapshot into a versioned protein annotation vocabulary. Stable, namespaced term identifiers are the biological identity. Integer tokens are specific to this vocabulary and grammar version. Release summary Field Value Vocabulary version 2026_02-support10-v1 Grammar version 1 Swiss-Prot release 2026_02 Swiss-Prot release date 2026-06-10 Build date 2026-08-26… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/SwissProt-Annotation-Vocabulary.tabularfeature-extraction1M<n<10M1 likes281 downloads29d agoHugging Face04MEYNG /sango-vocabulary Sango Vocabulary Dataset Dataset Description An open, structured, machine-readable trilingual vocabulary dataset for Sango (ISO 639-1: sg, ISO 639-3: sag), the co-official language of the Central African Republic (with French) and its most widely spoken language. Sango is a creole language with over 5 million speakers, yet it remains severely underrepresented in NLP research and digital resources. This dataset provides trilingual vocabulary entries… See the full description on the dataset page: https://huggingface.co/datasets/MEYNG/sango-vocabulary.texttranslationn<1K1 likes237 downloads20d agoHugging Face05Synthyra /ESMC-6B-SAE-Annotation-Vocabulary-Features Vocabulary interpretations of ESMC-6B SAE features One row for every one of the 16,384 features of biohub/ESMC-6B-sae-layer60-k64-codebook16384, giving the protein annotation vocabulary term that best identifies what the feature detects, together with how well that identification holds on proteins the assignment never saw. This is the counterpart to biohub/ESMC-SAE-Features, produced without a language model. Where that release gives a free-text hypothesis per feature, this… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ESMC-6B-SAE-Annotation-Vocabulary-Features.tabularfeature-extraction100K<n<1M0 likes150 downloads29d agoHugging Face06wannaphong /korean-vocabulary-5000 Koko Korean 5K — Multilingual Vocabulary Dataset 5,000 carefully curated Korean vocabulary entries with English translations, romanization, contextual usage notes, and example sentences. Each entry is also translated into 9 additional languages, giving researchers and developers a high-quality parallel corpus of 50,000 aligned vocabulary records anchored to Korean. The dataset reflects real conversation patterns from K-dramas, K-pop, and everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/korean-vocabulary-5000.texttranslation10K<n<100K0 likes136 downloads5mo agoHugging Face07Lots-of-LoRAs /task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation.texttext-generationn<1K1 likes127 downloads2y agoHugging Face08jaylee8864 /korean-vocabulary-5000 Koko Korean 5K — Multilingual Vocabulary Dataset 5,000 carefully curated Korean vocabulary entries with English translations, romanization, contextual usage notes, and example sentences. Each entry is also translated into 9 additional languages, giving researchers and developers a high-quality parallel corpus of 50,000 aligned vocabulary records anchored to Korean. The dataset reflects real conversation patterns from K-dramas, K-pop, and everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/jaylee8864/korean-vocabulary-5000.texttranslation10K<n<100K0 likes117 downloads5mo agoHugging Face09ryanjosephkamp /ars-magna-vocabulary Ars Magna Vocabulary The exact vocabulary Ars Magna judges words against, so that its claim to find every anagram of your letters can be checked rather than taken on trust. It is three things: English OpenList at one pinned revision, a short, public list of the site's own additions, and a short list of the site's listed forms, the contractions whose letters the search knows. Nothing else. A word the site accepts is in one of them. Beside the words are the term classes: the short… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/ars-magna-vocabulary.textn<1K0 likes106 downloads2d agoHugging Face10lego573402 /bible-vocabulary-difficulty Bible vocabulary-difficulty metrics, 12 translations Per-verse reading-difficulty metrics plus a cross-language book-name table keyed on USFM codes. Produced by bible-reader — see scripts/export_dataset.py. No verse text This dataset contains references and derived metrics only, never verse text. That is deliberate: it keeps translations under copyright (NASB) publishable as derived data, and it keeps the download small. Fetch the texts themselves from their own… See the full description on the dataset page: https://huggingface.co/datasets/lego573402/bible-vocabulary-difficulty.tabulartext-classification100K<n<1M0 likes68 downloads2mo agoHugging Face11wordlevel /toefl-essential-vocabulary-1k 🎓 TOEFL Essential Vocabulary Dataset (AI-Enriched) A meticulously curated, AI-enriched dataset of 1,000 high-frequency academic words essential for the TOEFL iBT, IELTS, and advanced English comprehension. 🌟 Why This Dataset? This dataset is specifically engineered for NLP applications, language learning platforms, and academic research. Each entry includes: Academic Theme: The specific field (e.g., Biology, Sociology) where the word frequently appears. Exact Synonyms:… See the full description on the dataset page: https://huggingface.co/datasets/wordlevel/toefl-essential-vocabulary-1k.texttext-classification1K<n<10K1 likes54 downloads5mo agoHugging Face12lotdpbc /word-orb-vocabulary Word Orb Vocabulary Intelligence Structured vocabulary intelligence for AI agents, educators, and researchers. 162,253 words with pronunciation, etymology, age-appropriate definitions, translations across 47 languages, and ethical context. Dataset Description Word Orb is the world's most comprehensive structured vocabulary dataset designed for AI agents and education technology. Each word entry includes: IPA pronunciation for text-to-speech and phonetics research… See the full description on the dataset page: https://huggingface.co/datasets/lotdpbc/word-orb-vocabulary.tabulartext-classification100K<n<1M0 likes47 downloads6mo agoHugging Face13kalixlouiis /lokaniti-vocabulary-pairs Lokaniti Pali-Burmese Vocabulary Pairs 📖 About the Dataset This dataset provides a clean, deduplicated collection of Pali and Burmese word/phrase pairs extracted directly from the Nissaya breakdowns of the Lokaniti text. It is designed to support lexicography, alignment tasks, and machine translation experiments involving classical Pali and Burmese languages. Total Unique Pairs: 1,640 Data Integrity: 0 null values. 👤 Who Created This dataset and… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/lokaniti-vocabulary-pairs.texttranslation1K<n<10K1 likes43 downloads2mo agoHugging Face14MichaelAnthony /lemonseed-vocabulary lemonseed-vocabulary LemonSeed — WordNet vocabulary Q&A with chain-of-thought definitions. Contents vocab.jsonl (4000 rows) Format JSON Lines (.jsonl), one example per line. Provenance Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella. textquestion-answering1K<n<10K0 likes42 downloads29d agoHugging Face15SwiftieJerry /english-vocabulary-materialsgated English Vocabulary Teaching Materials(雅思与初高中词汇教学资料) 中学段的英语词汇教学资料:雅思分级词汇(预备班 / 一阶 / 二阶 / 三阶)的词汇本、词测本、配套听力录音与听说读讲义,外加初高中词表。 原始材料是 PDF、MP3 和 Excel —— 词表分散在 Excel 的多张工作表里,音频按中文文件名散落各目录,PDF 里的词测没法检索。这份仓库做了两件事:60 个原始文件原样归档不做改动,另外从 Excel 抽出 8257 条结构化词条存成 CSV/JSONL,可以直接读进来做背诵、默写、出题或全文检索。 ⚠️ 版权提醒 这批材料整理自绿新的教研资料,不是原创数据集,著作权归原权利人所有。仓库采用 CC BY-NC-ND 4.0 并开启 gated access:禁止商业使用、再分发、公开镜像与演绎;研究用途允许,但须按第 7 节格式署名。完整条款见第 6 节。 1. 数据总览 指标 数值 清单内文件 85(另有… See the full description on the dataset page: https://huggingface.co/datasets/SwiftieJerry/english-vocabulary-materials.audiofill-mask10K<n<100K1 likes42 downloads27d agoHugging Face16yukiarimo /english-vocabularygatedDataset contains all English words from the dictionary! texttext-classification1M<n<10M7 likes40 downloads2y agoHugging Face17Manoj2702 /Sanskrit-to-English-Vocabulary-v1 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: Manoj, Nandish, Mayank, Abhiram Language(s) (NLP): Sanskrit, English License: MIT License Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More… See the full description on the dataset page: https://huggingface.co/datasets/Manoj2702/Sanskrit-to-English-Vocabulary-v1.text100K<n<1M1 likes27 downloads2y agoHugging Face18ndamulelonemakh /za_vocabularytext10K<n<100K0 likes20 downloads2y agoHugging Face19Lots-of-LoRAs /task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation.texttext-generationn<1K0 likes18 downloads2y agoHugging Face20lemon-mint /advanced_vocabulary_ko_en_v0.1text10K<n<100K0 likes14 downloads2y agoHugging Face21BioMedTok /vocabulary_nachos_lowercasedtext1M<n<10M0 likes13 downloads3y agoHugging Face22KomeijiForce /llama3_vocabulary_clusterThis dataset contains the clusters discovered in the vocabulary embeddings of the llama3-8b-instruct model. The 128256 vocabulary embeddings are separated into 1024 clusters by k-means, which show pattern correlations probably undesirable for diverse generation. We also prompt GPT-4o to summarize the commonality of vocabularies in the same cluster, which can be used for further analysis. This dataset is a part of the work on diverse LLM generation. [Paper], [Github] textsummarization1K<n<10K0 likes13 downloads2y agoHugging Face23supergoose /flan_combined_task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generationtext1K<n<10K0 likes13 downloads2y agoHugging Face24abdelhaqueidali /Amazigh_Researchers_Vocabulary Mohammed Lchger Vocabulary Dataset This dataset contains a collection of vocabulary compiled by Mohamed Lachgar, the owner of the Amazigh researchers blog which its dataset can be find here. Dataset Details Content: 2,000 scientific related terms and general vocabulary. Languages: Amazigh (zgh, ber) and English. Script: Tifinagh. Unique Feature: It is unique especially in its inclusion of scientific vocabulary. Acknowledgments Thanks to Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh_Researchers_Vocabulary.texttext-classification1K<n<10K1 likes13 downloads3mo agoHugging Face25fontan /anyfeature_vocabularytext1M<n<10M0 likes12 downloads2y agoHugging Face26KomeijiForce /olmo_vocabulary_clusterThis dataset contains the clusters discovered in the vocabulary embeddings of the olmo-7b-sft model. The 50280 vocabulary embeddings are separated into 512 clusters by k-means, which show pattern correlations probably undesirable for diverse generation. We also prompt GPT-4o to summarize the commonality of vocabularies in the same cluster, which can be used for further analysis. This dataset is a part of the work on diverse LLM generation. [Paper], [Github] textn<1K0 likes11 downloads2y agoHugging Face270009-0004-2896-8766 /vocabulary-storetextn<1K0 likes10 downloads9mo agoHugging Face28Gargaz /vocabularytext1K<n<10K1 likes9 downloads2y agoHugging Face29omar95 /wikimedia_es_vocabularytext1M<n<10M0 likes8 downloads11mo agoHugging Face30Oichin /wellness-vocabularyWellness & Healthcare Terminology Dataset This dataset contains a curated list of wellness terms, product categories, and SEO-optimized keywords related to high-quality healthcare products. Purpose: To help AI models understand the nuances of premium Japanese wellness standards and improve translation accuracy for the Vietnamese market. Maintained by oichin.net textn<1K0 likes8 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.