CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01moumeneb1 /large_vocabulary_datasettext100M<n<1B0 likes732 downloads5y agoHugging Face02egcortes /asr-jargon-specialized-vocabulary A Dataset for Evaluating ASR on Specialized Vocabulary Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026). Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code Configs Config Language Description synthetic_terms_en English Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms synthetic_terms_pt Portuguese Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.audioautomatic-speech-recognition10K<n<100K0 likes632 downloads3mo agoHugging Face03Synthyra /SwissProt-Annotation-Vocabulary Swiss-Prot Annotation Vocabulary 2026_02 This release converts a pinned Swiss-Prot snapshot into a versioned protein annotation vocabulary. Stable, namespaced term identifiers are the biological identity. Integer tokens are specific to this vocabulary and grammar version. Release summary Field Value Vocabulary version 2026_02-support10-v1 Grammar version 1 Swiss-Prot release 2026_02 Swiss-Prot release date 2026-06-10 Build date 2026-08-26… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/SwissProt-Annotation-Vocabulary.tabularfeature-extraction1M<n<10M1 likes281 downloads29d agoHugging Face04MEYNG /sango-vocabulary Sango Vocabulary Dataset Dataset Description An open, structured, machine-readable trilingual vocabulary dataset for Sango (ISO 639-1: sg, ISO 639-3: sag), the co-official language of the Central African Republic (with French) and its most widely spoken language. Sango is a creole language with over 5 million speakers, yet it remains severely underrepresented in NLP research and digital resources. This dataset provides trilingual vocabulary entries… See the full description on the dataset page: https://huggingface.co/datasets/MEYNG/sango-vocabulary.texttranslationn<1K1 likes237 downloads20d agoHugging Face05Synthyra /ESMC-6B-SAE-Annotation-Vocabulary-Features Vocabulary interpretations of ESMC-6B SAE features One row for every one of the 16,384 features of biohub/ESMC-6B-sae-layer60-k64-codebook16384, giving the protein annotation vocabulary term that best identifies what the feature detects, together with how well that identification holds on proteins the assignment never saw. This is the counterpart to biohub/ESMC-SAE-Features, produced without a language model. Where that release gives a free-text hypothesis per feature, this… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ESMC-6B-SAE-Annotation-Vocabulary-Features.tabularfeature-extraction100K<n<1M0 likes150 downloads28d agoHugging Face06YangCaoCS /Open-Vocabulary-ScanNetThe Open-Vocabulary ScanNet datasets from CoDA and CoDAv2. If the dataset is helpful, please cite: @inproceedings{dai2017scannet, title={ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes}, author={Dai, Angela and Chang, Angel X. and Savva, Manolis and Halber, Maciej and Funkhouser, Thomas and Nie{\ss}ner, Matthias}, booktitle = {Proc. Computer Vision and Pattern Recognition (CVPR), IEEE}, year = {2017} } @inproceedings{cao2023coda, title={CoDA: Collaborative… See the full description on the dataset page: https://huggingface.co/datasets/YangCaoCS/Open-Vocabulary-ScanNet.0 likes146 downloads7mo agoHugging Face07wannaphong /korean-vocabulary-5000 Koko Korean 5K — Multilingual Vocabulary Dataset 5,000 carefully curated Korean vocabulary entries with English translations, romanization, contextual usage notes, and example sentences. Each entry is also translated into 9 additional languages, giving researchers and developers a high-quality parallel corpus of 50,000 aligned vocabulary records anchored to Korean. The dataset reflects real conversation patterns from K-dramas, K-pop, and everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/korean-vocabulary-5000.texttranslation10K<n<100K0 likes136 downloads5mo agoHugging Face08Lots-of-LoRAs /task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation.texttext-generationn<1K1 likes127 downloads2y agoHugging Face09jaylee8864 /korean-vocabulary-5000 Koko Korean 5K — Multilingual Vocabulary Dataset 5,000 carefully curated Korean vocabulary entries with English translations, romanization, contextual usage notes, and example sentences. Each entry is also translated into 9 additional languages, giving researchers and developers a high-quality parallel corpus of 50,000 aligned vocabulary records anchored to Korean. The dataset reflects real conversation patterns from K-dramas, K-pop, and everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/jaylee8864/korean-vocabulary-5000.texttranslation10K<n<100K0 likes117 downloads5mo agoHugging Face10ryanjosephkamp /ars-magna-vocabulary Ars Magna Vocabulary The exact vocabulary Ars Magna judges words against, so that its claim to find every anagram of your letters can be checked rather than taken on trust. It is three things: English OpenList at one pinned revision, a short, public list of the site's own additions, and a short list of the site's listed forms, the contractions whose letters the search knows. Nothing else. A word the site accepts is in one of them. Beside the words are the term classes: the short… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/ars-magna-vocabulary.textn<1K0 likes106 downloads2d agoHugging Face11aborasheed /Multilingual-Core-Vocabulary Multilingual Core Vocabulary 🌍 This dataset contains millions of frequency-sorted, highly accurate words across 19 languages. It is designed to be the ultimate resource for building cross-lingual applications, AI similarity agents, and translation models. Dataset Structure This repository contains two variations of the dataset: Massive_Dataset: Contains over 6 million words. The words were extracted and frequency-sorted from FastText, cleaned from internet noise… See the full description on the dataset page: https://huggingface.co/datasets/aborasheed/Multilingual-Core-Vocabulary.translation0 likes101 downloads3mo agoHugging Face12mustafaalkanxgmail /Multilingual-Core-Vocabulary Multilingual Core Vocabulary 🌍 This dataset contains millions of frequency-sorted, highly accurate words across 19 languages. It is designed to be the ultimate resource for building cross-lingual applications, AI similarity agents, and translation models. Dataset Structure This repository contains two variations of the dataset: Massive_Dataset: Contains over 6 million words. The words were extracted and frequency-sorted from FastText, cleaned from internet noise… See the full description on the dataset page: https://huggingface.co/datasets/mustafaalkanxgmail/Multilingual-Core-Vocabulary.translation0 likes78 downloads2mo agoHugging Face13ITOTII /tavern-canvas-vocabulary Tavern Canvas Vocabulary Packages Prebuilt vocabulary packages for the Tavern Canvas SillyTavern extension. The extension reads catalog.json to list installable packages; each package under packages/<package_id>/ contains a manifest.json plus MessagePack+gzip index shards, verified per-shard by SHA-256 at install time. Contents package_id records languages bytes note tavern-canvas-baseline 50,000 en, zh-CN 3.9 MB bundled with the extension, listed here… See the full description on the dataset page: https://huggingface.co/datasets/ITOTII/tavern-canvas-vocabulary.0 likes73 downloads2mo agoHugging Face14lego573402 /bible-vocabulary-difficulty Bible vocabulary-difficulty metrics, 12 translations Per-verse reading-difficulty metrics plus a cross-language book-name table keyed on USFM codes. Produced by bible-reader — see scripts/export_dataset.py. No verse text This dataset contains references and derived metrics only, never verse text. That is deliberate: it keeps translations under copyright (NASB) publishable as derived data, and it keeps the download small. Fetch the texts themselves from their own… See the full description on the dataset page: https://huggingface.co/datasets/lego573402/bible-vocabulary-difficulty.tabulartext-classification100K<n<1M0 likes68 downloads2mo agoHugging Face15wordlevel /toefl-essential-vocabulary-1k 🎓 TOEFL Essential Vocabulary Dataset (AI-Enriched) A meticulously curated, AI-enriched dataset of 1,000 high-frequency academic words essential for the TOEFL iBT, IELTS, and advanced English comprehension. 🌟 Why This Dataset? This dataset is specifically engineered for NLP applications, language learning platforms, and academic research. Each entry includes: Academic Theme: The specific field (e.g., Biology, Sociology) where the word frequently appears. Exact Synonyms:… See the full description on the dataset page: https://huggingface.co/datasets/wordlevel/toefl-essential-vocabulary-1k.texttext-classification1K<n<10K1 likes54 downloads5mo agoHugging Face16lotdpbc /word-orb-vocabulary Word Orb Vocabulary Intelligence Structured vocabulary intelligence for AI agents, educators, and researchers. 162,253 words with pronunciation, etymology, age-appropriate definitions, translations across 47 languages, and ethical context. Dataset Description Word Orb is the world's most comprehensive structured vocabulary dataset designed for AI agents and education technology. Each word entry includes: IPA pronunciation for text-to-speech and phonetics research… See the full description on the dataset page: https://huggingface.co/datasets/lotdpbc/word-orb-vocabulary.tabulartext-classification100K<n<1M0 likes47 downloads6mo agoHugging Face17kalixlouiis /lokaniti-vocabulary-pairs Lokaniti Pali-Burmese Vocabulary Pairs 📖 About the Dataset This dataset provides a clean, deduplicated collection of Pali and Burmese word/phrase pairs extracted directly from the Nissaya breakdowns of the Lokaniti text. It is designed to support lexicography, alignment tasks, and machine translation experiments involving classical Pali and Burmese languages. Total Unique Pairs: 1,640 Data Integrity: 0 null values. 👤 Who Created This dataset and… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/lokaniti-vocabulary-pairs.texttranslation1K<n<10K1 likes43 downloads2mo agoHugging Face18MichaelAnthony /lemonseed-vocabulary lemonseed-vocabulary LemonSeed — WordNet vocabulary Q&A with chain-of-thought definitions. Contents vocab.jsonl (4000 rows) Format JSON Lines (.jsonl), one example per line. Provenance Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella. textquestion-answering1K<n<10K0 likes42 downloads29d agoHugging Face19SwiftieJerry /english-vocabulary-materialsgated English Vocabulary Teaching Materials(雅思与初高中词汇教学资料) 中学段的英语词汇教学资料:雅思分级词汇(预备班 / 一阶 / 二阶 / 三阶)的词汇本、词测本、配套听力录音与听说读讲义,外加初高中词表。 原始材料是 PDF、MP3 和 Excel —— 词表分散在 Excel 的多张工作表里,音频按中文文件名散落各目录,PDF 里的词测没法检索。这份仓库做了两件事:60 个原始文件原样归档不做改动,另外从 Excel 抽出 8257 条结构化词条存成 CSV/JSONL,可以直接读进来做背诵、默写、出题或全文检索。 ⚠️ 版权提醒 这批材料整理自绿新的教研资料,不是原创数据集,著作权归原权利人所有。仓库采用 CC BY-NC-ND 4.0 并开启 gated access:禁止商业使用、再分发、公开镜像与演绎;研究用途允许,但须按第 7 节格式署名。完整条款见第 6 节。 1. 数据总览 指标 数值 清单内文件 85(另有… See the full description on the dataset page: https://huggingface.co/datasets/SwiftieJerry/english-vocabulary-materials.audiofill-mask10K<n<100K1 likes42 downloads27d agoHugging Face20yukiarimo /english-vocabularygatedDataset contains all English words from the dictionary! texttext-classification1M<n<10M7 likes40 downloads2y agoHugging Face21YangCaoCS /Open-Vocabulary-SUN-RGBDThe Open-Vocabulary SUN-RGBD datasets from CoDA and CoDAv2. If the dataset is helpful, please cite: @inproceedings{song2015sun, title={Sun rgb-d: A rgb-d scene understanding benchmark suite}, author={Song, Shuran and Lichtenberg, Samuel P and Xiao, Jianxiong}, booktitle={CVPR}, year={2015} } @inproceedings{cao2023coda, title={CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection}, author={Cao, Yang and Zeng, Yihan and Xu, Hang… See the full description on the dataset page: https://huggingface.co/datasets/YangCaoCS/Open-Vocabulary-SUN-RGBD.0 likes36 downloads8mo agoHugging Face22xqt /jlpt_n5_vocabulary Jisho JLPT-N5 Relational Dataset This dataset provides a comprehensive and highly normalized collection of Japanese vocabulary for the JLPT-N5 level, scraped from Jisho.org. Unlike flat datasets, this version uses a relational schema to separate headwords, metadata (JLPT/WaniKani levels), and individual English definitions. 🏗 Scraper Architecture The data was generated using a custom Python scraper following a robust state-machine logic. Logic Highlights:… See the full description on the dataset page: https://huggingface.co/datasets/xqt/jlpt_n5_vocabulary.translation1K<n<10K0 likes31 downloads6mo agoHugging Face23xqt /jlpt_n5_vocabulary_tagged Jisho JLPT-N5 Relational Dataset This dataset provides a comprehensive and highly normalized collection of Japanese vocabulary for the JLPT-N5 level, scraped from Jisho.org and enriched with DeepSeek-V3 AI tagging. 🏗 Scraper & AI Architecture The data was generated using a custom Python scraper and post-processed using a batched AI tagging pipeline. Logic Highlights: Normalization: Every English sense is a unique row with a specific meaning_no.… See the full description on the dataset page: https://huggingface.co/datasets/xqt/jlpt_n5_vocabulary_tagged.translation1K<n<10K0 likes31 downloads6mo agoHugging Face24Manoj2702 /Sanskrit-to-English-Vocabulary-v1 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: Manoj, Nandish, Mayank, Abhiram Language(s) (NLP): Sanskrit, English License: MIT License Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More… See the full description on the dataset page: https://huggingface.co/datasets/Manoj2702/Sanskrit-to-English-Vocabulary-v1.text100K<n<1M1 likes27 downloads2y agoHugging Face25FormosanBank /ePark_xue_xi_ci_biao_learning_vocabulary FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum. This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary.audioautomatic-speech-recognition10K<n<100K0 likes26 downloads2mo agoHugging Face26ndamulelonemakh /za_vocabularytext10K<n<100K0 likes20 downloads2y agoHugging Face27Lots-of-LoRAs /task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation.texttext-generationn<1K0 likes18 downloads2y agoHugging Face28dibo /hatevolution-vocabulary-expansion Dataset Info The hatevolution-vocabulary-expansion dataset contains the data used for Experiment 2 in the paper Hatevolution: What Static Benchmarks Don't Tell Us (Di Bonaventura et al., 2025). It is built using the NeoBench dataset (Zheng et al., 2024), further annotated for hate speech detection. text-classificationn<1K0 likes17 downloads1y agoHugging Face29lemon-mint /advanced_vocabulary_ko_en_v0.1text10K<n<100K0 likes14 downloads2y agoHugging Face30BioMedTok /vocabulary_nachos_lowercasedtext1M<n<10M0 likes13 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.