CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01palshub /phonemizer-dicts Phonemizer Dicts Pre-generated IPA dictionaries for GPL-free text-to-phonemes lookup. Files en-us.tsv — 124K English (US) words, tab-separated word<TAB>IPA Provenance Generated by running espeak-ng over an English wordlist. The TSV is program output; espeak-ng source (GPL-3.0) is not redistributed here. Regeneration See scripts/generate-espeak-dict.py in the tts-rd-team repo. text100K<n<1M0 likes23k downloads5mo agoHugging Face02TigreGotico /barranquenho-ipa-dict-synthetic Barranquenho IPA Pronunciation Dictionary The first and only IPA pronunciation dictionary of Barranquenho — the Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal), a mixed system born of centuries of Portuguese–Spanish (Extremaduran / Andalusian) contact on the raia. Every headword is written in the Convenção Ortográfica do Barranquenho (2025) orthography and paired with a broad-phonemic IPA transcription plus Portuguese and Spanish glosses. This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.texttext-to-speech1K<n<10K0 likes592 downloads2mo agoHugging Face03fraug-library /synonyms_dictionnaries Description Apache OpenOffice dictionnaries text1M<n<10M3 likes530 downloads2y agoHugging Face04taresco /afri-dict Afri-Dict Dataset Summary afri-dict is a bilingual dictionary dataset for four major African languages: Hausa, Igbo, Swahili, and Yoruba. Entries include a headword, part-of-speech tag, and definition in English or the target African language. This dataset can serve as a foundational resource for machine translation systems, language learning tools, spell checkers, cross-lingual search, and other NLP applications for African languages. Languages… See the full description on the dataset page: https://huggingface.co/datasets/taresco/afri-dict.texttranslation10K<n<100K0 likes448 downloads1mo agoHugging Face05binjang /NIKL-korean-english-dictionary Column Name Type Description 설명 Form str Registered word entry 단어 Part of Speech str or None Part of speech of the word in Korean 품사 Korean Definition List[str] Definition of the word in Korean 해당 단어의 한글 정의 English Definition List[str] or None Definition of the word in English 한글 정의의 영문 번역본 Usages List[str] or None Sample sentence or dialogue 해당 단어의 예문 (문장 또는 대화 형식) Vocabulary Level str or None Difficulty of the word (3 levels) 단어의 난이도 ('초급', '중급', '고급') Semantic… See the full description on the dataset page: https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary.texttranslation10K<n<100K7 likes219 downloads3y agoHugging Face06kavyamanohar /Pronunciation-dictionary-malayalam Malayalam Pronunciation Dictionary This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here It gives Phonemic transcription of Malayalam words in IPA format. Dataset Details Dataset Description This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon] (https://pypi.org/project/mlphon/) Python library. Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.text100K<n<1M2 likes151 downloads2y agoHugging Face07obaydata /ths-quant-factor-dictionary THS Quant Factor Dictionary (同花顺量化因子字典) Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors. These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data. Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.tabularn<1K0 likes134 downloads6mo agoHugging Face08SmartQHSE /hse-acronym-dictionary-2026 Canonical landing page: https://www.smartqhse.com/datasets/hse-acronym-dictionary-2026 HSE Acronym Dictionary 2026 Authoritative reference of 150+ HSE / EHS / occupational-safety acronyms with one-line definitions. Covers metrics (TRIR, LTIFR, DART, EMR, WBGT), regulations (OSHA, RIDDOR, COSHH, CDM, COMAH, PSM, OSHAD-SF), methodologies (HAZOP, HAZID, LOPA, SIL, FMEA, RCA, BBS, ICAM, TapRooT, Tripod Beta), bodies (NEBOSH, IOSH, IIRSM, IOGP, ACGIH, NIOSH, ANSI, ASSP), and… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-acronym-dictionary-2026.textn<1K0 likes131 downloads4mo agoHugging Face09fa0311 /warabi-dictionary-extended-gpl IME Dictionary Extended GPL Japanese IME dictionary pack. The SQLite database and TSV files contain the same conversion and prediction records. See NOTICE.md before redistribution. Compilation license: GPL-3.0-only SQLite tables: entries, entry_sources, predictions, prediction_sources TSV exports: entries.tsv, predictions.tsv Kaomoji flag The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter with WHERE NOT kaomoji for a face-free dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-extended-gpl.text1M<n<10M0 likes98 downloads1mo agoHugging Face10fa0311 /warabi-dictionary-core IME Dictionary Core Japanese IME dictionary pack. The SQLite database and TSV files contain the same conversion and prediction records. See NOTICE.md before redistribution. Compilation license: CC-BY-4.0 SQLite tables: entries, entry_sources, predictions, prediction_sources TSV exports: entries.tsv, predictions.tsv Kaomoji flag The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter with WHERE NOT kaomoji for a face-free dictionary. A candidate… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-core.text1M<n<10M0 likes89 downloads1mo agoHugging Face11samuraijun /shitto-mania-dic English Summary This dataset accompanies our NLP2026 study on language resource design in the RAG era. Using a 321-episode Japanese dataset derived from the essay series Shitto Mania, we found that structured metadata substantially outperformed full text in the tested reference-retrieval setting. Key finding: structured metadata alone achieves 11.1× better retrieval performance than full text in TF-IDF retrieval (59.0% vs 5.3% Recall@10, STRUCT queries, n=300). This advantage… See the full description on the dataset page: https://huggingface.co/datasets/samuraijun/shitto-mania-dic.tabulartext-retrieval1K<n<10K0 likes86 downloads1mo agoHugging Face12mmi01 /DICE DICE: Dataset for Controlled Evaluation of Idiomatic Expressions Summary DICE is a corpus for testing models' understanding of potentially idiomatic expressions (PIEs) in contexts, and potential effects of memorisation. PIEs are expressions that can be interpreted with a figurative meaning or a literal meaning, depending on the context in which they occur. For instance, "let the cat out of the bag" could mean "revealing a secret", it could also hold a more literal… See the full description on the dataset page: https://huggingface.co/datasets/mmi01/DICE.text1K<n<10K1 likes80 downloads11mo agoHugging Face13i-Lang /iLang-Dict iLang Dictionary v5.0 88 verbs. 13 Greek aliases. 29 core modifiers plus a 20-key media profile. 25 entities (17 addressable, 8 role). 32 declarations. Includes the v5.0 judgment vocabulary and the v4.1 media profile for image, video and audio. The complete verb dictionary for iLang v5.0, the native language of artificial intelligence. It reduces semantic loss between human intent and machine execution. iLang is the first protocol to formally map Greek mathematical symbols (Σ… See the full description on the dataset page: https://huggingface.co/datasets/i-Lang/iLang-Dict.textn<1K0 likes80 downloads15h agoHugging Face14DusunDictionary /dusun-dictionary-corpus Dusun-English-Malay Dictionary and Corpus A trilingual dataset containing Dusun, English, and Malay words, phrases, and sentences compiled for linguistic research, dictionary development, and machine translation. This dataset is based on the Dusun language as spoken by the Dusun ethnic group of Sabah, Malaysia. While the Dusun dialect in the dataset shares approximately 99% similarity with standardized Kadazandusun, there may be minor differences in vocabulary, spelling, and… See the full description on the dataset page: https://huggingface.co/datasets/DusunDictionary/dusun-dictionary-corpus.texttranslation10K<n<100K3 likes74 downloads1mo agoHugging Face15agenticx /DICTranktext1K<n<10K0 likes65 downloads1y agoHugging Face16nhatnguyet /64-que-kinh-dich 64 quẻ Kinh Dịch The 64 hexagrams of the I Ching 1. Mô tả · Description Đủ 64 quẻ theo thứ tự Chu Dịch, kèm tên Việt, tên Hán, quẻ thượng, quẻ hạ và tượng quẻ. All 64 hexagrams in King Wen order, with Vietnamese and Han names, upper and lower trigrams, and the image. Số dòng · Rows: 64 Phiên bản · Version: 1.0.0 (2026-09-16) Mã hoá · Encoding: UTF-8 không BOM 2. Cấu trúc · Structure Cột · Column Kiểu · Type Ý nghĩa · Meaning id string… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/64-que-kinh-dich.tabularn<1K0 likes62 downloads2d agoHugging Face17MAKILINGDING /english_dictionarytext100K<n<1M16 likes61 downloads3y agoHugging Face18aaljuhani /TCGA-SARC-dict-tumortabular10M<n<100M0 likes58 downloads9mo agoHugging Face19MedicalSAM3 /sam3-low-dice-2d-nnunet SAM3 low-Dice 2D datasets for nnU-Net Private research export of two small 2D datasets on which the balanced-finish SAM3 LoRA validation Dice was below 0.5. The purpose is to test whether a dataset-specific nnU-Net can fit these data and to distinguish data/training limitations from inference bugs. Dataset SAM3 Dice SAM3 IoU Evaluated validation images Actual SAM3 training images DRIVE 0.212233 0.118717 2 14 RAVIR 0.224709 0.128455 2 16 The two-image validation… See the full description on the dataset page: https://huggingface.co/datasets/MedicalSAM3/sam3-low-dice-2d-nnunet.imageimage-segmentationn<1K0 likes54 downloads15d agoHugging Face20Tamajeq1286 /Tamajaq-English-dictionarytext10K<n<100K1 likes53 downloads2mo agoHugging Face21YuviBitts /marathi-dictionary Marathi Dictionary Marathi-to-Marathi Dictionary Dataset Description Synthetically generated meanings in marathi for over 38k words text10K<n<100K0 likes50 downloads1y agoHugging Face22VedatTurkkal /english_career_occupation_terminology_dictionarytext100K<n<1M0 likes47 downloads20d agoHugging Face23TigreGotico /dicionario_barranquenho Dicionário de Barranquenho Structured lexical dataset derived from the first published dictionary of Barranquenho, a Romance contact language spoken in Barrancos, Portugal. Contains 1,680 entries with Portuguese and Spanish glosses, grammatical categories, semantic fields, source attributions, and synonym cross-references. Language Barranquenho (glottocode: barr1245; no ISO 639-3 code assigned at time of publication) is a contact language spoken in the municipality of… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/dicionario_barranquenho.documenttranslation1K<n<10K0 likes46 downloads5mo agoHugging Face24cyberali32112 /healthcare-data-dictionary Healthcare Data Dictionary — ISO-11179 Standard Terms A sample dataset of 5,000 standardized healthcare data column names, abbreviations, and definitions for data engineers building on Snowflake, Databricks, BigQuery, and dbt. Dataset Description This dataset is a sample from the mdatool Healthcare Data Dictionary — the most comprehensive ISO-11179 compliant healthcare data dictionary available for data engineers. What is ISO-11179? ISO-11179 is… See the full description on the dataset page: https://huggingface.co/datasets/cyberali32112/healthcare-data-dictionary.text1K<n<10K0 likes39 downloads5d agoHugging Face25afrilang-edu /afrilang-dicttextmultiple-choice100K<n<1M0 likes35 downloads1y agoHugging Face26mrrtmob /english-khmer-dictionary 📖 English–Khmer Dictionary Dataset A comprehensive bilingual English–Khmer (ភាសាខ្មែរ) dictionary dataset in CSV format containing 170,000+ entries. Each entry includes the original English word, its Khmer translation, part of speech, full definitions in both languages, and example sentences — making it one of the richer English–Khmer lexical resources available for NLP and language learning. Dataset Description This dataset provides structured dictionary entries pairing… See the full description on the dataset page: https://huggingface.co/datasets/mrrtmob/english-khmer-dictionary.tabulartranslation100K<n<1M1 likes33 downloads7mo agoHugging Face27W4ashabii /root_affix_dictionary Nepali Root-Affix Dictionary A word-level morphological segmentation dataset for Nepali, mapping surface words to their root and affix components (e.g. अँकाइनु → अँका + इनु). Dataset Details Dataset Description This dataset was built by scraping the 10th edition of the Nepali dictionary, parsing the scraped entries into structured word/root pairs, and then applying a heuristic rule engine to extend root-affix coverage to surface forms not explicitly… See the full description on the dataset page: https://huggingface.co/datasets/W4ashabii/root_affix_dictionary.texttext-generation10K<n<100K0 likes33 downloads2mo agoHugging Face28researchaudio /mac-dictation-privacy-matrix Mac Dictation Privacy Matrix An open, source-reviewed dataset comparing the documented privacy boundaries of 18 Mac dictation products. The matrix separates questions that are often collapsed into one label: where microphone audio becomes a transcript; whether an optional cloud, cleanup, assistant, or agent path exists; whether the documented speech path works offline after setup; what the publisher says it retains; which account, subscription, license, or provider boundary… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/mac-dictation-privacy-matrix.textn<1K0 likes33 downloads2mo agoHugging Face29OTAR3088 /CellTissue-dict CellTissue-dict This dictionary is comprised of a merge of two source dictionaries (from Bioportal and ChEMBL), used to tag cell and tissue descriptions in text. Dataset Details Dataset Sources This dictionary was produced by a combination of quick data sourcing, followed by partial manual categorising of terms. Repository: https://github.com/ML4LitS/OTAR3088 Direct Use This data may be used to tag phrases in natural language, categorising… See the full description on the dataset page: https://huggingface.co/datasets/OTAR3088/CellTissue-dict.texttext-classification1K<n<10K0 likes32 downloads5mo agoHugging Face30wannaphong /thai-mm-dict 🗂️ Dataset Card: Thai-Myanmar Dictionary (2025) 📝 Dataset Summary The Thai-Myanmar Dictionary (2025) is a high-quality bilingual lexical dataset created by Htet Myet Lynn.It provides direct word-to-word and phrase mappings between Thai and Myanmar (Burmese), supporting both linguistic use and machine learning applications. The dataset is released under the MIT License, allowing free usage, modification, redistribution, and integration into both academic and commercial… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-mm-dict.text10K<n<100K0 likes32 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.