datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phonemizer-dicts
Phonemizer Dicts
Pre-generated IPA dictionaries for GPL-free text-to-phonemes lookup.
Files
en-us.tsv — 124K English (US) words, tab-separated word<TAB>IPA
Provenance
Generated by running espeak-ng over an English wordlist. The TSV is program output; espeak-ng source (GPL-3.0) is not redistributed here.
Regeneration
See scripts/generate-espeak-dict.py in the tts-rd-team repo.
barranquenho-ipa-dict-synthetic
Barranquenho IPA Pronunciation Dictionary
The first and only IPA pronunciation dictionary of Barranquenho — the
Ibero-Romance contact variety spoken in Barrancos (Baixo Alentejo, Portugal),
a mixed system born of centuries of Portuguese–Spanish (Extremaduran /
Andalusian) contact on the raia. Every headword is written in the
Convenção Ortográfica do Barranquenho (2025) orthography and paired with a
broad-phonemic IPA transcription plus Portuguese and Spanish glosses.
This… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/barranquenho-ipa-dict-synthetic.synonyms_dictionnaries
Description
Apache OpenOffice dictionnaries
afri-dict
Afri-Dict
Dataset Summary
afri-dict is a bilingual dictionary dataset for four major African languages: Hausa, Igbo, Swahili, and Yoruba.
Entries include a headword, part-of-speech tag, and definition in English or the target African language.
This dataset can serve as a foundational resource for machine translation systems, language learning tools, spell checkers, cross-lingual search, and other NLP applications for African languages.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/taresco/afri-dict.NIKL-korean-english-dictionary
Column Name
Type
Description
설명
Form
str
Registered word entry
단어
Part of Speech
str or None
Part of speech of the word in Korean
품사
Korean Definition
List[str]
Definition of the word in Korean
해당 단어의 한글 정의
English Definition
List[str] or None
Definition of the word in English
한글 정의의 영문 번역본
Usages
List[str] or None
Sample sentence or dialogue
해당 단어의 예문 (문장 또는 대화 형식)
Vocabulary Level
str or None
Difficulty of the word (3 levels)
단어의 난이도 ('초급', '중급', '고급')
Semantic… See the full description on the dataset page: https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary.Pronunciation-dictionary-malayalam
Malayalam Pronunciation Dictionary
This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here
It gives Phonemic transcription of Malayalam words in IPA format.
Dataset Details
Dataset Description
This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon]
(https://pypi.org/project/mlphon/) Python library.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.ths-quant-factor-dictionary
THS Quant Factor Dictionary (同花顺量化因子字典)
Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors.
These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data.
Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.hse-acronym-dictionary-2026
Canonical landing page: https://www.smartqhse.com/datasets/hse-acronym-dictionary-2026
HSE Acronym Dictionary 2026
Authoritative reference of 150+ HSE / EHS / occupational-safety acronyms with one-line definitions. Covers metrics (TRIR, LTIFR, DART, EMR, WBGT), regulations (OSHA, RIDDOR, COSHH, CDM, COMAH, PSM, OSHAD-SF), methodologies (HAZOP, HAZID, LOPA, SIL, FMEA, RCA, BBS, ICAM, TapRooT, Tripod Beta), bodies (NEBOSH, IOSH, IIRSM, IOGP, ACGIH, NIOSH, ANSI, ASSP), and… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-acronym-dictionary-2026.warabi-dictionary-extended-gpl
IME Dictionary Extended GPL
Japanese IME dictionary pack. The SQLite database and TSV files contain the
same conversion and prediction records. See NOTICE.md before redistribution.
Compilation license: GPL-3.0-only
SQLite tables: entries, entry_sources, predictions, prediction_sources
TSV exports: entries.tsv, predictions.tsv
Kaomoji flag
The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter
with WHERE NOT kaomoji for a face-free dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-extended-gpl.warabi-dictionary-core
IME Dictionary Core
Japanese IME dictionary pack. The SQLite database and TSV files contain the
same conversion and prediction records. See NOTICE.md before redistribution.
Compilation license: CC-BY-4.0
SQLite tables: entries, entry_sources, predictions, prediction_sources
TSV exports: entries.tsv, predictions.tsv
Kaomoji flag
The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter
with WHERE NOT kaomoji for a face-free dictionary. A candidate… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-core.shitto-mania-dic
English Summary
This dataset accompanies our NLP2026 study on language resource design in the RAG era.
Using a 321-episode Japanese dataset derived from the essay series Shitto Mania, we found that structured metadata substantially outperformed full text in the tested reference-retrieval setting.
Key finding: structured metadata alone achieves 11.1× better retrieval performance than full text in TF-IDF retrieval (59.0% vs 5.3% Recall@10, STRUCT queries, n=300). This advantage… See the full description on the dataset page: https://huggingface.co/datasets/samuraijun/shitto-mania-dic.DICE
DICE: Dataset for Controlled Evaluation of Idiomatic Expressions
Summary
DICE is a corpus for testing models' understanding of potentially idiomatic expressions (PIEs) in contexts, and potential effects of memorisation.
PIEs are expressions that can be interpreted with a figurative meaning or a literal meaning, depending on the context in which they occur.
For instance, "let the cat out of the bag" could mean "revealing a secret", it could also hold a more literal… See the full description on the dataset page: https://huggingface.co/datasets/mmi01/DICE.iLang-Dict
iLang Dictionary v5.0
88 verbs. 13 Greek aliases. 29 core modifiers plus a 20-key media profile. 25 entities (17 addressable, 8 role). 32 declarations. Includes the v5.0 judgment vocabulary and the v4.1 media profile for image, video and audio.
The complete verb dictionary for iLang v5.0, the native language of artificial intelligence. It reduces semantic loss between human intent and machine execution. iLang is the first protocol to formally map Greek mathematical symbols (Σ… See the full description on the dataset page: https://huggingface.co/datasets/i-Lang/iLang-Dict.dusun-dictionary-corpus
Dusun-English-Malay Dictionary and Corpus
A trilingual dataset containing Dusun, English, and Malay words, phrases, and sentences compiled for linguistic research, dictionary development, and machine translation.
This dataset is based on the Dusun language as spoken by the Dusun ethnic group of Sabah, Malaysia. While the Dusun dialect in the dataset shares approximately 99% similarity with standardized Kadazandusun, there may be minor differences in vocabulary, spelling, and… See the full description on the dataset page: https://huggingface.co/datasets/DusunDictionary/dusun-dictionary-corpus.DICTrank64-que-kinh-dich
64 quẻ Kinh Dịch
The 64 hexagrams of the I Ching
1. Mô tả · Description
Đủ 64 quẻ theo thứ tự Chu Dịch, kèm tên Việt, tên Hán, quẻ thượng, quẻ hạ và tượng quẻ.
All 64 hexagrams in King Wen order, with Vietnamese and Han names, upper and lower trigrams, and the image.
Số dòng · Rows: 64
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type
Ý nghĩa · Meaning
id
string… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/64-que-kinh-dich.english_dictionaryTCGA-SARC-dict-tumorsam3-low-dice-2d-nnunet
SAM3 low-Dice 2D datasets for nnU-Net
Private research export of two small 2D datasets on which the balanced-finish
SAM3 LoRA validation Dice was below 0.5. The purpose is to test whether a
dataset-specific nnU-Net can fit these data and to distinguish data/training
limitations from inference bugs.
Dataset
SAM3 Dice
SAM3 IoU
Evaluated validation images
Actual SAM3 training images
DRIVE
0.212233
0.118717
2
14
RAVIR
0.224709
0.128455
2
16
The two-image validation… See the full description on the dataset page: https://huggingface.co/datasets/MedicalSAM3/sam3-low-dice-2d-nnunet.Tamajaq-English-dictionarymarathi-dictionary
Marathi Dictionary
Marathi-to-Marathi Dictionary
Dataset Description
Synthetically generated meanings in marathi for over 38k words
english_career_occupation_terminology_dictionarydicionario_barranquenho
Dicionário de Barranquenho
Structured lexical dataset derived from the first published dictionary of Barranquenho, a Romance contact language spoken in Barrancos, Portugal. Contains 1,680 entries with Portuguese and Spanish glosses, grammatical categories, semantic fields, source attributions, and synonym cross-references.
Language
Barranquenho (glottocode: barr1245; no ISO 639-3 code assigned at time of publication) is a contact language spoken in the municipality of… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/dicionario_barranquenho.healthcare-data-dictionary
Healthcare Data Dictionary — ISO-11179 Standard Terms
A sample dataset of 5,000 standardized healthcare data column names,
abbreviations, and definitions for data engineers building on
Snowflake, Databricks, BigQuery, and dbt.
Dataset Description
This dataset is a sample from the mdatool Healthcare Data Dictionary —
the most comprehensive ISO-11179 compliant healthcare data dictionary
available for data engineers.
What is ISO-11179?
ISO-11179 is… See the full description on the dataset page: https://huggingface.co/datasets/cyberali32112/healthcare-data-dictionary.afrilang-dictenglish-khmer-dictionary
📖 English–Khmer Dictionary Dataset
A comprehensive bilingual English–Khmer (ភាសាខ្មែរ) dictionary dataset in CSV format containing 170,000+ entries. Each entry includes the original English word, its Khmer translation, part of speech, full definitions in both languages, and example sentences — making it one of the richer English–Khmer lexical resources available for NLP and language learning.
Dataset Description
This dataset provides structured dictionary entries pairing… See the full description on the dataset page: https://huggingface.co/datasets/mrrtmob/english-khmer-dictionary.root_affix_dictionary
Nepali Root-Affix Dictionary
A word-level morphological segmentation dataset for Nepali, mapping surface
words to their root and affix components (e.g. अँकाइनु → अँका + इनु).
Dataset Details
Dataset Description
This dataset was built by scraping the 10th edition of the Nepali
dictionary, parsing the scraped entries into structured word/root pairs,
and then applying a heuristic rule engine to extend root-affix coverage
to surface forms not explicitly… See the full description on the dataset page: https://huggingface.co/datasets/W4ashabii/root_affix_dictionary.mac-dictation-privacy-matrix
Mac Dictation Privacy Matrix
An open, source-reviewed dataset comparing the documented privacy boundaries of
18 Mac dictation products.
The matrix separates questions that are often collapsed into one label:
where microphone audio becomes a transcript;
whether an optional cloud, cleanup, assistant, or agent path exists;
whether the documented speech path works offline after setup;
what the publisher says it retains;
which account, subscription, license, or provider boundary… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/mac-dictation-privacy-matrix.CellTissue-dict
CellTissue-dict
This dictionary is comprised of a merge of two source dictionaries (from Bioportal and ChEMBL), used to tag cell and tissue descriptions in text.
Dataset Details
Dataset Sources
This dictionary was produced by a combination of quick data sourcing, followed by partial manual categorising of terms.
Repository: https://github.com/ML4LitS/OTAR3088
Direct Use
This data may be used to tag phrases in natural language, categorising… See the full description on the dataset page: https://huggingface.co/datasets/OTAR3088/CellTissue-dict.thai-mm-dict
🗂️ Dataset Card: Thai-Myanmar Dictionary (2025)
📝 Dataset Summary
The Thai-Myanmar Dictionary (2025) is a high-quality bilingual lexical dataset created by Htet Myet Lynn.It provides direct word-to-word and phrase mappings between Thai and Myanmar (Burmese), supporting both linguistic use and machine learning applications.
The dataset is released under the MIT License, allowing free usage, modification, redistribution, and integration into both academic and commercial… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-mm-dict.
