datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NIKL-korean-english-dictionary
Column Name
Type
Description
설명
Form
str
Registered word entry
단어
Part of Speech
str or None
Part of speech of the word in Korean
품사
Korean Definition
List[str]
Definition of the word in Korean
해당 단어의 한글 정의
English Definition
List[str] or None
Definition of the word in English
한글 정의의 영문 번역본
Usages
List[str] or None
Sample sentence or dialogue
해당 단어의 예문 (문장 또는 대화 형식)
Vocabulary Level
str or None
Difficulty of the word (3 levels)
단어의 난이도 ('초급', '중급', '고급')
Semantic… See the full description on the dataset page: https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary.Pronunciation-dictionary-malayalam
Malayalam Pronunciation Dictionary
This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here
It gives Phonemic transcription of Malayalam words in IPA format.
Dataset Details
Dataset Description
This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon]
(https://pypi.org/project/mlphon/) Python library.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.ths-quant-factor-dictionary
THS Quant Factor Dictionary (同花顺量化因子字典)
Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors.
These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data.
Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.warabi-dictionary-extended-gpl
IME Dictionary Extended GPL
Japanese IME dictionary pack. The SQLite database and TSV files contain the
same conversion and prediction records. See NOTICE.md before redistribution.
Compilation license: GPL-3.0-only
SQLite tables: entries, entry_sources, predictions, prediction_sources
TSV exports: entries.tsv, predictions.tsv
Kaomoji flag
The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter
with WHERE NOT kaomoji for a face-free dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-extended-gpl.warabi-dictionary-core
IME Dictionary Core
Japanese IME dictionary pack. The SQLite database and TSV files contain the
same conversion and prediction records. See NOTICE.md before redistribution.
Compilation license: CC-BY-4.0
SQLite tables: entries, entry_sources, predictions, prediction_sources
TSV exports: entries.tsv, predictions.tsv
Kaomoji flag
The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter
with WHERE NOT kaomoji for a face-free dictionary. A candidate… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-core.dusun-dictionary-corpus
Dusun-English-Malay Dictionary and Corpus
A trilingual dataset containing Dusun, English, and Malay words, phrases, and sentences compiled for linguistic research, dictionary development, and machine translation.
This dataset is based on the Dusun language as spoken by the Dusun ethnic group of Sabah, Malaysia. While the Dusun dialect in the dataset shares approximately 99% similarity with standardized Kadazandusun, there may be minor differences in vocabulary, spelling, and… See the full description on the dataset page: https://huggingface.co/datasets/DusunDictionary/dusun-dictionary-corpus.english_dictionaryTamajaq-English-dictionarymarathi-dictionary
Marathi Dictionary
Marathi-to-Marathi Dictionary
Dataset Description
Synthetically generated meanings in marathi for over 38k words
english_career_occupation_terminology_dictionaryroot_affix_dictionary
Nepali Root-Affix Dictionary
A word-level morphological segmentation dataset for Nepali, mapping surface
words to their root and affix components (e.g. अँकाइनु → अँका + इनु).
Dataset Details
Dataset Description
This dataset was built by scraping the 10th edition of the Nepali
dictionary, parsing the scraped entries into structured word/root pairs,
and then applying a heuristic rule engine to extend root-affix coverage
to surface forms not explicitly… See the full description on the dataset page: https://huggingface.co/datasets/W4ashabii/root_affix_dictionary.hse-acronym-dictionary-2026
Canonical landing page: https://www.smartqhse.com/datasets/hse-acronym-dictionary-2026
HSE Acronym Dictionary 2026
Authoritative reference of 150+ HSE / EHS / occupational-safety acronyms with one-line definitions. Covers metrics (TRIR, LTIFR, DART, EMR, WBGT), regulations (OSHA, RIDDOR, COSHH, CDM, COMAH, PSM, OSHAD-SF), methodologies (HAZOP, HAZID, LOPA, SIL, FMEA, RCA, BBS, ICAM, TapRooT, Tripod Beta), bodies (NEBOSH, IOSH, IIRSM, IOGP, ACGIH, NIOSH, ANSI, ASSP), and… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-acronym-dictionary-2026.healthcare-data-dictionary
Healthcare Data Dictionary — ISO-11179 Standard Terms
A sample dataset of 5,000 standardized healthcare data column names,
abbreviations, and definitions for data engineers building on
Snowflake, Databricks, BigQuery, and dbt.
Dataset Description
This dataset is a sample from the mdatool Healthcare Data Dictionary —
the most comprehensive ISO-11179 compliant healthcare data dictionary
available for data engineers.
What is ISO-11179?
ISO-11179 is… See the full description on the dataset page: https://huggingface.co/datasets/cyberali32112/healthcare-data-dictionary.isan-phonetic-dictionary
Isan Phonetic Dictionary Dataset
Summary
This dataset is a phonetic dictionary focused on Isan (Northeastern Thai) pronunciations. It is structured to handle linguistic complexities such as:
Phonetic Variations (เสียงแปร): Words that have multiple valid pronunciations without changing the meaning.
Homographs (คำพ้องรูป): Words that are spelled the same but have different pronunciations and meanings depending on the context.
The data is provided in TSV (Tab-Separated… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/isan-phonetic-dictionary.english-khmer-dictionary
📖 English–Khmer Dictionary Dataset
A comprehensive bilingual English–Khmer (ភាសាខ្មែរ) dictionary dataset in CSV format containing 170,000+ entries. Each entry includes the original English word, its Khmer translation, part of speech, full definitions in both languages, and example sentences — making it one of the richer English–Khmer lexical resources available for NLP and language learning.
Dataset Description
This dataset provides structured dictionary entries pairing… See the full description on the dataset page: https://huggingface.co/datasets/mrrtmob/english-khmer-dictionary.Turkish-Spelling-Dictionaryvietnamese-dictionary
Vietnamese Dictionary dataset
Dataset Details
Vietnamese dictionary, crawled from tudientv.com which can use to train word embedding, feature extraction,..
Dataset Source
Website: https://tudientv.com
DEZD-Dictionary-Kirmancki
DEZD Language Dataset (Inisiatifê Projea Zonê Ma Projekt)
Über das Projekt
Dieses Dataset ist das digitale Rückgrat des DEZD-Modells (Dimili-Elewi-Zaza-Dersimi). Es dient der wissenschaftlichen Dokumentation und dem digitalen Erhalt der Kirmancki-Sprache (Zonê Ma / Zazaki/ Dimili ). Das Projekt wird durch die Initiative Inisiatifê Projea Zonê Ma und das Online-Wörterbuch qesebend-sozluk.de realisiert.
Struktur des DEZD-Modells
Das Akronym DEZD repräsentiert… See the full description on the dataset page: https://huggingface.co/datasets/Manidar/DEZD-Dictionary-Kirmancki.languages_datasetThis dataset contains a set of 8612 languages from across the world as well as data such as Glottocode, ISO-639-3 codes, names, language families etc.
Original source: https://glottolog.org/glottolog/language
turkish-legal-terms-dictionary
Turkish Legal Terms Dictionary / Türkçe Hukuki Terimler Sözlüğü
🇹🇷 Türkçe
Bu proje, Türkçe hukuki terimlerin ve anlamlarının bulunduğu kapsamlı bir sözlük içermektedir.
📋 İçerik
Toplam terim sayısı: 3.000+ hukuki terim
Format: CSV
Dil: Türkçe
Kapsam: Genel hukuk terimleri, medeni hukuk, ticaret hukuku, ceza hukuku ve diğer hukuk dalları
📁 Dosya Yapısı
├── turkish-legal-terms-dictionary.csv # Ana sözlük dosyası
└── README.md # Bu… See the full description on the dataset page: https://huggingface.co/datasets/istaken/turkish-legal-terms-dictionary.healthcare-data-dictionary
Healthcare Data Dictionary — ISO-11179 Standard Terms
A sample dataset of 5,000 standardized healthcare data column names,
abbreviations, and definitions for data engineers building on
Snowflake, Databricks, BigQuery, and dbt.
Dataset Description
This dataset is a sample from the mdatool Healthcare Data Dictionary —
the most comprehensive ISO-11179 compliant healthcare data dictionary
available for data engineers.
What is ISO-11179?
ISO-11179 is… See the full description on the dataset page: https://huggingface.co/datasets/mdatool/healthcare-data-dictionary.AI-Dictionary
AI Dictionary Dataset
Welcome to the AI Dictionary dataset on HuggingFace. This dataset is a comprehensive tool comprised of 16,665 unique key phrases that describe the whole domain of Artificial Intelligence (AI). It serves both the research community and industry domains, aiding in the identification of radical innovations and uncovering applications of AI in new domains.
This dataset is the result of the research paper "The AI Dictionary: The Foundation for a Text-Based Tool to… See the full description on the dataset page: https://huggingface.co/datasets/J0nasW/AI-Dictionary.Amazigh-English-Dictionarybulgarian-dictionary-2024
Bulgarian Dictionary 2024
Dataset Summary
This is a dictionary of single-word Bulgarian tokens, tagged by the approproate part-of-speech tag.
Supported Tasks
token-classification: The dataset can be used to train a model for token classification, which consists in applying a class to each token in a sequence.
Languages
bg: Only Bulgarian is supported by this dataset.
Dataset Structure
Data Instances
Each instance contains… See the full description on the dataset page: https://huggingface.co/datasets/thebogko/bulgarian-dictionary-2024.scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.vietnamese-emoticon-dictionaryconcept-to-root-dictionary
🌿 Concept-to-Root Dictionary
A mapping of universal concepts to Arabic triliteral roots for semantic compression
📖 Overview
This dataset provides mappings between universal semantic concepts and Arabic triliteral roots, designed for use as a compression layer in Large Language Models.
What are Arabic Roots?
Arabic uses a root-and-pattern morphological system where most words derive from 3-letter roots:
Root
Core Meaning
Derived Words… See the full description on the dataset page: https://huggingface.co/datasets/root-semantic-research/concept-to-root-dictionary.isan-phonetic-dictionary
Isan Phonetic Dictionary Dataset
Summary
This dataset is a phonetic dictionary focused on Isan (Northeastern Thai) pronunciations. It is structured to handle linguistic complexities such as:
Phonetic Variations (เสียงแปร): Words that have multiple valid pronunciations without changing the meaning.
Homographs (คำพ้องรูป): Words that are spelled the same but have different pronunciations and meanings depending on the context.
The data is provided in TSV (Tab-Separated… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/isan-phonetic-dictionary.parsed_hindi_dictionaryParsed Hindi Dictionary
scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.
