datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-dictionary
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-dictionary.italian-dictionary
Italian Dictionary
Introduction
This dataset contains most of the words in the Italian dictionary. They were obtained from Wiktionary and the license is the same as its contents CC BY-SA 4.0
License
You are free to:
Share — copy and redistribute the material in any medium or format for any purpose, even commercially.
Adapt — remix, transform, and build upon the material for any purpose, even commercially.
The licensor cannot revoke these freedoms… See the full description on the dataset page: https://huggingface.co/datasets/mik3ml/italian-dictionary.opengloss-v1.1-dictionary
OpenGloss Dictionary v1.1 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,637 lexemes
7,701,312 semantic edges… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-dictionary.opengloss-v1.3-dictionary
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
205,988 lexemes
8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.ksl-pose-dictionary-poc
KSL Pose Dictionary (PoC)
한국수어(KSL) text-to-pose 시제품용 keypoint 데이터셋.
docent_AI_sign_research_02 프로젝트에서 생성. Neural Sign Actors (CVPR 2024) 접근법을 KSL에 적용하는 Path B (Dictionary-based) 시제품의 핵심 데이터셋.
개요
자산
갯수
키포인트
sldict keypoint (국립국어원 한국수어사전)
1,444 단어
OpenPose 137 (RTMW-DW-L-M 추출)
NIASL2021 gloss segmentation keypoint (재난 안전 도메인)
2,287 base gloss
OpenPose 137 (NIASL 원본)
Hybrid sign index
4,511 unique signs
단어 → keypoint 경로 매핑
Stage 1 학습 corpus
20,085 samples… See the full description on the dataset page: https://huggingface.co/datasets/Trotquonalize/ksl-pose-dictionary-poc.myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.Dictionary-MKG
Dictionary-MKG: An LLM-Generated Multilingual Dictionary for Language Learners
Dictionary-MKG is a next-generation multilingual dataset designed to bridge the gap between static dictionaries and dynamic language learning. Generated using state-of-the-art LLMs (currently gemini-3-flash-preview), this project aims to provide structured, high-quality learning resources for language pairs that are historically under-served (e.g., learning Korean through Spanish).
You can find an… See the full description on the dataset page: https://huggingface.co/datasets/kenantang/Dictionary-MKG.alpaca-bulgarian-dictionary
Bulgarian Dictionary Dataset
This dataset is a collection of Bulgarian words along with their linguistic details, derived forms, synonyms, incorrect usages, and more. It is designed for use in natural language processing (NLP) tasks, such as training language models, building dictionaries, or enhancing word prediction systems.
Source Information
The data is sourced from the Читанка Речник, a free online dictionary. The Читанка Речник project aims to preserve and make… See the full description on the dataset page: https://huggingface.co/datasets/vislupus/alpaca-bulgarian-dictionary.VAM-Dictionary-Datasettw-dictionary
Dataset Card for tw-dictionary
本資料集整合中華民國公開之三部繁體中文辭典/成語典:
dictionary-of-chinese-idioms:成語典
mandarin-chinese-mini-dictionary:國語辭典簡編本
revised-mandarin-chinese-dictionary:重編國語辭典修訂本
可作為繁體中文模型在用詞、成語、字義、注音等基礎語言知識上的補強語料。
Dataset Details
Dataset Description
這幾部辭典/成語典是繁體中文最具權威性的公開字/詞知識庫之一,內容涵蓋字音、字義、詞性、例句、出處典故等。本資料集將三部辭典分別作為獨立 config,方便依需求單獨使用或合併。
可用於:
增強模型對繁中字詞、成語、典故的覆蓋。
訓練字音/注音相關任務(部分 config 含注音資訊)。
作為文言/古文 / 成語使用情境之教學素材。
Curated by: Huang Liang Hsun
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-dictionary.
