datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlel_wd_dictionaryXLEL-WD is a multilingual event linking dataset. This sub-dataset contains a dictionary of events from Wikidata. The multilingual descriptions for Wikidata event items are taken from the corresponding Wikipedia articles.opengloss-v1.3-dictionary
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-dictionary.Pulaar_Dictionary
Saggitorde — Dictionnaire Pulaar / Français / Anglais
Dictionnaire multilingue Pulaar ↔ Français ↔ Anglais extrait et enrichi à partir du Saggitorde de Ceerno Abuu Sih, enrichi par les terminologies de l'ARPRIM, ILN, MFS et d'autres sources spécialisées.
Statistiques
Statistique
Valeur
Nombre total d'entrées
1862
Nombre de domaines
24
Nombre de sources
24
Langues
Pulaar (ff), Français (fr), Anglais (en)
Structure des données… See the full description on the dataset page: https://huggingface.co/datasets/ARPRIM/Pulaar_Dictionary.gemma-2b-dictionary-embeddings-all-layers
Gemma-2B Dictionary Embeddings - All Layers
This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers.
Dataset Structure
metadata.json: Contains dataset metadata (model info, dimensions, word count)
embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26)
Usage
import pickle
from huggingface_hub import hf_hub_download
# Download a specific layer
layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.catalan-dictionary
Dataset Card for ca-text-corpus
Descripció (ca)
En aquest repositori s'apleguen llistes de paraules etiquetades amb la categoria gramatical, usades per a construir eines com correctors ortogràfics i gramaticals.
Dataset Summary
Catalan word lists with part of speech labeling curated by humans. Contains 1 180 773 forms including verbs, nouns, adjectives, names or toponyms. These word lists are used to build applications like Catalan spellcheckers or… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-dictionary.sign-dictionary-isl
Dataset Card for Sign Dictionary Dataset
This dataset contains Indian sign language videos with one gloss per video. There are 3077 seperate lex items or glosses included.
The dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0).
Dataset Details
There is a total of 2.5 hours of sign videos.
How to use
import webdataset as wds
import numpy as np
import json
import tempfile
import os
import cv2
def… See the full description on the dataset page: https://huggingface.co/datasets/bridgeconn/sign-dictionary-isl.opengloss-dictionary-definitions
OpenGloss Dictionary (Definition-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the definitions-level view where each record represents one sense definition.
Key Statistics
536,829 sense definitions across 150,101 English lexemes
9.1… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-dictionary-definitions.NIKL-korean-english-dictionary
Column Name
Type
Description
설명
Form
str
Registered word entry
단어
Part of Speech
str or None
Part of speech of the word in Korean
품사
Korean Definition
List[str]
Definition of the word in Korean
해당 단어의 한글 정의
English Definition
List[str] or None
Definition of the word in English
한글 정의의 영문 번역본
Usages
List[str] or None
Sample sentence or dialogue
해당 단어의 예문 (문장 또는 대화 형식)
Vocabulary Level
str or None
Difficulty of the word (3 levels)
단어의 난이도 ('초급', '중급', '고급')
Semantic… See the full description on the dataset page: https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary.opengloss-dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English lexemes
9.1… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-dictionary.iraqi-dictionaryuyghur-dictionary-dataset
维吾尔语多语言词典数据集
维吾尔语-汉语-英语多语言词典数据集,适用于大型语言模型(LLM)微调训练。
数据集统计
数据集
条目数
大小
语言方向
ug-cn.jsonl
3,079,016
456 MB
维吾尔语 ⟷ 汉语
ug-en.jsonl
685,836
108 MB
维吾尔语 ⟷ 英语
en-ug.jsonl
705,534
112 MB
英语 ⟷ 维吾尔语
ug-ug.jsonl
139,920
46 MB
维吾尔语释义
cn-cn.jsonl
127,646
39 MB
汉语释义
总计: 4,737,952 条(双向)/ 约 236 万唯一对
数据格式
{
"instruction": "请翻译以下维吾尔语词汇",
"input": "مەركىزى",
"output": "中心的,中央的"
}
字段说明:
instruction: 任务指令
input: 输入文本
output: 输出文本
使用示例… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-dictionary-dataset.robogate-failure-dictionary
RoboGate Failure Dictionary
50,000+ Physics-Validated Pick & Place Failure Patterns across 4 Robots (Franka Panda, UR5e, UR3e, UR10e)
A structured database of robot AI failure patterns collected from NVIDIA Isaac Sim physical simulations using Two-Stage Adaptive Sampling. Each experiment records the exact conditions under which a robot succeeded or failed at Pick & Place tasks.
Quick Stats
Franka Uniform
Franka Boundary
UR5e
UR3e
UR10e
Combined… See the full description on the dataset page: https://huggingface.co/datasets/liveplex/robogate-failure-dictionary.Pronunciation-dictionary-malayalam
Malayalam Pronunciation Dictionary
This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here
It gives Phonemic transcription of Malayalam words in IPA format.
Dataset Details
Dataset Description
This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon]
(https://pypi.org/project/mlphon/) Python library.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.french-dictionary
French Dictionary
A ready-to-use offline French language dictionary derived from the French Wiktionary. Available in two formats to suit different use cases: SQLite for desktop applications and real-time querying, and Parquet for data science and machine learning pipelines.
Contains nearly 900,000 distinct word forms including conjugated verb forms, with structured definitions, usage examples, and rich linguistic metadata.
Acknowledgements
This dataset would not… See the full description on the dataset page: https://huggingface.co/datasets/Kartmaan/french-dictionary.Malay-Dialect-Dictionary
Malay Dialect Dictionary
This is non official Malay Dialect Dictionary gathered from multiple sources in internet.
gemma-2b-dictionary-embeddingsths-quant-factor-dictionary
THS Quant Factor Dictionary (同花顺量化因子字典)
Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors.
These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data.
Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.hse-acronym-dictionary-2026
Canonical landing page: https://www.smartqhse.com/datasets/hse-acronym-dictionary-2026
HSE Acronym Dictionary 2026
Authoritative reference of 150+ HSE / EHS / occupational-safety acronyms with one-line definitions. Covers metrics (TRIR, LTIFR, DART, EMR, WBGT), regulations (OSHA, RIDDOR, COSHH, CDM, COMAH, PSM, OSHAD-SF), methodologies (HAZOP, HAZID, LOPA, SIL, FMEA, RCA, BBS, ICAM, TapRooT, Tripod Beta), bodies (NEBOSH, IOSH, IIRSM, IOGP, ACGIH, NIOSH, ANSI, ASSP), and… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-acronym-dictionary-2026.khmer-dictionary-44k
RAC Khmer Dictionary 2022
Data was extracted from Khmer Dictionary 2022 by Royal Academy of Cambodia. This is for research purpose only! Not for commercial use.
vi-en-mathematics-dictionaryVietAlpha English–Vietnamese Mathematics Dictionary
Research page ·
VietAlpha Lab ·
Source scan
The VietAlpha English–Vietnamese Mathematics Dictionary turns a 709-page printed reference work into a machine-readable bilingual lexicon. It contains 26,205 English and Vietnamese mathematics entries digitized from Cung Kim Tiến's Từ Điển Toán Học Anh – Việt, Việt – Anh and organized as JSON Lines.
What is in the dataset
Direction
Entries
English to Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/vi-en-mathematics-dictionary.myanmar_typeset_dictionary_OCR
Myanmar Typeset Dictionary OCR Dataset
This is a synthetically generated, realistically formatted dataset modeling a Myanmar-Myanmar dictionary. It is designed for training and validating OCR models, Document Layout Analysis (DLA) pipelines, and structural key-value extraction models.
The dataset contains a highly diverse set of pages containing multiple column flows, tabular glossaries, running headers/footers, realistic backgrounds, and dynamic typography (four fonts paired… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_typeset_dictionary_OCR.edmt-dictionary-english-khmer-llm-translated
EDMT English-Khmer Dictionary Dataset (LLM-Translated via Gemini 3.7 & Quality Evaluated)
A comprehensive, high-coverage English-to-Khmer bilingual dictionary dataset containing 176,064 entries and 110,504 distinct English headwords, built upon the open-source EDMT Dictionary Database (Webster's Revised Unabridged Dictionary).
Every word definition and translation has been translated and adapted into natural, grammatically sound Khmer using Gemini 3.7 Flash. In addition, both… See the full description on the dataset page: https://huggingface.co/datasets/namae101/edmt-dictionary-english-khmer-llm-translated.warabi-dictionary-extended-gpl
IME Dictionary Extended GPL
Japanese IME dictionary pack. The SQLite database and TSV files contain the
same conversion and prediction records. See NOTICE.md before redistribution.
Compilation license: GPL-3.0-only
SQLite tables: entries, entry_sources, predictions, prediction_sources
TSV exports: entries.tsv, predictions.tsv
Kaomoji flag
The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter
with WHERE NOT kaomoji for a face-free dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-extended-gpl.fr-vi-mathematics-dictionaryVietAlpha French–Vietnamese Mathematics Dictionary
Research page ·
VietAlpha Lab
The VietAlpha French–Vietnamese Mathematics Dictionary is a machine-readable edition of Danh-từ Toán-học Pháp-Việt, compiled in Saigon in 1964 by the Mathematics Committee of the National Committee for the Compilation of Specialized Dictionaries. The release contains 4,095 dictionary entries and a 1,369-item Vietnamese index reconstructed from the printed volume.
This dataset records how a Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/fr-vi-mathematics-dictionary.khmer-khmer-dictionaryurban_dictionarywarabi-dictionary-core
IME Dictionary Core
Japanese IME dictionary pack. The SQLite database and TSV files contain the
same conversion and prediction records. See NOTICE.md before redistribution.
Compilation license: CC-BY-4.0
SQLite tables: entries, entry_sources, predictions, prediction_sources
TSV exports: entries.tsv, predictions.tsv
Kaomoji flag
The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter
with WHERE NOT kaomoji for a face-free dictionary. A candidate… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-core.task123_conala_sort_dictionary
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task123_conala_sort_dictionary
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task123_conala_sort_dictionary.italian-dictionary
Italian Dictionary
Introduction
This dataset contains most of the words in the Italian dictionary. They were obtained from Wiktionary and the license is the same as its contents CC BY-SA 4.0
License
You are free to:
Share — copy and redistribute the material in any medium or format for any purpose, even commercially.
Adapt — remix, transform, and build upon the material for any purpose, even commercially.
The licensor cannot revoke these freedoms… See the full description on the dataset page: https://huggingface.co/datasets/mik3ml/italian-dictionary.dusun-dictionary-corpus
Dusun-English-Malay Dictionary and Corpus
A trilingual dataset containing Dusun, English, and Malay words, phrases, and sentences compiled for linguistic research, dictionary development, and machine translation.
This dataset is based on the Dusun language as spoken by the Dusun ethnic group of Sabah, Malaysia. While the Dusun dialect in the dataset shares approximately 99% similarity with standardized Kadazandusun, there may be minor differences in vocabulary, spelling, and… See the full description on the dataset page: https://huggingface.co/datasets/DusunDictionary/dusun-dictionary-corpus.
