datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-dictionary
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-dictionary.handy-dictation-editing
Handy dictation-editing corpus
Turns a raw dictated transcript into the text the speaker meant to write.
in : um so the meeting is uh moved to friday no wait thursday at three
out: The meeting is Thursday at three.
Three jobs at once, because they are not separable in speech: drop filler words,
repair punctuation and capitalisation, and — the hard one — when the speaker
changes their mind mid-sentence, delete the wording they abandoned and keep only
what they settled on.
Built… See the full description on the dataset page: https://huggingface.co/datasets/MagicNoThief/handy-dictation-editing.MathCOT-oss-vs-DeepSeek
Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
This dataset is the one used from the paper, available here 📄
This dataset consists of 242k math questions, with the verified generated answer (with reasoning) by both DeepSeek-R1-0528 and gpt-oss-120b.
The original prompts and the DeepSeek-R1-0528 traces were taken from NVIDIA's Nemotron-Post-Training-Dataset-v1.
Citation
If you found this dataset useful, please cite the paper below:… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/MathCOT-oss-vs-DeepSeek.DataClawEval
DataClawEval
An executable benchmark for end-to-end data-engineering agents in industrial environments.
DataClawEval measures an autonomous agent's ability to inspect data, implement and debug pipelines,
and materialize correct artifacts in realistic data-engineering workflows. It contains 100
production-grounded tasks across five execution engines: PySpark, MySQL, HiveSQL, PrestoSQL/Trino,
and FlinkSQL. Each task runs in an isolated Docker sandbox and is evaluated by a… See the full description on the dataset page: https://huggingface.co/datasets/dicemy/DataClawEval.DICE-BENCH
🎲 DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
🔗 Links for Reference
Repository: https://github.com/snuhcc/DICE-Bench
Paper: https://arxiv.org/abs/2506.22853
Project page: https://snuhcc.github.io/DICE-Bench/
Point of Contact: kyochul@snu.ac.kr
📖 Paper Description
DICE-BENCH is a benchmark that tests how well large language models can call external functions in realistic… See the full description on the dataset page: https://huggingface.co/datasets/OfficerChul/DICE-BENCH.italian-dictionary
Italian Dictionary
Introduction
This dataset contains most of the words in the Italian dictionary. They were obtained from Wiktionary and the license is the same as its contents CC BY-SA 4.0
License
You are free to:
Share — copy and redistribute the material in any medium or format for any purpose, even commercially.
Adapt — remix, transform, and build upon the material for any purpose, even commercially.
The licensor cannot revoke these freedoms… See the full description on the dataset page: https://huggingface.co/datasets/mik3ml/italian-dictionary.ShadowBench
ShadowBench
ShadowBench is a Lean 4 full autoformalization benchmark. Given a natural-language proof, a list of allowed Lean 4 imports, and formalization rules, produce a Lean 4 snippet that states the canonical theorem and proves it.
Splits
Split
Problems
Description
test
178
The problem set of the ShadowBench paper. Use this split for the public leaderboard.
icml
126
The earlier release used by the ICML 2026 AI4Math Workshop & Challenge 4 on… See the full description on the dataset page: https://huggingface.co/datasets/DicoTiar/ShadowBench.diccionario-psicologia-es
Diccionario de Psicología en Español para IA
Dataset léxico-conceptual de psicología en español, diseñado para entrenamiento y fine-tuning de modelos de lenguaje (LLMs/NLP). Contiene 3,002 términos únicos con 12 campos estructurados, cubriendo 18 áreas de la psicología.
Autor: Dr. Juan Moisés de la Serna · ORCID: 0000-0002-8401-8018 · UNIR
Estadísticas (v2.0.0)
Métrica
Valor
Términos únicos
3,002
Áreas de psicología
18
Campos por término
12
Formatos… See the full description on the dataset page: https://huggingface.co/datasets/juanmoisesdelas/diccionario-psicologia-es.opengloss-v1.1-dictionary
OpenGloss Dictionary v1.1 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,637 lexemes
7,701,312 semantic edges… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-dictionary.dictation-cleanup-examples
Dictation cleanup examples
A sample of the hand-written cases behind
SpeakoFlow Mini, published so the
conventions the model follows are inspectable rather than described.
Seven cases in each of fifteen categories, spread across short, medium and long transcripts.
Every case was written by hand. None of it is captured speech.
This is not a benchmark
Read that before using it for anything.
These cases are drawn from the training pool, not from the held-out set the… See the full description on the dataset page: https://huggingface.co/datasets/SpeakoFlow/dictation-cleanup-examples.opengloss-v1.3-dictionary
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
205,988 lexemes
8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.ksl-pose-dictionary-poc
KSL Pose Dictionary (PoC)
한국수어(KSL) text-to-pose 시제품용 keypoint 데이터셋.
docent_AI_sign_research_02 프로젝트에서 생성. Neural Sign Actors (CVPR 2024) 접근법을 KSL에 적용하는 Path B (Dictionary-based) 시제품의 핵심 데이터셋.
개요
자산
갯수
키포인트
sldict keypoint (국립국어원 한국수어사전)
1,444 단어
OpenPose 137 (RTMW-DW-L-M 추출)
NIASL2021 gloss segmentation keypoint (재난 안전 도메인)
2,287 base gloss
OpenPose 137 (NIASL 원본)
Hybrid sign index
4,511 unique signs
단어 → keypoint 경로 매핑
Stage 1 학습 corpus
20,085 samples… See the full description on the dataset page: https://huggingface.co/datasets/Trotquonalize/ksl-pose-dictionary-poc.myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.Dictionary-MKG
Dictionary-MKG: An LLM-Generated Multilingual Dictionary for Language Learners
Dictionary-MKG is a next-generation multilingual dataset designed to bridge the gap between static dictionaries and dynamic language learning. Generated using state-of-the-art LLMs (currently gemini-3-flash-preview), this project aims to provide structured, high-quality learning resources for language pairs that are historically under-served (e.g., learning Korean through Spanish).
You can find an… See the full description on the dataset page: https://huggingface.co/datasets/kenantang/Dictionary-MKG.alpaca-bulgarian-dictionary
Bulgarian Dictionary Dataset
This dataset is a collection of Bulgarian words along with their linguistic details, derived forms, synonyms, incorrect usages, and more. It is designed for use in natural language processing (NLP) tasks, such as training language models, building dictionaries, or enhancing word prediction systems.
Source Information
The data is sourced from the Читанка Речник, a free online dictionary. The Читанка Речник project aims to preserve and make… See the full description on the dataset page: https://huggingface.co/datasets/vislupus/alpaca-bulgarian-dictionary.VAM-Dictionary-Datasetmon_eng_dict_instructions
Mon-English Dictionary Instruction Dataset (Mon-AI Project)
📌 Project Overview
This dataset is a comprehensive, scalable, and high-quality Mon-English Instruction-Prompt Dataset designed specifically for supervised fine-tuning (SFT) of Large Language Models (LLMs).
The Mon language (ISO 639-3: mnw) is historically rich but classified as a low-resource language in the digital and AI landscape. The core mission of this project is to scale Mon linguistic resources… See the full description on the dataset page: https://huggingface.co/datasets/Nenemin95/mon_eng_dict_instructions.dickens_data_quality_checkstw-dictionary
Dataset Card for tw-dictionary
本資料集整合中華民國公開之三部繁體中文辭典/成語典:
dictionary-of-chinese-idioms:成語典
mandarin-chinese-mini-dictionary:國語辭典簡編本
revised-mandarin-chinese-dictionary:重編國語辭典修訂本
可作為繁體中文模型在用詞、成語、字義、注音等基礎語言知識上的補強語料。
Dataset Details
Dataset Description
這幾部辭典/成語典是繁體中文最具權威性的公開字/詞知識庫之一,內容涵蓋字音、字義、詞性、例句、出處典故等。本資料集將三部辭典分別作為獨立 config,方便依需求單獨使用或合併。
可用於:
增強模型對繁中字詞、成語、典故的覆蓋。
訓練字音/注音相關任務(部分 config 含注音資訊)。
作為文言/古文 / 成語使用情境之教學素材。
Curated by: Huang Liang Hsun
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-dictionary.mon_eng_dict_expanded
Mon-English Dictionary (Volume 2)
📌 Dataset Summary
This is a separate, dedicated Mon-English dictionary dataset structured for AI training, machine translation, and linguistic research.
Maintained by: Mon Community
Format: JSON
License: CC-BY-4.0
