CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mjbommar /opengloss-v1.3-dictionary See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M1 likes764 downloads17d agoHugging Face02MagicNoThief /handy-dictation-editing Handy dictation-editing corpus Turns a raw dictated transcript into the text the speaker meant to write. in : um so the meeting is uh moved to friday no wait thursday at three out: The meeting is Thursday at three. Three jobs at once, because they are not separable in speech: drop filler words, repair punctuation and capitalisation, and — the hard one — when the speaker changes their mind mid-sentence, delete the wording they abandoned and keep only what they settled on. Built… See the full description on the dataset page: https://huggingface.co/datasets/MagicNoThief/handy-dictation-editing.texttext-generation100K<n<1M1 likes220 downloads19d agoHugging Face03dicta-il /MathCOT-oss-vs-DeepSeek Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces This dataset is the one used from the paper, available here 📄 This dataset consists of 242k math questions, with the verified generated answer (with reasoning) by both DeepSeek-R1-0528 and gpt-oss-120b. The original prompts and the DeepSeek-R1-0528 traces were taken from NVIDIA's Nemotron-Post-Training-Dataset-v1. Citation If you found this dataset useful, please cite the paper below:… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/MathCOT-oss-vs-DeepSeek.texttext-generation100K<n<1M2 likes199 downloads10mo agoHugging Face04dicemy /DataClawEval DataClawEval An executable benchmark for end-to-end data-engineering agents in industrial environments. DataClawEval measures an autonomous agent's ability to inspect data, implement and debug pipelines, and materialize correct artifacts in realistic data-engineering workflows. It contains 100 production-grounded tasks across five execution engines: PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL. Each task runs in an isolated Docker sandbox and is evaluated by a… See the full description on the dataset page: https://huggingface.co/datasets/dicemy/DataClawEval.texttext-generationn<1K0 likes161 downloads2mo agoHugging Face05OfficerChul /DICE-BENCH 🎲 DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues 🔗 Links for Reference Repository: https://github.com/snuhcc/DICE-Bench Paper: https://arxiv.org/abs/2506.22853 Project page: https://snuhcc.github.io/DICE-Bench/ Point of Contact: kyochul@snu.ac.kr 📖 Paper Description DICE-BENCH is a benchmark that tests how well large language models can call external functions in realistic… See the full description on the dataset page: https://huggingface.co/datasets/OfficerChul/DICE-BENCH.texttext-generation1K<n<10K3 likes110 downloads1y agoHugging Face06mik3ml /italian-dictionary Italian Dictionary Introduction This dataset contains most of the words in the Italian dictionary. They were obtained from Wiktionary and the license is the same as its contents CC BY-SA 4.0 License You are free to: Share — copy and redistribute the material in any medium or format for any purpose, even commercially. Adapt — remix, transform, and build upon the material for any purpose, even commercially. The licensor cannot revoke these freedoms… See the full description on the dataset page: https://huggingface.co/datasets/mik3ml/italian-dictionary.texttext-generation100K<n<1M8 likes81 downloads2y agoHugging Face07DicoTiar /ShadowBench ShadowBench ShadowBench is a Lean 4 full autoformalization benchmark. Given a natural-language proof, a list of allowed Lean 4 imports, and formalization rules, produce a Lean 4 snippet that states the canonical theorem and proves it. Splits Split Problems Description test 178 The problem set of the ShadowBench paper. Use this split for the public leaderboard. icml 126 The earlier release used by the ICML 2026 AI4Math Workshop & Challenge 4 on… See the full description on the dataset page: https://huggingface.co/datasets/DicoTiar/ShadowBench.texttext-generationn<1K0 likes67 downloads8d agoHugging Face08juanmoisesdelas /diccionario-psicologia-es Diccionario de Psicología en Español para IA Dataset léxico-conceptual de psicología en español, diseñado para entrenamiento y fine-tuning de modelos de lenguaje (LLMs/NLP). Contiene 3,002 términos únicos con 12 campos estructurados, cubriendo 18 áreas de la psicología. Autor: Dr. Juan Moisés de la Serna · ORCID: 0000-0002-8401-8018 · UNIR Estadísticas (v2.0.0) Métrica Valor Términos únicos 3,002 Áreas de psicología 18 Campos por término 12 Formatos… See the full description on the dataset page: https://huggingface.co/datasets/juanmoisesdelas/diccionario-psicologia-es.texttext-generation1K<n<10K0 likes56 downloads6mo agoHugging Face09mjbommar /opengloss-v1.1-dictionary OpenGloss Dictionary v1.1 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,637 lexemes 7,701,312 semantic edges… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-dictionary.tabulartext-generation100K<n<1M0 likes54 downloads10mo agoHugging Face10SpeakoFlow /dictation-cleanup-examples Dictation cleanup examples A sample of the hand-written cases behind SpeakoFlow Mini, published so the conventions the model follows are inspectable rather than described. Seven cases in each of fifteen categories, spread across short, medium and long transcripts. Every case was written by hand. None of it is captured speech. This is not a benchmark Read that before using it for anything. These cases are drawn from the training pool, not from the held-out set the… See the full description on the dataset page: https://huggingface.co/datasets/SpeakoFlow/dictation-cleanup-examples.texttext-generationn<1K0 likes54 downloads26d agoHugging Face11Cloudadorablebearcloudbear /opengloss-v1.3-dictionary OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 205,988 lexemes 8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M0 likes49 downloads1mo agoHugging Face12Trotquonalize /ksl-pose-dictionary-poc KSL Pose Dictionary (PoC) 한국수어(KSL) text-to-pose 시제품용 keypoint 데이터셋. docent_AI_sign_research_02 프로젝트에서 생성. Neural Sign Actors (CVPR 2024) 접근법을 KSL에 적용하는 Path B (Dictionary-based) 시제품의 핵심 데이터셋. 개요 자산 갯수 키포인트 sldict keypoint (국립국어원 한국수어사전) 1,444 단어 OpenPose 137 (RTMW-DW-L-M 추출) NIASL2021 gloss segmentation keypoint (재난 안전 도메인) 2,287 base gloss OpenPose 137 (NIASL 원본) Hybrid sign index 4,511 unique signs 단어 → keypoint 경로 매핑 Stage 1 학습 corpus 20,085 samples… See the full description on the dataset page: https://huggingface.co/datasets/Trotquonalize/ksl-pose-dictionary-poc.tabulartext-generation10K<n<100K0 likes45 downloads4mo agoHugging Face13freococo /myanmar-english-pali-dictionary Myanmar–English–Pali Dictionary Dataset Summary This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein). It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary. The dataset is intended for research and educational purposes, including but not limited to: Natural Language Processing (NLP) Machine Translation (MT) Lexicography Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.texttranslation10K<n<100K1 likes41 downloads8mo agoHugging Face14DatarrX /pali-myanmar-dictionary-corpus Pali-Myanmar Dictionary Corpus (Instruction-Ready) Dataset Summary The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning. Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.texttranslation100K<n<1M6 likes38 downloads5mo agoHugging Face15kenantang /Dictionary-MKG Dictionary-MKG: An LLM-Generated Multilingual Dictionary for Language Learners Dictionary-MKG is a next-generation multilingual dataset designed to bridge the gap between static dictionaries and dynamic language learning. Generated using state-of-the-art LLMs (currently gemini-3-flash-preview), this project aims to provide structured, high-quality learning resources for language pairs that are historically under-served (e.g., learning Korean through Spanish). You can find an… See the full description on the dataset page: https://huggingface.co/datasets/kenantang/Dictionary-MKG.texttranslation1K<n<10K3 likes36 downloads8mo agoHugging Face16vislupus /alpaca-bulgarian-dictionary Bulgarian Dictionary Dataset This dataset is a collection of Bulgarian words along with their linguistic details, derived forms, synonyms, incorrect usages, and more. It is designed for use in natural language processing (NLP) tasks, such as training language models, building dictionaries, or enhancing word prediction systems. Source Information The data is sourced from the Читанка Речник, a free online dictionary. The Читанка Речник project aims to preserve and make… See the full description on the dataset page: https://huggingface.co/datasets/vislupus/alpaca-bulgarian-dictionary.texttext-generation100K<n<1M3 likes29 downloads2y agoHugging Face17abdelhaqueidali /VAM-Dictionary-Datasettexttext-classification1K<n<10K1 likes21 downloads4mo agoHugging Face18Nenemin95 /mon_eng_dict_instructions Mon-English Dictionary Instruction Dataset (Mon-AI Project) 📌 Project Overview This dataset is a comprehensive, scalable, and high-quality Mon-English Instruction-Prompt Dataset designed specifically for supervised fine-tuning (SFT) of Large Language Models (LLMs). The Mon language (ISO 639-3: mnw) is historically rich but classified as a low-resource language in the digital and AI landscape. The core mission of this project is to scale Mon linguistic resources… See the full description on the dataset page: https://huggingface.co/datasets/Nenemin95/mon_eng_dict_instructions.texttranslation100K<n<1M0 likes15 downloads3mo agoHugging Face19elsatch /dickens_data_quality_checkstextquestion-answeringn<1K0 likes12 downloads3y agoHugging Face20lianghsun /tw-dictionarygated Dataset Card for tw-dictionary 本資料集整合中華民國公開之三部繁體中文辭典/成語典: dictionary-of-chinese-idioms:成語典 mandarin-chinese-mini-dictionary:國語辭典簡編本 revised-mandarin-chinese-dictionary:重編國語辭典修訂本 可作為繁體中文模型在用詞、成語、字義、注音等基礎語言知識上的補強語料。 Dataset Details Dataset Description 這幾部辭典/成語典是繁體中文最具權威性的公開字/詞知識庫之一,內容涵蓋字音、字義、詞性、例句、出處典故等。本資料集將三部辭典分別作為獨立 config,方便依需求單獨使用或合併。 可用於: 增強模型對繁中字詞、成語、典故的覆蓋。 訓練字音/注音相關任務(部分 config 含注音資訊)。 作為文言/古文 / 成語使用情境之教學素材。 Curated by: Huang Liang Hsun Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-dictionary.texttext-generation100K<n<1M0 likes7 downloads5mo agoHugging Face21Nenemin95 /mon_eng_dict_expanded Mon-English Dictionary (Volume 2) 📌 Dataset Summary This is a separate, dedicated Mon-English dictionary dataset structured for AI training, machine translation, and linguistic research. Maintained by: Mon Community Format: JSON License: CC-BY-4.0 texttranslation1K<n<10K0 likes5 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.