CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mjbommar /opengloss-v1.3-dictionary See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M1 likes773 downloads15d agoHugging Face02mjbommar /opengloss-dictionary-definitions OpenGloss Dictionary (Definition-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the definitions-level view where each record represents one sense definition. Key Statistics 536,829 sense definitions across 150,101 English lexemes 9.1… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-dictionary-definitions.tabulartext-generation100K<n<1M1 likes228 downloads10mo agoHugging Face03mjbommar /opengloss-dictionary OpenGloss Dictionary (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,101 lexemes across 150,101 English lexemes 9.1… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-dictionary.tabulartext-generation100K<n<1M5 likes214 downloads10mo agoHugging Face04dagim /urban-dictionary-embeddings Dataset Card for "urban-dictionary-embeddings" More Information needed tabular1M<n<10M1 likes201 downloads3y agoHugging Face05LeeHarrold /gemma-2b-dictionary-embeddings-all-layers Gemma-2B Dictionary Embeddings - All Layers This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers. Dataset Structure metadata.json: Contains dataset metadata (model info, dimensions, word count) embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26) Usage import pickle from huggingface_hub import hf_hub_download # Download a specific layer layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.tabularn<1K0 likes198 downloads1y agoHugging Face06liveplex /robogate-failure-dictionary RoboGate Failure Dictionary 50,000+ Physics-Validated Pick & Place Failure Patterns across 4 Robots (Franka Panda, UR5e, UR3e, UR10e) A structured database of robot AI failure patterns collected from NVIDIA Isaac Sim physical simulations using Two-Stage Adaptive Sampling. Each experiment records the exact conditions under which a robot succeeded or failed at Pick & Place tasks. Quick Stats Franka Uniform Franka Boundary UR5e UR3e UR10e Combined… See the full description on the dataset page: https://huggingface.co/datasets/liveplex/robogate-failure-dictionary.tabularrobotics10K<n<100K0 likes166 downloads1mo agoHugging Face07Kartmaan /french-dictionary French Dictionary A ready-to-use offline French language dictionary derived from the French Wiktionary. Available in two formats to suit different use cases: SQLite for desktop applications and real-time querying, and Parquet for data science and machine learning pipelines. Contains nearly 900,000 distinct word forms including conjugated verb forms, with structured definitions, usage examples, and rich linguistic metadata. Acknowledgements This dataset would not… See the full description on the dataset page: https://huggingface.co/datasets/Kartmaan/french-dictionary.tabular1M<n<10M1 likes152 downloads6mo agoHugging Face08obaydata /ths-quant-factor-dictionary THS Quant Factor Dictionary (同花顺量化因子字典) Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors. These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data. Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.tabularn<1K0 likes133 downloads6mo agoHugging Face09namae101 /edmt-dictionary-english-khmer-llm-translated EDMT English-Khmer Dictionary Dataset (LLM-Translated via Gemini 3.7 & Quality Evaluated) A comprehensive, high-coverage English-to-Khmer bilingual dictionary dataset containing 176,064 entries and 110,504 distinct English headwords, built upon the open-source EDMT Dictionary Database (Webster's Revised Unabridged Dictionary). Every word definition and translation has been translated and adapted into natural, grammatically sound Khmer using Gemini 3.7 Flash. In addition, both… See the full description on the dataset page: https://huggingface.co/datasets/namae101/edmt-dictionary-english-khmer-llm-translated.tabular100K<n<1M0 likes106 downloads26d agoHugging Face10Cloudadorablebearcloudbear /opengloss-dictionary OpenGloss Dictionary (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-dictionary.tabulartext-generation100K<n<1M0 likes60 downloads1mo agoHugging Face11caioloures /personal_dictionary OpenGloss Dictionary (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/caioloures/personal_dictionary.tabulartext-generation100K<n<1M1 likes54 downloads2mo agoHugging Face12mjbommar /opengloss-v1.1-dictionary OpenGloss Dictionary v1.1 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,637 lexemes 7,701,312 semantic edges… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-dictionary.tabulartext-generation100K<n<1M0 likes51 downloads10mo agoHugging Face13Cloudadorablebearcloudbear /opengloss-v1.3-dictionary OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 205,988 lexemes 8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M0 likes49 downloads1mo agoHugging Face14Trotquonalize /ksl-pose-dictionary-poc KSL Pose Dictionary (PoC) 한국수어(KSL) text-to-pose 시제품용 keypoint 데이터셋. docent_AI_sign_research_02 프로젝트에서 생성. Neural Sign Actors (CVPR 2024) 접근법을 KSL에 적용하는 Path B (Dictionary-based) 시제품의 핵심 데이터셋. 개요 자산 갯수 키포인트 sldict keypoint (국립국어원 한국수어사전) 1,444 단어 OpenPose 137 (RTMW-DW-L-M 추출) NIASL2021 gloss segmentation keypoint (재난 안전 도메인) 2,287 base gloss OpenPose 137 (NIASL 원본) Hybrid sign index 4,511 unique signs 단어 → keypoint 경로 매핑 Stage 1 학습 corpus 20,085 samples… See the full description on the dataset page: https://huggingface.co/datasets/Trotquonalize/ksl-pose-dictionary-poc.tabulartext-generation10K<n<100K0 likes41 downloads4mo agoHugging Face15LukeEuser /docvqa_singledoc_train_val_dataset_dictionarytabular10K<n<100K0 likes40 downloads3y agoHugging Face16mrrtmob /english-khmer-dictionary 📖 English–Khmer Dictionary Dataset A comprehensive bilingual English–Khmer (ភាសាខ្មែរ) dictionary dataset in CSV format containing 170,000+ entries. Each entry includes the original English word, its Khmer translation, part of speech, full definitions in both languages, and example sentences — making it one of the richer English–Khmer lexical resources available for NLP and language learning. Dataset Description This dataset provides structured dictionary entries pairing… See the full description on the dataset page: https://huggingface.co/datasets/mrrtmob/english-khmer-dictionary.tabulartranslation100K<n<1M1 likes29 downloads7mo agoHugging Face17ultimate-dictionary /languages_datasetThis dataset contains a set of 8612 languages from across the world as well as data such as Glottocode, ISO-639-3 codes, names, language families etc. Original source: https://glottolog.org/glottolog/language tabular1K<n<10K1 likes23 downloads1y agoHugging Face18irisdewildt /docvqa_singledoc_train_val_dataset_dictionarytabular10K<n<100K0 likes19 downloads3y agoHugging Face19abdelhaqueidali /Tawalt-Amazigh-French-Dictionarytabular10K<n<100K0 likes18 downloads4mo agoHugging Face20abdelhaqueidali /Amazigh-English-Dictionarytabular10K<n<100K0 likes15 downloads4mo agoHugging Face21ScoutieAutoML /scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning Chinese, this dataset contains grammar, syntax, spelling and punctuation rules, as well as Chinese words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_chinese_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification1K<n<10K0 likes14 downloads2y agoHugging Face22ScoutieAutoML /scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification10K<n<100K0 likes11 downloads2y agoHugging Face23abdelhaqueidali /Tawalt-Amazigh-Arabic-Dictionarytabular10K<n<100K0 likes11 downloads4mo agoHugging Face24Zaanthai /balochi-dictionary Zaanth Balochi Dictionary — v0.1 (Latin script) The first release of the open Balochi dictionary by Zaanth, an open platform for Balochi & Brahui language data. 30 entries of the most frequent Balochi words (Latin script / Syáhag), each with an English meaning verified by a native Makrani Balochi speaker, corpus frequency, and a real example sentence with translation. Method Words ranked by frequency across 18,930 Balochi sentences; candidate meanings were… See the full description on the dataset page: https://huggingface.co/datasets/Zaanthai/balochi-dictionary.tabularn<1K0 likes9 downloads3mo agoHugging Face25irisdewildt /docvqa_singledoc_test_dataset_dictionary Dataset Card for "docvqa_singledoc_test_dataset_dictionary" More Information needed tabular1K<n<10K0 likes8 downloads3y agoHugging Face26metythorn /english-khmer-dictionary English-Khmer Dictionary This dataset is copied from mrrtmob/english-khmer-dictionary. Usage from datasets import load_dataset ds = load_dataset("metythorn/english-khmer-dictionary") print(ds) tabulartranslation100K<n<1M0 likes7 downloads5mo agoHugging Face27kowndinya23 /dipole-dictionary-data-v0tabular1M<n<10M1 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.