CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01adithya7 /xlel_wd_dictionaryXLEL-WD is a multilingual event linking dataset. This sub-dataset contains a dictionary of events from Wikidata. The multilingual descriptions for Wikidata event items are taken from the corresponding Wikipedia articles.text100K<n<1M3 likes933 downloads4y agoHugging Face02byliang /REALISTA_latent_dictionary REALISTA Latent Direction Dictionaries Pre-computed stage-2 latent perturbation direction dictionaries for REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations (ICML 2026). ⚠️ Warning: This method could be misused for malicious purposes. Provided for research and educational purposes only. What this is REALISTA attacks a target LLM by optimizing a sparse combination of latent editing directions, each corresponding to a… See the full description on the dataset page: https://huggingface.co/datasets/byliang/REALISTA_latent_dictionary.2 likes901 downloads3mo agoHugging Face03mjbommar /opengloss-v1.3-dictionary See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M1 likes773 downloads15d agoHugging Face04mjbommar /opengloss-v1.2-dictionary OpenGloss Dictionary v1.2 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 162,314 lexemes 7,798,653 semantic edges… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-dictionary.text-generation100K<n<1M0 likes540 downloads6mo agoHugging Face05ARPRIM /Pulaar_Dictionary Saggitorde — Dictionnaire Pulaar / Français / Anglais Dictionnaire multilingue Pulaar ↔ Français ↔ Anglais extrait et enrichi à partir du Saggitorde de Ceerno Abuu Sih, enrichi par les terminologies de l'ARPRIM, ILN, MFS et d'autres sources spécialisées. Statistiques Statistique Valeur Nombre total d'entrées 1862 Nombre de domaines 24 Nombre de sources 24 Langues Pulaar (ff), Français (fr), Anglais (en) Structure des données… See the full description on the dataset page: https://huggingface.co/datasets/ARPRIM/Pulaar_Dictionary.text1K<n<10K1 likes366 downloads12d agoHugging Face06softcatala /catalan-dictionary Dataset Card for ca-text-corpus Descripció (ca) En aquest repositori s'apleguen llistes de paraules etiquetades amb la categoria gramatical, usades per a construir eines com correctors ortogràfics i gramaticals. Dataset Summary Catalan word lists with part of speech labeling curated by humans. Contains 1 180 773 forms including verbs, nouns, adjectives, names or toponyms. These word lists are used to build applications like Catalan spellcheckers or… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-dictionary.texttext-generation1M<n<10M3 likes265 downloads2mo agoHugging Face07bridgeconn /sign-dictionary-isl Dataset Card for Sign Dictionary Dataset This dataset contains Indian sign language videos with one gloss per video. There are 3077 seperate lex items or glosses included. The dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0). Dataset Details There is a total of 2.5 hours of sign videos. How to use import webdataset as wds import numpy as np import json import tempfile import os import cv2 def… See the full description on the dataset page: https://huggingface.co/datasets/bridgeconn/sign-dictionary-isl.text1K<n<10K1 likes228 downloads11mo agoHugging Face08mjbommar /opengloss-dictionary-definitions OpenGloss Dictionary (Definition-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the definitions-level view where each record represents one sense definition. Key Statistics 536,829 sense definitions across 150,101 English lexemes 9.1… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-dictionary-definitions.tabulartext-generation100K<n<1M1 likes228 downloads10mo agoHugging Face09binjang /NIKL-korean-english-dictionary Column Name Type Description 설명 Form str Registered word entry 단어 Part of Speech str or None Part of speech of the word in Korean 품사 Korean Definition List[str] Definition of the word in Korean 해당 단어의 한글 정의 English Definition List[str] or None Definition of the word in English 한글 정의의 영문 번역본 Usages List[str] or None Sample sentence or dialogue 해당 단어의 예문 (문장 또는 대화 형식) Vocabulary Level str or None Difficulty of the word (3 levels) 단어의 난이도 ('초급', '중급', '고급') Semantic… See the full description on the dataset page: https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary.texttranslation10K<n<100K7 likes217 downloads3y agoHugging Face10mjbommar /opengloss-dictionary OpenGloss Dictionary (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,101 lexemes across 150,101 English lexemes 9.1… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-dictionary.tabulartext-generation100K<n<1M5 likes214 downloads10mo agoHugging Face11SaifSilverHand /iraqi-dictionarytextn<1K5 likes210 downloads2d agoHugging Face12dagim /urban-dictionary-embeddings Dataset Card for "urban-dictionary-embeddings" More Information needed tabular1M<n<10M1 likes201 downloads3y agoHugging Face13LeeHarrold /gemma-2b-dictionary-embeddings-all-layers Gemma-2B Dictionary Embeddings - All Layers This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers. Dataset Structure metadata.json: Contains dataset metadata (model info, dimensions, word count) embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26) Usage import pickle from huggingface_hub import hf_hub_download # Download a specific layer layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.tabularn<1K0 likes198 downloads1y agoHugging Face14liveplex /robogate-failure-dictionary RoboGate Failure Dictionary 50,000+ Physics-Validated Pick & Place Failure Patterns across 4 Robots (Franka Panda, UR5e, UR3e, UR10e) A structured database of robot AI failure patterns collected from NVIDIA Isaac Sim physical simulations using Two-Stage Adaptive Sampling. Each experiment records the exact conditions under which a robot succeeded or failed at Pick & Place tasks. Quick Stats Franka Uniform Franka Boundary UR5e UR3e UR10e Combined… See the full description on the dataset page: https://huggingface.co/datasets/liveplex/robogate-failure-dictionary.tabularrobotics10K<n<100K0 likes166 downloads1mo agoHugging Face15anke01 /uyghur-dictionary-dataset 维吾尔语多语言词典数据集 维吾尔语-汉语-英语多语言词典数据集,适用于大型语言模型(LLM)微调训练。 数据集统计 数据集 条目数 大小 语言方向 ug-cn.jsonl 3,079,016 456 MB 维吾尔语 ⟷ 汉语 ug-en.jsonl 685,836 108 MB 维吾尔语 ⟷ 英语 en-ug.jsonl 705,534 112 MB 英语 ⟷ 维吾尔语 ug-ug.jsonl 139,920 46 MB 维吾尔语释义 cn-cn.jsonl 127,646 39 MB 汉语释义 总计: 4,737,952 条(双向)/ 约 236 万唯一对 数据格式 { "instruction": "请翻译以下维吾尔语词汇", "input": "مەركىزى", "output": "中心的,中央的" } 字段说明: instruction: 任务指令 input: 输入文本 output: 输出文本 使用示例… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-dictionary-dataset.texttranslation1M<n<10M2 likes165 downloads7mo agoHugging Face16Kartmaan /french-dictionary French Dictionary A ready-to-use offline French language dictionary derived from the French Wiktionary. Available in two formats to suit different use cases: SQLite for desktop applications and real-time querying, and Parquet for data science and machine learning pipelines. Contains nearly 900,000 distinct word forms including conjugated verb forms, with structured definitions, usage examples, and rich linguistic metadata. Acknowledgements This dataset would not… See the full description on the dataset page: https://huggingface.co/datasets/Kartmaan/french-dictionary.tabular1M<n<10M1 likes152 downloads6mo agoHugging Face17mesolitica /Malay-Dialect-Dictionary Malay Dialect Dictionary This is non official Malay Dialect Dictionary gathered from multiple sources in internet. text1K<n<10K0 likes147 downloads1y agoHugging Face18kavyamanohar /Pronunciation-dictionary-malayalam Malayalam Pronunciation Dictionary This Dataset has an alternate name of Malayalam Phonetic Lexicon. It is curated from the original source here It gives Phonemic transcription of Malayalam words in IPA format. Dataset Details Dataset Description This is a collection of Malayalam words and their pronunciation described in IPA format. The pronunciations has been automatically generated using [Mlphon] (https://pypi.org/project/mlphon/) Python library. Curated… See the full description on the dataset page: https://huggingface.co/datasets/kavyamanohar/Pronunciation-dictionary-malayalam.text100K<n<1M2 likes144 downloads2y agoHugging Face19LeeHarrold /gemma-2b-dictionary-embeddingstext10K<n<100K0 likes139 downloads1y agoHugging Face20obaydata /ths-quant-factor-dictionary THS Quant Factor Dictionary (同花顺量化因子字典) Quantitative factor dictionaries from THS (同花顺/Tonghuashun), covering A-share and overseas markets. Includes alpha factors, Barra risk factors, sell-side consensus estimates, and real-time news factors. These dictionaries describe the schema and metadata of THS's quantitative factor database — they do not contain actual factor values, but serve as essential references for anyone working with THS quant data. Files… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/ths-quant-factor-dictionary.tabularn<1K0 likes133 downloads6mo agoHugging Face21freococo /myanmar_typeset_dictionary_OCR Myanmar Typeset Dictionary OCR Dataset This is a synthetically generated, realistically formatted dataset modeling a Myanmar-Myanmar dictionary. It is designed for training and validating OCR models, Document Layout Analysis (DLA) pipelines, and structural key-value extraction models. The dataset contains a highly diverse set of pages containing multiple column flows, tabular glossaries, running headers/footers, realistic backgrounds, and dynamic typography (four fonts paired… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_typeset_dictionary_OCR.imageobject-detectionn<1K2 likes120 downloads2mo agoHugging Face22seanghay /khmer-dictionary-44k RAC Khmer Dictionary 2022 Data was extracted from Khmer Dictionary 2022 by Royal Academy of Cambodia. This is for research purpose only! Not for commercial use. text10K<n<100K11 likes117 downloads2y agoHugging Face23BlazingCerulean /OpenLeaf-Dictionary A Lean English Dictionary A compressed, offline-ready English dictionary derived from the kaikki.org machine-readable extraction of Wiktionary. Packed into a single SQLite file (~53MB) so it can be bundled directly into an app instead of queried over a network. Why this exists I built this to power offline dictionary lookups in OpenLeaf, an Android e-reader app, tap a word while reading, get a definition, no network round-trip. OpenLeaf It's published here in case… See the full description on the dataset page: https://huggingface.co/datasets/BlazingCerulean/OpenLeaf-Dictionary.100K<n<1M0 likes112 downloads20d agoHugging Face24VietAlphaLabs /vi-en-mathematics-dictionaryVietAlpha English–Vietnamese Mathematics Dictionary Research page · VietAlpha Lab · Source scan The VietAlpha English–Vietnamese Mathematics Dictionary turns a 709-page printed reference work into a machine-readable bilingual lexicon. It contains 26,205 English and Vietnamese mathematics entries digitized from Cung Kim Tiến's Từ Điển Toán Học Anh – Việt, Việt – Anh and organized as JSON Lines. What is in the dataset Direction Entries English to Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/vi-en-mathematics-dictionary.texttranslation10K<n<100K1 likes108 downloads1d agoHugging Face25namae101 /edmt-dictionary-english-khmer-llm-translated EDMT English-Khmer Dictionary Dataset (LLM-Translated via Gemini 3.7 & Quality Evaluated) A comprehensive, high-coverage English-to-Khmer bilingual dictionary dataset containing 176,064 entries and 110,504 distinct English headwords, built upon the open-source EDMT Dictionary Database (Webster's Revised Unabridged Dictionary). Every word definition and translation has been translated and adapted into natural, grammatically sound Khmer using Gemini 3.7 Flash. In addition, both… See the full description on the dataset page: https://huggingface.co/datasets/namae101/edmt-dictionary-english-khmer-llm-translated.tabular100K<n<1M0 likes106 downloads26d agoHugging Face26gaurannggg7 /asl-dictionary SignLink ASL Video Dictionary Matrix This dataset contains a processed matrix of American Sign Language (ASL) video clips used as the core dictionary lookup engine for SignLink, an offline voice-to-sign inference pipeline. 🤝 Attribution & Data Source The video assets in this dataset are sourced directly from the Sign-Language-Mocap-Archive created by StudioGalt. Original Creator: StudioGalt Source Repository: StudioGalt/Sign-Language-Mocap-Archive License:… See the full description on the dataset page: https://huggingface.co/datasets/gaurannggg7/asl-dictionary.videotranslation1K<n<10K0 likes100 downloads4mo agoHugging Face27conwaychriscosmo /text-embedding-3-large-english-dictionaryREADME written by Claude inspired by Chris NLTK English Word Embeddings Dataset This dataset contains embeddings for every word in the English language according to the Natural Language Toolkit (NLTK). It provides a comprehensive resource for researchers, developers, and AI enthusiasts working on natural language processing tasks. The dataset is segmented into 7 parts based on alphabetic order. Dataset Overview Source: NLTK English vocabulary Embedding Model: OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/conwaychriscosmo/text-embedding-3-large-english-dictionary.1 likes99 downloads2y agoHugging Face28fa0311 /warabi-dictionary-extended-gpl IME Dictionary Extended GPL Japanese IME dictionary pack. The SQLite database and TSV files contain the same conversion and prediction records. See NOTICE.md before redistribution. Compilation license: GPL-3.0-only SQLite tables: entries, entry_sources, predictions, prediction_sources TSV exports: entries.tsv, predictions.tsv Kaomoji flag The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter with WHERE NOT kaomoji for a face-free dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-extended-gpl.text1M<n<10M0 likes96 downloads1mo agoHugging Face29fa0311 /warabi-dictionary-core IME Dictionary Core Japanese IME dictionary pack. The SQLite database and TSV files contain the same conversion and prediction records. See NOTICE.md before redistribution. Compilation license: CC-BY-4.0 SQLite tables: entries, entry_sources, predictions, prediction_sources TSV exports: entries.tsv, predictions.tsv Kaomoji flag The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter with WHERE NOT kaomoji for a face-free dictionary. A candidate… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-core.text1M<n<10M0 likes89 downloads1mo agoHugging Face30VietAlphaLabs /fr-vi-mathematics-dictionaryVietAlpha French–Vietnamese Mathematics Dictionary Research page · VietAlpha Lab The VietAlpha French–Vietnamese Mathematics Dictionary is a machine-readable edition of Danh-từ Toán-học Pháp-Việt, compiled in Saigon in 1964 by the Mathematics Committee of the National Committee for the Compilation of Specialized Dictionaries. The release contains 4,095 dictionary entries and a 1,369-item Vietnamese index reconstructed from the printed volume. This dataset records how a Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/fr-vi-mathematics-dictionary.texttranslation1K<n<10K0 likes88 downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.