datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-dictionary
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-dictionary.dichvucong-gov-vn
Vietnam Administrative Procedures — full structured detail (dichvucong.gov.vn)
🇻🇳 Tóm tắt. 3,927 thủ tục hành chính từ Cổng Dịch vụ công
Quốc gia, mỗi thủ tục kèm toàn bộ nội dung có cấu trúc: trình tự, cách
thức, thành phần hồ sơ, phí/lệ phí, căn cứ pháp lý, kết quả, cơ quan thực
hiện. Kèm embedding + toạ độ UMAP/PCA/t-SNE và một báo cáo phân tích sâu.
🇬🇧 Summary. 3,927 Vietnamese administrative procedures from the
National Public Service Portal, each with the full… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/dichvucong-gov-vn.opengloss-dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English lexemes
9.1… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-dictionary.opengloss-dictionary-definitions
OpenGloss Dictionary (Definition-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the definitions-level view where each record represents one sense definition.
Key Statistics
536,829 sense definitions across 150,101 English lexemes
9.1… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-dictionary-definitions.opengloss-v1.2-dictionary
OpenGloss Dictionary v1.2 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
162,314 lexemes
7,798,653 semantic edges… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-dictionary.woori_spring_dict
Dataset Card for "woori_spring_dict"
This dataset is a NLP learnable form of woori mal saem(우리말샘) a Korean collaborative open source dictionary.
It follows the original copyright policy (cc-by-sa-2.0)
This version is built from xls_20230602
우리말샘을 학습 가능한 형태로 처리한 데이터입니다.
우리말샘의 저작권을 따릅니다.
xls_20230602으로부터 생성되었습니다.
More Information needed
basic_korean_dict
Dataset Card for "basic_korean_dict"
This dataset is a NLP learnable form of Korean Basic Dictionary(한국어기초사전).
It follows the original copyright policy (cc-by-sa-2.0)
Some words have usage examples in other languages, effectively rendering this into a parallel corpus.
This version is built from xls_20230601
한국어 기초 사전을 학습 가능한 형태로 처리한 데이터입니다.
한국어 기초 사전의 저작권을 따릅니다.
여러 언어로 이루어진 표제어들이 있어 병렬 말뭉치의 기능이 있습니다.
xls_20230601으로부터 생성되었습니다.
opengloss-dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-dictionary.opengloss-v1.1-dictionary
OpenGloss Dictionary v1.1 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,637 lexemes
7,701,312 semantic edges… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-dictionary.personal_dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/caioloures/personal_dictionary.R550-ROS2-Graph-Dictionary
R550 ROS2 Graph Dictionary
Part of RabbitRobot VLN & Open Robot Assets: public demos and datasets distilled from real ROS 2 robot engineering.
这是 lijinghai 在 RabbitRobot / R550 机器人开发中整理出的公开 ROS 2 图谱字典数据集。
原始项目是一个本地 Web 字典工具:把真实机器人启动后的功能域、节点、Topic、Service、Action、launch 命令和工程说明整理到浏览器里,方便调试、检索和复盘。这个 Hugging Face 版本只发布脱敏后的结构化数据,不发布真实内网地址、访问凭据、完整运行日志或私有机器人文件。
What is included
data/r550_ros2_features.json: 14 个功能域的结构化 ROS 2 字典。
data/feature_summary.csv:… See the full description on the dataset page: https://huggingface.co/datasets/lijinghai/R550-ROS2-Graph-Dictionary.opengloss-v1.3-dictionary
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
205,988 lexemes
8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.dichvucong-gov-vn
Vietnam Administrative Procedures — full structured detail (dichvucong.gov.vn)
🇻🇳 Tóm tắt. 3,927 thủ tục hành chính từ Cổng Dịch vụ công
Quốc gia, mỗi thủ tục kèm toàn bộ nội dung có cấu trúc: trình tự, cách
thức, thành phần hồ sơ, phí/lệ phí, căn cứ pháp lý, kết quả, cơ quan thực
hiện. Kèm embedding + toạ độ UMAP/PCA/t-SNE và một báo cáo phân tích sâu.
🇬🇧 Summary. 3,927 Vietnamese administrative procedures from the
National Public Service Portal, each with the full… See the full description on the dataset page: https://huggingface.co/datasets/kmt74/dichvucong-gov-vn.opengloss-v1.1-dictionary
OpenGloss Dictionary v1.1 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,637 lexemes
7,701,312 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.1-dictionary.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.aiysha-diction-500
AIySha: yShade.AI AI Agent
This is the formatted dataset for customizing the diction of the bot backed by llama-2-7b-chat model.
The dataset fits the prompt template for the chat model and is ready to be used for fine tuning purposes.
The goal of the dataset is to train the model to be specialized as a beauty advisor.
mon_eng_dict_instructions
Mon-English Dictionary Instruction Dataset (Mon-AI Project)
📌 Project Overview
This dataset is a comprehensive, scalable, and high-quality Mon-English Instruction-Prompt Dataset designed specifically for supervised fine-tuning (SFT) of Large Language Models (LLMs).
The Mon language (ISO 639-3: mnw) is historically rich but classified as a low-resource language in the digital and AI landscape. The core mission of this project is to scale Mon linguistic resources… See the full description on the dataset page: https://huggingface.co/datasets/Nenemin95/mon_eng_dict_instructions.dickens_data_quality_checksMMVP
MMVP Benchmark Datacard
Basic Information
Title: MMVP Benchmark
Description: The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying “CLIP-blind pairs” – images that are perceived as similar by CLIP despite having clear visual differences. MMVP benchmarks the performance of state-of-the-art systems, including GPT-4V, across nine basic visual patterns. It highlights the challenges these systems face in answering straightforward questions, often leading to… See the full description on the dataset page: https://huggingface.co/datasets/DickMan42/MMVP.dictionary-embeddingsEmbeddings generated from the model multi-qa-mpnet-base-dot-v1 being trained on MAKILINGDING/english_dictionary
aiysha-diction
AIySha: yShade.AI AI Agent
This is the base dataset for customizing the diction of the bot backed by llama-2-7b-chat model.
The dataset needs to be reformatted to fit the prompt template for the chat model in order to use for fine tuning purposes.
The goal of the dataset is to train the model to be specialized as a beauty advisor.
dictionaryПроект по запуску бота
в реальном времени
полезная нагрузка 8541
Открытый код python
быстрое внедрение
удобное взаимодействие
автотестирование
Надежность
