datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
large_vocabulary_datasetasr-jargon-specialized-vocabulary
A Dataset for Evaluating ASR on Specialized Vocabulary
Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026).
Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code
Configs
Config
Language
Description
synthetic_terms_en
English
Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms
synthetic_terms_pt
Portuguese
Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.SwissProt-Annotation-Vocabulary
Swiss-Prot Annotation Vocabulary 2026_02
This release converts a pinned Swiss-Prot snapshot into a versioned protein annotation vocabulary. Stable, namespaced term identifiers are the biological identity. Integer tokens are specific to this vocabulary and grammar version.
Release summary
Field
Value
Vocabulary version
2026_02-support10-v1
Grammar version
1
Swiss-Prot release
2026_02
Swiss-Prot release date
2026-06-10
Build date
2026-08-26… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/SwissProt-Annotation-Vocabulary.sango-vocabulary
Sango Vocabulary Dataset
Dataset Description
An open, structured, machine-readable trilingual vocabulary dataset for Sango (ISO 639-1: sg, ISO 639-3: sag), the co-official language of the Central African Republic (with French) and its most widely spoken language. Sango is a creole language with over 5 million speakers, yet it remains severely underrepresented in NLP research and digital resources.
This dataset provides trilingual vocabulary entries… See the full description on the dataset page: https://huggingface.co/datasets/MEYNG/sango-vocabulary.ESMC-6B-SAE-Annotation-Vocabulary-Features
Vocabulary interpretations of ESMC-6B SAE features
One row for every one of the 16,384 features of
biohub/ESMC-6B-sae-layer60-k64-codebook16384, giving the protein annotation vocabulary term
that best identifies what the feature detects, together with how well that identification holds on
proteins the assignment never saw.
This is the counterpart to biohub/ESMC-SAE-Features, produced without a language model. Where that
release gives a free-text hypothesis per feature, this… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ESMC-6B-SAE-Annotation-Vocabulary-Features.Open-Vocabulary-ScanNetThe Open-Vocabulary ScanNet datasets from CoDA and CoDAv2.
If the dataset is helpful, please cite:
@inproceedings{dai2017scannet,
title={ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes},
author={Dai, Angela and Chang, Angel X. and Savva, Manolis and Halber, Maciej and Funkhouser, Thomas and Nie{\ss}ner, Matthias},
booktitle = {Proc. Computer Vision and Pattern Recognition (CVPR), IEEE},
year = {2017}
}
@inproceedings{cao2023coda,
title={CoDA: Collaborative… See the full description on the dataset page: https://huggingface.co/datasets/YangCaoCS/Open-Vocabulary-ScanNet.korean-vocabulary-5000
Koko Korean 5K — Multilingual Vocabulary Dataset
5,000 carefully curated Korean vocabulary entries with English translations,
romanization, contextual usage notes, and example sentences. Each entry is
also translated into 9 additional languages, giving researchers and
developers a high-quality parallel corpus of 50,000 aligned vocabulary
records anchored to Korean.
The dataset reflects real conversation patterns from K-dramas, K-pop, and
everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/korean-vocabulary-5000.task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation.korean-vocabulary-5000
Koko Korean 5K — Multilingual Vocabulary Dataset
5,000 carefully curated Korean vocabulary entries with English translations,
romanization, contextual usage notes, and example sentences. Each entry is
also translated into 9 additional languages, giving researchers and
developers a high-quality parallel corpus of 50,000 aligned vocabulary
records anchored to Korean.
The dataset reflects real conversation patterns from K-dramas, K-pop, and
everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/jaylee8864/korean-vocabulary-5000.ars-magna-vocabulary
Ars Magna Vocabulary
The exact vocabulary Ars Magna judges words against, so that its
claim to find every anagram of your letters can be checked rather than taken on trust.
It is three things: English OpenList at one pinned revision, a short, public list of
the site's own additions, and a short list of the site's listed forms, the contractions
whose letters the search knows. Nothing else. A word the site accepts is in one of them. Beside
the words are the term classes: the short… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/ars-magna-vocabulary.Multilingual-Core-Vocabulary
Multilingual Core Vocabulary 🌍
This dataset contains millions of frequency-sorted, highly accurate words across 19 languages. It is designed to be the ultimate resource for building cross-lingual applications, AI similarity agents, and translation models.
Dataset Structure
This repository contains two variations of the dataset:
Massive_Dataset: Contains over 6 million words. The words were extracted and frequency-sorted from FastText, cleaned from internet noise… See the full description on the dataset page: https://huggingface.co/datasets/aborasheed/Multilingual-Core-Vocabulary.Multilingual-Core-Vocabulary
Multilingual Core Vocabulary 🌍
This dataset contains millions of frequency-sorted, highly accurate words across 19 languages. It is designed to be the ultimate resource for building cross-lingual applications, AI similarity agents, and translation models.
Dataset Structure
This repository contains two variations of the dataset:
Massive_Dataset: Contains over 6 million words. The words were extracted and frequency-sorted from FastText, cleaned from internet noise… See the full description on the dataset page: https://huggingface.co/datasets/mustafaalkanxgmail/Multilingual-Core-Vocabulary.tavern-canvas-vocabulary
Tavern Canvas Vocabulary Packages
Prebuilt vocabulary packages for the Tavern Canvas SillyTavern extension. The extension reads catalog.json to list installable packages; each package under packages/<package_id>/ contains a manifest.json plus MessagePack+gzip index shards, verified per-shard by SHA-256 at install time.
Contents
package_id
records
languages
bytes
note
tavern-canvas-baseline
50,000
en, zh-CN
3.9 MB
bundled with the extension, listed here… See the full description on the dataset page: https://huggingface.co/datasets/ITOTII/tavern-canvas-vocabulary.bible-vocabulary-difficulty
Bible vocabulary-difficulty metrics, 12 translations
Per-verse reading-difficulty metrics plus a cross-language book-name table
keyed on USFM codes. Produced by bible-reader — see
scripts/export_dataset.py.
No verse text
This dataset contains references and derived metrics only, never verse
text. That is deliberate: it keeps translations under copyright (NASB)
publishable as derived data, and it keeps the download small. Fetch the texts
themselves from their own… See the full description on the dataset page: https://huggingface.co/datasets/lego573402/bible-vocabulary-difficulty.toefl-essential-vocabulary-1k
🎓 TOEFL Essential Vocabulary Dataset (AI-Enriched)
A meticulously curated, AI-enriched dataset of 1,000 high-frequency academic words essential for the TOEFL iBT, IELTS, and advanced English comprehension.
🌟 Why This Dataset?
This dataset is specifically engineered for NLP applications, language learning platforms, and academic research. Each entry includes:
Academic Theme: The specific field (e.g., Biology, Sociology) where the word frequently appears.
Exact Synonyms:… See the full description on the dataset page: https://huggingface.co/datasets/wordlevel/toefl-essential-vocabulary-1k.word-orb-vocabulary
Word Orb Vocabulary Intelligence
Structured vocabulary intelligence for AI agents, educators, and researchers. 162,253 words with pronunciation, etymology, age-appropriate definitions, translations across 47 languages, and ethical context.
Dataset Description
Word Orb is the world's most comprehensive structured vocabulary dataset designed for AI agents and education technology. Each word entry includes:
IPA pronunciation for text-to-speech and phonetics research… See the full description on the dataset page: https://huggingface.co/datasets/lotdpbc/word-orb-vocabulary.lokaniti-vocabulary-pairs
Lokaniti Pali-Burmese Vocabulary Pairs
📖 About the Dataset
This dataset provides a clean, deduplicated collection of Pali and Burmese word/phrase pairs extracted directly from the Nissaya breakdowns of the Lokaniti text. It is designed to support lexicography, alignment tasks, and machine translation experiments involving classical Pali and Burmese languages.
Total Unique Pairs: 1,640
Data Integrity: 0 null values.
👤 Who Created
This dataset and… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/lokaniti-vocabulary-pairs.lemonseed-vocabulary
lemonseed-vocabulary
LemonSeed — WordNet vocabulary Q&A with chain-of-thought definitions.
Contents
vocab.jsonl (4000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
english-vocabulary-materials
English Vocabulary Teaching Materials(雅思与初高中词汇教学资料)
中学段的英语词汇教学资料:雅思分级词汇(预备班 / 一阶 / 二阶 / 三阶)的词汇本、词测本、配套听力录音与听说读讲义,外加初高中词表。
原始材料是 PDF、MP3 和 Excel —— 词表分散在 Excel 的多张工作表里,音频按中文文件名散落各目录,PDF 里的词测没法检索。这份仓库做了两件事:60 个原始文件原样归档不做改动,另外从 Excel 抽出 8257 条结构化词条存成 CSV/JSONL,可以直接读进来做背诵、默写、出题或全文检索。
⚠️ 版权提醒
这批材料整理自绿新的教研资料,不是原创数据集,著作权归原权利人所有。仓库采用 CC BY-NC-ND 4.0 并开启 gated access:禁止商业使用、再分发、公开镜像与演绎;研究用途允许,但须按第 7 节格式署名。完整条款见第 6 节。
1. 数据总览
指标
数值
清单内文件
85(另有… See the full description on the dataset page: https://huggingface.co/datasets/SwiftieJerry/english-vocabulary-materials.english-vocabularyDataset contains all English words from the dictionary!
Open-Vocabulary-SUN-RGBDThe Open-Vocabulary SUN-RGBD datasets from CoDA and CoDAv2.
If the dataset is helpful, please cite:
@inproceedings{song2015sun,
title={Sun rgb-d: A rgb-d scene understanding benchmark suite},
author={Song, Shuran and Lichtenberg, Samuel P and Xiao, Jianxiong},
booktitle={CVPR},
year={2015}
}
@inproceedings{cao2023coda,
title={CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection},
author={Cao, Yang and Zeng, Yihan and Xu, Hang… See the full description on the dataset page: https://huggingface.co/datasets/YangCaoCS/Open-Vocabulary-SUN-RGBD.jlpt_n5_vocabulary
Jisho JLPT-N5 Relational Dataset
This dataset provides a comprehensive and highly normalized collection of Japanese vocabulary for the JLPT-N5 level, scraped from Jisho.org. Unlike flat datasets, this version uses a relational schema to separate headwords, metadata (JLPT/WaniKani levels), and individual English definitions.
🏗 Scraper Architecture
The data was generated using a custom Python scraper following a robust state-machine logic.
Logic Highlights:… See the full description on the dataset page: https://huggingface.co/datasets/xqt/jlpt_n5_vocabulary.jlpt_n5_vocabulary_tagged
Jisho JLPT-N5 Relational Dataset
This dataset provides a comprehensive and highly normalized collection of Japanese vocabulary for the JLPT-N5 level, scraped from Jisho.org and enriched with DeepSeek-V3 AI tagging.
🏗 Scraper & AI Architecture
The data was generated using a custom Python scraper and post-processed using a batched AI tagging pipeline.
Logic Highlights:
Normalization: Every English sense is a unique row with a specific meaning_no.… See the full description on the dataset page: https://huggingface.co/datasets/xqt/jlpt_n5_vocabulary_tagged.Sanskrit-to-English-Vocabulary-v1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: Manoj, Nandish, Mayank, Abhiram
Language(s) (NLP): Sanskrit, English
License: MIT License
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Manoj2702/Sanskrit-to-English-Vocabulary-v1.ePark_xue_xi_ci_biao_learning_vocabulary
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary.za_vocabularytask104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation.hatevolution-vocabulary-expansion
Dataset Info
The hatevolution-vocabulary-expansion dataset contains the data used for Experiment 2 in the paper Hatevolution: What Static Benchmarks Don't Tell Us (Di Bonaventura et al., 2025).
It is built using the NeoBench dataset (Zheng et al., 2024), further annotated for hate speech detection.
advanced_vocabulary_ko_en_v0.1vocabulary_nachos_lowercased
