datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
large_vocabulary_datasetasr-jargon-specialized-vocabulary
A Dataset for Evaluating ASR on Specialized Vocabulary
Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026).
Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code
Configs
Config
Language
Description
synthetic_terms_en
English
Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms
synthetic_terms_pt
Portuguese
Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.SwissProt-Annotation-Vocabulary
Swiss-Prot Annotation Vocabulary 2026_02
This release converts a pinned Swiss-Prot snapshot into a versioned protein annotation vocabulary. Stable, namespaced term identifiers are the biological identity. Integer tokens are specific to this vocabulary and grammar version.
Release summary
Field
Value
Vocabulary version
2026_02-support10-v1
Grammar version
1
Swiss-Prot release
2026_02
Swiss-Prot release date
2026-06-10
Build date
2026-08-26… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/SwissProt-Annotation-Vocabulary.sango-vocabulary
Sango Vocabulary Dataset
Dataset Description
An open, structured, machine-readable trilingual vocabulary dataset for Sango (ISO 639-1: sg, ISO 639-3: sag), the co-official language of the Central African Republic (with French) and its most widely spoken language. Sango is a creole language with over 5 million speakers, yet it remains severely underrepresented in NLP research and digital resources.
This dataset provides trilingual vocabulary entries… See the full description on the dataset page: https://huggingface.co/datasets/MEYNG/sango-vocabulary.ESMC-6B-SAE-Annotation-Vocabulary-Features
Vocabulary interpretations of ESMC-6B SAE features
One row for every one of the 16,384 features of
biohub/ESMC-6B-sae-layer60-k64-codebook16384, giving the protein annotation vocabulary term
that best identifies what the feature detects, together with how well that identification holds on
proteins the assignment never saw.
This is the counterpart to biohub/ESMC-SAE-Features, produced without a language model. Where that
release gives a free-text hypothesis per feature, this… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ESMC-6B-SAE-Annotation-Vocabulary-Features.korean-vocabulary-5000
Koko Korean 5K — Multilingual Vocabulary Dataset
5,000 carefully curated Korean vocabulary entries with English translations,
romanization, contextual usage notes, and example sentences. Each entry is
also translated into 9 additional languages, giving researchers and
developers a high-quality parallel corpus of 50,000 aligned vocabulary
records anchored to Korean.
The dataset reflects real conversation patterns from K-dramas, K-pop, and
everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/korean-vocabulary-5000.task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation.korean-vocabulary-5000
Koko Korean 5K — Multilingual Vocabulary Dataset
5,000 carefully curated Korean vocabulary entries with English translations,
romanization, contextual usage notes, and example sentences. Each entry is
also translated into 9 additional languages, giving researchers and
developers a high-quality parallel corpus of 50,000 aligned vocabulary
records anchored to Korean.
The dataset reflects real conversation patterns from K-dramas, K-pop, and
everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/jaylee8864/korean-vocabulary-5000.ars-magna-vocabulary
Ars Magna Vocabulary
The exact vocabulary Ars Magna judges words against, so that its
claim to find every anagram of your letters can be checked rather than taken on trust.
It is three things: English OpenList at one pinned revision, a short, public list of
the site's own additions, and a short list of the site's listed forms, the contractions
whose letters the search knows. Nothing else. A word the site accepts is in one of them. Beside
the words are the term classes: the short… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/ars-magna-vocabulary.bible-vocabulary-difficulty
Bible vocabulary-difficulty metrics, 12 translations
Per-verse reading-difficulty metrics plus a cross-language book-name table
keyed on USFM codes. Produced by bible-reader — see
scripts/export_dataset.py.
No verse text
This dataset contains references and derived metrics only, never verse
text. That is deliberate: it keeps translations under copyright (NASB)
publishable as derived data, and it keeps the download small. Fetch the texts
themselves from their own… See the full description on the dataset page: https://huggingface.co/datasets/lego573402/bible-vocabulary-difficulty.toefl-essential-vocabulary-1k
🎓 TOEFL Essential Vocabulary Dataset (AI-Enriched)
A meticulously curated, AI-enriched dataset of 1,000 high-frequency academic words essential for the TOEFL iBT, IELTS, and advanced English comprehension.
🌟 Why This Dataset?
This dataset is specifically engineered for NLP applications, language learning platforms, and academic research. Each entry includes:
Academic Theme: The specific field (e.g., Biology, Sociology) where the word frequently appears.
Exact Synonyms:… See the full description on the dataset page: https://huggingface.co/datasets/wordlevel/toefl-essential-vocabulary-1k.word-orb-vocabulary
Word Orb Vocabulary Intelligence
Structured vocabulary intelligence for AI agents, educators, and researchers. 162,253 words with pronunciation, etymology, age-appropriate definitions, translations across 47 languages, and ethical context.
Dataset Description
Word Orb is the world's most comprehensive structured vocabulary dataset designed for AI agents and education technology. Each word entry includes:
IPA pronunciation for text-to-speech and phonetics research… See the full description on the dataset page: https://huggingface.co/datasets/lotdpbc/word-orb-vocabulary.lokaniti-vocabulary-pairs
Lokaniti Pali-Burmese Vocabulary Pairs
📖 About the Dataset
This dataset provides a clean, deduplicated collection of Pali and Burmese word/phrase pairs extracted directly from the Nissaya breakdowns of the Lokaniti text. It is designed to support lexicography, alignment tasks, and machine translation experiments involving classical Pali and Burmese languages.
Total Unique Pairs: 1,640
Data Integrity: 0 null values.
👤 Who Created
This dataset and… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/lokaniti-vocabulary-pairs.lemonseed-vocabulary
lemonseed-vocabulary
LemonSeed — WordNet vocabulary Q&A with chain-of-thought definitions.
Contents
vocab.jsonl (4000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
english-vocabulary-materials
English Vocabulary Teaching Materials(雅思与初高中词汇教学资料)
中学段的英语词汇教学资料:雅思分级词汇(预备班 / 一阶 / 二阶 / 三阶)的词汇本、词测本、配套听力录音与听说读讲义,外加初高中词表。
原始材料是 PDF、MP3 和 Excel —— 词表分散在 Excel 的多张工作表里,音频按中文文件名散落各目录,PDF 里的词测没法检索。这份仓库做了两件事:60 个原始文件原样归档不做改动,另外从 Excel 抽出 8257 条结构化词条存成 CSV/JSONL,可以直接读进来做背诵、默写、出题或全文检索。
⚠️ 版权提醒
这批材料整理自绿新的教研资料,不是原创数据集,著作权归原权利人所有。仓库采用 CC BY-NC-ND 4.0 并开启 gated access:禁止商业使用、再分发、公开镜像与演绎;研究用途允许,但须按第 7 节格式署名。完整条款见第 6 节。
1. 数据总览
指标
数值
清单内文件
85(另有… See the full description on the dataset page: https://huggingface.co/datasets/SwiftieJerry/english-vocabulary-materials.english-vocabularyDataset contains all English words from the dictionary!
Sanskrit-to-English-Vocabulary-v1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: Manoj, Nandish, Mayank, Abhiram
Language(s) (NLP): Sanskrit, English
License: MIT License
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Manoj2702/Sanskrit-to-English-Vocabulary-v1.za_vocabularytask104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation.advanced_vocabulary_ko_en_v0.1vocabulary_nachos_lowercasedllama3_vocabulary_clusterThis dataset contains the clusters discovered in the vocabulary embeddings of the llama3-8b-instruct model.
The 128256 vocabulary embeddings are separated into 1024 clusters by k-means, which show pattern correlations probably undesirable for diverse generation.
We also prompt GPT-4o to summarize the commonality of vocabularies in the same cluster, which can be used for further analysis.
This dataset is a part of the work on diverse LLM generation. [Paper], [Github]
flan_combined_task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generationAmazigh_Researchers_Vocabulary
Mohammed Lchger Vocabulary Dataset
This dataset contains a collection of vocabulary compiled by Mohamed Lachgar, the owner of the Amazigh researchers blog which its dataset can be find here.
Dataset Details
Content: 2,000 scientific related terms and general vocabulary.
Languages: Amazigh (zgh, ber) and English.
Script: Tifinagh.
Unique Feature: It is unique especially in its inclusion of scientific vocabulary.
Acknowledgments
Thanks to Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh_Researchers_Vocabulary.anyfeature_vocabularyolmo_vocabulary_clusterThis dataset contains the clusters discovered in the vocabulary embeddings of the olmo-7b-sft model.
The 50280 vocabulary embeddings are separated into 512 clusters by k-means, which show pattern correlations probably undesirable for diverse generation.
We also prompt GPT-4o to summarize the commonality of vocabularies in the same cluster, which can be used for further analysis.
This dataset is a part of the work on diverse LLM generation. [Paper], [Github]
vocabulary-storevocabularywikimedia_es_vocabularywellness-vocabularyWellness & Healthcare Terminology Dataset
This dataset contains a curated list of wellness terms, product categories, and SEO-optimized keywords related to high-quality healthcare products.
Purpose: To help AI models understand the nuances of premium Japanese wellness standards and improve translation accuracy for the Vietnamese market.
Maintained by oichin.net
