Vocabulary
large_vocabulary_datasetasr-jargon-specialized-vocabulary
A Dataset for Evaluating ASR on Specialized Vocabulary
Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026).
Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code
Configs
Config
Language
Description
synthetic_terms_en
English
Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms
synthetic_terms_pt
Portuguese
Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.SwissProt-Annotation-Vocabulary
Swiss-Prot Annotation Vocabulary 2026_02
This release converts a pinned Swiss-Prot snapshot into a versioned protein annotation vocabulary. Stable, namespaced term identifiers are the biological identity. Integer tokens are specific to this vocabulary and grammar version.
Release summary
Field
Value
Vocabulary version
2026_02-support10-v1
Grammar version
1
Swiss-Prot release
2026_02
Swiss-Prot release date
2026-06-10
Build date
2026-08-26… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/SwissProt-Annotation-Vocabulary.sango-vocabulary
Sango Vocabulary Dataset
Dataset Description
An open, structured, machine-readable trilingual vocabulary dataset for Sango (ISO 639-1: sg, ISO 639-3: sag), the co-official language of the Central African Republic (with French) and its most widely spoken language. Sango is a creole language with over 5 million speakers, yet it remains severely underrepresented in NLP research and digital resources.
This dataset provides trilingual vocabulary entries… See the full description on the dataset page: https://huggingface.co/datasets/MEYNG/sango-vocabulary.ESMC-6B-SAE-Annotation-Vocabulary-Features
Vocabulary interpretations of ESMC-6B SAE features
One row for every one of the 16,384 features of
biohub/ESMC-6B-sae-layer60-k64-codebook16384, giving the protein annotation vocabulary term
that best identifies what the feature detects, together with how well that identification holds on
proteins the assignment never saw.
This is the counterpart to biohub/ESMC-SAE-Features, produced without a language model. Where that
release gives a free-text hypothesis per feature, this… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ESMC-6B-SAE-Annotation-Vocabulary-Features.Open-Vocabulary-ScanNetThe Open-Vocabulary ScanNet datasets from CoDA and CoDAv2.
If the dataset is helpful, please cite:
@inproceedings{dai2017scannet,
title={ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes},
author={Dai, Angela and Chang, Angel X. and Savva, Manolis and Halber, Maciej and Funkhouser, Thomas and Nie{\ss}ner, Matthias},
booktitle = {Proc. Computer Vision and Pattern Recognition (CVPR), IEEE},
year = {2017}
}
@inproceedings{cao2023coda,
title={CoDA: Collaborative… See the full description on the dataset page: https://huggingface.co/datasets/YangCaoCS/Open-Vocabulary-ScanNet.
