datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VietMed-NER
Medical Spoken Named Entity Recognition (NAACL 2025)
Description:
Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our knowledge, our Vietnamese real-world dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types.… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/VietMed-NER.CV18-NER
CV-18 NER
CV-18 NER is the first publicly available dataset for Named Entity Recognition (NER) from Arabic speech. It was created by augmenting the Arabic Common Voice 18 corpus with manual NER annotations following the fine-grained Wojood schema, which covers 21 entity types.
The dataset provides a benchmark for evaluating both pipeline systems (ASR + text NER) and end-to-end speech NER models. It is particularly valuable for research in low-resource settings and morphologically… See the full description on the dataset page: https://huggingface.co/datasets/Elyadata/CV18-NER.
