datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
few-nerd
Dataset Card for "Few-NERD"
#dataset-description)
Dataset Summary
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Annotations
Personal and Sensitive InformationConsiderations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information
Contributions
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/few-nerd.gutenberg_spacy-ner
Dataset Card for "gutenberg_spacy-ner"
More Information needed
open-ner-standardized
Dataset Card for OpenNER 1.0
OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets.
OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies.
We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-standardized.sample-ner
Dataset Card for "sample-ner"
More Information needed
gutenberg_spacy-neropen-ner-core-types
Dataset Card for OpenNER 1.0
OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets.
OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies.
We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-core-types.WIESP2022-NER
Dataset for the first Workshop on Information Extraction from Scientific Publications (WIESP/2022).
Dataset Description
Datasets with text fragments from astrophysics papers, provided by the NASA Astrophysical Data System with manually tagged astronomical facilities and other entities of interest (e.g., celestial objects).Datasets are in JSON Lines format (each line is a json dictionary).The datasets are formatted similarly to the CONLL2003 format. Each token is… See the full description on the dataset page: https://huggingface.co/datasets/adsabs/WIESP2022-NER.nereltaqpol_insilico_dms
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: Yulia E. Tomilova,
Nikolai E. Russkikh,
Igor M. Yi,
Elizaveta V. Shaburova,
Viktor N. Tomilov,
Galina B. Pyrinova,
Svetlana O. Brezhneva,
Olga S. Tikhonyuk,
Nadezhda S. Gololobova,
Dmitriy V. Popichenko,
Maxim O. Arkhipov,
Leonid O. Bryzgalov,
Evgeny V. Brenner… See the full description on the dataset page: https://huggingface.co/datasets/nerusskikh/taqpol_insilico_dms.universal_nerThis is an exact duplicate of https://huggingface.co/datasets/universalner/universal_ner, which is not compatible with modern versions of datasets anymore because loading data via a custom script is no longer supported. All credit goes to the original creators. Original README below.
Dataset Card for Universal NER
Dataset Summary
Universal NER (UNER) is an open, community-driven initiative aimed at creating gold-standard benchmarks for Named Entity Recognition (NER)… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/universal_ner.instruct_uie_nerfingpt-ner
Dataset Card for "fingpt-ner"
More Information needed
Pile-NER-biomed-IOBperfume-dataset
Perfume Dataset
Hugging Face-friendly export of crawled perfume records from parfumo_tidytuesday.
Files
data.parquet: primary tabular artifact for datasets.load_dataset(...)
data.jsonl: JSON Lines version of the same split
Suggested usage
from datasets import load_dataset
dataset = load_dataset("your-org/perfume-dataset", split="train")
print(dataset[0])
Schema overview
Each row corresponds to one crawled perfume record and preserves the raw… See the full description on the dataset page: https://huggingface.co/datasets/junyeong-nero/perfume-dataset.clue-ner
CLUE-NER 命名实体识别数据集
字段说明
text: 文本
entities: 文本中包含的实体
id: 实体 id
entity: 实体对应的字符串
start_offset: 实体开始位置
end_offset: 实体结束位置的下一位
label: 实体对应的开始位置
jeju-dialect-to-standardThis dataset was created by extracting only labeled text from the Jeju dialect utterance dataset available on AIHub.
by extracting only labeling text.
FSCM_Flood_nerfVietMed-NER
Medical Spoken Named Entity Recognition (NAACL 2025)
Description:
Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our knowledge, our Vietnamese real-world dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types.… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/VietMed-NER.ner-augmentation
Impresso HIPE-2022 NER dataset
Named-entity-recognition data derived from the
HIPE-2022 shared task hipe2020 bundle(s), prepared
for the Impresso NER models. Train/dev are taken from revision
v2.1; the test split from v2.1-test-all-unmasked (the
post-evaluation release of the withheld test gold labels).
Languages: de, en, fr
Bundle(s): hipe2020
Config (variant): hipe2020
Label scheme: IOB2 over the NE-COARSE-LIT column.
The loadable parquet config carries tokens / ner_tags (as… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/ner-augmentation.swedish_medical_ner
Dataset Card for swedish_medical_ner
Dataset Summary
SwedMedNER is Named Entity Recognition dataset on medical text in Swedish. It consists three subsets which are in turn derived from three different sources respectively: the Swedish Wikipedia (a.k.a. wiki), Läkartidningen (a.k.a. lt), and 1177 Vårdguiden (a.k.a. 1177). While the Swedish Wikipedia and Läkartidningen subsets in total contains over 790000 sequences with 60 characters each, the 1177 Vårdguiden subset is… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/swedish_medical_ner.archaeo_ner_dutch
Dutch Archaeology NER Dataset
A selection of Dutch archaeology field reports, annotated by archaeology students from Leiden University.
Labels
The following labels are included:
ART, artefacts ('bijl', 'pijlpunt')
MAT, materials ('vuursteen', 'ijzer')
PER, time periods ('Middeleeuwen', '400 v. Chr.')
CON, archaeological contexts ('greppel','beerput')
LOC, locations ('Amsterdam', 'Oss')
SPE, species ('Betula nana', 'koe')
Folds
The reason I supply 5 folds is… See the full description on the dataset page: https://huggingface.co/datasets/alexbrandsen/archaeo_ner_dutch.english_news_weak_ner
Large Weak Labelled NER corpus
Dataset Summary
The dataset is generated through weak labelling of the scraped and preprocessed news corpus (bloomberg's news). so, only to research purpose.
In order of the tokenization, news were splitted into sentences using nltk.PunktSentenceTokenizer (so, sometimes, tokenization might be not perfect)
Usage
from datasets import load_dataset
articles_ds = load_dataset("imvladikon/english_news_weak_ner", "articles") # just… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/english_news_weak_ner.squeeze3d_rf_nerfmaeThis dataset is for the paper Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor. It contains latent representations of 3D data used for training and evaluating the Squeeze3D model.
Project page
Code
nero-move-allThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"tcp.x",
"tcp.y",
"tcp.z",
"tcp.r1",
"tcp.r2",
"tcp.r3",
"tcp.r4",
"tcp.r5",
"tcp.r6",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushine/nero-move-all.nero-allThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"tcp.x",
"tcp.y",
"tcp.z",
"tcp.r1",
"tcp.r2",
"tcp.r3",
"tcp.r4",
"tcp.r5",
"tcp.r6",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushine/nero-all.ancora-ca-ner
Dataset Card for AnCora-Ca-NER
Dataset Summary
This is a dataset for Named Entity Recognition (NER) in Catalan. It adapts AnCora corpus for Machine Learning and Language Model evaluation purposes.
This dataset was developed by BSC TeMU as part of the Projecte AINA, to enrich the Catalan Language Understanding Benchmark (CLUB).
Supported Tasks and Leaderboards
Named Entities Recognition, Language Model
Languages
The dataset is in Catalan… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ancora-ca-ner.nero_bag_box
nero_bag_box — dual-arm bag/box manipulation (LeRobot v2.0)
Teleoperated demonstrations on the AgileX NERO dual-arm rig (two 7-DoF
S-R-S arms), collected with a Pico 4 WebXR teleop rig and cut into
single-task clips. Ready for π0.5 / openpi-style fine-tuning.
Tasks
Put the bag into the box
Pour the bag out of the box
Features
key
shape
notes
observation.state
float32[16]
left joint1–7 (rad), left gripper (0–1, measured), right joint1–7… See the full description on the dataset page: https://huggingface.co/datasets/NoahWeiss/nero_bag_box.msra-nerresume-ner-training-data
Resume NER Training Dataset
This dataset contains training data for Named Entity Recognition (NER) on resume text. It's used to train the yashpwr/resume-ner-bert model.
Dataset Summary
Task: Token Classification (NER)
Language: English
Domain: Resume/CV text
Size: 22855 examples
Format: JSONL with BIO tagging
Entity Types
The dataset includes the following entity types commonly found in resumes:
PERSON: Names of individuals
ORG: Organizations, companies… See the full description on the dataset page: https://huggingface.co/datasets/yashpwr/resume-ner-training-data.saraiki-ner-dataset
