datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Voxpopuli_NER
VoxPopuli_NER
VoxPopuli-NER is derived from the VoxPopuli corpus and specifically enhanced for
Named Entity Recognition (NER) tasks focusing on political and geographical entities.
It includes 879 audio samples, annotated with 2469 unique entity types. The dataset consists of the English part of the test set of VoxPopuli.
See full details in the WhisperNER paper.
citation
If you find this usful, please cite the following works:
@article{ayache2024whisperner… See the full description on the dataset page: https://huggingface.co/datasets/aiola/Voxpopuli_NER.vc-deal-flow-signal
Startup GitHub Engineering Velocity Panel
A longitudinal dataset of public GitHub engineering-activity signals for venture-backed startups. It is published under CC BY 4.0 for reproducible research, data journalism, and analysis of alternative data in venture capital.
219 startup-period observations
55 unique startups
18 sectors
4 quarterly periods: Q3 2025, Q4 2025, Q1 2026, and Q2 2026
No missing values in the primary table
Version: 1.0.0
The 219 rows are startup-period… See the full description on the dataset page: https://huggingface.co/datasets/the-data-nerd/vc-deal-flow-signal.ner_datasetfever-ner
FEVER Entity Retrieval Benchmark
Frozen benchmark for evaluating retrieval methods on the BEIR FEVER dataset (5.4M Wikipedia articles, 6,666 test queries). All data is pre-built so you can test a new method without re-running BM25 or dense retrieval.
Files
Core benchmark data (for testing new methods)
File
Size
What it is
beir_pool.json
31 MB
BM25 top-100 candidate pool (k1=1.2, b=0.75). 6,666 queries, each with 100 candidate docids +… See the full description on the dataset page: https://huggingface.co/datasets/Kim-el/fever-ner.hard-nerve-dcb861
hard-nerve-dcb861
Synthetic sensors test data: 53 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/PaulSanchez/hard-nerve-dcb861.nervous-section-ccc5e3
nervous-section-ccc5e3
Synthetic products test data: 32 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Indigo-Owen/nervous-section-ccc5e3.nervous-crew-48b38a
nervous-crew-48b38a
Synthetic weather test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Meridian-Sable/nervous-crew-48b38a.nervous-budget-86eac7
nervous-budget-86eac7
Synthetic products test data: 59 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/velvetThomas/nervous-budget-86eac7.uzbek_NER
Uzbek NER Gold
Uzbek NER Gold is a token-level named entity recognition dataset for Uzbek. The dataset is distributed as a UTF-8 TSV file and uses BIO tagging for named entities.
Dataset Summary
Dataset ID: uznlp-uz/uzbek_NER
Language: Uzbek (uz)
Rows: 59,569 token rows
Columns: 5
Sentences: 4,176
Split: train
Format: UTF-8 TSV
Data file: Uzbek_NER_Gold.tsv
License: CC BY 4.0
Data Fields
Field
Description
Sentence
Sentence identifier.… See the full description on the dataset page: https://huggingface.co/datasets/uznlp-uz/uzbek_NER.nervous-radio-8ab78b
nervous-radio-8ab78b
Synthetic products test data: 45 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Velvet-Michael/nervous-radio-8ab78b.eval-gliner2-ner-fin-dice-soft-20260709ner_quechua_iic
Dataset Card for WikiANN
Dataset Summary
NER_Quechua_IIC is a named entity recognition dataset consisting of dictionary texts provided by the Peruvian Ministry of Education, annotated with LOC (location), PER (person) and ORG (organization) tags in the IOB2 format.
Supported Tasks and Leaderboards
named-entity-recognition: The dataset can be used to train a model for named entity recognition in Quechua languages.
eval-gliner2-ner-bionlp2004-boundary-smoothing-validationeval-gliner2-ner-ontonotes5-boundary-smoothing-testeval-gliner2-ner-fin-boundary-smoothing-validationrecipes_for_dishes_and_food_with_vectors_sentiment_ners
Description in English:
The dataset is collected from Russian-language Telegram channels with various food recipes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/recipes_for_dishes_and_food_with_vectors_sentiment_ners.eval-gliner2-ner-bionlp2004-boundary-smoothing-testeval-gliner2-ner-ncbi_disease-boundary-smoothing-validationeval-gliner2-ner-wnut2017-boundary-smoothing-testeval-gliner2-ner-conll2003-boundary-smoothing-testeval-gliner2-ner-fin-boundary-smoothing-testeval-gliner2-ner-bc5cdr-boundary-smoothing-validationeval-gliner2-ner-ontonotes5-boundary-smoothing-validationaugmented_dataset_llm_generated_NER
📚 Augmented LLM-Generated NER Dataset for Scholarly Text
🧠 Dataset Summary
This dataset contains synthetically generated academic text tailored for Named Entity Recognition (NER) in the software engineering domain. The synthetic data augments scholarly writing using large language models (LLMs), with entity consistency maintained via token preservation.
The dataset is generated by merging and rephrasing pairs of annotated sentences from scholarly papers using… See the full description on the dataset page: https://huggingface.co/datasets/psresearch/augmented_dataset_llm_generated_NER.eval-gliner2-ner-wnut2017-boundary-smoothing-validationeval-gliner2-ner-mit_restaurant-boundary-smoothing-testai-human-text-classification
AI vs Human Sentence Classification Dataset
Dataset Summary
sentence_dataset is a sentence-level binary classification dataset containing approximately 9.84 million sentences labelled as either AI-generated (1) or human-written (0).
It was constructed by extracting individual sentences from two source datasets and merging them:
Dataset 1 — ai_vs_human_content_v2_20000.csv: 20,000 rows of short text and code snippets with rich metadata (prompt, topic, source… See the full description on the dataset page: https://huggingface.co/datasets/Nerdy37/ai-human-text-classification.eval-gliner2-ner-ontonotes5-affine-boundary-smoothing-testeval-gliner2-ner-mit_restaurant-affine-boundary-smoothing-validationeval-gliner2-ner-ncbi_disease-boundary-smoothing-test
