CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aiola /Voxpopuli_NER VoxPopuli_NER VoxPopuli-NER is derived from the VoxPopuli corpus and specifically enhanced for Named Entity Recognition (NER) tasks focusing on political and geographical entities. It includes 879 audio samples, annotated with 2469 unique entity types. The dataset consists of the English part of the test set of VoxPopuli. See full details in the WhisperNER paper. citation If you find this usful, please cite the following works: @article{ayache2024whisperner… See the full description on the dataset page: https://huggingface.co/datasets/aiola/Voxpopuli_NER.tabulartext-classification1K<n<10K1 likes227 downloads2y agoHugging Face02the-data-nerd /vc-deal-flow-signal Startup GitHub Engineering Velocity Panel A longitudinal dataset of public GitHub engineering-activity signals for venture-backed startups. It is published under CC BY 4.0 for reproducible research, data journalism, and analysis of alternative data in venture capital. 219 startup-period observations 55 unique startups 18 sectors 4 quarterly periods: Q3 2025, Q4 2025, Q1 2026, and Q2 2026 No missing values in the primary table Version: 1.0.0 The 219 rows are startup-period… See the full description on the dataset page: https://huggingface.co/datasets/the-data-nerd/vc-deal-flow-signal.tabulartabular-classificationn<1K0 likes131 downloads1mo agoHugging Face03Zoe10 /ner_datasettabularn<1K3 likes126 downloads5y agoHugging Face04Kim-el /fever-ner FEVER Entity Retrieval Benchmark Frozen benchmark for evaluating retrieval methods on the BEIR FEVER dataset (5.4M Wikipedia articles, 6,666 test queries). All data is pre-built so you can test a new method without re-running BM25 or dense retrieval. Files Core benchmark data (for testing new methods) File Size What it is beir_pool.json 31 MB BM25 top-100 candidate pool (k1=1.2, b=0.75). 6,666 queries, each with 100 candidate docids +… See the full description on the dataset page: https://huggingface.co/datasets/Kim-el/fever-ner.tabular1K<n<10K0 likes87 downloads4mo agoHugging Face05PaulSanchez /hard-nerve-dcb861 hard-nerve-dcb861 Synthetic sensors test data: 53 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/PaulSanchez/hard-nerve-dcb861.tabularn<1K0 likes46 downloads12d agoHugging Face06Indigo-Owen /nervous-section-ccc5e3 nervous-section-ccc5e3 Synthetic products test data: 32 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Indigo-Owen/nervous-section-ccc5e3.tabularn<1K0 likes46 downloads12d agoHugging Face07Meridian-Sable /nervous-crew-48b38a nervous-crew-48b38a Synthetic weather test data: 46 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Meridian-Sable/nervous-crew-48b38a.tabularn<1K0 likes38 downloads12d agoHugging Face08velvetThomas /nervous-budget-86eac7 nervous-budget-86eac7 Synthetic products test data: 59 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/velvetThomas/nervous-budget-86eac7.tabularn<1K0 likes35 downloads12d agoHugging Face09uznlp-uz /uzbek_NER Uzbek NER Gold Uzbek NER Gold is a token-level named entity recognition dataset for Uzbek. The dataset is distributed as a UTF-8 TSV file and uses BIO tagging for named entities. Dataset Summary Dataset ID: uznlp-uz/uzbek_NER Language: Uzbek (uz) Rows: 59,569 token rows Columns: 5 Sentences: 4,176 Split: train Format: UTF-8 TSV Data file: Uzbek_NER_Gold.tsv License: CC BY 4.0 Data Fields Field Description Sentence Sentence identifier.… See the full description on the dataset page: https://huggingface.co/datasets/uznlp-uz/uzbek_NER.tabulartoken-classification10K<n<100K1 likes31 downloads3mo agoHugging Face10Velvet-Michael /nervous-radio-8ab78b nervous-radio-8ab78b Synthetic products test data: 45 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Velvet-Michael/nervous-radio-8ab78b.tabularn<1K0 likes30 downloads12d agoHugging Face11AITeamUIT /eval-gliner2-ner-fin-dice-soft-20260709tabularn<1K0 likes29 downloads3mo agoHugging Face12Llamacha /ner_quechua_iic Dataset Card for WikiANN Dataset Summary NER_Quechua_IIC is a named entity recognition dataset consisting of dictionary texts provided by the Peruvian Ministry of Education, annotated with LOC (location), PER (person) and ORG (organization) tags in the IOB2 format. Supported Tasks and Leaderboards named-entity-recognition: The dataset can be used to train a model for named entity recognition in Quechua languages. tabulartoken-classification10K<n<100K1 likes27 downloads4y agoHugging Face13AITeamUIT /eval-gliner2-ner-bionlp2004-boundary-smoothing-validationtabularn<1K0 likes27 downloads3mo agoHugging Face14AITeamUIT /eval-gliner2-ner-ontonotes5-boundary-smoothing-testtabularn<1K0 likes25 downloads3mo agoHugging Face15AITeamUIT /eval-gliner2-ner-fin-boundary-smoothing-validationtabularn<1K0 likes25 downloads3mo agoHugging Face16ScoutieAutoML /recipes_for_dishes_and_food_with_vectors_sentiment_ners Description in English: The dataset is collected from Russian-language Telegram channels with various food recipes,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/recipes_for_dishes_and_food_with_vectors_sentiment_ners.tabulartext-classification10K<n<100K2 likes24 downloads2y agoHugging Face17AITeamUIT /eval-gliner2-ner-bionlp2004-boundary-smoothing-testtabularn<1K0 likes23 downloads3mo agoHugging Face18AITeamUIT /eval-gliner2-ner-ncbi_disease-boundary-smoothing-validationtabularn<1K0 likes23 downloads3mo agoHugging Face19AITeamUIT /eval-gliner2-ner-wnut2017-boundary-smoothing-testtabularn<1K0 likes22 downloads3mo agoHugging Face20AITeamUIT /eval-gliner2-ner-conll2003-boundary-smoothing-testtabularn<1K0 likes21 downloads3mo agoHugging Face21AITeamUIT /eval-gliner2-ner-fin-boundary-smoothing-testtabularn<1K0 likes21 downloads3mo agoHugging Face22AITeamUIT /eval-gliner2-ner-bc5cdr-boundary-smoothing-validationtabularn<1K0 likes21 downloads3mo agoHugging Face23AITeamUIT /eval-gliner2-ner-ontonotes5-boundary-smoothing-validationtabularn<1K0 likes19 downloads3mo agoHugging Face24psresearch /augmented_dataset_llm_generated_NER 📚 Augmented LLM-Generated NER Dataset for Scholarly Text 🧠 Dataset Summary This dataset contains synthetically generated academic text tailored for Named Entity Recognition (NER) in the software engineering domain. The synthetic data augments scholarly writing using large language models (LLMs), with entity consistency maintained via token preservation. The dataset is generated by merging and rephrasing pairs of annotated sentences from scholarly papers using… See the full description on the dataset page: https://huggingface.co/datasets/psresearch/augmented_dataset_llm_generated_NER.tabulartoken-classification1K<n<10K0 likes18 downloads1y agoHugging Face25AITeamUIT /eval-gliner2-ner-wnut2017-boundary-smoothing-validationtabularn<1K0 likes18 downloads3mo agoHugging Face26AITeamUIT /eval-gliner2-ner-mit_restaurant-boundary-smoothing-testtabularn<1K0 likes17 downloads3mo agoHugging Face27Nerdy37 /ai-human-text-classification AI vs Human Sentence Classification Dataset Dataset Summary sentence_dataset is a sentence-level binary classification dataset containing approximately 9.84 million sentences labelled as either AI-generated (1) or human-written (0). It was constructed by extracting individual sentences from two source datasets and merging them: Dataset 1 — ai_vs_human_content_v2_20000.csv: 20,000 rows of short text and code snippets with rich metadata (prompt, topic, source… See the full description on the dataset page: https://huggingface.co/datasets/Nerdy37/ai-human-text-classification.tabulartext-classification1M<n<10M0 likes16 downloads3mo agoHugging Face28AITeamUIT /eval-gliner2-ner-ontonotes5-affine-boundary-smoothing-testtabularn<1K0 likes16 downloads3mo agoHugging Face29AITeamUIT /eval-gliner2-ner-mit_restaurant-affine-boundary-smoothing-validationtabularn<1K0 likes16 downloads3mo agoHugging Face30AITeamUIT /eval-gliner2-ner-ncbi_disease-boundary-smoothing-testtabularn<1K0 likes15 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.