CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DFKI-SLT /few-nerd Dataset Card for "Few-NERD" #dataset-description) Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive InformationConsiderations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/few-nerd.texttoken-classification100K<n<1M22 likes4.4k downloads11mo agoHugging Face02argilla /gutenberg_spacy-ner Dataset Card for "gutenberg_spacy-ner" More Information needed textn<1K4 likes2.4k downloads3y agoHugging Face03bltlab /open-ner-standardized Dataset Card for OpenNER 1.0 OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies. We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-standardized.texttoken-classification100K<n<1M2 likes1.4k downloads9mo agoHugging Face04Prikshit7766 /sample-ner Dataset Card for "sample-ner" More Information needed texttoken-classification1K<n<10K0 likes1.2k downloads2y agoHugging Face05rubrix /gutenberg_spacy-nertextn<1K1 likes591 downloads5y agoHugging Face06bltlab /open-ner-core-types Dataset Card for OpenNER 1.0 OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies. We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-core-types.texttoken-classification100K<n<1M2 likes517 downloads9mo agoHugging Face07adsabs /WIESP2022-NER Dataset for the first Workshop on Information Extraction from Scientific Publications (WIESP/2022). Dataset Description Datasets with text fragments from astrophysics papers, provided by the NASA Astrophysical Data System with manually tagged astronomical facilities and other entities of interest (e.g., celestial objects).Datasets are in JSON Lines format (each line is a json dictionary).The datasets are formatted similarly to the CONLL2003 format. Each token is… See the full description on the dataset page: https://huggingface.co/datasets/adsabs/WIESP2022-NER.texttoken-classification1K<n<10K10 likes470 downloads3y agoHugging Face08danasone /nereltextn<1K0 likes429 downloads7mo agoHugging Face09nerusskikh /taqpol_insilico_dms Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: Yulia E. Tomilova, Nikolai E. Russkikh, Igor M. Yi, Elizaveta V. Shaburova, Viktor N. Tomilov, Galina B. Pyrinova, Svetlana O. Brezhneva, Olga S. Tikhonyuk, Nadezhda S. Gololobova, Dmitriy V. Popichenko, Maxim O. Arkhipov, Leonid O. Bryzgalov, Evgeny V. Brenner… See the full description on the dataset page: https://huggingface.co/datasets/nerusskikh/taqpol_insilico_dms.tabular10M<n<100M0 likes426 downloads2y agoHugging Face10BramVanroy /universal_nerThis is an exact duplicate of https://huggingface.co/datasets/universalner/universal_ner, which is not compatible with modern versions of datasets anymore because loading data via a custom script is no longer supported. All credit goes to the original creators. Original README below. Dataset Card for Universal NER Dataset Summary Universal NER (UNER) is an open, community-driven initiative aimed at creating gold-standard benchmarks for Named Entity Recognition (NER)… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/universal_ner.texttoken-classification10K<n<100K0 likes415 downloads10mo agoHugging Face11bentrevett /instruct_uie_nertext1M<n<10M0 likes392 downloads1y agoHugging Face12FinGPT /fingpt-ner Dataset Card for "fingpt-ner" More Information needed textn<1K5 likes382 downloads3y agoHugging Face13disi-unibo-nlp /Pile-NER-biomed-IOBtext10K<n<100K0 likes359 downloads1y agoHugging Face14junyeong-nero /perfume-dataset Perfume Dataset Hugging Face-friendly export of crawled perfume records from parfumo_tidytuesday. Files data.parquet: primary tabular artifact for datasets.load_dataset(...) data.jsonl: JSON Lines version of the same split Suggested usage from datasets import load_dataset dataset = load_dataset("your-org/perfume-dataset", split="train") print(dataset[0]) Schema overview Each row corresponds to one crawled perfume record and preserves the raw… See the full description on the dataset page: https://huggingface.co/datasets/junyeong-nero/perfume-dataset.tabular10K<n<100K1 likes359 downloads6mo agoHugging Face15xusenlin /clue-ner CLUE-NER 命名实体识别数据集 字段说明 text: 文本 entities: 文本中包含的实体 id: 实体 id entity: 实体对应的字符串 start_offset: 实体开始位置 end_offset: 实体结束位置的下一位 label: 实体对应的开始位置 text10K<n<100K12 likes319 downloads4y agoHugging Face16junyeong-nero /jeju-dialect-to-standardThis dataset was created by extracting only labeled text from the Jeju dialect utterance dataset available on AIHub. by extracting only labeling text. text100K<n<1M0 likes317 downloads2y agoHugging Face17ParkSY /FSCM_Flood_nerfimage10K<n<100K0 likes298 downloads1y agoHugging Face18leduckhai /VietMed-NER Medical Spoken Named Entity Recognition (NAACL 2025) Description: Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our knowledge, our Vietnamese real-world dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types.… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/VietMed-NER.audiotoken-classification1K<n<10K9 likes296 downloads1y agoHugging Face19impresso-project /ner-augmentation Impresso HIPE-2022 NER dataset Named-entity-recognition data derived from the HIPE-2022 shared task hipe2020 bundle(s), prepared for the Impresso NER models. Train/dev are taken from revision v2.1; the test split from v2.1-test-all-unmasked (the post-evaluation release of the withheld test gold labels). Languages: de, en, fr Bundle(s): hipe2020 Config (variant): hipe2020 Label scheme: IOB2 over the NE-COARSE-LIT column. The loadable parquet config carries tokens / ner_tags (as… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/ner-augmentation.texttoken-classification100K<n<1M0 likes278 downloads14d agoHugging Face20community-datasets /swedish_medical_ner Dataset Card for swedish_medical_ner Dataset Summary SwedMedNER is Named Entity Recognition dataset on medical text in Swedish. It consists three subsets which are in turn derived from three different sources respectively: the Swedish Wikipedia (a.k.a. wiki), Läkartidningen (a.k.a. lt), and 1177 Vårdguiden (a.k.a. 1177). While the Swedish Wikipedia and Läkartidningen subsets in total contains over 790000 sequences with 60 characters each, the 1177 Vårdguiden subset is… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/swedish_medical_ner.texttoken-classification100K<n<1M4 likes253 downloads2y agoHugging Face21alexbrandsen /archaeo_ner_dutch Dutch Archaeology NER Dataset A selection of Dutch archaeology field reports, annotated by archaeology students from Leiden University. Labels The following labels are included: ART, artefacts ('bijl', 'pijlpunt') MAT, materials ('vuursteen', 'ijzer') PER, time periods ('Middeleeuwen', '400 v. Chr.') CON, archaeological contexts ('greppel','beerput') LOC, locations ('Amsterdam', 'Oss') SPE, species ('Betula nana', 'koe') Folds The reason I supply 5 folds is… See the full description on the dataset page: https://huggingface.co/datasets/alexbrandsen/archaeo_ner_dutch.texttoken-classification100K<n<1M1 likes241 downloads3y agoHugging Face22imvladikon /english_news_weak_ner Large Weak Labelled NER corpus Dataset Summary The dataset is generated through weak labelling of the scraped and preprocessed news corpus (bloomberg's news). so, only to research purpose. In order of the tokenization, news were splitted into sentences using nltk.PunktSentenceTokenizer (so, sometimes, tokenization might be not perfect) Usage from datasets import load_dataset articles_ds = load_dataset("imvladikon/english_news_weak_ner", "articles") # just… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/english_news_weak_ner.texttoken-classification1M<n<10M5 likes220 downloads3y agoHugging Face23rishitdagli /squeeze3d_rf_nerfmaeThis dataset is for the paper Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor. It contains latent representations of 3D data used for training and evaluating the Squeeze3D model. Project page Code image-to-3d1K<n<10K0 likes213 downloads1y agoHugging Face24zjushine /nero-move-allThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "tcp.x", "tcp.y", "tcp.z", "tcp.r1", "tcp.r2", "tcp.r3", "tcp.r4", "tcp.r5", "tcp.r6", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushine/nero-move-all.tabularrobotics10K<n<100K0 likes210 downloads17d agoHugging Face25zjushine /nero-allThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "tcp.x", "tcp.y", "tcp.z", "tcp.r1", "tcp.r2", "tcp.r3", "tcp.r4", "tcp.r5", "tcp.r6", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushine/nero-all.tabularrobotics10K<n<100K0 likes205 downloads18d agoHugging Face26projecte-aina /ancora-ca-ner Dataset Card for AnCora-Ca-NER Dataset Summary This is a dataset for Named Entity Recognition (NER) in Catalan. It adapts AnCora corpus for Machine Learning and Language Model evaluation purposes. This dataset was developed by BSC TeMU as part of the Projecte AINA, to enrich the Catalan Language Understanding Benchmark (CLUB). Supported Tasks and Leaderboards Named Entities Recognition, Language Model Languages The dataset is in Catalan… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ancora-ca-ner.text10K<n<100K3 likes203 downloads2y agoHugging Face27NoahWeiss /nero_bag_box nero_bag_box — dual-arm bag/box manipulation (LeRobot v2.0) Teleoperated demonstrations on the AgileX NERO dual-arm rig (two 7-DoF S-R-S arms), collected with a Pico 4 WebXR teleop rig and cut into single-task clips. Ready for π0.5 / openpi-style fine-tuning. Tasks Put the bag into the box Pour the bag out of the box Features key shape notes observation.state float32[16] left joint1–7 (rad), left gripper (0–1, measured), right joint1–7… See the full description on the dataset page: https://huggingface.co/datasets/NoahWeiss/nero_bag_box.tabularrobotics10K<n<100K0 likes202 downloads1mo agoHugging Face28PassbyGrocer /msra-nertext10K<n<100K0 likes200 downloads2y agoHugging Face29yashpwr /resume-ner-training-data Resume NER Training Dataset This dataset contains training data for Named Entity Recognition (NER) on resume text. It's used to train the yashpwr/resume-ner-bert model. Dataset Summary Task: Token Classification (NER) Language: English Domain: Resume/CV text Size: 22855 examples Format: JSONL with BIO tagging Entity Types The dataset includes the following entity types commonly found in resumes: PERSON: Names of individuals ORG: Organizations, companies… See the full description on the dataset page: https://huggingface.co/datasets/yashpwr/resume-ner-training-data.texttoken-classification10K<n<100K1 likes199 downloads1y agoHugging Face30themohal /saraiki-ner-datasetgatedtextn<1K0 likes193 downloads6h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.