CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DFKI-SLT /few-nerd Dataset Card for "Few-NERD" #dataset-description) Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive InformationConsiderations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/few-nerd.texttoken-classification100K<n<1M22 likes4.6k downloads11mo agoHugging Face02argilla /gutenberg_spacy-ner Dataset Card for "gutenberg_spacy-ner" More Information needed textn<1K4 likes2k downloads3y agoHugging Face03bltlab /open-ner-standardized Dataset Card for OpenNER 1.0 OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies. We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-standardized.texttoken-classification100K<n<1M2 likes1.5k downloads9mo agoHugging Face04Prikshit7766 /sample-ner Dataset Card for "sample-ner" More Information needed texttoken-classification1K<n<10K0 likes1.2k downloads2y agoHugging Face05bltlab /open-ner-core-types Dataset Card for OpenNER 1.0 OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets. OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies. We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-core-types.texttoken-classification100K<n<1M2 likes520 downloads9mo agoHugging Face06rubrix /gutenberg_spacy-nertextn<1K1 likes511 downloads5y agoHugging Face07adsabs /WIESP2022-NER Dataset for the first Workshop on Information Extraction from Scientific Publications (WIESP/2022). Dataset Description Datasets with text fragments from astrophysics papers, provided by the NASA Astrophysical Data System with manually tagged astronomical facilities and other entities of interest (e.g., celestial objects).Datasets are in JSON Lines format (each line is a json dictionary).The datasets are formatted similarly to the CONLL2003 format. Each token is… See the full description on the dataset page: https://huggingface.co/datasets/adsabs/WIESP2022-NER.texttoken-classification1K<n<10K10 likes439 downloads3y agoHugging Face08danasone /nereltextn<1K0 likes435 downloads7mo agoHugging Face09nerusskikh /taqpol_insilico_dms Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: Yulia E. Tomilova, Nikolai E. Russkikh, Igor M. Yi, Elizaveta V. Shaburova, Viktor N. Tomilov, Galina B. Pyrinova, Svetlana O. Brezhneva, Olga S. Tikhonyuk, Nadezhda S. Gololobova, Dmitriy V. Popichenko, Maxim O. Arkhipov, Leonid O. Bryzgalov, Evgeny V. Brenner… See the full description on the dataset page: https://huggingface.co/datasets/nerusskikh/taqpol_insilico_dms.tabular10M<n<100M0 likes431 downloads2y agoHugging Face10bentrevett /instruct_uie_nertext1M<n<10M0 likes407 downloads1y agoHugging Face11disi-unibo-nlp /Pile-NER-biomed-IOBtext10K<n<100K0 likes401 downloads1y agoHugging Face12BramVanroy /universal_nerThis is an exact duplicate of https://huggingface.co/datasets/universalner/universal_ner, which is not compatible with modern versions of datasets anymore because loading data via a custom script is no longer supported. All credit goes to the original creators. Original README below. Dataset Card for Universal NER Dataset Summary Universal NER (UNER) is an open, community-driven initiative aimed at creating gold-standard benchmarks for Named Entity Recognition (NER)… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/universal_ner.texttoken-classification10K<n<100K0 likes364 downloads10mo agoHugging Face13junyeong-nero /perfume-dataset Perfume Dataset Hugging Face-friendly export of crawled perfume records from parfumo_tidytuesday. Files data.parquet: primary tabular artifact for datasets.load_dataset(...) data.jsonl: JSON Lines version of the same split Suggested usage from datasets import load_dataset dataset = load_dataset("your-org/perfume-dataset", split="train") print(dataset[0]) Schema overview Each row corresponds to one crawled perfume record and preserves the raw… See the full description on the dataset page: https://huggingface.co/datasets/junyeong-nero/perfume-dataset.tabular10K<n<100K1 likes359 downloads6mo agoHugging Face14FinGPT /fingpt-ner Dataset Card for "fingpt-ner" More Information needed textn<1K5 likes351 downloads3y agoHugging Face15ParkSY /FSCM_Flood_nerfimage10K<n<100K0 likes324 downloads1y agoHugging Face16xusenlin /clue-ner CLUE-NER 命名实体识别数据集 字段说明 text: 文本 entities: 文本中包含的实体 id: 实体 id entity: 实体对应的字符串 start_offset: 实体开始位置 end_offset: 实体结束位置的下一位 label: 实体对应的开始位置 text10K<n<100K12 likes315 downloads4y agoHugging Face17leduckhai /VietMed-NER Medical Spoken Named Entity Recognition (NAACL 2025) Description: Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our knowledge, our Vietnamese real-world dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types.… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/VietMed-NER.audiotoken-classification1K<n<10K9 likes272 downloads1y agoHugging Face18community-datasets /swedish_medical_ner Dataset Card for swedish_medical_ner Dataset Summary SwedMedNER is Named Entity Recognition dataset on medical text in Swedish. It consists three subsets which are in turn derived from three different sources respectively: the Swedish Wikipedia (a.k.a. wiki), Läkartidningen (a.k.a. lt), and 1177 Vårdguiden (a.k.a. 1177). While the Swedish Wikipedia and Läkartidningen subsets in total contains over 790000 sequences with 60 characters each, the 1177 Vårdguiden subset is… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/swedish_medical_ner.texttoken-classification100K<n<1M4 likes245 downloads2y agoHugging Face19alexbrandsen /archaeo_ner_dutch Dutch Archaeology NER Dataset A selection of Dutch archaeology field reports, annotated by archaeology students from Leiden University. Labels The following labels are included: ART, artefacts ('bijl', 'pijlpunt') MAT, materials ('vuursteen', 'ijzer') PER, time periods ('Middeleeuwen', '400 v. Chr.') CON, archaeological contexts ('greppel','beerput') LOC, locations ('Amsterdam', 'Oss') SPE, species ('Betula nana', 'koe') Folds The reason I supply 5 folds is… See the full description on the dataset page: https://huggingface.co/datasets/alexbrandsen/archaeo_ner_dutch.texttoken-classification100K<n<1M1 likes238 downloads3y agoHugging Face20zjushine /nero-move-allThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "tcp.x", "tcp.y", "tcp.z", "tcp.r1", "tcp.r2", "tcp.r3", "tcp.r4", "tcp.r5", "tcp.r6", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushine/nero-move-all.tabularrobotics10K<n<100K0 likes226 downloads19d agoHugging Face21zjushine /nero-allThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "tcp.x", "tcp.y", "tcp.z", "tcp.r1", "tcp.r2", "tcp.r3", "tcp.r4", "tcp.r5", "tcp.r6", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushine/nero-all.tabularrobotics10K<n<100K0 likes223 downloads20d agoHugging Face22imvladikon /english_news_weak_ner Large Weak Labelled NER corpus Dataset Summary The dataset is generated through weak labelling of the scraped and preprocessed news corpus (bloomberg's news). so, only to research purpose. In order of the tokenization, news were splitted into sentences using nltk.PunktSentenceTokenizer (so, sometimes, tokenization might be not perfect) Usage from datasets import load_dataset articles_ds = load_dataset("imvladikon/english_news_weak_ner", "articles") # just… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/english_news_weak_ner.texttoken-classification1M<n<10M5 likes220 downloads3y agoHugging Face23NoahWeiss /nero_bag_box nero_bag_box — dual-arm bag/box manipulation (LeRobot v2.0) Teleoperated demonstrations on the AgileX NERO dual-arm rig (two 7-DoF S-R-S arms), collected with a Pico 4 WebXR teleop rig and cut into single-task clips. Ready for π0.5 / openpi-style fine-tuning. Tasks Put the bag into the box Pour the bag out of the box Features key shape notes observation.state float32[16] left joint1–7 (rad), left gripper (0–1, measured), right joint1–7… See the full description on the dataset page: https://huggingface.co/datasets/NoahWeiss/nero_bag_box.tabularrobotics10K<n<100K0 likes217 downloads1mo agoHugging Face24themohal /saraiki-ner-datasetgatedtextn<1K0 likes210 downloads5h agoHugging Face25rishitdagli /squeeze3d_rf_nerfmaeThis dataset is for the paper Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor. It contains latent representations of 3D data used for training and evaluating the Squeeze3D model. Project page Code image-to-3d1K<n<10K0 likes209 downloads1y agoHugging Face26yashpwr /resume-ner-training-data Resume NER Training Dataset This dataset contains training data for Named Entity Recognition (NER) on resume text. It's used to train the yashpwr/resume-ner-bert model. Dataset Summary Task: Token Classification (NER) Language: English Domain: Resume/CV text Size: 22855 examples Format: JSONL with BIO tagging Entity Types The dataset includes the following entity types commonly found in resumes: PERSON: Names of individuals ORG: Organizations, companies… See the full description on the dataset page: https://huggingface.co/datasets/yashpwr/resume-ner-training-data.texttoken-classification10K<n<100K1 likes209 downloads1y agoHugging Face27PassbyGrocer /msra-nertext10K<n<100K0 likes204 downloads2y agoHugging Face28projecte-aina /ancora-ca-ner Dataset Card for AnCora-Ca-NER Dataset Summary This is a dataset for Named Entity Recognition (NER) in Catalan. It adapts AnCora corpus for Machine Learning and Language Model evaluation purposes. This dataset was developed by BSC TeMU as part of the Projecte AINA, to enrich the Catalan Language Understanding Benchmark (CLUB). Supported Tasks and Leaderboards Named Entities Recognition, Language Model Languages The dataset is in Catalan… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ancora-ca-ner.text10K<n<100K3 likes186 downloads2y agoHugging Face29nwu-ctext /afrikaans_ner_corpus Dataset Card for Afrikaans Ner Corpus Dataset Summary The Afrikaans Ner Corpus is an Afrikaans dataset developed by The Centre for Text Technology (CTexT), North-West University, South Africa. The data is based on documents from the South African goverment domain and crawled from gov.za websites. It was created to support NER task for Afrikaans language. The dataset uses CoNLL shared task annotation standards. Supported Tasks and Leaderboards [More… See the full description on the dataset page: https://huggingface.co/datasets/nwu-ctext/afrikaans_ner_corpus.texttoken-classification1K<n<10K8 likes181 downloads3y agoHugging Face30ljvmiranda921 /tlunified-ner 🪐 spaCy Project: TLUnified-NER Corpus Homepage: Github Repository: Github Point of Contact: ljvmiranda@gmail.com Dataset Summary This dataset contains the annotated TLUnified corpora from Cruz and Cheng (2021). It is a curated sample of around 7,000 documents for the named entity recognition (NER) task. The majority of the corpus are news reports in Tagalog, resembling the domain of the original ConLL 2003. There are three entity types: Person (PER), Organization… See the full description on the dataset page: https://huggingface.co/datasets/ljvmiranda921/tlunified-ner.texttoken-classification1K<n<10K6 likes181 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.