datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
named_entity_recognition_document_contexttask960_ancora-ca-ner_named_entity_recognition
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task960_ancora-ca-ner_named_entity_recognition
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task960_ancora-ca-ner_named_entity_recognition.flan_combined_task1544_conll2002_named_entity_recognition_answer_generationamharic-named-entity-recognition
Amharic Named Entity Recognition Dataset
This dataset can be used to train models for Named Entity Recognition.
Dataset Source
https://github.com/uhh-lt/ethiopicmodels/blob/master/am/data/NER/train.txt
Finetuned Models
The following transformer models were finetuned using this dataset. The reported precision, recall, and f1 metrics are macro averages.
Model
Size (# params)
Precision
Recall
F1
bert-medium-amharic
40.5M
0.64
0.73
0.68… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-named-entity-recognition.entity_recognition
Data for the NERMuD shared task (Evalita 2023)
This data is the one used for the NERMuD shared task organized
at Evalita 2023.
The dataset contains the Wikinews, fiction, and De Gasperi subsets of KIND, where test data is used for development.
Content of the dataset
Split
Sentences
wn_train
10,912
wn_dev
2,594
wn_test
2,088
fic_train11,423
fic_dev
1,051
fic_test
1,517
adg_train
5,147
adg_dev
1,122
adg_test
521
Set
Sentences… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/entity_recognition.sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600Dataset for sarvam's entity normalisation task. More detailed information can be found here, in the main model repo: Hugging Face
Detailed Report (Writeup): Google Drive
It also has a gguf variant, with certain additional gguf based innstructions: Hugging Face
Model inference script can be found here: Colab
Model predictions can be found in this dataset and both the repo files. named as:
eval_data_001_predictions.csv and eval_data_001_predictions_excel.csv.
train_data_001_predictions.csvand… See the full description on the dataset page: https://huggingface.co/datasets/Tasmay-Tib/sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600.NVR-Entity-Recognition-Experiment
NVR Entity Recognition Experiment
Overview
This repository contains a training dataset designed for entity recognition in Network Video Recorder (NVR) applications, specifically focused on newborn safety monitoring. The dataset uses a stuffed animal as a privacy-conscious substitute for actual newborn footage, enabling the development of computer vision models that can identify critical safety scenarios in nursery environments.
Purpose
The primary goal of this… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/NVR-Entity-Recognition-Experiment.named_entity_recognitionnamed-entity-recognitionnlp.6.named_entity_recognition
Dataset Card for "nlp.6.named_entity_recognition"
More Information needed
task1544_conll2002_named_entity_recognition_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1544_conll2002_named_entity_recognition_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1544_conll2002_named_entity_recognition_answer_generation.Named_entity_recognitionrequirements-entity-recognitionlaw_entity_recognition
Dataset Card for Dataset Name
The dataset transforms complex legal passages into structured outputs, detailing entities, their interrelationships, and claims, providing a foundation for a legal knowledge graph to facilitate advanced analysis and applications.
Dataset Details
Dataset Description
The dataset in question is a specialized collection designed for legal text analysis, where each input is a passage of legal text—ranging from case law to statutory… See the full description on the dataset page: https://huggingface.co/datasets/rubenamtz0/law_entity_recognition.named_entity_recognitionAPIS_OEBL__Named_Entity_RecognitionJSON file of 6,941 sentences of historical biographies, annotated with "PER" (Person), "ORG" (Organisation), "LOC" (Location).
source
The original data was extracted from the Austrian Biographical Lexicon (ÖBL) in the context of the Austrian Prosopographical Information System (APIS) project.
From there, samples were randomly pulled and annotated for Named Entity Recognition tasks, which form this dataset.
The texts concern numerous smaller biographies in the time period between… See the full description on the dataset page: https://huggingface.co/datasets/SteffRhes/APIS_OEBL__Named_Entity_Recognition.german_legal_entity_recognitionGreek_name_entity_recognitionSlovenian_name_entity_recognitionnamed_entity_recognitionEnglish_name_entity_recognition1_1_legal_entity_recognition_promptshallucinated_entity_recognition_HER_datasetTest
Italian_name_entity_recognitionbible_entity_recognitionproject-EntityRecognition-f9a1b253-0e97-4ecb-be54-fa89fa1f62caproject-EntityRecognition-9c28eaf0-236b-450b-a255-c54b06410864project-EntityRecognition-5cde751a-826c-4653-a162-13de02f805adPolish_name_entity_recognitionflan_combined_task960_ancora-ca-ner_named_entity_recognition
