datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
named_entity_recognition_document_contextcs_czech-named-entity-corpus_2.0
Dataset Card for Czech Named Entity Corpus 2.0
Dataset Description
The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test)
Dataset Features
Each sample contains:
text: source sentence
entities: list of selected entities. Each entity contains:
category_id: string identifier of the entity category
category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.task960_ancora-ca-ner_named_entity_recognition
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task960_ancora-ca-ner_named_entity_recognition
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task960_ancora-ca-ner_named_entity_recognition.flan_combined_task1544_conll2002_named_entity_recognition_answer_generationamharic-named-entity-recognition
Amharic Named Entity Recognition Dataset
This dataset can be used to train models for Named Entity Recognition.
Dataset Source
https://github.com/uhh-lt/ethiopicmodels/blob/master/am/data/NER/train.txt
Finetuned Models
The following transformer models were finetuned using this dataset. The reported precision, recall, and f1 metrics are macro averages.
Model
Size (# params)
Precision
Recall
F1
bert-medium-amharic
40.5M
0.64
0.73
0.68… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-named-entity-recognition.named_entity_recognitionnamed-entity-recognitionNamed_entity_recognitiontask1544_conll2002_named_entity_recognition_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1544_conll2002_named_entity_recognition_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1544_conll2002_named_entity_recognition_answer_generation.named_entity_recognitionNamedEntityRecognition_SLUE-VoxPopuliNamedEntityLocalization_SLUE-VoxPopuliKorean-Math-Named-Entity-Translated
목적
더이상 Hermitian Matrix를 "허미트 행렬", "허미트리안 행렬", "헤르미트 행렬" 이라고 부르지 않는 모델을 만들기 위해
특징
수학 관련 영어 웹 데이터를 기반으로 수학자, 지명, 수학자 이름/지명이 포함된 수학 공식이나 이론명 등의 (비교적)올바른 한국어 명칭을 생성
다만 웹 데이터 기반이므로, 수학자 이름 외에도 "답변자 ID"와 같은 일반인 이름, 프로그램 코드의 함수명으로 추측되는 텍스트도 Named Entity로 포함되어 있음
성과 이름이 잘못 매칭된 경우도 있음
Sonnet 3.7 기반이므로 잘못된 번역도 존재함
예시: Zu Chongzhi - 주 충즈 (올바른 번역: 조충지 -> Sonnet은 대부분의 고대 중국인 이름을 중국어 발음 그대로 번역함)
기반 데이터 셋
ajibawa-2023/Maths-College 에서 instruction 부분만 추출
수학 관련 웹 데이터… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/Korean-Math-Named-Entity-Translated.flan_combined_task960_ancora-ca-ner_named_entity_recognition
