nerel
Datasets
All datasets matching “nerel”nerelNEREL
NEREL dataset
Dataset Description
NEREL dataset (https://doi.org/10.48550/arXiv.2108.13112) is
a Russian dataset for named entity recognition and relation extraction.
NEREL is significantly larger than existing Russian datasets:
to date it contains 56K annotated named entities and 39K annotated relations.
Its important difference from previous datasets is annotation of nested named
entities, as well as relations within nested entities and at the discourse
level. NEREL can… See the full description on the dataset page: https://huggingface.co/datasets/iluvvatar/NEREL.NEREL_bench
NEREL-bench
Summary
NEREL-bench is a benchmark dataset designed to evaluate the capabilities of Large Language Models (LLMs) in performing knowledge graph construction tasks on Russian-language texts. The dataset focuses on three fundamental tasks essential for building knowledge graphs: named entity recognition, relation extraction between entities, and generation of contextual definitions for both entities and relations. These tasks are critical for evaluating whether… See the full description on the dataset page: https://huggingface.co/datasets/bond005/NEREL_bench.NEREL_instruct
NEREL-instruct
NEREL-instruct is an instruction-based dataset derived from the NEREL corpus — a large Russian dataset annotated with nested named entities, relations, and events. The original NEREL annotations (texts + manual entity/relation markup) were converted into a structured instruction-following format using Qwen2.5-32B-Instruct. The result is a semi‑synthetic dataset designed for fine‑tuning large language models (LLMs) on a variety of information extraction tasks.
The… See the full description on the dataset page: https://huggingface.co/datasets/bond005/NEREL_instruct.nerel_dataset
TeSla NeReL Dataset
Скуп за обучавање модела за обележавање и повезивање именованих ентитета (NER+NEL)
Преко 150.000 реченица анотираних реченица из различитих домена
Named Entity Recognition and Linking (NER+NEL) Model Training Set for Serbian
Over 150,000 annotated sentences from various domains
Editor
Milica Ikonić Nešić
@MilicaIK
Editor… See the full description on the dataset page: https://huggingface.co/datasets/te-sla/nerel_dataset.nerel_short
About DataSet
The dataset based on NEREL corpus.
For more information about original data, please visit this source
Example of preparing original data illustrated in <Prepare_original_data.ipynb>
Additional info
The dataset consist 29 entities, each of them can be as beginner part of entity "B-" as inner "I-".
Frequency for each entity:
I-AGE: 284
B-AGE: 247
B-AWARD: 285
I-AWARD: 466
B-CITY: 1080
I-CITY: 39
B-COUNTRY: 2378
I-COUNTRY: 128
B-CRIME: 214
I-CRIME: 372… See the full description on the dataset page: https://huggingface.co/datasets/surdan/nerel_short.
