datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NEREL_bench
NEREL-bench
Summary
NEREL-bench is a benchmark dataset designed to evaluate the capabilities of Large Language Models (LLMs) in performing knowledge graph construction tasks on Russian-language texts. The dataset focuses on three fundamental tasks essential for building knowledge graphs: named entity recognition, relation extraction between entities, and generation of contextual definitions for both entities and relations. These tasks are critical for evaluating whether… See the full description on the dataset page: https://huggingface.co/datasets/bond005/NEREL_bench.NEREL_instruct
NEREL-instruct
NEREL-instruct is an instruction-based dataset derived from the NEREL corpus — a large Russian dataset annotated with nested named entities, relations, and events. The original NEREL annotations (texts + manual entity/relation markup) were converted into a structured instruction-following format using Qwen2.5-32B-Instruct. The result is a semi‑synthetic dataset designed for fine‑tuning large language models (LLMs) on a variety of information extraction tasks.
The… See the full description on the dataset page: https://huggingface.co/datasets/bond005/NEREL_instruct.
