datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task1390_wscfixed_coreference
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1390_wscfixed_coreference
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1390_wscfixed_coreference.flan2021_coreference_raw
Flan 2021 Coreference Tasks
Project: https://github.com/google-research/FLAN/tree/main/flan/v2
Data source: DataProvenanceInitiative/flan2021_submix_original
Details
This dataset contains all coreference examples that were included in the Flan 2022 collection which were orignally included in Flan 2021.
The data is copied from the preprocessed Flan2021 dataset at DataProvenanceInitiative/flan2021_submix_original.
COREFERENCE_TASK_NAMES = {… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/flan2021_coreference_raw.task891_gap_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task891_gap_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task891_gap_coreference_resolution.task893_gap_fill_the_blank_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task893_gap_fill_the_blank_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task893_gap_fill_the_blank_coreference_resolution.task892_gap_reverse_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task892_gap_reverse_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task892_gap_reverse_coreference_resolution.id_coreference_resolutionWe built Indonesian coreference resolution that solves not only pronoun referenced to proper noun, but also proper noun to proper noun and pronoun to pronoun.
The differences with the available Indonesian coreference resolution lay on the problem scope and features.
We conducted experiments using various features (lexical and shallow syntactic features) such as appositive feature, nearest candidate feature, direct sentence feature, previous and next word feature, and a lexical feature of first person.
We also modified the method to build the training set by selecting the negative examples by cross pairing every single markable that appear between antecedent and anaphor.
Compared with two available methods to build the training set, we conducted experiments using C45 algorithm.
Using 200 news sentences, the best experiment achieved 71.6% F-Measure score.coreference-dataset-ua
Silver Ukrainian Coreference Dataset
Dataset Description
Dataset Summary
A silver coreference resolution dataset for the Ukrainian language. The dataset was generated automatically with the usage of the word alignment method from the following English dataset: https://github.com/d5555/Coreference-dataset.
The word alignment method was implemented by Andrii Kursin (aqrsn@ukr.net).
Languages
Ukrainian
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/artemkramov/coreference-dataset-ua.coreference-challenge
PI-LLM Bench: The Core Retrieval Challenge Behind MRCR
Update: Accepted to COLM 2026 (San Francisco).
AAAI 2026 Worshop Oral: LaMAS (LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems) Jan/2026 Singapole
ICML 2025 Long-Context Foundation Models Workshop Accepted.
A simple context interference evaluation.
Update: This dataset is integrated into Moonshot AI(Kimi)'s internal benchmarking framework for assessing ** tracking capacity and… See the full description on the dataset page: https://huggingface.co/datasets/giantfish-fly/coreference-challenge.dfm10-mimir-event-coreference-sft
dfm10-mimir-event-coreference-sft
Controlled event-coreference and temporal-state supervision grounded in licensed passages.
Contents
Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
Schema: chat messages, optional condition and tools, plus provenance
Shards: 2
Rows: 300,000
Category: Commonsense reasoning
Upstream material
Novel event-coreference and temporal-state synthetic tasks
Processing
Gemma 4 31B generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-mimir-event-coreference-sft.wino_x_fr_prompt_coreference
wino_x_fr_prompt_coreference
Summary
wino_x_fr_prompt_coreference is a subset of the Dataset of French Prompts (DFP).It contains 27,930 rows that can be used for a coreference task.The original data (without prompts) comes from the dataset wino_x by Emelin et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by Muennighoff et al.… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/wino_x_fr_prompt_coreference.coreference-resolutionxwinograd_fr_prompt_coreference
xwinograd_fr_prompt_coreference
Summary
xwinograd_fr_prompt_coreference is a subset of the Dataset of French Prompts (DFP).It contains 830 rows that can be used for a coreference task.The original data (without prompts) comes from the dataset xwinograd by Muennighoff where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by Muennighoff et al.… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/xwinograd_fr_prompt_coreference.flan_combined_task892_gap_reverse_coreference_resolutionflan_combined_task891_gap_coreference_resolutionflan2021-coreferenceflan_combined_task893_gap_fill_the_blank_coreference_resolutionrlvr_task1390_wscfixed_coreferencepronominal_coreference_resolutionCo-Reference-Nepalimultimodal_dialogue_coreference_resolution
멀티 모달 대화 모델을 위한 패션 지식 대화 데이터 셋 :복합대화 연구용 데이터셋 V.2 (데이터)
목적 및 소개
목적 : 텍스트로 이루어진 대화 뿐만 아니라 이미지까지 포함된 대화도 언어모델이 처리할 수 있도록하는 것이 목적.
도메인 : 이미지와 텍스트 모두에 대한 이해가 있어야 대화가 가능한 패션으로 도메인을 설정.
대화 내용 : 패션 주제를 정해 놓고 이에대해 system과 user가 서로 대화를 나누는 내용으로 구성. 주로 user가 패션에 대한 지식이나 이미지를 요청하고 system이 그에대한 답변으로 관련 지식이나 이미지를 찾아주는 형태를 가짐.
특징
의도 및 감정 라벨링 : user 및 system의 발화와 함께 발화가 가지는 의도 (intent), 감정 (sentiment)을 함께 라벨링함.
패션 속성 리스트 라벨링 : user가 어떤 종류의 패션 이미지를 원하는지 패션… See the full description on the dataset page: https://huggingface.co/datasets/KETI-NLP/multimodal_dialogue_coreference_resolution.qwen3_0.6b-rlvr_task1390_wscfixed_coreferenceniv2_coreference_raw
Natural Instructions v2 Coreference Tasks
Project: https://github.com/allenai/natural-instructions
Data source: DataProvenanceInitiative/niv2_submix_original
Details
This dataset contains all coreference examples that were included in the Flan 2022 collection which were orignally published in Super-Natural-Instructions.
The data is copied from the preprocessed Natural Instructions v2 dataset at DataProvenanceInitiative/niv2_submix_original.
These tasks are:… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/niv2_coreference_raw.Co-reference_BenchmarkKorean-CoreferenceResolutionflan2021-coreference-held-out
