datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task249_enhanced_wsc_pronoun_disambiguation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task249_enhanced_wsc_pronoun_disambiguation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task249_enhanced_wsc_pronoun_disambiguation.bbeh-disambiguation-qa
Reference
@article{kazemi2025big,
title={Big-bench extra hard},
author={Kazemi, Mehran and Fatemi, Bahare and Bansal, Hritik and Palowitch, John and Anastasiou, Chrysovalantis and Mehta, Sanket Vaibhav and Jain, Lalit K and Aglietti, Virginia and Jindal, Disha and Chen, Peter and others},
journal={arXiv preprint arXiv:2502.19187},
year={2025}
}
LLM-Chinese-Textual-Disambiguation
Chinese Textual Ambiguity Dataset
This dataset is the accompanying dataset for the paper:
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
Paper (arXiv): https://arxiv.org/abs/2507.23121
Project repository: https://github.com/ictup/LLM-Chinese-Textual-Disambiguation
Dataset Summary
This release contains 925 Chinese textual ambiguity records collected and annotated for research on ambiguity detection, ambiguity understanding… See the full description on the dataset page: https://huggingface.co/datasets/pip1237/LLM-Chinese-Textual-Disambiguation.Tool_Selection_Disambiguation
🇰🇿 Kazakh Tool Selection and Disambiguation Dataset
Dataset Summary
Kazakh Tool Selection and Disambiguation Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI scenarios that require choosing the most appropriate tool from multiple available options.
The dataset focuses on tool-selection reasoning, where the assistant must understand the user’s intent, compare available tools, avoid unnecessary… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Tool_Selection_Disambiguation.
