datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM-Chinese-Textual-Disambiguation
Chinese Textual Ambiguity Dataset
This dataset is the accompanying dataset for the paper:
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
Paper (arXiv): https://arxiv.org/abs/2507.23121
Project repository: https://github.com/ictup/LLM-Chinese-Textual-Disambiguation
Dataset Summary
This release contains 925 Chinese textual ambiguity records collected and annotated for research on ambiguity detection, ambiguity understanding… See the full description on the dataset page: https://huggingface.co/datasets/pip1237/LLM-Chinese-Textual-Disambiguation.task775_pawsx_chinese_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task775_pawsx_chinese_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task775_pawsx_chinese_text_modification.
