kenhktsui/squad_v2_factuality_v1
squad_v2_factuality_v1 This dataset is derived from "squad_v2" training "context" with the following steps. NER is run to extract entities. Lexicon of person's name, date, organisation name and location are collected. 20% of the time, one of the text attribute (person's name, date, organisation name and location) is randomly replaced. For consistency of context, all other place with the same name is also replaced. Purpose of the Dataset The purpose of this… See the full description on the dataset page: https://huggingface.co/datasets/kenhktsui/squad_v2_factuality_v1.
squadv2factuality_v1
This dataset is derived from "squad_v2" training "context" with the following steps.
- NER is run to extract entities.
- Lexicon of person's name, date, organisation name and location are collected.
- 20% of the time, one of the text attribute (person's name, date, organisation name and location) is randomly replaced. For consistency of context, all other place with the same name is also replaced.
# Purpose of the Dataset The purpose of this dataset is to assess if a language model could detect factuality.
