CoolFace
Datasetpublic

juice-cultural-eval/JuICE

JuICE Sources Repository: https://anonymous.4open.science/r/JuICE HuggingFace: juice-cultural-eval/JuiCE About We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh)… See the full description on the dataset page: https://huggingface.co/datasets/juice-cultural-eval/JuICE.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
1likes25downloads
Dataset Card

JuICE

Sources

About

We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both English and their countries' main languages.

Data Description

  • —qa_id (key): id of query-response pair
  • —country: country
  • —language_type: english or native
  • —language: language code (en, ko, id, bn)
  • —query_text
  • —response_text
  • —error_groups
  • —error_group_id (key)
  • —paragraph_idx_list
  • —thick_categories
  • —errors
  • —error_id (key)
  • —paragraph_idx_list: index of paragraph in response_text where span is contained
  • —detector_id
  • —detector_type: human or model
  • —span
  • —explanation
  • —raw_annotation
  • —thick_category