juice-cultural-eval/JuICE
JuICE Sources Repository: https://anonymous.4open.science/r/JuICE HuggingFace: juice-cultural-eval/JuiCE About We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh)… See the full description on the dataset page: https://huggingface.co/datasets/juice-cultural-eval/JuICE.
JuICE
Sources
- Repository: https://anonymous.4open.science/r/JuICE
- HuggingFace: juice-cultural-eval/JuiCE
About
We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both English and their countries' main languages.
Data Description
qa_id(key): id of query-response paircountry: countrylanguage_type:englishornativelanguage: language code (en,ko,id,bn)query_textresponse_texterror_groupserror_group_id(key)paragraph_idx_listthick_categorieserrorserror_id(key)paragraph_idx_list: index of paragraph inresponse_textwherespanis containeddetector_iddetector_type:humanormodelspanexplanationraw_annotationthick_category
