CoolFace
Datasetpublic

HiTZ/PAWS-eu

Dataset Card for PAWS-eu Point of Contact: hitz@ehu.eus Dataset Description Dataset Summary PAWS-eu is the professional translation to Basque of the PAWS dataset (Zhang et al., 2019), in the spirit of the PAWS-X effort (Yang et al., 2019). PAWS consist of sentence pairs that have high lexical overlap but that may or may not be paraphrases. Languages eu-ES Dataset Structure Data Fields id (str): A… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/PAWS-eu.

sourceHugging Faceotherupdated 2y agoView on Hugging Face
0likes61downloads
Dataset Card

Dataset Card for PAWS-eu

Dataset Description

Dataset Summary

PAWS-eu is the professional translation to Basque of the PAWS dataset (Zhang et al., 2019), in the spirit of the PAWS-X effort (Yang et al., 2019). PAWS consist of sentence pairs that have high lexical overlap but that may or may not be paraphrases.

Languages

  • eu-ES

Dataset Structure

Data Fields

  • id (str): A unique id for each pair.
  • sentence1 (str): The first sentence.
  • sentence2 (str): The second sentence.
  • label (int): (noisy) label for each pair; 0 indicates that the pair has different meaning, while 1 indicates the pair is a paraphrase.

Data Splits

nametest
default2000

Dataset Creation

This dataset is a professional translation of the English PAWS dataset into Basque, commissioned by HiTZ (UPV/EHU) within the ILENIA project. For more information on how PAWS was created, please refer to their articles (see above).

Additional Information

This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project ILENIA with reference 2022/TL22/00215335.

Licensing Information

Original PAWS-X License:

The dataset may be freely used for any purpose, although acknowledgement of Google LLC ("Google") as the data source would be appreciated. The dataset is provided "AS IS" without any warranty, express or implied. Google disclaims all liability for any damages, direct or indirect, resulting from the use of the dataset.

Citation Information

@inproceedings{baucells-etal-2025-iberobench,
    title = "{I}bero{B}ench: A Benchmark for {LLM} Evaluation in {I}berian Languages",
    author = "Baucells, Irene  and
      Aula-Blasco, Javier  and
      de-Dios-Flores, Iria  and
      Paniagua Su{\'a}rez, Silvia  and
      Perez, Naiara  and
      Salles, Anna  and
      Sotelo Docio, Susana  and
      Falc{\~a}o, J{\'u}lia  and
      Saiz, Jose Javier  and
      Sepulveda Torres, Robiert  and
      Barnes, Jeremy  and
      Gamallo, Pablo  and
      Gonzalez-Agirre, Aitor  and
      Rigau, German  and
      Villegas, Marta",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.699/",
    pages = "10491--10519",
}