datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
top-hits-spotifyxnli-eu
Dataset Card for XNLIeu
XNLIeu is an extension of XNLI translated from English to Basque. It has been designed as a cross-lingual dataset for the Natural Language Inference task, a text-classification task that consists on classifying pairs of sentences, a premise and a hypothesis, according to their semantic relation out of three possible labels: entailment, contradiction and neutral.
Dataset Details
Dataset Description
XNLI is a popular Natural… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/xnli-eu.CROQ
🌍🏺 CROQ: Culture-Related Open Questions
A multilingual benchmark for evaluating cultural and regional biases in large language models through open-ended cultural questions.
CROQ (Culture-Related Open Questions) is a multilingual dataset designed to uncover cultural and regional biases in large language models (LLMs). Unlike traditional cultural benchmarks based on multiple-choice or factual questions, CROQ focuses on open-ended cultural questions that have no single correct… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CROQ.casimedicos-arg
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
CasiMedicos-Arg is, to the best of our knowledge, the first
multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are
enriched with a natural language explanation written by doctors.
The casimedicos-exp have been manually annotated with
argument components (i.e., premise, claim) and argument… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-arg.social-behavior-emotionsCONAN-EUSContent Warning: This dataset contains examples of offensive language that do not reflect the authors’ views
CONAN-EUS: Basque and Spanish Parallel Counter Narratives Dataset
CONAN-EUS was created by professionally translating all 6654 English HS-CN pairs of the original CONAN dataset into
Basque and Spanish. For experimentation we generated train, validation and test splits in a way that no HS-CN pairs occurred across them.
CONAN-EUS Splits
Total HS-CN… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/CONAN-EUS.spotify-hit-prediction-analysis
Your browser does not support the video tag.
🎵 Spotify Hit Prediction - Exploratory Data Analysis (EDA)
Project Overview
This project analyzes audio features from Spotify to predict track popularity. Using a sample of 2,000 tracks, I explored how technical attributes like energy and danceability relate to a song's success.
🔍 Research Questions & Insights
I addressed several key questions during the EDA:
Is the data balanced? I analyzed the ratio of… See the full description on the dataset page: https://huggingface.co/datasets/Ohad777/spotify-hit-prediction-analysis.Parallel-XNLIvarBrief dataset description:
Native: the native partition of XNLIeu (Heredia et al., 2024) adapted into three Basque dialects.
Test: the test partition of XNLI adapted into three Basque dialects.
All_dialects _together is a train/dev/test split that includes both native and test instances, stratified according to dialects.
basqueparl
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions
This repository contains BasqueParl, a bilingual corpus for political discourse analysis. It covers transcriptions from the Parliament of
the Basque Autonomous Community for eight years and two legislative terms (2012-2020), and its main characteristic is the presence of Basque-Spanish
code-switching speeches.
📖 Paper: BasqueParl A Bilingual Corpus of Basque Parliamentary Transcriptions In LREC 2022.… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/basqueparl.EN2CSWork in progress.
gemini_testEuskanolDS
Dataset Card for EuskañolDS
EuskañolDS is a naturally sourced corpus for Basque-Spanish code-switching, created by filtering publicly available corpora in Basque and Spanish.
Dataset Details
Dataset Description
Code-switching (CS) remains a significant challenge in Natural Language Processing (NLP), mainly due a lack of relevant data. In the context of the contact between the Basque and Spanish languages in the north of the Iberian Peninsula, CS frequently… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/EuskanolDS.flores_plus_gender
FLORES+Gender
This dataset builds on the FLORES+ benchmark, developed by Meta to assess machine translation (MT) systems for low-resource languages. FLORES+Gender is designed to assess gender bias in MT. While the typical approach examines bias by translating from a genderless language into a gendered one, this dataset follows the methodology of Costa-jussà et al. (2023) and reverses the direction to analyse whether translation quality is affected by the predominant grammatical… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/flores_plus_gender.generatedpredictive_maintenance-dataset
Predictive Maintenance Dataset
Description
This dataset contains engine sensor readings used to predict engine health.
Features
Engine RPM
Lub Oil Pressure
Fuel Pressure
Coolant Pressure
Lub Oil Temperature
Coolant Temperature
Target
Engine Condition
0 = Healthy
1 = Maintenance Required
Number of Records
19535
Machine Learning Task
Binary Classification
Source
PGP Capstone Project… See the full description on the dataset page: https://huggingface.co/datasets/hiteshsharma/predictive_maintenance-dataset.VStarBench_hit_1QA_Filmfood_analysishate-speech-en{
"label": {
0: "normal",
1: "offensive",
2: "hateful",
3: "abusive",
4: "fearful",
5: "disrespectful",
99: "unknown"
},
"tweet": <string>
}
ALIA_scientic_articles
Dataset Description
This dataset contains a high-quality, real (non-synthetic) parallel corpus in the scientific and academic domain, covering Basque (eu), Spanish (es), and English (en).
Unlike synthetically generated datasets, this corpus is built from genuine, human-authored trilingual abstracts extracted from ADDI (Archivo Digital de Docencia e Investigación), the institutional digital repository of the University of the Basque Country (UPV/EHU). It is specifically designed… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ALIA_scientic_articles.XNLIvar
XNLIvar: Basque and Spanish variation-inclusive NLI
This repository contains the data and code used in the paper Lost in Variation? Evaluating NLI Performance in Basque and Spanish Geographical Variants
This paper evaluates the capacity of current language technologies to understand Basque and Spanish language varieties. We use Natural Language Inferenc (NLI) as a pivot task and introduce a novel, manually-curated parallel dataset in Basque and Spanish and their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/XNLIvar.film_qa_pairs_datasetQA_AVOhate-speech-frVir-Pat-2024-Intents
VirPat-2024
In this repository, you will find the corpus of Intents used in the following paper: https://aclanthology.org/2024.lrec-main.182/
Introduction
Virtual patients (VPs) have emerged as potent educational tools in medical training and healthcare simulation. They offer medical students the opportunity to simulate authentic clinical consultations, allowing them to practice a wide range of scenarios and gain valuable experience before engaging with real patients… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Vir-Pat-2024-Intents.Film_QA_Pairshate-speech-fr-v2NCDB_colisionhitl-gold-finalhitobot
