CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BSC-LT /multi_lmentry Multi-LMentry This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?". Dataset Details Dataset Description Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/multi_lmentry.textquestion-answering100K<n<1M12 likes960 downloads5mo agoHugging Face02BSC-LT /IFEval_es Dataset Card for IFEval_es IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.textquestion-answeringn<1K1 likes107 downloads10mo agoHugging Face03BSC-LT /EQ-bench_es Dataset Card for EQ Bench Dataset (Spanish Version) This dataset card documents the Spanish adaptation of the EQ-Bench benchmark. The original dataset was designed to evaluate emotional reasoning in language models through dialogue-based prompts. Dataset Details Dataset Description EQ-Bench (Spanish Version) is a translated and linguistically adapted version of the original EQ-Bench dataset. Its design responds to the need to adapt the emotional detection… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EQ-bench_es.textquestion-answeringn<1K0 likes85 downloads1y agoHugging Face04BSC-LT /LexBOE Dataset Card for LexBOE Dataset summary LexBOE is a Spanish legal text classification dataset built from articles extracted from the Boletín Oficial del Estado (BOE), the official source of legislation and administrative acts in Spain. The articles included in the dataset were published between 2022 and 2024. LexBOE reflects contemporary legal-administrative language and is intended for the training and evaluation of language models on legal text classification tasks.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/LexBOE.texttext-classification10K<n<100K0 likes85 downloads6mo agoHugging Face05BSC-LT /EQ-bench_ca Dataset Card for EQ Bench Dataset (Catalan Version) This dataset card documents the Catalan adaptation of the EQ-Bench benchmark. The original dataset was designed to evaluate emotional reasoning in language models through dialogue-based prompts. Dataset Details Dataset Description EQ-Bench (Catalan Version) is a translated and linguistically adapted version of the original EQ-Bench dataset. Its design responds to the need to adapt the emotional detection… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EQ-bench_ca.textquestion-answeringn<1K0 likes70 downloads1y agoHugging Face06BSC-LT /AbSanitas Dataset Card for AbSanitas Dataset summary AbSanitas is a Spanish biomedical information retrieval dataset built from biomedical texts collected from official academic repositories and open-access sources. This dataset is designed to support the training and evaluation of encoder models on biomedical retrieval and semantic matching tasks in Spanish. Curated by: Barcelona Supercomputing Center (BSC) Funded by: ALIA Language(s) (NLP): Spanish (es) License: CC BY-NC-ND 4.0… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/AbSanitas.texttext-retrieval10K<n<100K0 likes69 downloads6mo agoHugging Face07BSC-LT /AbScientia Dataset Card for AbScientia Dataset summary AbScientia is a Spanish STEM scientific text classification dataset built from scientific abstracts collected from official academic repositories and open-access sources. The dataset focuses on Science, Technology, Engineering, and Mathematics (STEM) disciplines and reflects domain-specific scientific language in Spanish. This dataset is designed to support the training and evaluation of encoder models on STEM scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/AbScientia.texttext-classification10K<n<100K0 likes52 downloads7mo agoHugging Face08BSC-LT /SIQA_es Dataset Card for SIQA (Spanish Version) Dataset summary This dataset provides the Spanish translation and adaptation of the SIQA (Social Interaction Question Answering) validation set. The original dataset was designed to evaluate social commonsense reasoning in LLMs by presenting a collection of questions based on everyday social situations, with the ultimate goal of challenging models to infer motivations, reactions and social implications behind human actions. This… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/SIQA_es.textquestion-answering1K<n<10K0 likes42 downloads5mo agoHugging Face09open-llm-leaderboard /BSC-LT__salamandra-7b-detailsgated Dataset Card for Evaluation run of BSC-LT/salamandra-7b Dataset automatically created during the evaluation run of model BSC-LT/salamandra-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BSC-LT__salamandra-7b-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face10BSC-LT /XitXatTools Dataset Card for XitXat Tools XitXat Tools is a dataset comprising simulated Catalan call center conversations. Each conversation is annotated with structured tool calls, making it suitable for training and evaluating language models with function-calling capabilities. Dataset Details Dataset Sources Repository: XitXat Uses The dataset can be utilized for: Training language models to handle function-calling scenarios… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/XitXatTools.texttext-generationn<1K0 likes39 downloads4mo agoHugging Face11open-llm-leaderboard /BSC-LT__salamandra-7b-instruct-detailsgated Dataset Card for Evaluation run of BSC-LT/salamandra-7b-instruct Dataset automatically created during the evaluation run of model BSC-LT/salamandra-7b-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BSC-LT__salamandra-7b-instruct-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face12BSC-LT /cabreu_dolly_summarizationtext1K<n<10K0 likes37 downloads3y agoHugging Face13BSC-LT /InstrucatQA Dataset Card for Dataset Name Instructional dataset to finetune models used for RAG applications Dataset Details Dataset Description This dataset is a merge from QA instructions from InstruCAT (ca), SQUAC (es), SQUAD (en), plus generalists CA and ES MENTOR datasets to provide a cognitive background for generating responses. Contains splits of 66139 (train) and 11674 (validation) instructions Curated by: [More Information Needed] Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/InstrucatQA.textquestion-answering10K<n<100K0 likes37 downloads3y agoHugging Face14BSC-LT /MULTI_corpus Dataset Card for MULTI-Corpus Dataset Summary This corpus was compiled as part of the TRAIN project (Traducción Automática para la Inclusión, Automatic Translation for Inclusion), funded by MCIN/AEI and ERDF. It aggregates 15,191,441 sentence-level entries covering four extremely low-resource languages: Tamazight/Amazigh (ZGH), Pashto (PS), Wolof (WO), and Romani (ROM), paired with one or more high-resource counterparts (English, Spanish, French), plus monolingual data… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/MULTI_corpus.texttranslation100K<n<1M0 likes25 downloads4mo agoHugging Face15BSC-LT /aguila7b-private-inferencegated Aguila7b Private Inference Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/aguila7b-private-inference.textn<1K0 likes2 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.