CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HPAI-BSC /Aloe-Beta-General-Collection Aloe-Beta-Medical-Collection Collection of curated general datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including: Coding, math, data analysis, STEM, etc. Function calling Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.textquestion-answering10K<n<100K2 likes146 downloads10mo agoHugging Face02BSC-LT /salamandra-guard-dataset Salamandra Guard Dataset Dataset Description The Salamandra Guard dataset is a comprehensive multilingual safety classification corpus designed for training and evaluating content moderation systems in Catalan, Spanish. It consists of 21,335 carefully curated conversational examples annotated across a hierarchical safety taxonomy. This dataset represents a significant advancement in culturally-grounded safety data, with particular emphasis on Catalan—a language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/salamandra-guard-dataset.documenttext-generation10K<n<100K3 likes136 downloads14d agoHugging Face03BSC-LT /IFEval_es Dataset Card for IFEval_es IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.textquestion-answeringn<1K1 likes117 downloads10mo agoHugging Face04HPAI-BSC /Aloe-Beta-DPO Aloe-Beta-Medical-Collection Collection of curated DPO datasets used to align Aloe-Beta. Dataset Details Dataset Description The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data: Medical preference data: TsinghuaC3I/UltraMedical-Preference General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.textquestion-answering100K<n<1M2 likes109 downloads1y agoHugging Face05HPAI-BSC /NotSoTiny-25-12 NotSoTiny: A Large, Living Benchmark for RTL Code Generation Summary NotSoTiny is a large, structurally rich, and "living" benchmark designed to assess Large Language Models (LLMs) on the generation of context-aware RTL (Register-Transfer Level) code. Built from hundreds of real hardware designs produced by the Tiny Tapeout community, this benchmark overcomes the limitations of prior static… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/NotSoTiny-25-12.texttext-generation1K<n<10K0 likes70 downloads8mo agoHugging Face06BSC-LT /ALIA-2606-DPO-safety Dataset Card for BSC Multilingual Synthetic Safety Preferences Dataset Summary This dataset consists of synthetic safety preference data generated to align language models across five languages: Catalan, Spanish, English, Basque, and Galician. Building on the PKU-SafeRLHF and Tulu 3/Ultrafeedback methodologies for creating preference data, this dataset leverages an LLM-as-a-judge approach to automatically score and pair model responses to a massive pool of safety… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-DPO-safety.texttext-generation10K<n<100K0 likes59 downloads2mo agoHugging Face07HPAI-BSC /NotSoTiny-26-07NotSoTiny: A Large, Living Benchmark for RTL Code Generation Summary NotSoTiny is a large, structurally rich, and "living" benchmark designed to assess Large Language Models (LLMs) on the generation of context-aware RTL (Register-Transfer Level) code. Built from hundreds of real hardware designs produced by the Tiny Tapeout community, this benchmark overcomes the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/NotSoTiny-26-07.texttext-generation1K<n<10K0 likes45 downloads2mo agoHugging Face08BSC-LT /InstrucatQA Dataset Card for Dataset Name Instructional dataset to finetune models used for RAG applications Dataset Details Dataset Description This dataset is a merge from QA instructions from InstruCAT (ca), SQUAC (es), SQUAD (en), plus generalists CA and ES MENTOR datasets to provide a cognitive background for generating responses. Contains splits of 66139 (train) and 11674 (validation) instructions Curated by: [More Information Needed] Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/InstrucatQA.textquestion-answering10K<n<100K0 likes37 downloads3y agoHugging Face09BSC-LT /XitXatTools Dataset Card for XitXat Tools XitXat Tools is a dataset comprising simulated Catalan call center conversations. Each conversation is annotated with structured tool calls, making it suitable for training and evaluating language models with function-calling capabilities. Dataset Details Dataset Sources Repository: XitXat Uses The dataset can be utilized for: Training language models to handle function-calling scenarios… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/XitXatTools.texttext-generationn<1K0 likes34 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.