CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sapinsapin /filipinospeechcorpus Filipino Speech Corpus (FSC) Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet. 313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and hand/machine transcribed with Transcriber. This repo repackages the original .wav + .trs volumes as segment-level Parquet with inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.audioautomatic-speech-recognition100K<n<1M3 likes650 downloads1mo agoHugging Face02sapinsapin /pld Philippine Language Dataset (PLD) Ten Philippine languages, 980 speakers, 448 hours of prompted speech — one of the largest multilingual Philippine speech collections available as Parquet. 334,268 utterances · 448.2 hours · 980 speakers · 10 languages · 16kHz mono ▶ Try the models in your browser — transcribe, synthesize, or convert a voice in any of the ten languages, from your microphone or the preloaded clips. Collected by the University of the Philippines Diliman… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/pld.audioautomatic-speech-recognition100K<n<1M0 likes570 downloads2d agoHugging Face03sapienzanlp /dromedario-3-sft-dataset 🐪 Dataset Card for Dromedario 3 📋 Dataset Summary Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.texttext-generation100K<n<1M9 likes502 downloads8d agoHugging Face04sapinsapin /halo-hil halo-hil Web text in hil, re-filtered by language and prepared for pretraining. What changed, and why it had to The earlier version of this dataset was labelled hil by the crawler's own language detection, and that label was never verified. An audit on 2026-09-22 found that most of it was not hil: over a random sample of 1,499 sentences, GlotLID v3 called 44 % English, 22 % Filipino/Tagalog and only 12 % Hiligaynon — much of the corpus was Tagalog news copy and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.tabulartext-generationn<1K0 likes357 downloads2d agoHugging Face05sapienzanlp /wic Word in Context (WIC) Original Paper: https://wic-ita.github.io/ This dataset comes from EVALITA-2023. Word in Context task consists of establishing if a word w occurring in two different sentences s1 and s2 has the same meaning or not. We repropose this task to test generative LLMs defining a specific prompting strategy comparing the perplexities of possible continuations to understand the models' capabilities. Example Here you can see the structure of the single… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/wic.tabular1K<n<10K0 likes286 downloads2y agoHugging Face06viridono /CF-MS_Homo_sapiens_PPI CF-MS Elution Profile PPI Dataset Proteins typically function as part of larger complexes, and co-fractionation mass spectrometry (CF-MS) identifies these complexes by tracking which proteins "co-elute" — separate into the same fractions — during chromatography, since interacting proteins show highly correlated abundance patterns across fractions. These correlations are conventionally scored with a linear metric (Pearson correlation), but non-linear relationships in the elution… See the full description on the dataset page: https://huggingface.co/datasets/viridono/CF-MS_Homo_sapiens_PPI.text10M<n<100M2 likes273 downloads13d agoHugging Face07songlab /gpn-msa-sapiens-dataset Training windows for GPN-MSA-Sapiens For more information check out our paper and repository. Path in Snakemake: results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001 tabular1M<n<10M0 likes241 downloads2y agoHugging Face08sapienzanlp /INDAQA_CALAMITA Dataset Card for INDAQA 2 INDAQA 2 (CALAMITA update) is a large-scale Italian reading-comprehension and question-answering benchmark built from classic narrative works. The dataset is designed to support research in Italian NLP, reading comprehension, information retrieval, and language model evaluation on medium- and long-context narratives and it is released as part of the CALAMITA 2026 edition. Dataset Details Dataset Description INDAQA 2 is a… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/INDAQA_CALAMITA.textn<1K2 likes71 downloads7mo agoHugging Face09sapienzanlp-course-materials /hw-mnlp-2026 Dataset for Multilingual Natural Language Processing (MNLP) Homeworks This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course. Homework 1 - Semantic Search In the first homework, you are asked to build semantic search systems. You must only use the following variables: query: A single question in natural language. query_id: The question (query) identifier. candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.tabularsentence-similarity10K<n<100K0 likes71 downloads6mo agoHugging Face10sapienzanlp /ITALIC-gen Dataset Card for ITALICGEN ITALICGEN is an adaptation of ITALIC (a Multiple-choice QA (MCQA) benchmark focused on the Italian culture) to a generative, Open-ended (OE) setting.Note: The sample in the figure is a direct translation; the original questions are in Italian. Dataset Details Dataset Description ITALICGEN is entirely based on ITALIC; for in-depth details, refer to the original publication (Seveso et al., 2025, ITALIC: An Italian… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ITALIC-gen.textquestion-answering1K<n<10K5 likes64 downloads10mo agoHugging Face11sapienzanlp /quandho QUANDHO: QUestion ANswering Data for italian HistOry Original Paper: https://aclanthology.org/L16-1069.pdf QUANDHO (QUestion ANswering Data for italian HistOry) is an Italian question answering dataset created to cover the history of Italy in the first half of the XX century. Starting from QUANDHO we defined a Multi-choice QA dataset, with a correct answer and four different distractors. Data and Distractors Generation We relied on the original data, to create this… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/quandho.text1K<n<10K1 likes61 downloads2y agoHugging Face12sapienzanlp /prelearn Prerequisite RElation LEARNing (PRELEARN) Original Paper: https://ceur-ws.org/Vol-2765/paper164.pdf This dataset contains a collection of binary-labelled concept pairs (A,B) extracted from textbooks on four domains: data mining, geometry, physics and precalculus. Then, domain experts were asked to manually annotate if pairs of concepts showed a prerequisite relation or not, therefore the dataset consists of both positive and negative concept pairs. We obtained the data from the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/prelearn.text1K<n<10K1 likes60 downloads2y agoHugging Face13sapienzanlp /ReTraceQA Dataset Card for ReTraceQA Dataset Summary ReTraceQA is a dataset designed to evaluate the reasoning traces of Small Language Models (SLMs) on commonsense reasoning tasks. It includes model-generated traces across four benchmark datasets: CommonsenseQA, OpenBookQA, QASC, and StrategyQA. During the construction of ReTraceQA, only correct instances from the original benchmarks were retained, and erroneous instances were manually removed to ensure data quality. Each item in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ReTraceQA.tabular1K<n<10K1 likes56 downloads6mo agoHugging Face14sapienzanlp /MMLU-Adversarial Dataset Card for MMLU-Adversarial Dataset Summary MMLU-Adversarial is a diagnostic dataset designed to evaluate the ability of current LLM-based answer extraction techniques to detect instances in which the model produces invalid answers due to hallucinated or flawed reasoning. Each instance in the dataset includes a reasoning chain that undermines the validity of the final selected answer, and as such, should be labeled as invalid (e.g., [No Valid Answer]). The flawed… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/MMLU-Adversarial.text1K<n<10K1 likes51 downloads1y agoHugging Face15sapinsapin /BantayWika BantayWika A FineWeb-compatible pretraining text corpus for Philippine languages, derived from the Bantay-Wika corpus collected by the University of the Philippines Sentro ng Wikang Filipino (UP-SWF) and the UP Digital Signal Processing (DSP) Laboratory. The Bantay-Wika (Language Watch) project was started in 1994 by UP-SWF to track how the Philippine national language is used and develops, particularly in Philippine media. The first phase (1994–2004) involved manual collection and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/BantayWika.texttext-generation10K<n<100K0 likes46 downloads7mo agoHugging Face16sapienzanlp /ami Automatic Misogyny Identification (AMI) Original Paper: https://amievalita2020.github.io Task presented at EVALITA-2020 This task consists of tweet classification, specifically, categorization of the level of misogyny in a given text. We taken both subtasks, raw_dataset uploaded as Behaviour (3 class classification) and synthetic uploaded as Synth (2 class classification). Example Here you can see the structure of the single sample in the present dataset.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ami.text10K<n<100K0 likes38 downloads2y agoHugging Face17sapinsapin /halo-tgl halo-tgl Dataset Summary halo-tgl is a web-scraped tgl text corpus assembled for LLM pre-training. It contains documents from news sites, blogs, academic journals, and other web sources. Cleaning Pipeline The raw text column contains web-scraped content with significant noise. A cleaning pipeline produces the text_cleaned column by: Dropping navigation menus, markdown tables, bare URLs, image markdown Removing WordPress, Blogger, Scribd, and SlideShare… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-tgl.text1K<n<10K0 likes37 downloads6mo agoHugging Face18sapienzanlp /pretens Presupposed Taxonomies: Evaluating Neural Network Semantics (PreTENS) Original Paper: https://aclanthology.org/2022.semeval-1.29.pdf This dataset comes from SemEVAL-2022 shared tasks. The PreTENS task aims at focusing on semantic competence with specific attention on the evaluation of language models with respect to the recognition of appropriate taxonomic relations between two nominal arguments. We collected the Italian part of the original dataset, and more specifically only the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/pretens.text10K<n<100K1 likes35 downloads2y agoHugging Face19sapienzanlp /nermud Named-Entities Recognition on Multi-Domain Documents (NERMUD) Original paper: https://iris.unitn.it/retrieve/d833b9e4-e997-4ee4-b6aa-f5144a85f708/paper42.pdf NERMuD is a task presented at EVALITA 2023 consisting in the extraction and classification of named-entities in a document, such as persons, organizations, and locations. This dataset comes as a word level classification setting, we decided to reframe the task to be prompted to a generative LLMs as a multiclass classification… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/nermud.text10K<n<100K0 likes35 downloads2y agoHugging Face20sapienzanlp /ghigliottinai ghigliottinAI MCQA References: https://ghigliottin-ai.github.io/ https://nlp4fun.github.io/ Starting from two different EVALITA tasks, nlp4fun (EVALITA 2018) and ghigliottin-AI (EVALITA 2020), we collected cc. 600 different games extracted from TV show and from BOARDGAME of "L'Eredità". "La Ghigliottina" is a complex game, to be solved, it needs a very large comprehension of the italian cultural knowledge. It consists in: given five different, uncorrelated words, the solution is a… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ghigliottinai.textn<1K0 likes34 downloads2y agoHugging Face21sapinsapin /halo-bcl halo-bcl Dataset Summary halo-bcl is a web-scraped bcl text corpus assembled for LLM pre-training. It contains documents from news sites, blogs, academic journals, and other web sources. Cleaning Pipeline The raw text column contains web-scraped content with significant noise. A cleaning pipeline produces the text_cleaned column by: Dropping navigation menus, markdown tables, bare URLs, image markdown Removing WordPress, Blogger, Scribd, and SlideShare… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-bcl.text1K<n<10K0 likes34 downloads6mo agoHugging Face22sapienzanlp /indaqa INDAQA - Italian Narrative Dataset for Long-document Question-Answering INDAQA is the first Italian question-answering dataset specifically designed for long-context Italian narrative texts. The dataset contains 362 documents paired with reading comprehension questions and reference answers based on Italian literary works sourced from Wikisource. Questions and answers were automatically generated using Gemini and subsequently underwent both automatic filtering and… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/indaqa.tabularquestion-answeringn<1K3 likes33 downloads9mo agoHugging Face23sapienzanlp /BBH_italian BBH - Italian (IT) This dataset is an Italian translation of the BBH dataset. BBH Bench dataset and consists of 23 tasks that are particularly hard for current generation of language models. The dataset is called Big Bench Hard. Boolean Expressions: Evaluate the truth value of a random Boolean expression consisting of Boolean constants (True, False) and basic Boolean operators (and, or, not). Causal Judgment: Given a short story (involving moral, intentional, or counterfactual… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/BBH_italian.text1K<n<10K0 likes32 downloads10mo agoHugging Face24sapienzanlp /Concept-10kimage10K<n<100K2 likes30 downloads11mo agoHugging Face25sapienzanlp /it-Magpie-Llama-3.1-Pro-300K-Filtered-easy Dataset Card: Magpie-Llama-3.1-Pro-300K-Filtered (Italian Translation) Dataset Summary This dataset is the Italian translation of the Magpie-Llama-3.1-Pro-300K-Filtered dataset, designed specifically for instruction tuning in Italian. We translated the instances whose 'difficulty' value is 'easy'. The translation was carried out using TowerInstruct-Mistral-7B-v0.2. Languages: Italian (translated from English) Purpose: Instruction tuning in Italian Train Size: 59070 New… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/it-Magpie-Llama-3.1-Pro-300K-Filtered-easy.tabular10K<n<100K2 likes29 downloads2y agoHugging Face26sapienzanlp /MATH_hard_italian MATH - Italian (IT) This dataset is an Italian translation of the MATH dataset. MATH-hard dataset consists of hard problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more, across different mathematical domains. Dataset Details The dataset consists of open-ended questions, where each question is associated with a gold solution which contains the correct answer inside \boxed tag. Number of samples: 1320. Languages This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/MATH_hard_italian.text1K<n<10K0 likes29 downloads10mo agoHugging Face27sapienzanlp /discotex Assessing DIScourse COherence in Italian TEXts (DISCOTEX) Original Paper: https://sites.google.com/view/discotex/ Task presented at EVALITA-2023 The original task is about modelling discourse coherence for Italian texts. We focalized only on the first sub-task: Last Sentence Classification: given a short paragraph, and an individual sentence (target), the model will be asked to classify whether the target follows or not the paragraph. To assess the capability of a Language Model to… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/discotex.text10K<n<100K0 likes28 downloads2y agoHugging Face28sapirharary /PrefixNLItext100K<n<1M0 likes27 downloads11mo agoHugging Face29sapirharary /RAGTruthPrefixestext100K<n<1M0 likes26 downloads11mo agoHugging Face30sapienzanlp /MUSR_italian MuSR - Italian (IT) This dataset is an Italian translation of the MUSR dataset. MuSR dataset require multi-step reasoning with commonsense to answer questions aligned with narrative texts. Dataset Details The dataset consists of multiple-choice questions and narrative texts, where each question is associated with a set of answer choices. The task is to choose the correct answer choice based on the context provided in the texts. The dataset includes three QAs domains:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/MUSR_italian.textn<1K0 likes25 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.