CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01VillanovaAI /Multi-CoSyn-400Kgated Multi-CoSyn-400K Overview Multi-CoSyn-400K is a multilingual extension of the original CoSyn-400K dataset from AllenAI, which is in turn an extended and improved version of the PixMo-Docs dataset. The original CoSyn-400K dataset consists of question–answer pairs with reasoning on text-rich images, where both the images and the question-answer pairs with reasoning were generated using code-based rendering tools and text-only LLMs. Example of rendering tools are:… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-CoSyn-400K.imagevisual-question-answering100K<n<1M0 likes107 downloads19h agoHugging Face02VillanovaAI /multi-pixmo-capgated Multi-PixMo-Cap Overview Multi-PixMo-Cap is a multilingual extension of the original PixMo-Cap dataset from AllenAI.The original PixMo-Cap dataset was created by recording annotators speaking freely about an image for 60–90 seconds, then transforming the resulting audio transcripts into detailed captions using Claude (see the PixMo paper). Multi-PixMo-Cap follows the same multimodal concept, but all examples were re-generated from human captions using a… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-pixmo-cap.imageimage-to-text100K<n<1M2 likes75 downloads19h agoHugging Face03VillanovaAI /multi-pixmo-ask-model-anythinggated Multi-PixMo-AskModelAnything Overview Multi-PixMo-AskModelAnything is a multilingual extension of the original PixMo-AskModelAnything dataset from AllenAI, part of the PixMo series of multimodal resources. The original PixMo-AskModelAnything dataset consists of image-based question–answer pairs, where annotators authored freeform questions about an image, and answers were generated through a pipeline combining OCR output, dense captions, and a language-only LLM.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-pixmo-ask-model-anything.imagevisual-question-answering100K<n<1M1 likes66 downloads19h agoHugging Face04VillanovaAI /villanova-sft-2603gated Villanova-SFT-2603 Villanova-SFT-2603 is a large-scale, multilingual supervised fine-tuning (SFT) collection of datasets. It contains 1,711,114 instruction-response conversations spanning five European languages, covering chat, instruction following, reasoning, code, knowledge, and safety tasks. This dataset was used to train the Villanova-2B-2603 model family. All data has been processed through a rigorous curation pipeline that enforces schema normalization, hash-based… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/villanova-sft-2603.texttext-generation1M<n<10M3 likes63 downloads2mo agoHugging Face05VillanovaAI /Eurostat_Tourism_STS_Dataset_Turnover_in_Servicesgated Eurostat Tourism STS Dataset – Turnover in Services (Monthly) This repository contains data extracted from the Eurostat Short-Term Statistics (STS) domain, with a focus on: tour_sts – Tourism industries short-term indicators sts_setu_m – Turnover in services (monthly data) These datasets measure monthly turnover and sales volume indices across tourism-related industries following the NACE Rev.2 classification. Source: https://ec.europa.eu/eurostat/web/tourism/databaseLicense: CC… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Eurostat_Tourism_STS_Dataset_Turnover_in_Services.tabular10K<n<100K2 likes28 downloads10mo agoHugging Face06VillanovaAI /Multi-SciRIFFgated Multi-SciRIFF A multilingual adaptation of SciRIFF extending a filtered subset of the original English-only instruction-following scientific literature dataset to five languages with permissively licensed synthetic translations. The original SciRIFF dataset, by AllenAI, includes ~137 K instruction-following demonstrations for 54 scientific literature understanding tasks, organized with rich metadata describing domains, task families, and context. It was developed as a benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-SciRIFF.text10K<n<100K1 likes25 downloads10mo agoHugging Face07VillanovaAI /Temporal_Spatial_Tracking_Datasetgated Dataset Overview This dataset contains time-stamped spatial tracking records collected from tagged entities (e.g., wearable tags, assets, or devices) operating within a monitored environment.Each row represents a single localization event captured at a precise moment in time, including 3D position coordinates and device status information. The dataset is inherently temporal and spatial, making it suitable for trajectory reconstruction, movement analysis, and time-based behavioral… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Temporal_Spatial_Tracking_Dataset.tabular100K<n<1M0 likes24 downloads6mo agoHugging Face08VillanovaAI /Time-Series-Donationsgated Time-Series Donations Dataset Overview This repository provides a time-series dataset of donation dynamics over time.It is intended for experiments in: Time-series forecasting Trend and seasonality analysis Anomaly detection on donation flows Benchmarking classical and deep time-series models The data are organized in a tabular time-series format, with each row representing a time step and each column representing a numerical or categorical feature related to donations.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Time-Series-Donations.tabularn<1K0 likes23 downloads6mo agoHugging Face09VillanovaAI /multi-dialoguesgated Multilingual Dialogues Dataset Overview Multilingual Dialogues is a multilingual dataset of synthetically generated everyday conversations between two fictitious people. Most dialogues consist of 6 to 8 turns between the characters. Multilingual Dialogues is generated using the SODAverse pipeline from AllenAI. The original SODA dataset is generated starting from the Atomic 10x knowledge graph. Triples regarding social interactions are extracted and contextualized to get a… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-dialogues.text10K<n<100K2 likes19 downloads10mo agoHugging Face10VillanovaAI /Multi-FLAN-NIv2gated Multi-FLAN-NIv2 Overview This dataset is a multilingual subset of Natural Instructions v2 (NIv2) as included in the FLAN collection. The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information). In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-NIv2.text10K<n<100K0 likes13 downloads9mo agoHugging Face11VillanovaAI /fineweb-edu-sample-15k-metagated fineweb-edu-sample-15k-meta Overview This dataset is a metadata-enriched sample of the original FineWeb-Edu dataset. FineWeb-Edu is a large-scale web text corpus derived from Common Crawl, curated to emphasize educational value and instructional relevance. The dataset is constructed through targeted filtering and classification steps designed to surface content suitable for learning, teaching, and knowledge transfer, while preserving the diversity and scale of web-sourced… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/fineweb-edu-sample-15k-meta.tabular10K<n<100K0 likes12 downloads8mo agoHugging Face12VillanovaAI /Small-Scale-Fisheries-Tracking-Datasetgated Small-Scale Fisheries Tracking Dataset Raw GPS data from the study "Addressing gaps in small-scale fisheries: a low-cost tracking system" This dataset contains the raw GPS vessel-tracking data originally released on HuggingFace as part of a study focused on developing a low-cost monitoring solution for small-scale fisheries. It is provided here to support transparency and reproducibility of the original work. 📦 Dataset Contents File Description… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Small-Scale-Fisheries-Tracking-Dataset.tabular1K<n<10K0 likes9 downloads6mo agoHugging Face13VillanovaAI /fineweb-2-sample-60k-metagated fineweb-2-sample-60k-meta Overview This dataset is a metadata-enriched multilingual sample of the original FineWeb-2 dataset. FineWeb-2 is a large-scale, high-quality web text corpus derived from Common Crawl, built through extensive filtering, deduplication, and language identification steps to support the training of modern large language models. It emphasizes text quality, diversity, and transparency, and includes rich crawl-level metadata inherited from Common Crawl.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/fineweb-2-sample-60k-meta.tabular10K<n<100K0 likes9 downloads9mo agoHugging Face14VillanovaAI /Multi-FLAN-CoTgated Multi-FLAN-CoT Overview This dataset is a multilingual subset of the Chain-of-Thought (CoT) component of the original FLAN collection. The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information). In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-CoT.text10K<n<100K0 likes8 downloads9mo agoHugging Face15VillanovaAI /PharmaQA.ITgated PharmaQA.IT PharmaQA.IT is an Italian extractive question-answering dataset built from the Riassunti delle Caratteristiche del Prodotto (RCP), the official medicine leaflets issued by the Italian Medicines Agency (AIFA) and collected in PharmaER.IT. It contains 861 expert-validated question–answer pairs covering indications, contraindications, dosage, warnings, interactions, and pharmacological properties. The pairs were generated semi-automatically: a multimodal LLM prompted… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/PharmaQA.IT.textn<1K0 likes8 downloads18h agoHugging Face16VillanovaAI /finepdfs-sample-75k-metagated finepdfs-sample-75k-meta Overview This dataset is a metadata-enriched multilingual sample of the original FinePDFs dataset. FinePDFs is a large-scale collection of document-level texts extracted from PDF files, sourced primarily from Common Crawl. The dataset emphasizes high-quality document extraction, structural coherence, and large-scale coverage of technical, scientific, educational, and administrative content commonly distributed in PDF form. This release contains a… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/finepdfs-sample-75k-meta.tabular10K<n<100K0 likes7 downloads9mo agoHugging Face17VillanovaAI /Multi-FLAN-P3gated Multi-FLAN-P3 Overview This dataset is a multilingual subset of the P3 component of the original FLAN collection. The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information). In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the flan_v2_converted dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-P3.text1K<n<10K0 likes7 downloads9mo agoHugging Face18VillanovaAI /OpenData-Benchmark-ITAgated OpenData-Benchmark-ITA Overview OpenData-Benchmark-ITA is a multiple-choice benchmark dataset designed to evaluate the capability of Large Language Models (LLMs) to understand, retrieve, and reason over public Open Data published by European government portals. The current release focuses exclusively on Italian Open Data and is based on datasets published on the official Italian government portal, data.gov.it. Future releases will extend the benchmark to include… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/OpenData-Benchmark-ITA.textn<1K0 likes6 downloads9mo agoHugging Face19VillanovaAI /Multi-FLAN-Flan2021gated Multi-FLAN-Flan2021 Overview This dataset is a multilingual subset of the Flan2021 component of the original FLAN collection. The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information). In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the flan_v2_converted… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-Flan2021.text1K<n<10K0 likes5 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.