datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-CoSyn-400K
Multi-CoSyn-400K
Overview
Multi-CoSyn-400K is a multilingual extension of the original CoSyn-400K dataset from AllenAI, which is in turn an extended and improved version of the PixMo-Docs dataset.
The original CoSyn-400K dataset consists of question–answer pairs with reasoning on text-rich images, where both the images and the question-answer pairs with reasoning were generated using code-based rendering tools and text-only LLMs. Example of rendering tools are:… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-CoSyn-400K.multi-pixmo-cap
Multi-PixMo-Cap
Overview
Multi-PixMo-Cap is a multilingual extension of the original PixMo-Cap dataset from AllenAI.The original PixMo-Cap dataset was created by recording annotators speaking freely about an image for 60–90 seconds, then transforming the resulting audio transcripts into detailed captions using Claude (see the PixMo paper).
Multi-PixMo-Cap follows the same multimodal concept, but all examples were re-generated from human captions using a… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-pixmo-cap.multi-pixmo-ask-model-anything
Multi-PixMo-AskModelAnything
Overview
Multi-PixMo-AskModelAnything is a multilingual extension of the original PixMo-AskModelAnything dataset from AllenAI, part of the PixMo series of multimodal resources.
The original PixMo-AskModelAnything dataset consists of image-based question–answer pairs, where annotators authored freeform questions about an image, and answers were generated through a pipeline combining OCR output, dense captions, and a language-only LLM.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-pixmo-ask-model-anything.villanova-sft-2603
Villanova-SFT-2603
Villanova-SFT-2603 is a large-scale, multilingual supervised fine-tuning (SFT) collection of datasets. It contains 1,711,114 instruction-response conversations spanning five European languages, covering chat, instruction following, reasoning, code, knowledge, and safety tasks. This dataset was used to train the Villanova-2B-2603 model family.
All data has been processed through a rigorous curation pipeline that enforces schema normalization, hash-based… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/villanova-sft-2603.Eurostat_Tourism_STS_Dataset_Turnover_in_Services
Eurostat Tourism STS Dataset – Turnover in Services (Monthly)
This repository contains data extracted from the Eurostat Short-Term Statistics (STS) domain, with a focus on:
tour_sts – Tourism industries short-term indicators
sts_setu_m – Turnover in services (monthly data)
These datasets measure monthly turnover and sales volume indices across tourism-related industries following the NACE Rev.2 classification.
Source: https://ec.europa.eu/eurostat/web/tourism/databaseLicense: CC… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Eurostat_Tourism_STS_Dataset_Turnover_in_Services.Multi-SciRIFF
Multi-SciRIFF
A multilingual adaptation of SciRIFF extending a filtered subset of the original English-only instruction-following scientific literature dataset to five languages with permissively licensed synthetic translations.
The original SciRIFF dataset, by AllenAI, includes ~137 K instruction-following demonstrations for 54 scientific literature understanding tasks, organized with rich metadata describing domains, task families, and context. It was developed as a benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-SciRIFF.Temporal_Spatial_Tracking_Dataset
Dataset Overview
This dataset contains time-stamped spatial tracking records collected from tagged entities (e.g., wearable tags, assets, or devices) operating within a monitored environment.Each row represents a single localization event captured at a precise moment in time, including 3D position coordinates and device status information.
The dataset is inherently temporal and spatial, making it suitable for trajectory reconstruction, movement analysis, and time-based behavioral… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Temporal_Spatial_Tracking_Dataset.Time-Series-Donations
Time-Series Donations Dataset
Overview
This repository provides a time-series dataset of donation dynamics over time.It is intended for experiments in:
Time-series forecasting
Trend and seasonality analysis
Anomaly detection on donation flows
Benchmarking classical and deep time-series models
The data are organized in a tabular time-series format, with each row representing a time step and each column representing a numerical or categorical feature related to donations.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Time-Series-Donations.multi-dialogues
Multilingual Dialogues Dataset
Overview
Multilingual Dialogues is a multilingual dataset of synthetically generated everyday conversations between two fictitious people. Most dialogues consist of 6 to 8 turns between the characters.
Multilingual Dialogues is generated using the SODAverse pipeline from AllenAI. The original SODA dataset is generated starting from the Atomic 10x knowledge graph. Triples regarding social interactions are extracted and contextualized to get a… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-dialogues.Multi-FLAN-NIv2
Multi-FLAN-NIv2
Overview
This dataset is a multilingual subset of Natural Instructions v2 (NIv2) as included in the FLAN collection.
The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information).
In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-NIv2.fineweb-edu-sample-15k-meta
fineweb-edu-sample-15k-meta
Overview
This dataset is a metadata-enriched sample of the original FineWeb-Edu dataset.
FineWeb-Edu is a large-scale web text corpus derived from Common Crawl, curated to emphasize educational value and instructional relevance. The dataset is constructed through targeted filtering and classification steps designed to surface content suitable for learning, teaching, and knowledge transfer, while preserving the diversity and scale of web-sourced… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/fineweb-edu-sample-15k-meta.Small-Scale-Fisheries-Tracking-Dataset
Small-Scale Fisheries Tracking Dataset
Raw GPS data from the study "Addressing gaps in small-scale fisheries: a low-cost tracking system"
This dataset contains the raw GPS vessel-tracking data originally released on HuggingFace as part of a study focused on developing a low-cost monitoring solution for small-scale fisheries.
It is provided here to support transparency and reproducibility of the original work.
📦 Dataset Contents
File
Description… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Small-Scale-Fisheries-Tracking-Dataset.fineweb-2-sample-60k-meta
fineweb-2-sample-60k-meta
Overview
This dataset is a metadata-enriched multilingual sample of the original FineWeb-2 dataset.
FineWeb-2 is a large-scale, high-quality web text corpus derived from Common Crawl, built through extensive filtering, deduplication, and language identification steps to support the training of modern large language models. It emphasizes text quality, diversity, and transparency, and includes rich crawl-level metadata inherited from Common Crawl.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/fineweb-2-sample-60k-meta.Multi-FLAN-CoT
Multi-FLAN-CoT
Overview
This dataset is a multilingual subset of the Chain-of-Thought (CoT) component of the original FLAN collection.
The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information).
In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-CoT.PharmaQA.IT
PharmaQA.IT
PharmaQA.IT is an Italian extractive question-answering dataset built from the Riassunti delle Caratteristiche del Prodotto (RCP), the official medicine leaflets issued by the Italian Medicines Agency (AIFA) and collected in PharmaER.IT.
It contains 861 expert-validated question–answer pairs covering indications, contraindications, dosage, warnings, interactions, and pharmacological properties. The pairs were generated semi-automatically: a multimodal LLM prompted… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/PharmaQA.IT.finepdfs-sample-75k-meta
finepdfs-sample-75k-meta
Overview
This dataset is a metadata-enriched multilingual sample of the original FinePDFs dataset.
FinePDFs is a large-scale collection of document-level texts extracted from PDF files, sourced primarily from Common Crawl. The dataset emphasizes high-quality document extraction, structural coherence, and large-scale coverage of technical, scientific, educational, and administrative content commonly distributed in PDF form.
This release contains a… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/finepdfs-sample-75k-meta.Multi-FLAN-P3
Multi-FLAN-P3
Overview
This dataset is a multilingual subset of the P3 component of the original FLAN collection.
The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information).
In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the flan_v2_converted dataset.
The… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-P3.OpenData-Benchmark-ITA
OpenData-Benchmark-ITA
Overview
OpenData-Benchmark-ITA is a multiple-choice benchmark dataset designed to evaluate the capability of Large Language Models (LLMs) to understand, retrieve, and reason over public Open Data published by European government portals.
The current release focuses exclusively on Italian Open Data and is based on datasets published on the official Italian government portal, data.gov.it. Future releases will extend the benchmark to include… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/OpenData-Benchmark-ITA.Multi-FLAN-Flan2021
Multi-FLAN-Flan2021
Overview
This dataset is a multilingual subset of the Flan2021 component of the original FLAN collection.
The original FLAN collection is extremely large and aggregates many instruction-following datasets across tasks and domains (see the original FLAN v2 repo for more information).
In contrast, this release contains a selected subset of the original FLAN data. The selection strategy mirrors the filtering and sampling used in the flan_v2_converted… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-FLAN-Flan2021.
