villanova
Multi-CoSyn-400K
Multi-CoSyn-400K
Overview
Multi-CoSyn-400K is a multilingual extension of the original CoSyn-400K dataset from AllenAI, which is in turn an extended and improved version of the PixMo-Docs dataset.
The original CoSyn-400K dataset consists of question–answer pairs with reasoning on text-rich images, where both the images and the question-answer pairs with reasoning were generated using code-based rendering tools and text-only LLMs. Example of rendering tools are:… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-CoSyn-400K.multi-pixmo-cap
Multi-PixMo-Cap
Overview
Multi-PixMo-Cap is a multilingual extension of the original PixMo-Cap dataset from AllenAI.The original PixMo-Cap dataset was created by recording annotators speaking freely about an image for 60–90 seconds, then transforming the resulting audio transcripts into detailed captions using Claude (see the PixMo paper).
Multi-PixMo-Cap follows the same multimodal concept, but all examples were re-generated from human captions using a… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-pixmo-cap.multi-pixmo-ask-model-anything
Multi-PixMo-AskModelAnything
Overview
Multi-PixMo-AskModelAnything is a multilingual extension of the original PixMo-AskModelAnything dataset from AllenAI, part of the PixMo series of multimodal resources.
The original PixMo-AskModelAnything dataset consists of image-based question–answer pairs, where annotators authored freeform questions about an image, and answers were generated through a pipeline combining OCR output, dense captions, and a language-only LLM.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-pixmo-ask-model-anything.villanova-sft-2603
Villanova-SFT-2603
Villanova-SFT-2603 is a large-scale, multilingual supervised fine-tuning (SFT) collection of datasets. It contains 1,711,114 instruction-response conversations spanning five European languages, covering chat, instruction following, reasoning, code, knowledge, and safety tasks. This dataset was used to train the Villanova-2B-2603 model family.
All data has been processed through a rigorous curation pipeline that enforces schema normalization, hash-based… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/villanova-sft-2603.Eurostat_Tourism_STS_Dataset_Turnover_in_Services
Eurostat Tourism STS Dataset – Turnover in Services (Monthly)
This repository contains data extracted from the Eurostat Short-Term Statistics (STS) domain, with a focus on:
tour_sts – Tourism industries short-term indicators
sts_setu_m – Turnover in services (monthly data)
These datasets measure monthly turnover and sales volume indices across tourism-related industries following the NACE Rev.2 classification.
Source: https://ec.europa.eu/eurostat/web/tourism/databaseLicense: CC… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Eurostat_Tourism_STS_Dataset_Turnover_in_Services.Multi-SciRIFF
Multi-SciRIFF
A multilingual adaptation of SciRIFF extending a filtered subset of the original English-only instruction-following scientific literature dataset to five languages with permissively licensed synthetic translations.
The original SciRIFF dataset, by AllenAI, includes ~137 K instruction-following demonstrations for 54 scientific literature understanding tasks, organized with rich metadata describing domains, task families, and context. It was developed as a benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Multi-SciRIFF.
