datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.encyclopaedia-britannica-lance
Encyclopaedia Britannica (1771-1860) - Lance Format
This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading.
Dataset Details
Total Pages: 155,388
Total Volumes: 195
Format: Lance (columnar format with blob storage for images)
Source: National Library of Scotland (NLS)
License: Public Domain (CC0)
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/encyclopaedia-britannica-lance.ViInfographicsVQA
Introduction
ViInfographicsVQA is a Vietnamese Visual Question Answering (VQA) dataset constructed from infographics sourced from 26 different news platforms. The dataset is designed to support research in multimodal learning by providing diverse questions and answers based on real-world visual data. The detailed distribution of sources is presented in the table below.
Figure 1: The number of infographics per news source.
Developed by: @Namronaldo2004, @Kiet2302… See the full description on the dataset page: https://huggingface.co/datasets/Namronaldo2004/ViInfographicsVQA.MultiExpArt
Dataset Card for Multilingual Explain Artworks: MultiExpArt
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Description
Dataset Summary
As the performance of Large-scale Vision Language Models (LVLMs) improves, they are increasingly capable of responding in multiple languages, and there is an expectation that the demand for explanations generated by LVLMs will grow. However, pre-training… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/MultiExpArt.nanog-cancer-data
NanoG - Cancer Foundation-Model Training Data
Multimodal cancer corpus for NanoG1 (generative multimodal pretraining). Literature, structured biology, imaging, and grounded <simulate> traces.
Hub: Abd0r/nanog-cancer-dataAuthor: Syed Abdur Rehman Ali (@Abd0r) · 17 · independent
Train exclusion (hard): NCI-60 is out of training. Skip records whose source / path / text refer to NCI-60. Prefer NCI-ALMANAC, TCGA-sim, Polymathic, PMC/PubMed, TCGA omics, imaging.
How… See the full description on the dataset page: https://huggingface.co/datasets/Abd0r/nanog-cancer-data.medical-history-of-british-india
A Medical History of British India Dataset
Dataset Description
This dataset contains digitiaed official publications documenting medical research and public health in British India from 1850-1950. The collection represents a crucial period in medical history, capturing the transition from humoral to biochemical medical traditions and documenting major breakthroughs in bacteriology, parasitology, and vaccine development. These documents provide invaluable insights into… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/medical-history-of-british-india.Britain-and-UK-Handbooks-Dataset
Britain and UK Handbooks Dataset
Dataset Description
This dataset contains digitised Britain and UK Handbooks from the National Library of Scotland's digital collections. These annual reference books were originally published by the British Information Service (starting in 1946) to provide overseas readers with comprehensive information about the United Kingdom.
Dataset Summary
Source: National Library of Scotland - Britain and UK Handbooks
Time Period:… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.
Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/Nakhoon/Nemotron-Personas-Korea.Scottish-School-Exam-Papers
Scottish School Exam Papers Dataset
Dataset Description
This dataset contains digitised Scottish school examination papers from the National Library of Scotland's (NLS) digital collections. The papers represent historical educational assessment materials that have been processed with Optical Character Recognition (OCR) to extract text content alongside the original page images.
Dataset Summary
Source: National Library of Scotland - Scottish School Exam Papers… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/Scottish-School-Exam-Papers.Spiritualist_Newspaper
Dataset Card for Spiritualist_Newspaper
Dataset Description
This resource includes plain text transcriptions of The Spiritualist Newspaper (1869), created using Transkribus (https://app.transkribus.org/) and manually corrected across the first serialisation's 50 pages. The transcriptions included were generated using a text extraction model, layout analysis model and text region classifier, all trained within the Transkribus environment. The former model is available… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/Spiritualist_Newspaper.exercise-api
Exercise API — Dataset
Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la
Exercise API. Cada ejercicio incluye grupo muscular,
equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración
masculina y femenina (208 imágenes en total).
Configuraciones
images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento,
músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal
(image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.
