datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
encyclopaedia-britannica-lance-test
Encyclopaedia Britannica (1771-1860) - Lance Format
This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading.
Dataset Details
Total Pages: 155,388
Total Volumes: 195
Format: Lance (columnar format with blob storage for images)
Source: National Library of Scotland (NLS)
License: Public Domain (CC0)
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test.encyclopaedia-britannica-lance-test2
Encyclopaedia Britannica (1771-1860) - Lance Format
This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading.
Dataset Details
Total Pages: 155,388
Total Volumes: 195
Format: Lance (columnar format with blob storage for images)
Source: National Library of Scotland (NLS)
License: Public Domain (CC0)
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test2.handwriting-test
handwriting-test
This dataset contains handwriting stroke data collected using a stylus (S Pen) on a tablet device.
Data Format
Each sample in data/train.jsonl contains:
Field
Description
id
Unique identifier (UUID)
text
The prompt text that was written
created_at
ISO timestamp of when the sample was created
device
Device information (user agent, platform, pixel ratio)
canvas
Canvas dimensions (width, height)
strokes
Array of strokes, each containing… See the full description on the dataset page: https://huggingface.co/datasets/finnbusse/handwriting-test.ChemVLM_test_dataUsing this dataset, please kindly cite:
@inproceedings{li2025chemvlm,
title={Chemvlm: Exploring the power of multimodal large language models in chemistry area},
author={Li, Junxian and Zhang, Di and Wang, Xunzhi and Hao, Zeying and Lei, Jingdi and Tan, Qian and Zhou, Cai and Liu, Wei and Yang, Yaotian and Xiong, Xinrui and others},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={39},
number={1},
pages={415--423},
year={2025}
}… See the full description on the dataset page: https://huggingface.co/datasets/Duke-de-Artois/ChemVLM_test_data.test0611testtest_bench
CAD Synthesis Benchmark — test_bench
Small-scale benchmark for image → CadQuery code generation.
Task: Given rendered views of an industrial part, generate CadQuery code that reproduces it.
Evaluation: Execute code → export STEP → compute IoU vs GT STEP.
Splits
Split
Families
Description
test-iid
54 train families, XY plane
In-distribution test
test-ood-family
19 held-out families
Family generalization
test-ood-plane
54 train families, XZ/YZ
View… See the full description on the dataset page: https://huggingface.co/datasets/Hula0401/test_bench.handwriting-test-v2
handwriting-test-v2
This dataset contains handwriting stroke data collected using a stylus (S Pen) on a tablet device.
Optimized for training RNNs (Recurrent Neural Networks) on handwriting generation/recognition tasks.
Data Format
Each row in the Parquet files represents a complete handwriting sample:
Column
Type
Description
id
string
Unique identifier (UUID)
text
string
The prompt text that was written
dx
string (JSON array)
Delta X offsets between… See the full description on the dataset page: https://huggingface.co/datasets/finnbusse/handwriting-test-v2.Dataset_Large_test
PPU-Bench
datikz_test_paired_desc
datikz_test_paired_desc — the DaTikZ v2/v3 official test split, with paired descriptions and reference images
The 984-item held-out test split used to evaluate text → TikZ models, with two descriptions of every
diagram so that description style can be isolated from model quality, plus the reference image for
image-space metrics.
This is deliberately a separate repo from the training corpus
(explcre/diagram_desc_datikz), which
contains no test items. Train and test cannot be… See the full description on the dataset page: https://huggingface.co/datasets/explcre/datikz_test_paired_desc.MSCOCO2014_testThis dataset is only a smaller version of JustinLeeCEO/MSCOCO2014, only used for testing.
It was converted from MSCOCO 2014 aiming at adapting the COCO dataset to EasyR1 using the following script.
NOTE:
This dataset only use COCO's segmentation data and caption data from its trainset and valset.
The first N_val samples of original valset act as new valset.
The last N_test samples of original valset act as new testset.
import os
import json
from datasets import Dataset, DatasetDict, Sequence… See the full description on the dataset page: https://huggingface.co/datasets/JustinLeeCEO/MSCOCO2014_test.wisconsin-test-dataset
Wisconsin Test Dataset
Generated by DocParserEngine.
Field
Value
Documents
1
Records
1
Schema
full
Usage
from datasets import load_dataset
ds = load_dataset("Remixonwin/wisconsin-test-dataset")
Phi4-Mini-P2T-4B-TestingTesting Results for USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B)
MySuperDataset-TestRepo
MySuperDataset
1. Introduction
MySuperDataset is a comprehensive multi-domain text corpus designed for training and evaluating language models. This dataset has been carefully curated through multiple iterations to ensure high quality, diversity, and representativeness across various domains including science, technology, arts, and everyday conversations.
The dataset includes over 10 million samples spanning 15 different quality metrics.… See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/MySuperDataset-TestRepo.test01test2323323STEM_train_edit_predictions_test
STEM Image Edit Predictions Dataset
This dataset contains AI-generated edit predictions for STEM images based on captions and edit commands.
Dataset Structure
The dataset is organized in batches:
Total batches: 1
Each batch is stored in a separate directory (batch_0000, batch_0001, etc.)
Fields
Each item contains:
layout_summary: Concise description of source image layout and key elements
edit_analysis: Specific elements to be edited and changes to be made… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/STEM_train_edit_predictions_test.
