groundhog
font-square-pretrain-20M
📚 Citation
If you use this dataset in your research, please cite these papers:
@article{pippi2023evaluating,
title={Evaluating Synthetic Pre-Training for Handwriting Processing Tasks},
author={Pippi, Vittorio and Cascianelli, Silvia and Baraldi, Lorenzo and Cucchiara, Rita},
journal={Pattern Recognition Letters},
year={2023},
publisher={Elsevier}
}
@InProceedings{pippi2025zeroshot,
author = {Pippi, Vittorio and Quattrini, Fabio and Cascianelli, Silvia and Tonioni… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-pretrain-20M.GROUNDHOG-RAWfont-square-v2
Accessing the font-square-v2 Dataset on Hugging Face
The font-square-v2 dataset is hosted on Hugging Face at blowing-up-groundhogs/font-square-v2. It is stored in WebDataset format, with tar files organized as follows:
tars/train/: Contains {000..499}.tar shards for the main training split.
tars/fine_tune/: Contains {000..049}.tar shards for fine-tuning.
Each tar file contains multiple samples, where each sample includes:
An RGB image (.rgb.png)
A black-and-white image (.bw.png)… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-v2.font-square-v2-pairs-vaefont_square_charactersACC-dataset
ACC: Agent Context Compilation Dataset
Overview
This dataset contains 10,802 compiled long-context QA pairs derived from multi-turn agent trajectories, introduced in the paper ACC: Compiling Agent Trajectories for Long-Context Training.
Standard agent SFT masks tool responses and only supervises turn-level tool selection, leaving scattered evidence signals unused. Agent Context Compilation (ACC) converts trajectories from Search, Software Engineering (SWE), and SQL agents… See the full description on the dataset page: https://huggingface.co/datasets/groundhogLLM/ACC-dataset.
