datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.sql-grimoire
Image generated by DALL-E.
Grimoire of SQL
Overview
Grimoire of SQLite is a comprehensive dataset tailored for training and evaluating text-to-SQL models. It consolidates and enhances multiple existing datasets, including Spider, BirdBench, and Gretel, by correcting errors, refining natural language queries, and validating SQL queries for runnability. The dataset is specifically designed to support high-quality fine-tuning of models like GPT-4 and its variants… See the full description on the dataset page: https://huggingface.co/datasets/data-maki/sql-grimoire.grimoire-figr8
Grimoire — FIGR-8 (processed / final training data)
The final, post-processed FIGR-8 training data for the Grimoire SVG pipeline.
Shipped processed because regenerating it from the raw FIGR-8 corpus (convert + deepsvg
normalize + tokenize) is painful.
⚠️ Split differs from the paper. The paper describes a 427K-sample, 1,000-class, 90/5/5 subset
(that split file is datasets/make_font/figr8_final_paper_split.csv in the code repo), but it is not
reproducible from code (the class… See the full description on the dataset page: https://huggingface.co/datasets/Potpov/grimoire-figr8.grimoire-emoji
Grimoire — Emoji (processed masks)
Processed Twitter-emoji (twemoji) data for the Grimoire SVG pipeline. Shipped
processed because regenerating the layer masks requires SAM (Segment Anything)
segmentation — painful.
Contents
preprocessed_v2.tar (5.7 GB) — training-ready per-emoji layer data consumed by CartoonDataset
(dataset key cartoons): color_masks/, bw_masks/, contour_masks/, color_info/, coords/,
merged/ (each split into train/ val/) + color_count.npy.… See the full description on the dataset page: https://huggingface.co/datasets/Potpov/grimoire-emoji.grimoire-mnist
Grimoire — MNIST (raw PNGs + ART tokens)
MNIST for the Grimoire SVG pipeline. The stage-1 (VSQ) input is shipped raw (PNGs) —
the heavy pre-tiled tensors (2.7 TB) regenerate easily from them, so they are not shipped.
The stage-2 (ART) tokens are shipped ready to use.
Contents
mnist_png.tar.gz (15 MB) — training/{0..9}/*.png + testing/{0..9}/*.png (standard MNIST).
mnist_tokenized.tar (3.7 GB) — all precomputed VSQ token sets for stage-2 (ART) (the paper variants:… See the full description on the dataset page: https://huggingface.co/datasets/Potpov/grimoire-mnist.The-Left-Hand_The-Cabal-Grimoire-of-Walking-in-DarknessGrimoiregrimoire-dataset
