datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma4-serving-bench-data
Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data
Test data, charts, and the running research log from an autonomous research
loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via
llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop
summarizes findings, proposes a goal, tests it end-to-end, documents success or
failure, and publishes here + to GitHub.
Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K
ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.pelican-svg-drawings
Pelican SVG: frontier model drawings, scored
139 SVGs produced by seven frontier models answering Simon Willison's prompt,
"generate an SVG of a pelican riding a bicycle", each with the score the
pelican_svg_env
OpenEnv environment gave it. Four of the seven are open weights and three are closed.
Simon has run that prompt against nearly every model release since early 2025, but the
results live as embedded images across 129 blog posts and scattered gists. This dataset
exists… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/pelican-svg-drawings.ConflictGUI
ConflictGUI
ConflictGUI is a benchmark for evaluating conflict awareness in GUI agents.
Dataset Splits
Split
Feasible
Conflict1
Conflict2
Total
Calibration
564
300
300
1,164
Test
1,800
822
874
3,496
Sources
ConflictGUI is constructed from AMEX, AndroidControl, and AITZ. The conflict instructions and labels are newly annotated, while screenshots and original instructions remain subject to their respective upstream terms.… See the full description on the dataset page: https://huggingface.co/datasets/serein356/ConflictGUI.sermon-index
SermonIndex — Sermon Metadata and Scripture Index
Metadata for the sermon archive at SermonIndex.net,
including the mapping between scripture passages and the sermons that expound them.
Table
Rows
What it holds
sermon
64,106
Title, speaker, summary, description, media links, URLs
sermon_scripture
281,754
Scripture references, one row per sermon–passage pair
sermon_topic
115,526
Topic assignments, one row per sermon–topic pair
speaker
2,188
Preachers and authors… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/sermon-index.turkish-comprehensive-movie-series-dataset
Beyazperde Film & Series Dataset
This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings.
Dataset Summary
Total Movies: 27,227
Total Series: 11,240
Total Entries: 38,467
File Size: ~222 MB
Format: JSONL (JSON Lines)
Language: Turkish
Source: Beyazperde.com
Data Structure
Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.ser-re-training-data
SER & RE Training Data
Datasets for Semantic Entity Recognition (SER) and Relation Extraction (RE) training.
Datasets
1. XFUND (~/data/ser_re/xfund/)
Source: https://github.com/doc-analysis/XFUND/releases/tag/v1.0
Languages: zh, ja, es, fr, it, de, pt (7 languages)
Documents: 149 train + 50 val per language = 1,393 total
Images: 1,393 JPG files ({lang}{split}{idx}.jpg)
Annotations: {lang}.train.json / {lang}.val.json
Format: Word-level bboxes, BIO labels… See the full description on the dataset page: https://huggingface.co/datasets/bluecopa/ser-re-training-data.
