datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-intuition-20260906-403-demo-10
math-intuition-20260906-403-demo-10
3,936 mathematics problems drawn from 403 problem families, each derived from a
distinct arXiv paper. Every problem is generated answer-first, so the answer is known by
construction and is checked by the family's own verify() before the row is written.
No row in this file is ungraded.
This is the demo rung — read this before using it
Each family exposes a four-rung ladder: demo, easy, medium, hard. This file samples
demo, which… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-demo-10.tb2.0_demo
Terminal-Bench 2.0 Demo Trajectories
A curated set of 8 terminal-bench style task trajectories, split into two complementary subsets:
short — 5 trajectories with < 40 agent steps (observed range 17–31)
long — 3 trajectories with > 40 agent steps (observed range 55–68)
Each entry contains a self-contained task definition, a fully reproducible Docker environment, and the agent's complete execution trajectory — all verified to pass every test under strict test isolation (reward = 1.0… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/tb2.0_demo.Terminal_trajactory_demo
Terminal Agent Trajectory Demo
Complete multi-turn conversation trajectories of an AI agent solving programming tasks in a Linux terminal environment. Designed for training and evaluating Terminal/CLI agents.
Overview
Item
Details
Samples
20 (ID 1441–1460)
Task Language
Chinese instructions + English code
Difficulty
Medium
Expert Time Estimate
15 min
Environment
Linux / Python 3.13 / Docker
Task Categories
Category
Sample IDs… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/Terminal_trajactory_demo.pdfsys-page-v2-demo
pdfsys.page/v2 — 格式演示数据集
pdfsys.page/v2 是 pdfsystem_mnbvc
的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。
这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、
三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。
来源提示:这里的 PDF 页来自 OmniDocBench
与 olmOCR-bench 两个公开
benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些
文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。
一句话设计
一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的;
页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强;
图像像素要么是裁剪图、要么是整页光栅,二选一。
里面有什么
config
行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.rewire-jobs-demo
REWIRE on HF Jobs — a small, open, reproducible re-run
A proof-of-concept reproduction of REWIRE ("Recycling the Web", Meta/FAIR, COLM 2025 — arXiv:2506.04689), built entirely on Hugging Face Jobs with datatrove. REWIRE takes low-quality web documents that pre-training filters would throw away and rewrites them with an LLM into something worth keeping:
discarded web doc ──▶ LLM rewrite ──▶ re-score, same filter ──▶ keep survivors
(median score 0.02) (paper's… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/rewire-jobs-demo.demonstration_glm52
demonstration_glm52
Hindsight reasoning-chain demonstrations for SDFT, written by GLM-5.2 on the
training share of the swmbench/swmbench jin10_attributed_filtered_en_daily
rolling 70/30 split.
Each row is a prediction-market forecasting example; output_text is a worked
reasoning chain that a skilled forecaster would have produced before knowing
the outcome, ending in a residual bucket <answer>Bx</answer>. The writer was
privately shown the realised move and instructed never to… See the full description on the dataset page: https://huggingface.co/datasets/skyyyyks/demonstration_glm52.physics-30k-demo
Computational & Quantitative Sciences Q&A — Multi-Level Explanations
24 question-answer pairs generated from recent papers (arXiv 2024–2026),
covering 6 subfields across 6 papers.
Each paper is explained at 4 depth levels, each as a focused Q/A pair:
Level
Description
L1
Intuitive / Phenomenological — what is happening, plain language, analogies, no equations
L2
Conceptual / Structural — key components, pipeline/steps, minimal formalism
L3
Mechanistic / Formal —… See the full description on the dataset page: https://huggingface.co/datasets/planetoid-reader/physics-30k-demo.vgrout-leetcode-teacher-demos
vGROUT LeetCode teacher demonstrations
Cached teacher demonstrations used to warm up the
vGROUT gradient-routing experiments on the
ariahw/rl-rewardhacking LeetCode
environment. Each row is a full problem-specific completion. The kind column gives the
two demonstration types:
hack (215 rows): verified exploits of the run_tests loophole (hacked=True,
gt_pass=False).
solve (126 rows): correct solutions verified against the ground-truth tests
(gt_pass=True).
Why fewer… See the full description on the dataset page: https://huggingface.co/datasets/wassname/vgrout-leetcode-teacher-demos.demo
Dataset Card for Dataset Name
This is a lot of info wow.
Dataset Details
Dataset Description
Just a demo
Curated by: Turtles
Funded by [optional]: Turtles
Shared by [optional]: Turtles
Language(s) (NLP): Enlish
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/andrew-noske/demo.demodemo
