datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lean-proof-compression
LeanPolish: A Kernel-Verified Dataset and Symbolic Compression Framework for Lean 4 Proofs
A dataset of Lean 4 proof rewrite pairs produced by LeanPolish,
a kernel-verified proof-shortening tool. Every accepted
(original, replacement) pair was kernel-checked under Lean 4.21.0
with Mathlib v4.21.0 before emission, and the rewritten file was
re-elaborated end-to-end by a separate out-of-process verifier.
The dataset is suitable for training models that learn to compress,
simplify… See the full description on the dataset page: https://huggingface.co/datasets/leanpolish-anon/lean-proof-compression.compression_test2compression_test3compression_testdolly-15k-prompt-compression
Dolly-15k Prompt Compression
This dataset contains compressed versions of the Databricks Dolly-15k prompts. Each prompt was compressed using the gpt-5-nano model to minimize input tokens while preserving all constraints. You can explore the downstream model that relies on this data in the companion Space: Very Small Prompt Compression Demo.
Compression model: gpt-5-nano
Source dataset: databricks/databricks-dolly-15k
Rows: 15,000
Aggregate token savings: 289,540 → 215,219 tokens… See the full description on the dataset page: https://huggingface.co/datasets/gravitee-io/dolly-15k-prompt-compression.er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42-rollouts
er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
hellaswagf1-quad-tyre-temp-brake-temp-pack-compression-reaction-delta-restart-position-loss-v0.1What this repo does
This dataset models restart instability in Formula One. It predicts when the interaction between tyre temperature readiness, brake temperature readiness, pack compression intensity, and reaction delay creates a high probability of losing positions at a safety car restart.
Core quad
tyre_temp_index
brake_temp_index
pack_compression_index
reaction_time_delta_s
Prediction target
label_restart_position_loss
Binary forward label predicting position loss across the restart phase… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/f1-quad-tyre-temp-brake-temp-pack-compression-reaction-delta-restart-position-loss-v0.1.aria_compressioncompression_data_e39383acee6abb3a424319cdbeed62dcleanforge-compression-evaltrimr-compression-v3trimr-compression-v2tokenlens-compression-benchmark
TokenLens Compression Benchmark
A benchmark dataset measuring LLM prompt compression quality across 100 Wikipedia articles in 5 categories, generated using TokenLens.
Dataset Description
This dataset contains compression quality measurements for 100 Wikipedia articles compressed at 9 different ratios (0.1 to 0.9) using extractive embedding-based compression. For each article and compression ratio, the dataset records tokens saved, semantic similarity, ROUGE-L… See the full description on the dataset page: https://huggingface.co/datasets/lavanyaashri/tokenlens-compression-benchmark.
