datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ImageNet-C-jpeg_compression-severity_5compression-pretraining-data
Dataset
Each example contains prompt (chat format) and target fields.
from datasets import load_dataset
ds = load_dataset("leonli66/compression-pretraining-data", "<config_name>")
lean-proof-compression
LeanPolish: A Kernel-Verified Dataset and Symbolic Compression Framework for Lean 4 Proofs
A dataset of Lean 4 proof rewrite pairs produced by LeanPolish,
a kernel-verified proof-shortening tool. Every accepted
(original, replacement) pair was kernel-checked under Lean 4.21.0
with Mathlib v4.21.0 before emission, and the rewritten file was
re-elaborated end-to-end by a separate out-of-process verifier.
The dataset is suitable for training models that learn to compress,
simplify… See the full description on the dataset page: https://huggingface.co/datasets/leanpolish-anon/lean-proof-compression.round-trip-code-compressionsentence-compression
Dataset Card for Sentence Compression
This dataset is a collection of text-simplified pairs from the Sentence Compression project. See Sentence Compression for additional information.
This dataset can be used directly with Sentence Transformers to train embedding models.
Dataset Subsets
pair subset
Columns: "text", "simplified"
Column types: str, str
Examples:{
'text': "The USHL completed an expansion draft on Monday as 10 players who were on the rosters of… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/sentence-compression.ImageNet-C-jpeg_compression-severity_4OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-3-CompressionRate-0.1compression_dataset
Compression Dataset — ER Training Subset (3,200)
The 3,200 examples selected by the Efficient-Reasoning training script, stored as Parquet for Hugging Face Dataset Viewer support. This is a subset of daman1209arora/compression_dataset, not the full upstream dataset.
Exact selection
run_rloo_deepseek_1.5B_compression.sh sets --max_samples 3200. The loader in openrlhf/utils/utils.py takes the first 3,200 source rows before shuffling with seed 42. This repository… See the full description on the dataset page: https://huggingface.co/datasets/zjhhhh/compression_dataset.ImageNet-C-jpeg_compression-severity_3OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-CompressionRate-0.9OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-2-CompressionRate-0.3OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-3-CompressionRate-0.6OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-3-CompressionRate-0.4Star-41K-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-CompressionRate-0.5OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-2-CompressionRate-0.6task1340_msr_text_compression_compression
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1340_msr_text_compression_compression
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1340_msr_text_compression_compression.TokenComplexitysentence-compression-pairsCompression_CoCOpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-2-CompressionRate-0.8hellaswagflan_combined_task1340_msr_text_compression_compressionOpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-3-CompressionRate-0.8OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-2-CompressionRate-0.5OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-CompressionRate-0.7Star-41K-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-CompressionRate-0.8OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-Level-2-CompressionRate-0.4leanforge-compression-evalStar-41K-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-CompressionRate-0.0OpenR1-DeepSeek-R1-Distill-Qwen-7BBlockwiseCompressedReasoniong-CompressionRate-0.3
