datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datasets-tests-compressionvision-token-compression-bench
OPTIC-Bench
Optical Text In-Context Benchmark: how reliably do LLMs consume text
delivered as rendered images versus plain text tokens?
In summary, the evaluation reported here finds that optical text compression
is effective only within a narrow and specific envelope. Delivering content
as rendered images genuinely reduces input tokens, by thirteen to
fifty-four per cent depending on the model and the language, but only when
the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.all-deletion-compressionswikipedia-deletion-compressionssentence-compression
Dataset Card for "sentence-compression"
Dataset Summary
Dataset with pairs of equivalent sentences.
The dataset is provided "AS IS" without any warranty, express or implied.
Google disclaims all liability for any damages, direct or indirect, resulting from using the dataset.
Disclaimer: The team releasing sentence-compression did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/sentence-compression.llm-compressionThis is the compression corpora dataset used in the paper "Compression Represents Intelligence Linearly".
We find that LLMs’ intelligence – reflected by benchmark scores – almost linearly correlates with their ability to compress external text corpora. We measure intelligence along three key abilities: knowledge and commonsense, coding, and mathematical reasoning, and provide the corresponding compression corpora here respectively named cc, python, and arxiv_math.
Load the data… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/llm-compression.cai-semantic-equivalence-benchmark
Contradish CAI-Bench
The semantic equivalence benchmark from Contradish
Do AI systems give the same answer when the wording changes but the meaning does not?
Contradish CAI-Bench measures semantic invariance: whether an AI system remains behaviorally consistent across prompts that express the same intent in different words.
This Hugging Face release contains 420 human-readable prompt pairs across 19 domains. Contradish is the official benchmark runner, scoring… See the full description on the dataset page: https://huggingface.co/datasets/compressionawareintelligence/cai-semantic-equivalence-benchmark.lace-semantic-compression
LACE — Latent Adaptive Compression Engine
Semantic Compression Under Physical Channel Constraints: Cognitive Phase Transitions Under Bandwidth Constraints
Théophile Lafargue · April 2026 · Patent FR2511116
What this is
198 operational tasks (defense, medical, industrial) used to study what emerges when you force discrete semantic compression under LoRa/SMS physical constraints.
v1 result: Domain clustering real (mean coherence 68.4%). Retrieval/inference separation not… See the full description on the dataset page: https://huggingface.co/datasets/ox-ox/lace-semantic-compression.mn-context-compression-dataset-v1
MN Context Compression Dataset v1
Author: Homer Quan
This dataset is used to train context-compression models for improving the context efficiency of multi-agent runtimes, especially MirrorNeuron and the broader work at mirrorneuron.io.
We use this dataset to train models such as homerquan/mn-context-engine-lora-v2, and later protected-fact-focused context engines. The data emphasizes exact protected-span retention, source-reference preservation, budget-conditioned compression, and… See the full description on the dataset page: https://huggingface.co/datasets/homerquan/mn-context-compression-dataset-v1.semantic-compression-sft
sematic-compression-sft
Dataset Summary
sematic-compression-sft is a synthetic supervised fine-tuning dataset for semantic compression.
The task is to convert verbose natural-language or code inputs into compact outputs that preserve reasoning-relevant information.
This dataset is designed for training compression models/adapters used before downstream LLM inference to reduce prompt size while retaining functional utility.
Goal
The objective is not generic… See the full description on the dataset page: https://huggingface.co/datasets/Sudhendra/semantic-compression-sft.er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42-rollouts
er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
Fathom-V0.4-RL-Compressionsentence-compression-translated-nlThis is a Dutch version of the Sentence Compression dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
SparsePlug-Compression-Benchmarks
SparsePlug Hardware Profiler & Compression Benchmarks
Author: Brad Wallace (coo@koba42.com)Source Code: github.com/tensorrent/prime-fheLicense: Apache-2.0
Component Description
Hardware-adaptive sparsity profiling, TrinityWasm compression, and zero-degradation motif throughput benchmarks at 0%, 75%, and 90% sparsity.
Verified Benchmark Performance
Metric
Value
Vector Dim
384
Step Latency Ns
1420
Full Evaluation Ms
0.55
Throughput… See the full description on the dataset page: https://huggingface.co/datasets/K42COO/SparsePlug-Compression-Benchmarks.p2-etf-compression-complexity-resultsSFT_cot_compressionhan-knowledge-compression-distillation-dataset-v1
Humanoid Knowledge Compression & Embedding Distillation Dataset
This dataset captures knowledge transfer
between high-capacity cognitive nodes
and lightweight humanoid edge agents.
It includes embedding compression,
distillation mappings, and performance delta tracking.
Objective
To enable scalable knowledge sharing
across heterogeneous humanoid agents.
Data Fields
source_model_id
target_agent_id
original_embedding_vector
compressed_embedding_vector… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-knowledge-compression-distillation-dataset-v1.compression_data_e39383acee6abb3a424319cdbeed62dchan-memory-compression-samples-v1
Humanoid Memory Compression Samples
This dataset shows how raw experiences
are reduced into compact memory units.
It supports scalable long-term intelligence.
Contents
Raw experience reference
Compressed memory token
Retention priority
Use Cases
Memory optimization
Storage efficiency
Cognitive scaling
Part of
Humanoid Network (HAN)
License
MIT
