datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
compact-alignments
compact-alignments — per-verse, per-book, content-addressed
The token-position companion to lexeme-alignments (which is
aggregated/type-level and can't tell you what happened in any one verse). This dataset restores
position: for a given edition's Bible book, which Hebrew/Greek content word aligned to which
target-text token, verse by verse.
The authoritative list of what's published is always manifest.json, not this file.
Original-language source editions (needed… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments.CompactDS-102GB
Introduction
CompactDS is a diverse, high-quality, web-scale datastore that achieves high retrieval accuracy and subsecond latency on a single-node deployment, making it suitable for academic use. Its core design combines a compact set of high-quality, diverse data sources with in-memory approximate nearest neighbor (ANN) retrieval and on-disk exact search. We release CompactDS and our retrieval pipeline as a fully reproducible alternative to commercial search, supporting future… See the full description on the dataset page: https://huggingface.co/datasets/alrope/CompactDS-102GB.Compact_OpenAIRE_citation_graph
📚 Compact OpenAIRE Citation Graph
Based on OpenAIRE Graph v11.1.1 (source on Zenodo).
The complete OpenAIRE citation graph, distilled into a handful of compact, analysis-ready files — the full scholarly citation network of the open-science ecosystem, small enough to actually work with.
Citation graphs at this scale are usually locked behind multi-terabyte dumps and heavyweight infrastructure. This dataset makes the entire OpenAIRE citation network loadable… See the full description on the dataset page: https://huggingface.co/datasets/Zmeos/Compact_OpenAIRE_citation_graph.FlexiSLM-Data-2M-s2s-compact
FlexiSLM-Data — Speech-to-Speech Part (2.43M filtered samples, 385G in size)
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in
WebDataset format.
Related data releases… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/FlexiSLM-Data-2M-s2s-compact.slm-parameter-audit
SLM card-vs-artifact parameter audit
An autonomous audit of small-language-model repos on the Hugging Face Hub. For each
in-scope model (independent builders training very small models from scratch, roughly
0.5M–500M parameters), the parameter count stated in the model card is compared against
the actual artifact: the safetensors header, config.json, and the training script where
present. A mismatch is recorded when the card's number does not match the artifact's
real parameter… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-parameter-audit.Gargantua-R1-Compact
Gargantua-R1 Distribution
Gargantua-R1-Compact(experimental purpose)
Gargantua-R1-Compact is a large-scale, high-quality reasoning dataset primarily designed for mathematical reasoning and STEM education. It contains approximately 6.67 million problems and solution traces, with a strong emphasis on mathematics (over 70%), as well as coverage of scientific domains, algorithmic challenges, and creative logic puzzles. The dataset is suitable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gargantua-R1-Compact.agent-memory-compaction-trajectories
Agent Memory Compaction Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agent-memory-compaction-trajectories.droid_1.0.1_v30_compact_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_5.CompactDS-102GB-raw-textsscc-compact-av
SSCC compact balanced multimodal subset
This private derived dataset contains 288 synchronized SSCC clips from 15 medium-load,
clean operating conditions at speeds 60, 80, and 100. It retains recorder FLAC audio,
four anti-aliased 25 kHz vibration channels in compressed float32 NPZ, and five sparse frames
from both iOS and Android videos. Five sample IDs retain both unchanged source MP4s for
presentation and loader tests.
The subset is balanced between normal and fault states… See the full description on the dataset page: https://huggingface.co/datasets/DesanSilva/sscc-compact-av.Gargantua-R1-Compact
Gargantua-R1 Distribution
Gargantua-R1-Compact(experimental purpose)
Gargantua-R1-Compact is a large-scale, high-quality reasoning dataset primarily designed for mathematical reasoning and STEM education. It contains approximately 6.67 million problems and solution traces, with a strong emphasis on mathematics (over 70%), as well as coverage of scientific domains, algorithmic challenges, and creative logic puzzles. The dataset is suitable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Gargantua-R1-Compact.CompactDS-102GB-retrieval-resultsdroid_1.0.1_v30_compact_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 10240,
"total_frames": 2988169,
"total_tasks": 6798,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:10240"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_3.iter267-apex-i-compact-routefix-encbookcorpus_compact_1024_shard8_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard8_of_10_meta"
More Information needed
cfb-playoff-compacts-2026pokebot-bc-compact-v5libero_spatial_noop_usd3_link_lerobot_compact_v30_workdirbookcorpus_compact_1024_shard7_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard7_of_10_meta"
More Information needed
droid_1.0.1_v30_compactThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95584,
"total_frames": 27607757,
"total_tasks": 49596,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95584"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact.compact-jailbreaks
Dataset Featurization: Extracting Compact Jailbreaks
This repository contains the datasets used in our case study on extracting compact representations of jailbreak tactics, demonstrating how our unsupervised featurization pipeline can effectively compress large sets of adversarial prompts while maintaining their effectiveness and diversity.
Featurization - WildTeaming
Access both the input dataset from WildTeaming and the evaluation stage outputs containing candidate… See the full description on the dataset page: https://huggingface.co/datasets/Bravansky/compact-jailbreaks.bookcorpus_compact_1024_shard9_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard9_of_10_meta"
More Information needed
apparel23-smartcat-google-intersection-priced-half-compact
Apparel23 Priced Half Compact Atlas Artifact
Compact Atlas retrieval artifact for flavianv/apparel23-smartcat-google-intersection.
Contents:
clothing_2023_smartcat_google_catalog_text.jsonl: priced Apparel23 catalog rows with only fields used by the outfit harness.
e5_smartcat_google_intersection_embeddings.pt: embedding rows aligned to item_order.json.
item_order.json: ordered item ids for the compact embedding tensor.
build_summary.json: build metadata and row counts.… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/apparel23-smartcat-google-intersection-priced-half-compact.longvideogen_wavespeed_compact_3_trial3bookcorpus_compact_256
Dataset Card for "bookcorpus_compact_256"
Num samples: 2,389,359
More Information needed
drctx-compaction-openresearcher-mk1gr1_pickup_compact_h264This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 150,
"total_frames": 7800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:150"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vedpatwardhan/gr1_pickup_compact_h264.b1k-joint-v5-compactbookcorpus_compact_1024_shard4_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard4_of_10_meta"
More Information needed
bookcorpus_compact_1024_shard0_of_10_meta
Dataset Card for "bookcorpus_compact_1024_shard0_meta"
132 hours to finish
num_examples: 61605
size: 1.5GB
More Information needed
