datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthesized_datasetai2d-no-maskvila-q-data-traindbpedia-entities-efficient-splade-100K
DBPedia SPLADE + OpenAI: 100,000 SPLADE Sparse Vectors + OpenAI Embedding
This dataset has both OpenAI and SPLADE vectors for 100,000 DBPedia entries. This adds SPLADE Vectors to KShivendu/dbpedia-entities-openai-1M/
Model id used to make these vectors:
model_id = "naver/efficient-splade-VI-BT-large-doc"
For processing the query, use this:
model_id = "naver/efficient-splade-VI-BT-large-query"
If you'd like to extract the indices and weights/values from the vectors, you can do so… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-efficient-splade-100K.Z1-Code-Reasoning-107K
Z1: Efficient Test-time Scaling with Code
Train Large Language Model to Reason with Shifted Thinking
[📜 Paper] •
[🤗 HF Models] •
[🐱 GitHub]
Details
Please refer to https://github.com/efficientscaling/Z1.
Usage
from datasets import load_dataset
ds = load_dataset("efficientscaling/Z1-Code-Reasoning-107K")["train"]
ds[0]
Citation
@misc{yu2025efficientscaling,
title={Z1: Efficient Test-time Scaling with Code}… See the full description on the dataset page: https://huggingface.co/datasets/efficientscaling/Z1-Code-Reasoning-107K.dpsk-r1-labelingdclm-train-1.64m-tsp-efficientLongLive2.0-Toy-Dataset
LongLive2.0 Toy Dataset
This dataset is a toy format-checking dataset for the LongLive2.0 release
code. It is intended to help users verify AR diffusion training, DMD
distillation, and prompt formatting before preparing a larger dataset.
Dataset placeholder:
https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset
Expected Layout
The released toy dataset will contain two separate training folders:
ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2.0-Toy-Dataset.Efficient_ToolCallingrepro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
tts-serving-benchmarkThis repository contains metadata-only versions (audio columns removed) of the following Hugging Face datasets:
https://huggingface.co/datasets/MikhailT/lj-speech
https://huggingface.co/datasets/MikhailT/hifi-tts
https://huggingface.co/datasets/mythicinfinity/libritts
The dataset is intended for benchmarking the performance of TTS serving systems. Example usage is available in this script:
https://github.com/vox-serve/vox-serve/blob/main/benchmark/goodput.py
sts-serving-benchmarkThis repository contains metadata-only versions (audio columns removed) of the following Hugging Face datasets:
https://huggingface.co/datasets/hlt-lab/voicebench
The dataset is intended for benchmarking the performance of STS serving systems. Example usage is available in this script:
https://github.com/vox-serve/vox-serve/blob/main/benchmark/goodput.py
simple_r1bosnian-corpus-v1
Bosnian Corpus v1.0 (cleaned)
This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian,
designed for quantitative analysis of language entropy, “language energy”,
and modern NLP tasks.
The canonical release of this corpus is published on Zenodo:
DOI: 10.5281/zenodo.17757098
Corpus composition
The corpus is built from three publicly available resources released via the CLARIN.SI repository:
Sarajevo Corpus of SMS Messages in Bosnian 1.1
Bosnian… See the full description on the dataset page: https://huggingface.co/datasets/hyper-efficient-system-llc/bosnian-corpus-v1.E2AM_EfficientNetV2_S
E2AM Ablation Results: EfficientNetV2-S
Energy-aware training ablation study for EfficientNetV2-S across three image-classification datasets: CIFAR-10, CIFAR-100, and Tiny-ImageNet.
Each dataset has 15 training variants (8 individual-method M0..M7, 7 cumulative ablation C0..C6) at 50 epochs, plus a 5-variant deployment pipeline (FP32 baseline, structured pruning, pruning+finetune, INT8 quantization, pruned+INT8).
Status: 45 completed variants, 0 partial.
Quick links… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/E2AM_EfficientNetV2_S.efficient-team-4a18b8
efficient-team-4a18b8
Synthetic weather test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/SableField/efficient-team-4a18b8.efficient-speech-codec
Efficient Speech Codec
Paper
Code
Demo Page
efficient-llm-papers
Efficient LLM Papers — FineSet
A research-paper dataset on Efficient LLM Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Efficient LLM Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/efficient-llm-papers.repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Efficient_ToolCalling_trainrepro-from-prior-to-pro-efficient-skill-mastery-via-distribution-contractive-rl-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Z1-Code-Reasoning-Shortest-90Klight_r1repro-velr-efficient-video-reward-feedback-via-ensemble-latent-reward-models-traces
Agent traces
Agent sessions published from a Trackio Logbook.
TER-Token_Efficient_Reasoning
Token Efficient Reasoning
The Token Efficient Reasoning dataset contains high-quality, expert-level reasoning demonstrations structured to capture how domain experts actually think through complex problems. Unlike traditional reasoning traces, TER features information-dense, concise reasoning paths that maintain full logical integrity without veering into "cryptic" shorthand territory.
Dataset Details
TER consists of high-quality Question/Answers from pretraining corpora… See the full description on the dataset page: https://huggingface.co/datasets/SynthData/TER-Token_Efficient_Reasoning.Z1-Code-Reasoning-Longest-33KDRQA_Efficient_Reasoning_CoTefficient_llm
Data V4 for NeurIPS LLM Challenge
Contains 70949 samples collected from Huggingface:
Math: 1273
gsm8k
math_qa
math-eval/TAL-SCQ5K
TAL-SCQ5K-EN
meta-math/MetaMathQA
TIGER-Lab/MathInstruct
Science: 42513
lighteval/mmlu - 'all', "split": 'auxiliary_train'
lighteval/bbq_helm - 'all'
openbookqa - 'main'
ComplexQA: 2940
ARC-Challenge
ARC-Easy
piqa
social_i_qa
Muennighoff/babi
Rowan/hellaswag
ComplexQA1: 2060
medmcqa
winogrande_xl,
winogrande_debiased
boolq
sciq
CNN: 2787… See the full description on the dataset page: https://huggingface.co/datasets/transZ/efficient_llm.efficientnet_b0_b7_comparison_4000_rowsefficientrag-labeler-training-data
EfficientRAG Labeler Training Data
Training data for the Labeler component of EfficientRAG.
Format
JSONL with fields:
question — query text
chunk — retrieved passage text
token_labels — per-word binary labels (1=useful, 0=useless)
tag — <CONTINUE>, <FINISH>, or <TERMINATE>
Statistics
Count
Total samples
30,818
CONTINUE
positive multi-hop chunks
FINISH
single-hop answer chunks
TERMINATE
hard negatives
Data Sources
Source… See the full description on the dataset page: https://huggingface.co/datasets/Necent/efficientrag-labeler-training-data.
