datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthesized_datasetvila-q-data-trainLongLive2.0-Toy-Dataset
LongLive2.0 Toy Dataset
This dataset is a toy format-checking dataset for the LongLive2.0 release
code. It is intended to help users verify AR diffusion training, DMD
distillation, and prompt formatting before preparing a larger dataset.
Dataset placeholder:
https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset
Expected Layout
The released toy dataset will contain two separate training folders:
ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2.0-Toy-Dataset.repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-from-prior-to-pro-efficient-skill-mastery-via-distribution-contractive-rl-traces
Agent traces
Agent sessions published from a Trackio Logbook.
efficient-llm-papers
Efficient LLM Papers — FineSet
A research-paper dataset on Efficient LLM Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Efficient LLM Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/efficient-llm-papers.repro-velr-efficient-video-reward-feedback-via-ensemble-latent-reward-models-traces
Agent traces
Agent sessions published from a Trackio Logbook.
TER-Token_Efficient_Reasoning
Token Efficient Reasoning
The Token Efficient Reasoning dataset contains high-quality, expert-level reasoning demonstrations structured to capture how domain experts actually think through complex problems. Unlike traditional reasoning traces, TER features information-dense, concise reasoning paths that maintain full logical integrity without veering into "cryptic" shorthand territory.
Dataset Details
TER consists of high-quality Question/Answers from pretraining corpora… See the full description on the dataset page: https://huggingface.co/datasets/SynthData/TER-Token_Efficient_Reasoning.efficientrag-filter-training-data
EfficientRAG Filter Training Data
Training data for the Filter component of EfficientRAG.
Format
JSONL with fields:
query_info — concatenation of original question + extracted useful tokens
token_labels — per-word binary labels (1=keep, 0=discard)
Statistics
Count
Total samples
5,691
Data Sources
Source
Language
Samples
Method
HotpotQA (5K questions)
EN
~5K
Heuristic labels
Dragon-derec multi-hop (690)
RU
~700… See the full description on the dataset page: https://huggingface.co/datasets/Necent/efficientrag-filter-training-data.efficientrag-labeler-training-data
EfficientRAG Labeler Training Data
Training data for the Labeler component of EfficientRAG.
Format
JSONL with fields:
question — query text
chunk — retrieved passage text
token_labels — per-word binary labels (1=useful, 0=useless)
tag — <CONTINUE>, <FINISH>, or <TERMINATE>
Statistics
Count
Total samples
30,818
CONTINUE
positive multi-hop chunks
FINISH
single-hop answer chunks
TERMINATE
hard negatives
Data Sources
Source… See the full description on the dataset page: https://huggingface.co/datasets/Necent/efficientrag-labeler-training-data.
