datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ethereum-gas-estimator-accuracy
Ethereum gas estimator comparisons
Ethereum mainnet gas suggestions from public RPC providers, recorded alongside provider-reported head blocks and subsequent block-fee statistics. The panel supports analysis of differences in gas recommendations and their relationship to observed block conditions.
Contents
Table
Record
ethereum_gas_recommendations
A provider's gas-price and priority-fee suggestions, with its reported head block
ethereum_block_fees
A… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/ethereum-gas-estimator-accuracy.bitcoin-fee-estimator-accuracy
Bitcoin fee estimator comparisons
Timestamped Bitcoin fee recommendations from public providers, paired with fee-rate percentiles from sampled transactions that were mined. The panel supports comparisons of provider recommendations across confirmation targets and network conditions.
Contents
Table
Record
bitcoin_fee_recommendations
A provider's fee recommendation at one observation time and confirmation target
bitcoin_fee_percentile_comparisons
A… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-fee-estimator-accuracy.grammar-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-completions.essay-vocab-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-trl-completions.grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.e3-math-medhard-zero-accuracytree-species-high-accuracy-psp
High-Accuracy LiDAR-backed PSP Dataset
High-accuracy PSP subset with LiDAR-backed patches.
Size
5,737 LiDAR-backed high-coordinate-accuracy PSP tree samples
Species counts
BA: 36
CW: 1,362
DR: 557
FD: 651
HW: 2,776
MB: 52
SS: 303
Coordinate reliability filtering
Coordinate tiers follow the source PSP coordinate provenance:
High: DGPS, PPP
Medium: RGPS, VGPS
Low: MAP, GIS, GE, PREVIOUS, INTENDED, UNKNOWN, missing, or unrecognized… See the full description on the dataset page: https://huggingface.co/datasets/dfichuk/tree-species-high-accuracy-psp.accuracytest2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 899,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RobEvans/accuracytest2.legal-eyewitness-confidence-accuracy-coherence-decay-v0.1What this dataset is
You get
witness confidence
identification conditions
post event influences
corroboration
an accuracy indicator
You label whether confidence remains coherent with likely accuracy.
Task
Answer coherent or incoherent only.
What it tests
Detection of high confidence under low reliability conditions.
Contamination signals
media exposure, police feedback, show up identification, co witness discussion.
Separation of confidence from accuracy.
Why this matters
Courts often treat… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/legal-eyewitness-confidence-accuracy-coherence-decay-v0.1.accuracytestThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 899,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RobEvans/accuracytest.e3-math-medhard-zero-accuracy-22026_01_22_results_accuracy_compareSee https://github.com/cephcyn/SteerEval for the main description.
Please cite our paper if you use this in your own work:
@misc{zhou2026steerevalframeworkevaluatingsteerability,
title={SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation},
author={Joyce Zhou and Weijie Zhou and Doug Turnbull and Thorsten Joachims},
year={2026},
eprint={2601.21105},
archivePrefix={arXiv},
primaryClass={cs.IR}… See the full description on the dataset page: https://huggingface.co/datasets/cephcyn/2026_01_22_results_accuracy_compare.accuracytest3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 899,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RobEvans/accuracytest3.summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757233813
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757233813"
More Information needed
summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757240890
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757240890"
More Information needed
summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757241106
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757241106"
More Information needed
summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757241465
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757241465"
More Information needed
summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757511400
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757511400"
More Information needed
D-reflection_accuracy_commonsenseQA_longmult_3dig_10exsummarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757478697
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757478697"
More Information needed
D-reflection_accuracy_commonsenseQA_longmult_3dig_100ex
