datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atlas-24-frozen-prefix-potential-shaping
ATLAS report 24: frozen-prefix potential shaping
1. Question and links
Read this first. This data root holds the first attempt of report 24 on the campaign's old harness (verl 0.7.1): the shaped training is complete and the unshaped training stopped at step 20 with a known problem (the subsection at the end of this section). The question was rerun on the runtime of report 25 with both trainings at 40 steps; that rerun's trajectories, exports, checkpoints and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-24-frozen-prefix-potential-shaping.atlas-23-prefix-curves-on-a-larger-selector
ATLAS report 23: prefix curves on a larger selector
1. Question and links
Read this first. GPQA ran in full on both surfaces (198 questions, prefix lengths k = 1 to 8, 1584 states each). LiveCodeBench was started and stopped by the user at 455 and 531 of its 1400 states per surface and is not read: every table and the report read GPQA only. A reader who cannot fetch files from the Hub finds this whole data root mirrored in the private GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-23-prefix-curves-on-a-larger-selector.prm800k_onpolicy_multiturn_rtg_prefix0.2_roll4_maxrev100pile_trigram_prefixes
trigram_prefixes
See https://confirmlabs.org/posts/catalog.html for details.
id0: the first token in the trigram
id1: the second token in the trigram
id2: the most common token following (id0, id1) in The Pile
sum_count: the number of times that (id0, id1) appears in The Pile.
max_count: the number of times that id2 appears after (id0, id1) in The Pile.
frac_max: max_count / sum_count
alfworld-expert-prefix-rollouts
ALFWorld Expert-Prefix Rollout Landscape
This dataset measures how a frozen language-model actor's probability of
solving an ALFWorld task changes after replaying different-length prefixes of a
successful expert trajectory.
The collection contains all 3,553 ALFWorld training tasks from the Agent-G2 SFT
data. Eight independent actor rollouts were sampled from the initial state for
every task. For the 2,307 low-signal tasks with at most one root success, eight
additional… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/alfworld-expert-prefix-rollouts.sffop_1706381144_410msft_relabel_pythia6.9b_logprobs_prefix_chosenprm800k-truncated-prefix-legacy-results
PRM800K truncated-prefix legacy results
Public archive of legacy experiment artifacts produced before complete PRM800K attempts were restored. Scientific interpretation and the corrected rerun are documented in the source repository.
Declared files: 60748
Declared bytes: 187013895604
Verification: exact remote paths, file count, and aggregate size at a pinned revision
Full remote content readback: not performed
Gleipnir-Prefix-Teacher-Cache
Gleipnir Prefix Teacher Cache
Research artifact containing 133,947 intermediate tool-trajectory predictions
from Qwen/Qwen3.5-27B-FP8. This is a prediction cache, not a self-contained
trajectory dataset. It contains no trajectory text, original hard labels,
privileged rationales, or Kimi K3 full-trajectory targets.
Important numerical limitation
The completed cache failed its numerical-agreement audit. On a fixed
64-prefix sample, fresh versus cached probability… See the full description on the dataset page: https://huggingface.co/datasets/Jazhyc/Gleipnir-Prefix-Teacher-Cache.prm800k_onpolicy_multiturn_rtgshape_prefix0.2_roll4_maxrev100prefixbench
PrefixBench JSONL Datasets
These datasets generate deterministic prompts for testing KV prefix caching behavior in LLM inference servers such as vLLM and SGLang. The prompts use controlled shared prefixes plus small unique suffixes so benchmark clients can compare cache reuse, latency, and throughput across server configurations.
Files
shared_schema_1k.jsonl: Simple shared-prefix benchmark. Every request reuses the same extraction instruction, JSON schema, and few-shot… See the full description on the dataset page: https://huggingface.co/datasets/jaytonde05/prefixbench.prm800k_onpolicy_multiturn_cumm_rew_prefix0.2_roll4_maxrev100prm800k_onpolicy_multiturn_seprew_prefix0.1_roll4_maxrev100prm800k_onpolicy_multiturn_cummrew_prefix0.2_roll4_maxrev100prm800k_onpolicy_multiturn_seprew_prefix0.2_roll4_maxrev100ogmath5_onpolicy_multiturn_seprew_prefix0.1_roll4_maxrev100ogmath5_onpolicy_multiturn_seprew_prefix0.2_roll4_maxrev100video_timestamps_prefixThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 3,
"total_frames": 180,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/video_timestamps_prefix.lm_eval_viewerprm800k_onpolicy_multiturn_cummrew_prefix0.1_roll4_maxrev100ogmath5_onpolicy_multiturn_cummrew_prefix0.2_roll4_maxrev100genvf-prefixes-v1-qwen35-only-suffixes_finalqwen3_sft_correct_saving_v6shaer-eval-raw-fanar-diwan-prefix
Shaer Evaluation Results
Models: fanar_2_diwan_prefix
Source dataset: Shaer-AI/shaer-sft-test-generations-k5
Rows: 3481
Validation passed: True
Scored rows included: True
Dataset repo: Shaer-AI/shaer-eval-raw-fanar-diwan-prefix
Files
generations.jsonl: raw generation rows
generations.csv: raw generation rows in CSV
generations_scored.jsonl: raw rows plus meter/count evaluation
validation.json: validation summary
generations_scored.csv: scored rows in CSV… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-fanar-diwan-prefix.ppo_pythia410m_tldr6.9b_rm410mdata_mergedsft_prefix_nokl_full_eval-datasetqwen3_sft_correct_v2rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft_prefix_eval-datasetqwen3_1.7B_sft_correct_v1_viewerqwen3_sft_correct_v6_saving_4096rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft_prefix_nokl_checkpoint-78_eval-datasetqwen3_sft_correct_v1
