datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
salabs-virtual-spatial-digitaltwin-v8
🌐 SALabs 10,000,000-Node 3D Virtual Spatial & Digital Twin Avatar Kinematics Dataset (v8.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,000 USD) & Instant 8.0GB Master DownloadInstant download of the complete 8.0GB master archive containing 10,000,000 verified 3D spatial nodes, 18-DoF avatar kinematics, B-spline 4D motion tensors, Laplace-Beltrami spectral resonance, and commercial license certificate.
🌟 Executive Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-virtual-spatial-digitaltwin-v8.lm-eval-results-bobofrut-ladybird-base-7B-v8-private
Dataset Card for Evaluation run of bobofrut/ladybird-base-7B-v8
Dataset automatically created during the evaluation run of model bobofrut/ladybird-base-7B-v8
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bobofrut-ladybird-base-7B-v8-private.lm-eval-results-zhengr-MixTAO-7Bx2-MoE-v8.1-private
Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1
Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-zhengr-MixTAO-7Bx2-MoE-v8.1-private.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details.zhengr__MixTAO-7Bx2-MoE-v8.1-details
Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1
Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zhengr__MixTAO-7Bx2-MoE-v8.1-details.pankajmathur__orca_mini_v8_1_70b-details
Dataset Card for Evaluation run of pankajmathur/orca_mini_v8_1_70b
Dataset automatically created during the evaluation run of model pankajmathur/orca_mini_v8_1_70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/pankajmathur__orca_mini_v8_1_70b-details.Lil-R__PRYMMAL-ECE-7B-SLERP-V8-details
Dataset Card for Evaluation run of Lil-R/PRYMMAL-ECE-7B-SLERP-V8
Dataset automatically created during the evaluation run of model Lil-R/PRYMMAL-ECE-7B-SLERP-V8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__PRYMMAL-ECE-7B-SLERP-V8-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.7-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.7
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.7
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.7-details.sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg
sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg.sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg
sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg.sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg
sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg.sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg
sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg.sokoban_hard_allstep_thinking_cot_v8_res256_stop_prompt
Sokoban hard allstep CoT v8 res256 stop-prompt rebuild
This local dataset rebuilds the HF hard allstep-thinking CoT source against the v8 stop-required 256x256 state-replay hard train source. Rows removed by v8 layout dedup are skipped. Output shards are sorted to match the v8 hard train batch filenames and row order.
Summary: metadata/rebuild_summary.json
sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k1_world_model_20260622_perseg.sokoban_easy_v8_cot_chunk_kinf_forward_20260710_fwdpure
sokoban_easy_v8_cot_chunk_kinf_forward_20260710_fwdpure
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_kinf_forward_20260710_fwdpure.sokoban_allstep_thinking_cot_v8_res256_stop_prompt_20260603_041108
Sokoban Allstep Thinking CoT V8 Res256 Stop Prompt
This repository contains the single-shard gzip package for the rebuilt Sokoban all-step CoT SFT data.
Packaging
Packaged/uploaded from: /data/home/raychai/hf_datasets/sokoban_allstep_thinking_cot_v8_res256_stop_prompt_20260603_041108
Single shard: training/sokoban_allstep_thinking_cot_v8_res256_stop_prompt_20260603_041108.jsonl.gz
Shard policy: concatenated gzip multi-member stream of training/batch_*.jsonl.gz… See the full description on the dataset page: https://huggingface.co/datasets/novastar113/sokoban_allstep_thinking_cot_v8_res256_stop_prompt_20260603_041108.ylalain__ECE-PRYMMAL-YL-1B-SLERP-V8-details
Dataset Card for Evaluation run of ylalain/ECE-PRYMMAL-YL-1B-SLERP-V8
Dataset automatically created during the evaluation run of model ylalain/ECE-PRYMMAL-YL-1B-SLERP-V8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ylalain__ECE-PRYMMAL-YL-1B-SLERP-V8-details.T145__KRONOS-8B-V8-details
Dataset Card for Evaluation run of T145/KRONOS-8B-V8
Dataset automatically created during the evaluation run of model T145/KRONOS-8B-V8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/T145__KRONOS-8B-V8-details.sokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k3_world_model_20260622_perseg.sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k5_world_model_20260622_perseg.jaspionjader__Kosmos-EVAA-v8-8B-details
Dataset Card for Evaluation run of jaspionjader/Kosmos-EVAA-v8-8B
Dataset automatically created during the evaluation run of model jaspionjader/Kosmos-EVAA-v8-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaspionjader__Kosmos-EVAA-v8-8B-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.5-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.5
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.5-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.6-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.6
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.6
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.6-details.sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg
sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_k10_world_model_20260622_perseg.Weyaxi__Einstein-v8-Llama3.2-1B-details
Dataset Card for Evaluation run of Weyaxi/Einstein-v8-Llama3.2-1B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v8-Llama3.2-1B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Weyaxi__Einstein-v8-Llama3.2-1B-details.mixtao__MixTAO-7Bx2-MoE-v8.1-details
Dataset Card for Evaluation run of mixtao/MixTAO-7Bx2-MoE-v8.1
Dataset automatically created during the evaluation run of model mixtao/MixTAO-7Bx2-MoE-v8.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mixtao__MixTAO-7Bx2-MoE-v8.1-details.sometimesanotion__Qwentinuum-14B-v8-details
Dataset Card for Evaluation run of sometimesanotion/Qwentinuum-14B-v8
Dataset automatically created during the evaluation run of model sometimesanotion/Qwentinuum-14B-v8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwentinuum-14B-v8-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.8-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.8
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.8-details.sokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg
sokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg
Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT
VLM-Gym feedback-interval study.
Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are
base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think>
then the committed action chunk; between chunks a loss-0 "Action executed." + real frame
(GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k1_world_model_20260622_perseg.
