datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
so100_popcorn_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 5,
"total_frames": 12356,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samanthalhy/so100_popcorn_1.PDM-Lite-DVS
PDM-Lite-DVS
PDM-Lite-DVS is an independently collected synthetic CARLA 0.9.15 event-camera
dataset generated with the rule-based PDM-Lite expert on route configurations
published by carla_garage. The release lineage is 5,545 public route XMLs →
5,503 recordings in the frozen local pool → 257 rejected recordings → 5,246
published recordings. “Public route release” refers only to the route
configurations: the sensor measurements are an independent collection. This
is not the… See the full description on the dataset page: https://huggingface.co/datasets/SamanthaZhang/PDM-Lite-DVS.LEAD-DVS
LEAD-DVS
LEAD expert-driving data collected in CARLA 0.9.15 with synchronized RGB,
depth, semantic/instance segmentation, LiDAR, radar, HD map, metadata, 3D
bounding boxes, and a forward-facing DVS event camera.
Code release
Dataset construction, preprocessing, training, and evaluation code will be
published in SamanthaZhang-stu/ReflexWorldModel.
This GitHub repository is the designated code-release location for the project.
Contents
1,579… See the full description on the dataset page: https://huggingface.co/datasets/SamanthaZhang/LEAD-DVS.so100_popcorn_2_transportThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 516,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samanthalhy/so100_popcorn_2_transport.eval_so100_smol_popcorn_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 2,
"total_frames": 9011,
"total_tasks": 1,
"total_videos": 8,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samanthalhy/eval_so100_smol_popcorn_1.so100_popcorn_2_gatherThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 767,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samanthalhy/so100_popcorn_2_gather.eval_so100_smol_popcorn_1_retryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 3972,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samanthalhy/eval_so100_smol_popcorn_1_retry.details_macadeliccc__Samantha-Qwen-2-7B
Dataset Card for Evaluation run of macadeliccc/Samantha-Qwen-2-7B
Dataset automatically created during the evaluation run of model macadeliccc/Samantha-Qwen-2-7B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_macadeliccc__Samantha-Qwen-2-7B.samantha-r01-recursive-reasoning-corpus
Samantha R01 Recursive Reasoning Corpus
Answer-only corpus for the first isolated Samantha silent-tick / recursive-latent-reasoning validation. It is normalized for pre_train_recursive_reasoning.py and intentionally contains no visible chain-of-thought or source rationale fields.
The repository is private because it combines sources with mixed or unspecified redistribution terms. Access does not supersede any upstream license.
Split policy
train: 50,000… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/samantha-r01-recursive-reasoning-corpus.representational_stability
Dataset Card for Representational Stability Fictional Data
Dataset Summary
The Representational Stability fictional dataset is made to supplement the
Trilemma of Truth dataset (here).
The Trilemma of Truth data contains three types of statements:
Factually true statements
Factually false statements
Synthetic, neither-valued statements generated to mimic statements unseen during LLM training
The Representational Stability fictional dataset adds new types of statements:… See the full description on the dataset page: https://huggingface.co/datasets/samanthadies/representational_stability.macadeliccc__Samantha-Qwen-2-7B-details
Dataset Card for Evaluation run of macadeliccc/Samantha-Qwen-2-7B
Dataset automatically created during the evaluation run of model macadeliccc/Samantha-Qwen-2-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/macadeliccc__Samantha-Qwen-2-7B-details.uukuguy__speechless-mistral-dolphin-orca-platypus-samantha-7b-details
Dataset Card for Evaluation run of uukuguy/speechless-mistral-dolphin-orca-platypus-samantha-7b
Dataset automatically created during the evaluation run of model uukuguy/speechless-mistral-dolphin-orca-platypus-samantha-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/uukuguy__speechless-mistral-dolphin-orca-platypus-samantha-7b-details.
