datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nt_multispecies_4096plantcad2-c4096distill-r1-qwen-1.5b-aime-24-4096-with-labels-prmdistill-r1-qwen-1.5b-hmmt-feb-25-4096-with-bt-model-with-sigmoiddistill-r1-qwen-1.5b-aime-24-4096-with-bt-model-wout-sigmoiddistill-r1-qwen-1.5b-hmmt-feb-24-4096-with-bt-model-wout-sigmoidnemotron_cc_v2_hq_packed4096_200shard
Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train)
Documents from nvidia/Nemotron-CC-v2
High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS
appended per document) and greedily packed into sequences of at most 4096 tokens.
A document is never split across a pack boundary; documents longer than 4096
are truncated to their own pack. Every pack ends on an EOS/document boundary.
Schema
index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.LoRa_4096blockfineweb-sample-100BT_over-4096-tokensgeometric-vocab-4096ddistill-r1-qwen-1.5b-aime-24-4096-with-bt-model-with-sigmoid4B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxivRULER-4096-llama-3.2-tokenizerRULER-4096-Qwen-3cut_4096_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.pointcloud": {
"dtype": "float32",
"shape": [
4096,
4
],
"names": [
"N_points",
"XYZI"
]
},
"observation.state": {
"dtype": "float32",
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/escapebirdy/cut_4096_v2.distill-r1-qwen-1.5b-aime-25-4096-with-bt-model-wout-sigmoid4096_1Mdistill-r1-qwen-1.5b-hmmt-feb-24-4096-with-bt-model-with-sigmoidcut_4096_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.pointcloud": {
"dtype": "float32",
"shape": [
4096,
4
],
"names": [
"N_points",
"XYZI"
]
},
"observation.state": {
"dtype": "float32",
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/escapebirdy/cut_4096_v3.nemotron_cc_v2_hq_packed4096
Nemotron-CC-v2 High-Quality, packed to 4096 tokens
5% subset of nvidia/Nemotron-CC-v2
High-Quality documents, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS
appended per document) and greedily packed into sequences of at most 4096 tokens.
A document is never split across a pack boundary; documents longer than 4096
are truncated to their own pack. Every pack ends on an EOS/document boundary.
Schema
index (int64): running pack id
input_ids… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096.llp-gold-37m-1.5m_T4096.0pick_and_place_mix_4096This dataset was created using LeRobot.
Dataset Description
QuestArmTeleop LeRobot v3 dataset. Point clouds are stored directly in Parquet as observation.pointcloud.xyz (float32 metres, shape [1024, 3], base_link), observation.pointcloud.rgb (uint8 RGB), and observation.pointcloud.valid_mask. They share the same row, frame index and timestamp as state/action.
Homepage: [More Information Needed]
Paper: [More Information Needed]
License: apache-2.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Keith-Luo/pick_and_place_mix_4096.pick_and_place_mix_append_4096This dataset was created using LeRobot.
Dataset Description
QuestArmTeleop LeRobot v3 dataset. Point clouds are stored directly in Parquet as observation.pointcloud.xyz (float32 metres, shape [1024, 3], base_link), observation.pointcloud.rgb (uint8 RGB), and observation.pointcloud.valid_mask. They share the same row, frame index and timestamp as state/action.
Homepage: [More Information Needed]
Paper: [More Information Needed]
License: apache-2.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Keith-Luo/pick_and_place_mix_append_4096.RULER-4096-llama-3.1-tokenizer-chat-templateselect_block_4096_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.pointcloud": {
"dtype": "float32",
"shape": [
4096,
4
],
"names": [
"N_points",
"XYZI"
]
},
"observation.state": {
"dtype": "float32",
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/escapebirdy/select_block_4096_v3.RULER-4096-Lamma3-InstructRULER-4096-Falcon-H1-3B-Base4B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxiv.qwen3-ids.kv-tags-explainedpolaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps
Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507
Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's
actual continuation for each, and the token accounting behind it.
The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is
decoded to text, the teacher is shown it under its own chat template, and the teacher's
reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.svae-freckles-4096-cifar10
SVAE Freckles 4096 — CIFAR-10 Omega Tokens
Precomputed spectral decomposition of CIFAR-10 through Freckles v41 (256×256), a frozen Spectral Variational Autoencoder trained exclusively on synthetic noise.
Each CIFAR-10 image is resized to 256×256, decomposed into 4096 patches (4×4 each), and passed through Freckles' encoder → SVD bottleneck. The 4 singular values per patch are stored as a (4, 64, 64) omega map — a 4-channel spatial representation of spectral energy.… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/svae-freckles-4096-cifar10.
