datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CUB-200-2011-Nelicense: apache-2.0
strategic_game_cube
Cube
This dataset contains 1.64 billion Rubik's Cube solves, totaling roughly 236.39 billion moves.it is generated by Fugaku using https://github.com/trincaog/magiccube
Each solve has two columns: 'Cube' and 'Actions',
'Cube': initial scrambled states of a 3-3-3 cube in string, such as:
WOWWYOBWOOGWRBYGGOGBBRRYOGRWORBBYYORYBWRYBOGBGYGWWGRRY
the visual state of this example is
NOTICE: Crambled Cube States are spread out into the above string, row by row.
'Actions': list of… See the full description on the dataset page: https://huggingface.co/datasets/laion/strategic_game_cube.cub200_dataset
Dataset Card for CUB_200_2011
Dataset Summary
The Caltech-UCSD Birds 200-2011 dataset (CUB-200-2011) is an extended version of the original CUB-200 dataset, featuring photos of 200 bird species primarily from North America. This 2011 version significantly expands its predecessor by doubling the number of images per class and introducing new part location annotations, alongside collecting detailed natural language descriptions for each image through Amazon Mechanical Turk… See the full description on the dataset page: https://huggingface.co/datasets/cassiekang/cub200_dataset.lewm-cube
LeWM Cube
An archive-free Lance conversion of quentinll/lewm-cube, released with LeWorldModel.
Open in the LeWM Dataset Visualizer.
Episodes
10,000
Timesteps
2,010,000
Episode length
201
Frames
224 × 224 RGB
Action dimensions
5
Files
train.lance: indexed training table (19.07 GiB)
train_episodes.lance: episode metadata
viewer/*.parquet: row-identical Hub-viewer mirror (12.54 GiB)
The viewer mirror repeats episode metadata on each… See the full description on the dataset page: https://huggingface.co/datasets/fracapuano/lewm-cube.cubeamericas_nli
Dataset Card for AmericasNLI
Dataset Summary
AmericasNLI is an extension of XNLI (Conneau et al., 2018) a natural language inference (NLI) dataset covering 15 high-resource languages to 10 low-resource indigenous languages spoken in the Americas: Ashaninka, Aymara, Bribri, Guarani, Nahuatl, Otomi, Quechua, Raramuri, Shipibo-Konibo, and Wixarika. As with MNLI, the goal is to predict textual entailment (does sentence A imply/contradict/neither sentence B) and is a… See the full description on the dataset page: https://huggingface.co/datasets/nala-cub/americas_nli.OpenCodeReasoning-2
OpenCodeReasoning-2: A Large-scale Dataset for Reasoning in Code Generation and Critique
Dataset Description
OpenCodeReasoning-2 is the largest reasoning-based synthetic dataset to date for coding, comprising 1.4M samples in Python and 1.1M samples in C++ across 34,799 unique competitive programming questions.
OpenCodeReasoning-2 is designed for supervised fine-tuning (SFT) tasks of code completion and code critique.
Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/cublya/OpenCodeReasoning-2.CUB
CUB (Image Retrieval)
This repository contains the raw images for the CUB dataset, formatted for use with Dino v3 retrieval applications.
Usage in Frontend
You can access these images directly via the Hugging Face resolve endpoint:
https://huggingface.co/datasets/<USERNAME>/CUB/resolve/main/<PATH_TO_FILE>
CUB_train
Dataset Card for "CUB_train"
More Information needed
cubicasa5k-yolo
CubiCase5K in YOLO format (walls / doors / windows)
Converted from CubiCase5K (CC BY-NC 4.0):
SVG Wall/Railing, Door, Window polygons rasterized on F1_scaled.png
and traced back to YOLO polygon labels. Swing arcs are ignored.
Class 0 wall
Class 1 door (opening polygon only)
Class 2 window
Splits: 90/5/5 train/val/test (random, seed 0). Labels are polygon format,
usable for both detection and instance segmentation in ultralytics.
ogb_cube_singlecub-druid
Dataset Card for DRUID
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the DRUID dataset by Hagström et al. (2024). For this version, we have sampled 4,500 DRUID entries for which a "true target" (the factcheck verdict) and a "new target" (the stance of the context) could be found.
Dataset Structure
Thus far, we use two versions of the dataset: gpt2-xl and pythia-6.9b with corresponding validation (200 samples) and test splits… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-druid.CUB_test
Dataset Card for "CUB_test"
More Information needed
cub-counterfact
Dataset Card for CounterFact
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the popular CounterFact dataset, originally proposed by Meng et al. (2022) and re-used in different variants by e.g. Ortu et al. (2024). For this version, the 899 CounterFact samples have been sampled based on the parametric memory of Pythia 6.9B, such that it contains samples for which the top model prediction without context is correct. We note that 546 samples in the… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-counterfact.octo-small-simpler-cube-stack-rollout-bank-50
Octo-Small SIMPLER cube-stack rollout bank
This bank contains exactly 50 deterministic Octo-Small
rollouts for StackGreenCubeOnYellowCubeBakedTexInScene-v1: 3
successes and 47 failures. Every episode has a complete
61-frame H.264 video, a seven-frame contact sheet, losslessly stored per-step
telemetry, derived phase/geometry metrics, an external-LLM diagnosis, exact
evidence values, and a future action-patch hypothesis for failures.
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/lsnu/octo-small-simpler-cube-stack-rollout-bank-50.cubemap-8kcub-nq
Dataset Card for NQ
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the popular NQ dataset, originally proposed by Kwiatkowski et al. (2019). For this version, NQ samples have been obtained based on whether we can recover the gold passage from the original Wikipedia page and for which there is one short answer (less than 5 words in length). The context used in these samples is the correct gold context annotated by the original NQ annotators… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-nq.cubemaps_padding_16px_captioned_40kspatialtunnel
SpatialTunnel
SpatialTunnel is a Blender-rendered diagnostic dataset for studying how vision-language models represent spatial relations internally. It was introduced in Why Far Looks Up: Probing Spatial Representation in Vision-Language Models (arXiv:2605.30161).
Resources
Project page
Contrastive-probing code
SpatialTunnel generation code
Dataset Configurations
Config
File
Rows
Description
phase_variation
phase_variation-*.parquet
12… See the full description on the dataset page: https://huggingface.co/datasets/cubec/spatialtunnel.piperx-put-cube-in-drawer-20260908-87ep
PiperX: put the cube into the drawer — 2026-09-08
87 recorded episodes, 54,463 frames, 30 FPS, approximately 30.26 minutes of recorded frames.
Collected on 2026-09-08 in Asia/Shanghai; source file times span approximately 03:37–05:06.
Task: put the cube into the drawer.
Contents
LeRobot v3.0 layout, dual PiperX arms, 14-dimensional action and state vectors.
Three RGB camera streams (top, left wrist, right wrist) and corresponding depth PNG ZIP archives.
Additional… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/piperx-put-cube-in-drawer-20260908-87ep.splatter-cube-pbmc3k
Splatter cube: controlled scRNA-seq variants with known cluster structure
180 simulated datasets: 60 parameter points x 3 seeds, each 2,000 cells x 16,085 genes with a known number of clusters.
Generated with splatter, with baseline
parameters estimated from real reference data rather than chosen by hand.
Files
data/pP_sS.h5ad -- counts (CSR). obs["Group"] holds the ground-truth cluster label.
metrics.csv -- one row per simulation: parameters, realised sparsity… See the full description on the dataset page: https://huggingface.co/datasets/btraven/splatter-cube-pbmc3k.cubert_ETHPy150Open
CuBERT ETH150 Open Benchmarks
This is an unofficial HuggingFace upload of the CuBERT ETH150 Open Benchmarks. This dataset was released along with Learning and Evaluating Contextual Embedding of Source Code.
Benchmarks and Fine-Tuned Models
Here we describe the 6 Python benchmarks we created. All 6 benchmarks were derived from ETH Py150 Open. All examples are stored as sharded text files. Each text line corresponds to a separate example encoded as a JSON object. For each… See the full description on the dataset page: https://huggingface.co/datasets/claudios/cubert_ETHPy150Open.CUB-200-2011piper_pick_cube_resize_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"joint_6.pos",
"gripper.pos"
],
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/GabrielNing/piper_pick_cube_resize_merged.ipfs_cuba_laws
Cuba Gaceta Oficial / Minjus Laws
Research snapshot of official national legislation from Gaceta Oficial / Ministerio de Justicia (gacetaoficial.gob.cu, minjus.gob.cu).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-18
Coverage
catalog-backed incomplete
Source
Gaceta Oficial / Ministerio de Justicia (gacetaoficial.gob.cu, minjus.gob.cu)
Collector… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_cuba_laws.ipfs_cuba_laws_ir
Cuba legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_cuba_laws (revision cd431e5d0272164d5109a328b18fe934ff6ad04f) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Cuba prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_cuba_laws_ir.cub200_retrievalCUB-200-2011cubicassa5k-coco
CubiCasa5K (COCO format)
Instance-segmentation dataset of residential floor plans, converted to a
COCO-style schema and packaged as Parquet with embedded images. Each image is
annotated with polygon masks for architectural elements (walls, doors, windows,
rooms, fixtures, etc.).
Repository: phungpx/cubicassa5k-coco
Source dataset: CubiCasa5K
Format: COCO instance segmentation (polygon segmentation + bbox)
Modalities: image + structured annotations
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/phungpx/cubicassa5k-coco.jam-alt-lines
Jam-ALT Lines
Jam-ALT Lines is a line-level version of the Jam-ALT lyrics transcription dataset.
Unlike Jam-ALT, this dataset contains one audio segment for each lyrics line, facilitating research that considers each line as a separate unit.
[!tip]
See the Jam-ALT project website for details and the JamendoLyrics community for related datasets.
Dataset flavors
Lyrics lines may overlap in time, which makes it impossible to have a one-to-one correspondence between… See the full description on the dataset page: https://huggingface.co/datasets/cublya/jam-alt-lines.
