CoolFace
Datasetpublic

cagataydev/neon-spatial-language-20k

neon-spatial-language-20k Spatial language grounding — 20K episodes of synthetic 14-DoF joint trajectories for Neon VLA training. Description Spatial reasoning with relative directions (left, right, above, near) Each episode contains: language_instruction: Natural language task description actions: JSON array of joint position trajectories (T × 14) length: Episode length (timesteps) Usage import pyarrow.parquet as pq table =… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/neon-spatial-language-20k.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes14downloads
Dataset Card

neon-spatial-language-20k

Spatial language grounding — 20K episodes of synthetic 14-DoF joint trajectories for Neon VLA training.

Description

Spatial reasoning with relative directions (left, right, above, near)

Each episode contains:

  • —language_instruction: Natural language task description
  • —actions: JSON array of joint position trajectories (T × 14)
  • —length: Episode length (timesteps)

Usage

python
import pyarrow.parquet as pq
table = pq.read_table("data.parquet")
df = table.to_pandas()
print(f"Episodes: {len(df)}, Columns: {list(df.columns)}")

Part of Neon VLA

Training data for Neon — open-source Vision-Language-Action model for humanoid whole-body control.

Total Neon dataset collection: 160K episodes across 8 datasets (~1.27GB)