datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cs336-basics-collection
CS336 Assignment 1 — Pre-tokenized Data & BPE Tokenizers
This repository contains preprocessing artifacts produced for Stanford CS336: Language Modeling from Scratch, Spring 2025 — Assignment 1: Basics.
It includes:
pre-tokenized TinyStories train/validation data,
pre-tokenized OpenWebText (OWT sample) train/validation data,
byte-level BPE vocabularies and merge tables for both datasets.
The main purpose of this repository is to avoid repeating the relatively expensive… See the full description on the dataset page: https://huggingface.co/datasets/victorhu493/cs336-basics-collection.xlerobot-final-v1-xlerobot_mobile_basics_h264_v1This dataset was created using LeRobot.
Dataset Description
XLeRobot Final v1 whole-body teleoperation smoke test
A public, real-hardware integration record for an XLeRobot-compatible 0.4.0
two-wheel configuration. This is an engineering smoke-test dataset, not a
household-task training benchmark.
Recorded system
Two Seeed Studio Pro SO-101 arms (leader/follower teleoperation)
Differential two-wheel base
Pan/tilt head
Three synchronized RGB… See the full description on the dataset page: https://huggingface.co/datasets/NOJIMA21/xlerobot-final-v1-xlerobot_mobile_basics_h264_v1.stock_market_basics
