CoolFace
Datasetpublic

TESS-Computer/tess-atari-15hz-384

TESS-Atari Stage 1 - Preprocessed (15Hz, 384x384) Training-ready version of the 15Hz dataset with images pre-resized to 384x384 (SmolVLM native resolution). Overview Metric Value Source TESS-Computer/atari-vla-stage1-15hz Samples 1,340,293 Image Size 384x384 (pre-resized) Action Rate 15 Hz (3 actions per observation) Format Lumine-style action tokens Why Preprocessed? Training VLMs requires resizing images to the model's… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/tess-atari-15hz-384.

sourceHugging Facemitupdated 9mo agoView on Hugging Face
0likes116downloads
Dataset Card

TESS-Atari Stage 1 - Preprocessed (15Hz, 384x384)

Training-ready version of the 15Hz dataset with images pre-resized to 384x384 (SmolVLM native resolution).

Overview

MetricValue
SourceTESS-Computer/atari-vla-stage1-15hz
Samples1,340,293
Image Size384x384 (pre-resized)
Action Rate15 Hz (3 actions per observation)
FormatLumine-style action tokens

Why Preprocessed?

Training VLMs requires resizing images to the model's native resolution. Doing this on-the-fly creates a CPU bottleneck. This dataset has images already resized, giving ~10x faster training:

Raw dataset:     160x210 → resize during training → slow (CPU bound)
Preprocessed:    384x384 → ready to use → fast (GPU saturated)

Action Format

<|action_start|> RIGHT ; RIGHT ; FIRE <|action_end|>
<|action_start|> LEFT ; LEFT ; LEFT <|action_end|>
<|action_start|> NOOP ; UP ; UPFIRE <|action_end|>

Schema

FieldTypeDescription
image_bytesbytesPNG at 384x384 (pre-resized)
actionstringLumine-style chunked action token
gamestringGame name
trial_idintHuman player trial number
frame_idxintFrame index in trial
image_sizeintAlways 384

Usage

python
from datasets import load_dataset
from PIL import Image
from io import BytesIO

# Load preprocessed dataset
ds = load_dataset("TESS-Computer/tess-atari-15hz-384", split="train")

# Images are already 384x384 - no resizing needed!
sample = ds[0]
img = Image.open(BytesIO(sample["image_bytes"]))
print(img.size)  # (384, 384)
print(sample["action"])  # <|action_start|> LEFT ; LEFT ; LEFT <|action_end|>

Training

bash
python scripts/train_v2.py \
    --preprocessed TESS-Computer/tess-atari-15hz-384 \
    --epochs 3 \
    --batch-size 4 \
    --grad-accum 32 \
    --wandb \
    --push-to-hub

Related

Citation

bibtex
@misc{tessatari2025,
  title={TESS-Atari: Vision-Language-Action Models for Atari Games},
  author={Lezzaik, Hussein},
  year={2025},
  url={https://github.com/HusseinLezzaik/TESS-Atari}
}

@misc{atarihead2019,
  title={Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset},
  author={Zhang, Ruohan and others},
  year={2019},
  url={https://zenodo.org/records/3451402}
}