tess
Datasets
All datasets matching “tess”minecraft-vla-stage1
Minecraft VLA Stage 1: Action Pretraining Data
Vision-Language-Action training data for Minecraft, processed from OpenAI's VPT contractor dataset.
Dataset Description
This dataset contains frame-action pairs from Minecraft gameplay, designed for training VLA models following the Lumine methodology.
Source
Original: OpenAI VPT Contractor Data (7.x subset)
Videos: 17,886 videos (330 hours of early-game gameplay)
Task: "Play Minecraft" with focus on first 30… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/minecraft-vla-stage1.details_migtissera__Tess-M-v1.3
Dataset Card for Evaluation run of migtissera/Tess-M-v1.3
Dataset automatically created during the evaluation run of model migtissera/Tess-M-v1.3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-M-v1.3.minecraft-vla-stage2
Minecraft VLA Stage 2: Instruction-Following Data
Stage 2 of the TESS-Minecraft Vision-Language-Action training pipeline.
Overview
This dataset adds task instructions to the Stage 1 visuomotor data, enabling instruction-following training.
Data Format
Field
Type
Description
id
string
Unique sample ID
video_id
string
Source video name
frame_idx
int
Frame index within video
instruction
string
Task instruction (empty for continuation frames)… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/minecraft-vla-stage2.mmu_tess_spoc
mmu_tess_spoc HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_tess_spoc.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_tess_spoc.TESSY-Code-80K
TESSY-Code-80K
📄 Paper Link
|
🔗 GitHub Repository
📣 Paper
🎉 Accepted at ICML 2026!
How to Fine-Tune a Reasoning Model? A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT Data
🚀 Overview
We construct a programming contest training dataset for Qwen3-8B by leveraging GPT-OSS-120B as the teacher model. The synthesized data preserves the strong reasoning capabilities of GPT-OSS-120B, while being aligned with the… See the full description on the dataset page: https://huggingface.co/datasets/CoopReason/TESSY-Code-80K.tess-agentnet
TESS AgentNet Dataset
Computer use trajectories for training Vision-Language-Action models.
Features
image: Screenshot (PIL Image)
instruction: Task description
action_type: 0=MOUSE, 1=KEYBOARD
mouse_x, mouse_y: Normalized coordinates [0,1]
click_type: 0-8 (NO_CLICK, LEFT_CLICK, etc.)
keyboard_text: Text with special tokens
os_type: ubuntu, windows_macos
episode_id, step_idx: Episode structure
Click Types
Index
Type
Description
0
NO_CLICK… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/tess-agentnet.
