datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_migtissera__Tess-M-v1.3
Dataset Card for Evaluation run of migtissera/Tess-M-v1.3
Dataset automatically created during the evaluation run of model migtissera/Tess-M-v1.3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-M-v1.3.mmu_tess_spoc
mmu_tess_spoc HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_tess_spoc.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_tess_spoc.atari-vla-stage1-5hz
TESS-Atari Stage 1 (5Hz)
Human gameplay demonstrations from Atari games, formatted for Vision-Language-Action (VLA) model training.
Overview
Metric
Value
Source
Atari-HEAD
Games
11 (overlapping with DIAMOND benchmark)
Samples
~4M
Action Rate
5 Hz (1 action per observation)
Format
Lumine-style action tokens
Games Included
Alien, Asterix, BankHeist, Breakout, DemonAttack, Freeway, Frostbite, Hero, MsPacman, RoadRunner, Seaquest… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/atari-vla-stage1-5hz.atari-vla-stage1-15hz
TESS-Atari Stage 1 (15Hz)
Human gameplay demonstrations from Atari games with action chunking, formatted for Vision-Language-Action (VLA) model training.
Overview
Metric
Value
Source
Atari-HEAD
Games
11 (overlapping with DIAMOND benchmark)
Samples
~1.3M
Observation Rate
5 Hz
Action Rate
15 Hz (3 actions per observation)
Format
Lumine-style action tokens
Why Action Chunking?
VLA models run at ~5 Hz inference speed, but Atari runs at… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/atari-vla-stage1-15hz.csgo-vla-stage1-5hz
CS:GO VLA Stage 1 Dataset (5Hz Chunked)
Vision-Language-Action dataset for Counter-Strike: Global Offensive with action chunking, converted from the TeaPearce CS:GO dataset.
Overview
Frame rate: 5Hz (every 3rd frame)
Action chunking: 3 actions per sample (~200ms coverage)
Total samples: ~1.8M chunks
Split: train / test following Diamond split
Map: Dust2 deathmatch
Action Format
<|action_start|> m1_x m1_y [keys1] ; m2_x m2_y [keys2] ; m3_x m3_y [keys3]… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/csgo-vla-stage1-5hz.mmu-norm-tess
TESS-SPOC light curves — reacquired L1 (release v1)
This L1 repository contains 31,058 per-sector light curves for 5,482 matched TIC objects. The data were reacquired from the TESS-SPOC HLSP at STScI (archive.stsci.edu/hlsps/tess-spoc), with every available sector stored separately alongside the native QUALITY bits and float64 times.
Schema (one row per star-sector)
Column
Meaning
tic
TIC identifier
sector
TESS sector — sectors are separate rows… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-tess.tess-atari-5hz-384tessera-experimentos
tessera-experimentos
Resultados dos experimentos sobre
tessera-extraidollm-gpt-5-mini-corrigido:
a mesma questão da OAB apresentada ao modelo com e sem a norma aplicável no contexto.
As condições
condição
o que vai no contexto
sem_lei
nada — a linha de base
com_lei_certa
só a norma da alternativa correta
com_lei_total
todas as normas transcritas na questão
com_lei_errada
a norma de outra questão — o controle
O controle é o que faz o resto… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/tessera-experimentos.tess-toi-candidates
TESS Objects of Interest (TOI) Planet Candidates
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Planet candidates identified by NASA's Transiting Exoplanet Survey Satellite (TESS), from the NASA Exoplanet Archive TOI catalog. Updated weekly.
TESS is a NASA space telescope launched in 2018 that surveys the entire sky for transiting exoplanets. When a star shows periodic brightness dips consistent with a planet… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/tess-toi-candidates.hyperliquid-ohlcv-1m
Tessera Analytics Hyperliquid Order-Flow OHLCV (1-minute) — free sample
A growing monthly sample of gold_ohlcv_1m, the
order-flow-enriched 1-minute OHLCV dataset served by Tessera Analytics:
open / high / low / close and volume for every minute, plus the order-flow context raw candles can't
show — the aggressor buy/sell volume split, cumulative volume delta (CVD), distinct taker counts,
taker fees and realized PnL, and the hourly forward-filled funding rate.
Coverage: BTC, ETH… See the full description on the dataset page: https://huggingface.co/datasets/tessera-analytics/hyperliquid-ohlcv-1m.tessera-resultados-tabelas
tessera — tabelas de resultado (com e sem RAG)
As tabelas do relatório em formato tabular, uma por arquivo, .parquet e .csv.
São os resultados de apresentar a mesma questão da OAB com e sem a norma aplicável
no contexto.
Gerações cruas, métricas por modelo e o relatório completo estão no dataset irmão:
juliadollis/tessera-experimentos.
Dataset de origem das questões:
juliadollis/tessera-extraidollm-gpt-5-mini-corrigido.
O desenho
366 questões × 9 modelos × 4… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/tessera-resultados-tabelas.tessera-8a92b237
tessera — corpus epoch 13, full sweep
Teacher-anchored SFT data harvested from every published Affine (Bittensor
SN120) duel scored against corpus epoch 13 — 230 duel records, chal-00760
through chal-01102, covering 2026-08-16 to 2026-08-24.
42,006 rows over 42,006 distinct turns (one row per turn), drawn from
4,981 trajectories and 3,751 strata. That is 70% of the 59,745-turn epoch-13
corpus, and 2.3× the 18,138 rows of
iamPi/tessera-77d11909,
which sampled a subset of the same… See the full description on the dataset page: https://huggingface.co/datasets/iamPi/tessera-8a92b237.tess-atari-15hz-384
TESS-Atari Stage 1 - Preprocessed (15Hz, 384x384)
Training-ready version of the 15Hz dataset with images pre-resized to 384x384 (SmolVLM native resolution).
Overview
Metric
Value
Source
TESS-Computer/atari-vla-stage1-15hz
Samples
1,340,293
Image Size
384x384 (pre-resized)
Action Rate
15 Hz (3 actions per observation)
Format
Lumine-style action tokens
Why Preprocessed?
Training VLMs requires resizing images to the model's native… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/tess-atari-15hz-384.tess---
description: 'TESS Light Curves From Full Frame Images ("TESS-SPOC")
'
homepage: https://archive.stsci.edu/hlsp/tess-spoc
version: 1.0.0
citation: "% % ACKNOWLEDGEMENT\n% % From: https://archive.stsci.edu/publishing/mission-acknowledgements\n\
% This paper includes data collected with the TESS mission, obtained from the MAST \ data archive at the Space Telescope Science Institute (STScI). Funding for the \ TESS mission is provided by the NASA Explorer Program. STScI is operated by… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/tess.CIVQA-TesseractOCR-LayoutLM
CIVQA TesseractOCR LayoutLM Dataset
The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR and encoded for the LayoutLM.
The pre-encoded dataset can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR
All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for processing the invoices.
Invoice number
Variable… See the full description on the dataset page: https://huggingface.co/datasets/SpringRollMonster/CIVQA-TesseractOCR-LayoutLM.csgo-vla-stage1-16hz
CS:GO VLA Stage 1 Dataset (16Hz)
Vision-Language-Action dataset for Counter-Strike: Global Offensive, converted from the TeaPearce CS:GO dataset.
Overview
Frame rate: 16Hz (native, 1 action per frame)
Total samples: ~5.5M frames
Split: train (5M) / test (500K) following Diamond split
Map: Dust2 deathmatch
Action Format
<|action_start|> mouse_x mouse_y [keys] <|action_end|>
Examples:
<|action_start|> 0 0 <|action_end|> # idle… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/csgo-vla-stage1-16hz.migtissera__Tess-v2.5.2-Qwen2-72B-details
Dataset Card for Evaluation run of migtissera/Tess-v2.5.2-Qwen2-72B
Dataset automatically created during the evaluation run of model migtissera/Tess-v2.5.2-Qwen2-72B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/migtissera__Tess-v2.5.2-Qwen2-72B-details.migtissera__Tess-3-Mistral-Nemo-12B-details
Dataset Card for Evaluation run of migtissera/Tess-3-Mistral-Nemo-12B
Dataset automatically created during the evaluation run of model migtissera/Tess-3-Mistral-Nemo-12B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/migtissera__Tess-3-Mistral-Nemo-12B-details.carla-simlingo-raw
SimLingo CARLA Dataset (Raw, 4Hz)
Raw driving data from CARLA simulator. No transformations or derived fields - all original measurements preserved as-is.
Dataset Summary
Source: SimLingo (CVPR 2025)
Scale: 228,757 frames (23 shards)
Frame Rate: 4 FPS
Resolution: 1024x512 RGB
Routes: Complete driving episodes (routes never split across shards)
Column Schema
Core Fields
Column
Type
Description
route_id
string
Route identifier
frame_idx… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/carla-simlingo-raw.tessera-77d11909details_migtissera__Tess-v2.5-Gemma-2-27B-alpha
Dataset Card for Evaluation run of migtissera/Tess-v2.5-Gemma-2-27B-alpha
Dataset automatically created during the evaluation run of model migtissera/Tess-v2.5-Gemma-2-27B-alpha.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-v2.5-Gemma-2-27B-alpha.details_migtissera__Tess-2.0-Llama-3-70B
Dataset Card for Evaluation run of migtissera/Tess-2.0-Llama-3-70B
Dataset automatically created during the evaluation run of model migtissera/Tess-2.0-Llama-3-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-2.0-Llama-3-70B.neopets-packdetails_migtissera__Tess-v2.5.2-Qwen2-72B
Dataset Card for Evaluation run of migtissera/Tess-v2.5.2-Qwen2-72B
Dataset automatically created during the evaluation run of model migtissera/Tess-v2.5.2-Qwen2-72B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-v2.5.2-Qwen2-72B.tess-atari-asterix-15hz-384
TESS-Atari: Asterix (15Hz, 384x384)
Single-game preprocessed dataset for VLA training.
Overview
Metric
Value
Game
Asterix
Samples
41,646
Image Size
384x384
Action Rate
15 Hz (3 actions per observation)
Format
Lumine-style action tokens
Filters Applied
score > 0 - Active gameplay only (no menus/idle)
No pure NOOP - Player actually taking actions
Action Format
<|action_start|> RIGHT ; UP ; UPRIGHT <|action_end|>
Usage… See the full description on the dataset page: https://huggingface.co/datasets/TESS-Computer/tess-atari-asterix-15hz-384.details_migtissera__Tess-3-Mistral-Nemo-12B
Dataset Card for Evaluation run of migtissera/Tess-3-Mistral-Nemo-12B
Dataset automatically created during the evaluation run of model migtissera/Tess-3-Mistral-Nemo-12B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-3-Mistral-Nemo-12B.tesstess---
description: 'TESS Light Curves From Full Frame Images ("TESS-SPOC")
'
homepage: https://archive.stsci.edu/hlsp/tess-spoc
version: 0.0.1
citation: "% % ACKNOWLEDGEMENT\n% % From: https://archive.stsci.edu/publishing/mission-acknowledgements\n\
% This paper includes data collected with the TESS mission, obtained from the MAST \ data archive at the Space Telescope Science Institute (STScI). Funding for the \ TESS mission is provided by the NASA Explorer Program. STScI is operated by… See the full description on the dataset page: https://huggingface.co/datasets/EiffL/tess.tessera-0f0e8791
tessera (cleaned)
Format-normalized derivative of iamPi/tessera-77d11909.
Same rows, same prompts; the completions have been repaired to the Affine
SWE turn contract and given a fixed action cue.
Completion format
Every completion is exactly:
</think>
THOUGHT: {thought}
The analysis is complete. Next command:
```bash
{command}
The prompt ends inside an open `<think>` block, so the completion opens with
`</think>`. `thought` is plain prose — no tags, no second… See the full description on the dataset page: https://huggingface.co/datasets/iamPi/tessera-0f0e8791.details_migtissera__Tess-v2.5-Qwen2-72B
Dataset Card for Evaluation run of migtissera/Tess-v2.5-Qwen2-72B
Dataset automatically created during the evaluation run of model migtissera/Tess-v2.5-Qwen2-72B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-v2.5-Qwen2-72B.
