datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Laion-Aesthetics-High-Resolution-GoT
Laion-Aesthetics-High-Resolution-GoT
Paper
Dataset Description
The Laion-Aesthetics-High-Resolution-GoT dataset is a collection of 3.77 million image-text pairs with rich grounding annotations. This dataset extends high-quality images from the LAION-Aesthetics collection with detailed text descriptions and object-level grounding information.
Key Features
Size: 3.77 million samples
Modalities: Image, Text, and Grounding Annotations
Image Resolution:… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT.kraken-trading-data
📈 Kraken Trading Data Collection
Overview
High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis.
This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling.
📊 Included Trading Pairs
Pair
Asset
Base Currency
Typical Daily Volume
XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.got-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown).
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
12
-
Prompts: 7660
Format version: 1.1
Load with lmprobe
from lmprobe import pull_dataset, load_activation_dataset
# Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.got-activations-qwen2.5-0.5b
Qwen/Qwen2.5-0.5B — Activation Dataset
Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987).
Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries.
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-23
896
-
1
-
logits_topk
-
k=100
last_token
1
1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.OmniEdit-GoTgotsf-ds
📶 Beam-Level (5G) Time-Series Dataset
📚 Citation
This dataset is released alongside the following paper:
Fechete, L., et al. “Goal-Oriented Time-Series Forecasting: Foundation Framework Design.” Proceedings of the AAAI Conference on Artificial Intelligence, 2026, Singapore.
If you use this dataset, please cite the above work.
This dataset introduces a novel multivariate time series specifically curated to support research in enabling accurate prediction of KPIs… See the full description on the dataset page: https://huggingface.co/datasets/netop/gotsf-ds.got-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Geometry of Truth curated dataset activations for Llama 3.1 70B base
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-79
8192
-
4
-
Prompts: 7660
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.go_to_lego_test21This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 234,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Baptiste-le-Beaudry/go_to_lego_test21.BallPickup
BallPickup
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
gotriple-pretraining-dataset
GoTriple Pretraining dataset
Summary
The GoTriple Pre-training Dataset is a multilingual corpus built from open-access research artefacts harvested via the GoTriple platform. It focuses on Social Sciences and Humanities (SSH) content, addressing their limited presence in standard LLM pre-training corpora.
Current release includes History, Sociology, Environmental Sciences, Psychology and Geography texts (~23.14B tokens).
Intended Use
Continuous… See the full description on the dataset page: https://huggingface.co/datasets/odoma/gotriple-pretraining-dataset.laions-got-talent-annotatedwarp_Research
Warp Research Dataset
Dataset Description
Dataset Summary
This dataset contains experimental results from warp field research, focusing on the relationship between warp factors, energy efficiency, and field characteristics.
Supported Tasks
Tabular Regression: Predict energy efficiency based on warp field parameters
Time Series Forecasting: Analyze temporal patterns in warp field behavior
Optimization: Identify optimal warp factor configurations for… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/warp_Research.go_to_lego_test1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 840,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Baptiste-le-Beaudry/go_to_lego_test1.go_to_lego_test16This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 3,
"total_frames": 766,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Baptiste-le-Beaudry/go_to_lego_test16.go_to_lego_test22This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 102,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Baptiste-le-Beaudry/go_to_lego_test22.VALUE_cola_got
Dataset Card for "VALUE_cola_got"
More Information needed
goto_pick_ver1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 153,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/khmin101/goto_pick_ver1.go_to_lego_test20This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 157,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Baptiste-le-Beaudry/go_to_lego_test20.go_to_lego_test79This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi_client",
"total_episodes": 2,
"total_frames": 1914,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Baptiste-le-Beaudry/go_to_lego_test79.goto_place_ver1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 307,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/khmin101/goto_place_ver1.gotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 50,
"total_frames": 87712,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yaswanth8390/got.VALUE_wikitext2_got
Dataset Card for "VALUE_wikitext2_got"
More Information needed
gothic_bible_got
Gothic Bible (4th Century)
Description
The Gothic Bible is the earliest known translation of the Bible into a Germanic language, made by Bishop Wulfila (Ulfilas, c. 311-383 AD) in the 4th century. Wulfila created the Gothic alphabet specifically for this translation. The surviving portions consist primarily of the Gospels (the Codex Argenteus) and parts of the Epistles and Old Testament. This is the oldest substantial text in any Germanic language and is of… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/gothic_bible_got.Got_Science_28kMULTI_VALUE_mnli_got_gotten
Dataset Card for "MULTI_VALUE_mnli_got_gotten"
More Information needed
GoToCompany__gemma2-9b-cpt-sahabatai-v1-instruct-details
Dataset Card for Evaluation run of GoToCompany/gemma2-9b-cpt-sahabatai-v1-instruct
Dataset automatically created during the evaluation run of model GoToCompany/gemma2-9b-cpt-sahabatai-v1-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/GoToCompany__gemma2-9b-cpt-sahabatai-v1-instruct-details.VALUE_wnli_got
Dataset Card for "VALUE_wnli_got"
More Information needed
MULTI_VALUE_mrpc_existential_got
Dataset Card for "MULTI_VALUE_mrpc_existential_got"
More Information needed
VALUE_qqp_got
Dataset Card for "VALUE_qqp_got"
More Information needed
MULTI_VALUE_wnli_existential_got
Dataset Card for "MULTI_VALUE_wnli_existential_got"
More Information needed
