datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laions_got_talent
LAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel.
Dataset Composition
The dataset includes:
Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.warp-speed
WarpSpeed Research Dataset
Dataset Summary
The WarpSpeed Research Dataset is a comprehensive collection of scientific research papers, experimental data, and theoretical materials focused on advanced propulsion concepts and physics principles inspired by Star Trek technologies. This dataset combines real-world physics research with theoretical frameworks to explore the possibilities of faster-than-light travel and advanced energy systems.
Data Collection and… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/warp-speed.laions_got_talent_enhanced_flash_annotations_and_long_captionsgot10klaions_got_talent_rawlaions_got_talent_enhanced_just_flash_annotationslaions_got_talent_pluslaions-got-talent-shuffled-with-long-captionslaions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.gotoubunnohanayome
Bangumi Image Base of Gotoubun No Hanayome
This is the image base of bangumi Gotoubun no Hanayome, we detected 134 characters, 16632 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/gotoubunnohanayome.chub-legacylaions_got_talent_enhanced_no_metadatalaions_got_talent_embs_only
laions_got_talent Whisper Embeddings (Embeddings + Metadata Only)
This dataset contains Whisper embeddings (NPY) and metadata (JSON). The original audio files (MP3) are NOT included.
Embeddings computed with: mkrausio/EmoWhisper-AnS-Small-v0.1
Includes original audio: No
Includes metadata: Yes (JSON)
Includes embeddings: Yes (NPY)
Creation date: 2025-05-11
laions_got_talent_german_bicodecLaion-Aesthetics-High-Resolution-GoT
Laion-Aesthetics-High-Resolution-GoT
Paper
Dataset Description
The Laion-Aesthetics-High-Resolution-GoT dataset is a collection of 3.77 million image-text pairs with rich grounding annotations. This dataset extends high-quality images from the LAION-Aesthetics collection with detailed text descriptions and object-level grounding information.
Key Features
Size: 3.77 million samples
Modalities: Image, Text, and Grounding Annotations
Image Resolution:… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT.laions-got-talent_wordlevel-annotationlaions_got_talent_clean_with_captionskraken-trading-data
📈 Kraken Trading Data Collection
Overview
High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis.
This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling.
📊 Included Trading Pairs
Pair
Asset
Base Currency
Typical Daily Volume
XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.got-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown).
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
12
-
Prompts: 7660
Format version: 1.1
Load with lmprobe
from lmprobe import pull_dataset, load_activation_dataset
# Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.sports-cards
Digital Card Magazine Dataset
This dataset contains sports card images and their associated metadata for training machine learning models in card recognition, text extraction, and value estimation.
Dataset Description
Dataset Summary
A comprehensive collection of sports card images and metadata, including:
Front and back card images
OCR-extracted text with confidence scores
AI-analyzed card attributes
Card details (player, team, year, etc.)
Vision API labels… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/sports-cards.fixtures_got_ocrGoT-10kgot-activations-qwen2.5-0.5b
Qwen/Qwen2.5-0.5B — Activation Dataset
Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987).
Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries.
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-23
896
-
1
-
logits_topk
-
k=100
last_token
1
1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.gotlaions_got_talent_previewOmniEdit-GoTgot10kJourneyDB-GoT
JourneyDB-GoT
Paper
Dataset Description
The JourneyDB-GoT dataset enriches the JourneyDB collection with rich grounding annotations added to text descriptions. This dataset combines high-quality AI-generated images from JourneyDB with detailed text descriptions and object-level grounding information.
Key Features
Modalities: Image, Text, and Grounding Annotations
Image Source: High-quality AI-generated images from JourneyDB
Text Descriptions: Each image has a… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/JourneyDB-GoT.STARGATE
🧠 STARGATE: CIA Remote Viewing Archive (Raw PDFs + Metadata)
Overview
STARGATE is the most comprehensive open-access archive of declassified CIA documents related to psychic research, remote viewing (RV), and anomalous cognition.
This dataset consolidates over 12,000 scanned PDF files, drawn from decades of classified government programs designed to investigate and operationalize extrasensory perception (ESP) in intelligence-gathering contexts.
Programs… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/STARGATE.gotsf-ds
📶 Beam-Level (5G) Time-Series Dataset
📚 Citation
This dataset is released alongside the following paper:
Fechete, L., et al. “Goal-Oriented Time-Series Forecasting: Foundation Framework Design.” Proceedings of the AAAI Conference on Artificial Intelligence, 2026, Singapore.
If you use this dataset, please cite the above work.
This dataset introduces a novel multivariate time series specifically curated to support research in enabling accurate prediction of KPIs… See the full description on the dataset page: https://huggingface.co/datasets/netop/gotsf-ds.
