datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TV-44kHz-Full
The "Thorsten-Voice" dataset
This truly open source (CC0 license) german (🇩🇪) voice dataset contains about 40 hours of transcribed voice recordings by Thorsten Müller,
a single male, native speaker in over 38.000 wave files.
Mono
Samplerate: 44.100Hz
Trimmed silence at begin/end
Denoised
Normalized to -24dB
Disclaimer
"Please keep in mind, I am not a professional speaker, just an open source speech technology enthusiast who donates his voice. I contribute my personal… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-44kHz-Full.sovereign-shadow-inference-bench
Sovereign Shadow Inference Bench
A public, versioned evidence surface for independent Hugging Face shadow inference beside Sovereign's primary OpenRouter/Revolver route.
What this dataset proves
The seed record in data/shadow_receipts.jsonl was produced by one real Hugging Face Inference Providers request. It records provider/model identity, request bounds, latency, hashes, literal-match outcome, source revision, and an immutable receipt hash.
What it… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-shadow-inference-bench.thor25
Thor25: THORChain Cross-Chain Data
This repository contains the Thor25 dataset released with LOCARD: An Agentic Framework for Blockchain Forensics, published at IEEE ICBC 2026, Brisbane, Australia, June 1-5, 2026.
Thor25 supports research on agentic blockchain forensics, cross-chain transaction tracing, and evidence-grounded tool use by LLM agents. It does not map cleanly to conventional NLP task categories such as question answering or text generation: solving each benchmark… See the full description on the dataset page: https://huggingface.co/datasets/xhyumiracle/thor25.PLI-parallax
PLI-Parallax
Predicted protein-ligand complexes are used as training data at considerable
scale today. Recent distillation sets contain several hundred thousand cofolded
BindingDB systems, and the filter applied to them is usually the predicting
model's own confidence score. That filter is a self-assessment, and cofolding
models have been shown to be confidently wrong in ways their own confidence does
not reveal.
This dataset provides the coordinates and derived distance labels… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/PLI-parallax.martial-arts-video-v1
Martial Arts Action Video Preview
This is a small preview of five short, staged-looking martial-arts action clips. The clips show guarded stances, upper-body strikes, close combat, footwork, falls/recovery, and apparent weapon exchanges.
Contents
videos/: five short MP4 clips
metadata.csv: one row per video with broad scene and action labels
segments.csv: coarse time windows with action descriptions
metadata.jsonl: the same metadata in JSON Lines format
frames/:… See the full description on the dataset page: https://huggingface.co/datasets/thordata/martial-arts-video-v1.theobroma
THEOBROMA v1.36
An aggregated open database of 1,132,805 natural products from 29 sources, with
per-compound license auditing, three-tier classification provenance, and
stereochemistry-aware deduplication.
Live instance: https://theobroma.l3s.uni-hannover.de
Archival record: https://doi.org/10.5281/zenodo.20443051 (concept DOI, resolves to latest)
This release: https://doi.org/10.5281/zenodo.22816330
Preprint: https://doi.org/10.64898/2026.06.12.731585
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/theobroma.mandarin-most-common-words-tr-en
Mandarin Most Common Words (TR-EN)
Overview
The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis.
This dataset was created by Stephanie Liu and Kamil Murat Yilmaz.
Dataset Content
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.motioncapture
Fight Scene Action Video Dataset (samples)
A VideoFolder-format dataset of fight / martial-arts action clips, provided by
Thordata. Each clip is annotated with a temporal action breakdown, per-segment
descriptions, an English caption, and ffmpeg-measured technical metadata.
Layout (HuggingFace VideoFolder)
data/
├── metadata.jsonl # one row per video; "file_name" links to the clip
└── train/
├── clip_01.mp4
└── ... clip_12.mp4
Load… See the full description on the dataset page: https://huggingface.co/datasets/Chris-davis-thor/motioncapture.sovereign-evidence-observatory
Sovereign Evidence Observatory Casebook
The interesting question is not “Which AI sounds smartest?” It is “What kind of evidence would make this claim true, false, or still undecidable?”
This public casebook is the first dataset for the Sovereign Evidence Observatory. It turns model agreement, disagreement and abstention into inspectable evidence objects instead of treating consensus as truth.
Five layers
Shadow Mesh — independent model/provider observations.… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-evidence-observatory.TV-24kHz-2025.12-Neutral-FT-Mini
Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini
Overview
This dataset is a small, high-quality fine-tuning dataset created specifically for speaker refinement and voice matching in Orpheus TTS models.
It consists of 60 newly recorded German speech samples, spoken in a neutral, relaxed, everyday style, closely reflecting the natural speaking voice of the original speaker.
This dataset is intended for:
Speaker adaptation and voice refinement
Fine-tuning Orpheus TTS models… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini.ball-sports-video-v1
Ball Sports Key Action Video Preview
This is a small preview of four sports videos from the Thordata Ball Sports Key Action product:
basketball: shooting
tennis: hitting
soccer: shooting
soccer: interception
The product listing states that the source collection is 1080p or higher, has no logos, subtitles, mosaics, watermarks, black borders, or noise, and includes metadata. The files in this package are the marketplace preview copies, which are lower-resolution preview encodes;… See the full description on the dataset page: https://huggingface.co/datasets/thordata/ball-sports-video-v1.THOR-Pretrain
THOR-Pretrain
Code to load data and details to come soon...
Release Notes
global_availability_index.parquet -> satellite/metadata.parquet
Layout
satellite/metadata.parquet: satellite sample metadata.
satellite/shards/: WebDataset tar shards.
satellite/tar_index/: byte-offset sidecar JSON for each satellite tar.
era5_land/metadata.parquet: ERA5-Land daily metadata.
era5_land/shards/: ERA5-Land daily WebDataset tar shards.
era5_land/tar_index/:… See the full description on the dataset page: https://huggingface.co/datasets/FM4CS/THOR-Pretrain.VR_Manipulatio
Dual-Arm VR Manipulation Video Dataset
A VideoFolder-format dataset of first-person (egocentric) dual-hand manipulation
clips captured with a VR headset, covering everyday bimanual tasks — handling a box,
carrying a tray, picking items into a cart, switching on a lamp, folding a shirt,
stacking blocks, capping a marker, loading a dishwasher, inserting pens, and placing
fruit onto a tray. Provided by Thordata. Each clip is annotated with a temporal
action breakdown, per-segment… See the full description on the dataset page: https://huggingface.co/datasets/Chris-davis-thor/VR_Manipulatio.so100stackingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 42,
"total_frames": 37264,
"total_tasks": 4,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:42"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thorns07/so100stacking.transportation
First-Person Mountain Driving Video Sample
Overview
This dataset contains a five-minute first-person driving video recorded on mountain roads. It is provided by ThorData for video understanding, scene analysis, preprocessing, and exploratory computer-vision research.
Files
videos/driver_pov_preview_0001.mp4: source video.
metadata.csv: video properties and the file reference used by the dataset viewer.
annotations.csv: ten fixed 30-second scene and… See the full description on the dataset page: https://huggingface.co/datasets/Chris-davis-thor/transportation.science-communication-video-v1
Science Communication Video Preview
This preview contains four short educational animation videos from the Thordata Multidisciplinary Science Communication Video Collection:
acetaldehyde oxidation
how volcanoes form
how typhoons form
solar wind and aurora
Each sample presents one focused knowledge topic through a coherent visual sequence. The videos are useful for demonstrating video captioning, visual question answering, and cross-modal reasoning.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/thordata/science-communication-video-v1.SPIDER-thorax
SPIDER-THORAX Dataset
SPIDER is a collection of supervised pathological datasets covering multiple organs, each with comprehensive class coverage. These datasets are professionally annotated by pathologists.
If you would like to support, sponsor, or obtain a commercial license for the SPIDER data and models, please contact us at models@hist.ai.
For a detailed description of SPIDER, methodology, and benchmark results, refer to our research paper:
📄 SPIDER: A Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/histai/SPIDER-thorax.protac-bench
PROTAC-Bench: A Cold-Target Benchmark for PROTAC Degradation Prediction
Dataset Description
PROTAC-Bench is a merged PROTAC degradation dataset containing 10,748 entries across 173 protein targets (9,359 unique SMILES). It combines data from PROTAC-DB 3.0, Ribes et al. (2024), and DegradeMaster with deduplication and canonical SMILES standardization. Each entry has a binary activity label (active: DC50 < 1 μM OR Dmax > 50%) along with target UniProt ID and E3 ligase type… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/protac-bench.rmh_subset_largeso100_test_diffusion_graspThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 25,
"total_frames": 11353,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thorns07/so100_test_diffusion_grasp.dual-arm-operation-video-v1
Dual-Arm Operation First-Person Video Preview
This preview contains four first-person operation videos from the Thordata Dual-Arm Operation VR Video Collection:
handling a remote-control battery
handling a food tray in a kitchen
inspecting and handling a packaged product in a supermarket
turning on a desk lamp
The product listing describes PICO 4 Ultra Enterprise capture at 1920x1080. The uploaded marketplace preview copies are lower-resolution encodes; the actual file… See the full description on the dataset page: https://huggingface.co/datasets/thordata/dual-arm-operation-video-v1.TV-24kHz-Neutral
Thorsten-Voice TV-24kHz-Neutral Dataset
This dataset is a resampled version of the "TV-2022.10-Neutral" configuration from the original Thorsten-Voice TV-44kHz-Full dataset, converted from 44.1kHz to 24kHz sampling rate.
Dataset Description
The Thorsten-Voice dataset contains German speech recordings by Thorsten Müller, suitable for text-to-speech (TTS) training and other speech synthesis tasks.
Changes from Original
Sample Rate: Converted from… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-Neutral.thor_eeg_mithor_dataset_16
thor_dataset_16
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1095,
"total_tasks":1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thorns07/so100_test.SABDAB_INDI_datasetshouseworkmkultra-foia-mineru-archive
MKULTRA FOIA MinerU Archive
The MKULTRA FOIA MinerU Archive is a public-interest research dataset
containing 21,237 pages of declassified MKULTRA and related-program
records.
The repository contains two separately preserved collections:
Collection
Pages
Source
Page export
Document export
Original PEERS FOIA collection
16,383
TIFF files obtained by PEERS through FOIA and converted to PNG
pages
documents
Supplemental authenticated declassified collection
4,854… See the full description on the dataset page: https://huggingface.co/datasets/Thorismund/mkultra-foia-mineru-archive.first-person-driving-video-v1
First-Person Perspective Driving Video Preview
This is a single public-preview sample from the Thordata First-Person Perspective Driving Video Collection. The product listing describes the source collection as driver-seat footage at 3840x2160 and approximately 24.6 Mb/s. The downloaded marketplace preview is a lower-resolution preview encode; its actual dimensions and frame rate are recorded in metadata.csv. It shows a roadside service area followed by open mountain roads and… See the full description on the dataset page: https://huggingface.co/datasets/thordata/first-person-driving-video-v1.datagen-thor-samples
Datagen Thor Samples
Multilingual JSONL samples generated by datagen-thor.
