CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Thorsten-Voice /TV-44kHz-Full The "Thorsten-Voice" dataset This truly open source (CC0 license) german (🇩🇪) voice dataset contains about 40 hours of transcribed voice recordings by Thorsten Müller, a single male, native speaker in over 38.000 wave files. Mono Samplerate: 44.100Hz Trimmed silence at begin/end Denoised Normalized to -24dB Disclaimer "Please keep in mind, I am not a professional speaker, just an open source speech technology enthusiast who donates his voice. I contribute my personal… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-44kHz-Full.audiotext-to-speech10K<n<100K10 likes649 downloads2y agoHugging Face02Thorsu /sovereign-shadow-inference-bench Sovereign Shadow Inference Bench A public, versioned evidence surface for independent Hugging Face shadow inference beside Sovereign's primary OpenRouter/Revolver route. What this dataset proves The seed record in data/shadow_receipts.jsonl was produced by one real Hugging Face Inference Providers request. It records provider/model identity, request bounds, latency, hashes, literal-match outcome, source revision, and an immutable receipt hash. What it… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-shadow-inference-bench.tabulartext-generationn<1K1 likes637 downloads1d agoHugging Face03xhyumiracle /thor25 Thor25: THORChain Cross-Chain Data This repository contains the Thor25 dataset released with LOCARD: An Agentic Framework for Blockchain Forensics, published at IEEE ICBC 2026, Brisbane, Australia, June 1-5, 2026. Thor25 supports research on agentic blockchain forensics, cross-chain transaction tracing, and evidence-grounded tool use by LLM agents. It does not map cleanly to conventional NLP task categories such as question answering or text generation: solving each benchmark… See the full description on the dataset page: https://huggingface.co/datasets/xhyumiracle/thor25.image100K<n<1M1 likes442 downloads1mo agoHugging Face04ThorKl /PLI-parallax PLI-Parallax Predicted protein-ligand complexes are used as training data at considerable scale today. Recent distillation sets contain several hundred thousand cofolded BindingDB systems, and the filter applied to them is usually the predicting model's own confidence score. That filter is a self-assessment, and cofolding models have been shown to be confidently wrong in ways their own confidence does not reveal. This dataset provides the coordinates and derived distance labels… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/PLI-parallax.tabular100M<n<1B0 likes272 downloads16d agoHugging Face05thordata /martial-arts-video-v1 Martial Arts Action Video Preview This is a small preview of five short, staged-looking martial-arts action clips. The clips show guarded stances, upper-body strikes, close combat, footwork, falls/recovery, and apparent weapon exchanges. Contents videos/: five short MP4 clips metadata.csv: one row per video with broad scene and action labels segments.csv: coarse time windows with action descriptions metadata.jsonl: the same metadata in JSON Lines format frames/:… See the full description on the dataset page: https://huggingface.co/datasets/thordata/martial-arts-video-v1.tabularn<1K0 likes216 downloads15d agoHugging Face06ThorKl /theobroma THEOBROMA v1.36 An aggregated open database of 1,132,805 natural products from 29 sources, with per-compound license auditing, three-tier classification provenance, and stereochemistry-aware deduplication. Live instance: https://theobroma.l3s.uni-hannover.de Archival record: https://doi.org/10.5281/zenodo.20443051 (concept DOI, resolves to latest) This release: https://doi.org/10.5281/zenodo.22816330 Preprint: https://doi.org/10.64898/2026.06.12.731585 Licensing… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/theobroma.tabular1M<n<10M0 likes186 downloads4d agoHugging Face07Thoria /mandarin-most-common-words-tr-en Mandarin Most Common Words (TR-EN) Overview The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis. This dataset was created by Stephanie Liu and Kamil Murat Yilmaz. Dataset Content The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.text1K<n<10K2 likes180 downloads5mo agoHugging Face08Chris-davis-thor /motioncapture Fight Scene Action Video Dataset (samples) A VideoFolder-format dataset of fight / martial-arts action clips, provided by Thordata. Each clip is annotated with a temporal action breakdown, per-segment descriptions, an English caption, and ffmpeg-measured technical metadata. Layout (HuggingFace VideoFolder) data/ ├── metadata.jsonl # one row per video; "file_name" links to the clip └── train/ ├── clip_01.mp4 └── ... clip_12.mp4 Load… See the full description on the dataset page: https://huggingface.co/datasets/Chris-davis-thor/motioncapture.videovideo-classificationn<1K0 likes160 downloads11d agoHugging Face09Thorsu /sovereign-evidence-observatory Sovereign Evidence Observatory Casebook The interesting question is not “Which AI sounds smartest?” It is “What kind of evidence would make this claim true, false, or still undecidable?” This public casebook is the first dataset for the Sovereign Evidence Observatory. It turns model agreement, disagreement and abstention into inspectable evidence objects instead of treating consensus as truth. Five layers Shadow Mesh — independent model/provider observations.… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-evidence-observatory.textn<1K1 likes152 downloads28d agoHugging Face10Thorsten-Voice /TV-24kHz-2025.12-Neutral-FT-Mini Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini Overview This dataset is a small, high-quality fine-tuning dataset created specifically for speaker refinement and voice matching in Orpheus TTS models. It consists of 60 newly recorded German speech samples, spoken in a neutral, relaxed, everyday style, closely reflecting the natural speaking voice of the original speaker. This dataset is intended for: Speaker adaptation and voice refinement Fine-tuning Orpheus TTS models… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini.audion<1K1 likes145 downloads9mo agoHugging Face11thordata /ball-sports-video-v1 Ball Sports Key Action Video Preview This is a small preview of four sports videos from the Thordata Ball Sports Key Action product: basketball: shooting tennis: hitting soccer: shooting soccer: interception The product listing states that the source collection is 1080p or higher, has no logos, subtitles, mosaics, watermarks, black borders, or noise, and includes metadata. The files in this package are the marketplace preview copies, which are lower-resolution preview encodes;… See the full description on the dataset page: https://huggingface.co/datasets/thordata/ball-sports-video-v1.tabularn<1K0 likes145 downloads14d agoHugging Face12FM4CS /THOR-Pretrain THOR-Pretrain Code to load data and details to come soon... Release Notes global_availability_index.parquet -> satellite/metadata.parquet Layout satellite/metadata.parquet: satellite sample metadata. satellite/shards/: WebDataset tar shards. satellite/tar_index/: byte-offset sidecar JSON for each satellite tar. era5_land/metadata.parquet: ERA5-Land daily metadata. era5_land/shards/: ERA5-Land daily WebDataset tar shards. era5_land/tar_index/:… See the full description on the dataset page: https://huggingface.co/datasets/FM4CS/THOR-Pretrain.textimage-feature-extraction10K<n<100K0 likes130 downloads3mo agoHugging Face13Chris-davis-thor /VR_Manipulatio Dual-Arm VR Manipulation Video Dataset A VideoFolder-format dataset of first-person (egocentric) dual-hand manipulation clips captured with a VR headset, covering everyday bimanual tasks — handling a box, carrying a tray, picking items into a cart, switching on a lamp, folding a shirt, stacking blocks, capping a marker, loading a dishwasher, inserting pens, and placing fruit onto a tray. Provided by Thordata. Each clip is annotated with a temporal action breakdown, per-segment… See the full description on the dataset page: https://huggingface.co/datasets/Chris-davis-thor/VR_Manipulatio.videovideo-classificationn<1K0 likes127 downloads8d agoHugging Face14Thorns07 /so100stackingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 42, "total_frames": 37264, "total_tasks": 4, "total_videos": 84, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:42" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thorns07/so100stacking.tabularrobotics10K<n<100K0 likes122 downloads1y agoHugging Face15Chris-davis-thor /transportation First-Person Mountain Driving Video Sample Overview This dataset contains a five-minute first-person driving video recorded on mountain roads. It is provided by ThorData for video understanding, scene analysis, preprocessing, and exploratory computer-vision research. Files videos/driver_pov_preview_0001.mp4: source video. metadata.csv: video properties and the file reference used by the dataset viewer. annotations.csv: ten fixed 30-second scene and… See the full description on the dataset page: https://huggingface.co/datasets/Chris-davis-thor/transportation.videovideo-classification0 likes117 downloads12d agoHugging Face16thordata /science-communication-video-v1 Science Communication Video Preview This preview contains four short educational animation videos from the Thordata Multidisciplinary Science Communication Video Collection: acetaldehyde oxidation how volcanoes form how typhoons form solar wind and aurora Each sample presents one focused knowledge topic through a coherent visual sequence. The videos are useful for demonstrating video captioning, visual question answering, and cross-modal reasoning. Contents… See the full description on the dataset page: https://huggingface.co/datasets/thordata/science-communication-video-v1.tabularn<1K1 likes117 downloads5d agoHugging Face17histai /SPIDER-thoraxgated SPIDER-THORAX Dataset SPIDER is a collection of supervised pathological datasets covering multiple organs, each with comprehensive class coverage. These datasets are professionally annotated by pathologists. If you would like to support, sponsor, or obtain a commercial license for the SPIDER data and models, please contact us at models@hist.ai. For a detailed description of SPIDER, methodology, and benchmark results, refer to our research paper: 📄 SPIDER: A Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/histai/SPIDER-thorax.image-classification100K<n<1M4 likes115 downloads2y agoHugging Face18ThorKl /protac-bench PROTAC-Bench: A Cold-Target Benchmark for PROTAC Degradation Prediction Dataset Description PROTAC-Bench is a merged PROTAC degradation dataset containing 10,748 entries across 173 protein targets (9,359 unique SMILES). It combines data from PROTAC-DB 3.0, Ribes et al. (2024), and DegradeMaster with deduplication and canonical SMILES standardization. Each entry has a binary activity label (active: DC50 < 1 μM OR Dmax > 50%) along with target UniProt ID and E3 ligase type… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/protac-bench.texttabular-classificationn<1K0 likes112 downloads4mo agoHugging Face19thorirhrafn /rmh_subset_largetext1M<n<10M0 likes111 downloads3y agoHugging Face20Thorns07 /so100_test_diffusion_graspThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 25, "total_frames": 11353, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:25" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thorns07/so100_test_diffusion_grasp.tabularrobotics10K<n<100K0 likes108 downloads1y agoHugging Face21thordata /dual-arm-operation-video-v1 Dual-Arm Operation First-Person Video Preview This preview contains four first-person operation videos from the Thordata Dual-Arm Operation VR Video Collection: handling a remote-control battery handling a food tray in a kitchen inspecting and handling a packaged product in a supermarket turning on a desk lamp The product listing describes PICO 4 Ultra Enterprise capture at 1920x1080. The uploaded marketplace preview copies are lower-resolution encodes; the actual file… See the full description on the dataset page: https://huggingface.co/datasets/thordata/dual-arm-operation-video-v1.tabularn<1K0 likes103 downloads7d agoHugging Face22Thorsten-Voice /TV-24kHz-Neutral Thorsten-Voice TV-24kHz-Neutral Dataset This dataset is a resampled version of the "TV-2022.10-Neutral" configuration from the original Thorsten-Voice TV-44kHz-Full dataset, converted from 44.1kHz to 24kHz sampling rate. Dataset Description The Thorsten-Voice dataset contains German speech recordings by Thorsten Müller, suitable for text-to-speech (TTS) training and other speech synthesis tasks. Changes from Original Sample Rate: Converted from… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-Neutral.audio10K<n<100K0 likes102 downloads1y agoHugging Face23NeuroBench /thor_eeg_mi1 likes90 downloads5mo agoHugging Face24alexchauncy /thor_dataset_16 thor_dataset_16 This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. videoroboticsn<1K0 likes84 downloads1y agoHugging Face25Thorns07 /so100_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 2, "total_frames": 1095, "total_tasks":1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Thorns07/so100_test.tabularrobotics10K<n<100K0 likes82 downloads1y agoHugging Face26ThorKl /SABDAB_INDI_datasets0 likes80 downloads1y agoHugging Face27Chris-davis-thor /houseworkvideon<1K0 likes77 downloads11d agoHugging Face28Thorismund /mkultra-foia-mineru-archive MKULTRA FOIA MinerU Archive The MKULTRA FOIA MinerU Archive is a public-interest research dataset containing 21,237 pages of declassified MKULTRA and related-program records. The repository contains two separately preserved collections: Collection Pages Source Page export Document export Original PEERS FOIA collection 16,383 TIFF files obtained by PEERS through FOIA and converted to PNG pages documents Supplemental authenticated declassified collection 4,854… See the full description on the dataset page: https://huggingface.co/datasets/Thorismund/mkultra-foia-mineru-archive.textimage-to-text10K<n<100K0 likes74 downloads2mo agoHugging Face29thordata /first-person-driving-video-v1 First-Person Perspective Driving Video Preview This is a single public-preview sample from the Thordata First-Person Perspective Driving Video Collection. The product listing describes the source collection as driver-seat footage at 3840x2160 and approximately 24.6 Mb/s. The downloaded marketplace preview is a lower-resolution preview encode; its actual dimensions and frame rate are recorded in metadata.csv. It shows a roadside service area followed by open mountain roads and… See the full description on the dataset page: https://huggingface.co/datasets/thordata/first-person-driving-video-v1.videon<1K0 likes69 downloads14d agoHugging Face30EtMmohammedHafsati /datagen-thor-samples Datagen Thor Samples Multilingual JSONL samples generated by datagen-thor. texttext-generation10K<n<100K0 likes64 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.