CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Thorsten-Voice /TV-44kHz-Full The "Thorsten-Voice" dataset This truly open source (CC0 license) german (🇩🇪) voice dataset contains about 40 hours of transcribed voice recordings by Thorsten Müller, a single male, native speaker in over 38.000 wave files. Mono Samplerate: 44.100Hz Trimmed silence at begin/end Denoised Normalized to -24dB Disclaimer "Please keep in mind, I am not a professional speaker, just an open source speech technology enthusiast who donates his voice. I contribute my personal… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-44kHz-Full.audiotext-to-speech10K<n<100K10 likes649 downloads2y agoHugging Face02Thorsu /sovereign-shadow-inference-bench Sovereign Shadow Inference Bench A public, versioned evidence surface for independent Hugging Face shadow inference beside Sovereign's primary OpenRouter/Revolver route. What this dataset proves The seed record in data/shadow_receipts.jsonl was produced by one real Hugging Face Inference Providers request. It records provider/model identity, request bounds, latency, hashes, literal-match outcome, source revision, and an immutable receipt hash. What it… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-shadow-inference-bench.tabulartext-generationn<1K1 likes637 downloads2d agoHugging Face03xhyumiracle /thor25 Thor25: THORChain Cross-Chain Data This repository contains the Thor25 dataset released with LOCARD: An Agentic Framework for Blockchain Forensics, published at IEEE ICBC 2026, Brisbane, Australia, June 1-5, 2026. Thor25 supports research on agentic blockchain forensics, cross-chain transaction tracing, and evidence-grounded tool use by LLM agents. It does not map cleanly to conventional NLP task categories such as question answering or text generation: solving each benchmark… See the full description on the dataset page: https://huggingface.co/datasets/xhyumiracle/thor25.image100K<n<1M1 likes442 downloads1mo agoHugging Face04ThorKl /PLI-parallax PLI-Parallax Predicted protein-ligand complexes are used as training data at considerable scale today. Recent distillation sets contain several hundred thousand cofolded BindingDB systems, and the filter applied to them is usually the predicting model's own confidence score. That filter is a self-assessment, and cofolding models have been shown to be confidently wrong in ways their own confidence does not reveal. This dataset provides the coordinates and derived distance labels… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/PLI-parallax.tabular100M<n<1B0 likes272 downloads16d agoHugging Face05thordata /martial-arts-video-v1 Martial Arts Action Video Preview This is a small preview of five short, staged-looking martial-arts action clips. The clips show guarded stances, upper-body strikes, close combat, footwork, falls/recovery, and apparent weapon exchanges. Contents videos/: five short MP4 clips metadata.csv: one row per video with broad scene and action labels segments.csv: coarse time windows with action descriptions metadata.jsonl: the same metadata in JSON Lines format frames/:… See the full description on the dataset page: https://huggingface.co/datasets/thordata/martial-arts-video-v1.tabularn<1K0 likes216 downloads16d agoHugging Face06ThorKl /theobroma THEOBROMA v1.36 An aggregated open database of 1,132,805 natural products from 29 sources, with per-compound license auditing, three-tier classification provenance, and stereochemistry-aware deduplication. Live instance: https://theobroma.l3s.uni-hannover.de Archival record: https://doi.org/10.5281/zenodo.20443051 (concept DOI, resolves to latest) This release: https://doi.org/10.5281/zenodo.22816330 Preprint: https://doi.org/10.64898/2026.06.12.731585 Licensing… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/theobroma.tabular1M<n<10M0 likes186 downloads5d agoHugging Face07Thoria /mandarin-most-common-words-tr-en Mandarin Most Common Words (TR-EN) Overview The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis. This dataset was created by Stephanie Liu and Kamil Murat Yilmaz. Dataset Content The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.text1K<n<10K2 likes180 downloads5mo agoHugging Face08Thorsu /sovereign-evidence-observatory Sovereign Evidence Observatory Casebook The interesting question is not “Which AI sounds smartest?” It is “What kind of evidence would make this claim true, false, or still undecidable?” This public casebook is the first dataset for the Sovereign Evidence Observatory. It turns model agreement, disagreement and abstention into inspectable evidence objects instead of treating consensus as truth. Five layers Shadow Mesh — independent model/provider observations.… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-evidence-observatory.textn<1K1 likes152 downloads28d agoHugging Face09Thorsten-Voice /TV-24kHz-2025.12-Neutral-FT-Mini Thorsten-Voice TV-24kHz-2025.12-Neutral-FT-Mini Overview This dataset is a small, high-quality fine-tuning dataset created specifically for speaker refinement and voice matching in Orpheus TTS models. It consists of 60 newly recorded German speech samples, spoken in a neutral, relaxed, everyday style, closely reflecting the natural speaking voice of the original speaker. This dataset is intended for: Speaker adaptation and voice refinement Fine-tuning Orpheus TTS models… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-2025.12-Neutral-FT-Mini.audion<1K1 likes145 downloads9mo agoHugging Face10thordata /ball-sports-video-v1 Ball Sports Key Action Video Preview This is a small preview of four sports videos from the Thordata Ball Sports Key Action product: basketball: shooting tennis: hitting soccer: shooting soccer: interception The product listing states that the source collection is 1080p or higher, has no logos, subtitles, mosaics, watermarks, black borders, or noise, and includes metadata. The files in this package are the marketplace preview copies, which are lower-resolution preview encodes;… See the full description on the dataset page: https://huggingface.co/datasets/thordata/ball-sports-video-v1.tabularn<1K0 likes145 downloads15d agoHugging Face11FM4CS /THOR-Pretrain THOR-Pretrain Code to load data and details to come soon... Release Notes global_availability_index.parquet -> satellite/metadata.parquet Layout satellite/metadata.parquet: satellite sample metadata. satellite/shards/: WebDataset tar shards. satellite/tar_index/: byte-offset sidecar JSON for each satellite tar. era5_land/metadata.parquet: ERA5-Land daily metadata. era5_land/shards/: ERA5-Land daily WebDataset tar shards. era5_land/tar_index/:… See the full description on the dataset page: https://huggingface.co/datasets/FM4CS/THOR-Pretrain.textimage-feature-extraction10K<n<100K1 likes130 downloads3mo agoHugging Face12thordata /science-communication-video-v1 Science Communication Video Preview This preview contains four short educational animation videos from the Thordata Multidisciplinary Science Communication Video Collection: acetaldehyde oxidation how volcanoes form how typhoons form solar wind and aurora Each sample presents one focused knowledge topic through a coherent visual sequence. The videos are useful for demonstrating video captioning, visual question answering, and cross-modal reasoning. Contents… See the full description on the dataset page: https://huggingface.co/datasets/thordata/science-communication-video-v1.tabularn<1K1 likes117 downloads6d agoHugging Face13ThorKl /protac-bench PROTAC-Bench: A Cold-Target Benchmark for PROTAC Degradation Prediction Dataset Description PROTAC-Bench is a merged PROTAC degradation dataset containing 10,748 entries across 173 protein targets (9,359 unique SMILES). It combines data from PROTAC-DB 3.0, Ribes et al. (2024), and DegradeMaster with deduplication and canonical SMILES standardization. Each entry has a binary activity label (active: DC50 < 1 μM OR Dmax > 50%) along with target UniProt ID and E3 ligase type… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/protac-bench.texttabular-classificationn<1K0 likes112 downloads4mo agoHugging Face14thorirhrafn /rmh_subset_largetext1M<n<10M0 likes111 downloads3y agoHugging Face15thordata /dual-arm-operation-video-v1 Dual-Arm Operation First-Person Video Preview This preview contains four first-person operation videos from the Thordata Dual-Arm Operation VR Video Collection: handling a remote-control battery handling a food tray in a kitchen inspecting and handling a packaged product in a supermarket turning on a desk lamp The product listing describes PICO 4 Ultra Enterprise capture at 1920x1080. The uploaded marketplace preview copies are lower-resolution encodes; the actual file… See the full description on the dataset page: https://huggingface.co/datasets/thordata/dual-arm-operation-video-v1.tabularn<1K0 likes103 downloads8d agoHugging Face16Thorsten-Voice /TV-24kHz-Neutral Thorsten-Voice TV-24kHz-Neutral Dataset This dataset is a resampled version of the "TV-2022.10-Neutral" configuration from the original Thorsten-Voice TV-44kHz-Full dataset, converted from 44.1kHz to 24kHz sampling rate. Dataset Description The Thorsten-Voice dataset contains German speech recordings by Thorsten Müller, suitable for text-to-speech (TTS) training and other speech synthesis tasks. Changes from Original Sample Rate: Converted from… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-24kHz-Neutral.audio10K<n<100K0 likes102 downloads1y agoHugging Face17Thorismund /mkultra-foia-mineru-archive MKULTRA FOIA MinerU Archive The MKULTRA FOIA MinerU Archive is a public-interest research dataset containing 21,237 pages of declassified MKULTRA and related-program records. The repository contains two separately preserved collections: Collection Pages Source Page export Document export Original PEERS FOIA collection 16,383 TIFF files obtained by PEERS through FOIA and converted to PNG pages documents Supplemental authenticated declassified collection 4,854… See the full description on the dataset page: https://huggingface.co/datasets/Thorismund/mkultra-foia-mineru-archive.textimage-to-text10K<n<100K0 likes74 downloads2mo agoHugging Face18EtMmohammedHafsati /datagen-thor-samples Datagen Thor Samples Multilingual JSONL samples generated by datagen-thor. texttext-generation10K<n<100K0 likes64 downloads4mo agoHugging Face19thorirhrafn /domar_datatext1K<n<10K0 likes57 downloads2y agoHugging Face20thorirhrafn /rmh_subset_medium2 Dataset Card for "rmh_subset_medium2" More Information needed text100K<n<1M0 likes53 downloads3y agoHugging Face21convoicon /Thoroughly_Engineered_Toxicitygated Thoroughly_Engineered_Toxicity Thoroughly Engineered Toxicity (TET) is a dataset created by filtering a set of prompts from Chat-Lmsys-1M dataset, each prompt has high potential of exposing the toxicity in Large Language models (LLMs). Related Links: Huggingface Dataset | Paper Please CITE our ACL 2024 paper when TET is used to help produce published results or is incorporated into other software: @inproceedings{luong-etal-2024-realistic, title = "Realistic Evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/convoicon/Thoroughly_Engineered_Toxicity.text1K<n<10K2 likes53 downloads2y agoHugging Face22thorirhrafn /rmh_subset_medium Dataset Card for "rmh_subset_medium" More Information needed text100K<n<1M0 likes52 downloads3y agoHugging Face23thorirhrafn /minigrid-wm-data2text100K<n<1M0 likes52 downloads11mo agoHugging Face24ThornZ /Search-R1-SFTtext10K<n<100K1 likes52 downloads8mo agoHugging Face25thorirhrafn /rmh_subset_large2 Dataset Card for "rmh_subset_large2" More Information needed text100K<n<1M0 likes47 downloads3y agoHugging Face26thorirhrafn /minigrid-wm-data3text100K<n<1M0 likes47 downloads11mo agoHugging Face27PEARLS-Lab /infini-thor-niehimage100K<n<1M0 likes43 downloads7mo agoHugging Face28PEARLS-Lab /infini-thor-traintext1K<n<10K0 likes34 downloads7mo agoHugging Face29thor1 /LLMForge This dataset contains roughly 30,000 inorganic solid-state synthesis recipes that were entirely generated by OpenAI GPT–4.1 to augment existing literature-mined protocols. It covers a wide range of compositions sampled from Materials Project for maximal chemical diversity, with complete precursor lists and two-step heat profiles (calcination and sintering conditions). By expanding the sparse experimental data over sixfold, this synthetic corpus enables more robust model pretraining and… See the full description on the dataset page: https://huggingface.co/datasets/thor1/LLMForge.tabular10K<n<100K1 likes32 downloads1y agoHugging Face30nthnael /rcv2-thoraximage1K<n<10K0 likes25 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.