thor
Datasets
All datasets matching “thor”TV-44kHz-Full
The "Thorsten-Voice" dataset
This truly open source (CC0 license) german (🇩🇪) voice dataset contains about 40 hours of transcribed voice recordings by Thorsten Müller,
a single male, native speaker in over 38.000 wave files.
Mono
Samplerate: 44.100Hz
Trimmed silence at begin/end
Denoised
Normalized to -24dB
Disclaimer
"Please keep in mind, I am not a professional speaker, just an open source speech technology enthusiast who donates his voice. I contribute my personal… See the full description on the dataset page: https://huggingface.co/datasets/Thorsten-Voice/TV-44kHz-Full.sovereign-shadow-inference-bench
Sovereign Shadow Inference Bench
A public, versioned evidence surface for independent Hugging Face shadow inference beside Sovereign's primary OpenRouter/Revolver route.
What this dataset proves
The seed record in data/shadow_receipts.jsonl was produced by one real Hugging Face Inference Providers request. It records provider/model identity, request bounds, latency, hashes, literal-match outcome, source revision, and an immutable receipt hash.
What it… See the full description on the dataset page: https://huggingface.co/datasets/Thorsu/sovereign-shadow-inference-bench.thor25
Thor25: THORChain Cross-Chain Data
This repository contains the Thor25 dataset released with LOCARD: An Agentic Framework for Blockchain Forensics, published at IEEE ICBC 2026, Brisbane, Australia, June 1-5, 2026.
Thor25 supports research on agentic blockchain forensics, cross-chain transaction tracing, and evidence-grounded tool use by LLM agents. It does not map cleanly to conventional NLP task categories such as question answering or text generation: solving each benchmark… See the full description on the dataset page: https://huggingface.co/datasets/xhyumiracle/thor25.PLI-parallax
PLI-Parallax
Predicted protein-ligand complexes are used as training data at considerable
scale today. Recent distillation sets contain several hundred thousand cofolded
BindingDB systems, and the filter applied to them is usually the predicting
model's own confidence score. That filter is a self-assessment, and cofolding
models have been shown to be confidently wrong in ways their own confidence does
not reveal.
This dataset provides the coordinates and derived distance labels… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/PLI-parallax.martial-arts-video-v1
Martial Arts Action Video Preview
This is a small preview of five short, staged-looking martial-arts action clips. The clips show guarded stances, upper-body strikes, close combat, footwork, falls/recovery, and apparent weapon exchanges.
Contents
videos/: five short MP4 clips
metadata.csv: one row per video with broad scene and action labels
segments.csv: coarse time windows with action descriptions
metadata.jsonl: the same metadata in JSON Lines format
frames/:… See the full description on the dataset page: https://huggingface.co/datasets/thordata/martial-arts-video-v1.theobroma
THEOBROMA v1.36
An aggregated open database of 1,132,805 natural products from 29 sources, with
per-compound license auditing, three-tier classification provenance, and
stereochemistry-aware deduplication.
Live instance: https://theobroma.l3s.uni-hannover.de
Archival record: https://doi.org/10.5281/zenodo.20443051 (concept DOI, resolves to latest)
This release: https://doi.org/10.5281/zenodo.22816330
Preprint: https://doi.org/10.64898/2026.06.12.731585
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/ThorKl/theobroma.
