datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenDebateEvidence-Anonymized
Dataset Card for OpenDebateEvidence (Anonymized)
A collection of evidence used in collegiate and high school debate competitions,
with all debater-identifying columns removed.
This is an anonymized redistribution of
Yusuf5/OpenCaselist. The
argumentative content is byte-for-byte unchanged. 26 of the original 45 columns
have been dropped. See Anonymization for exactly what was
removed and why.
Dataset Details
Dataset Description
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.reasoning-traces
Anonymous Reasoning Traces
This repository contains data accompanying an anonymous TMLR submission. It provides
192,000 sampled mathematical reasoning traces from 20 model configurations on four
30-question benchmarks. Each question has 80 sampled responses.
Contents
The repository provides two representations of the same attempts:
Configuration
Rows
Approximate size
Contents
meta
192,000
1.36 GiB
All models without token-level arrays
20 per-model… See the full description on the dataset page: https://huggingface.co/datasets/AnonymizedTMLRSubmission/reasoning-traces.reasoning-lite
Anonymous Reasoning Lite
This repository contains data accompanying an anonymous TMLR submission. It provides
1,211,520 sampled reasoning attempts from four model configurations on competition-math
and field-balanced multiple-choice questions. The data omits top-20 alternative-token
distributions while retaining realized-token log probabilities and ranks.
Contents
The same attempts are available in full and metadata-only representations:
Configuration group… See the full description on the dataset page: https://huggingface.co/datasets/AnonymizedTMLRSubmission/reasoning-lite.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.complete-voiceai-speech-dataset-anonymized
🎙️ Silencio Network: Voice AI Sample Dataset
📊 This is a sample. The full Silencio corpus contains 100,000+ hours across 170+ countries and 100+ languages.
📧 Contact: sofia@silencioai.com for custom datasets, bulk licensing, or specific language requests.
🌍 Why Silencio Data?
Silencio data is collected in the wild from a massive, opt-in community (2M+ contributors across 180+ countries), giving you:
✅ Real-world accents, dialects, devices, and… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/complete-voiceai-speech-dataset-anonymized.northwind_sales_anonymized_2023
Sales Transactions (Anonymized)
Anonymized sales transaction records generated internally by Northwind Analytics. No external source.
License: MIT.
real_human_vortexing_lerobot_anonymizedFace-anonymized copy of labos-sim/real_human_vortexing_lerobot.
Anonymization
This is a face-anonymized copy of labos-sim/real_human_vortexing_lerobot. Faces were blurred with deface (CenterFace detector, threshold 0.25, Gaussian blur, detection at 960x600, output at original resolution). Only detected face regions differ from the original dataset; all annotations, metadata, and non-face pixels are identical. Benchmark results reported for LabOS-Sim were computed on the original… See the full description on the dataset page: https://huggingface.co/datasets/labos-sim/real_human_vortexing_lerobot_anonymized.bsky-firehose-anonymized-dec-2025
Bluesky Firehose: Anonymized Posts (Dec 2025)
101,040 Bluesky posts collected via the AT Protocol firehose, December 2-25, 2025. All author DIDs, post URIs, and thread relationships are SHA-256 hashed. Includes sentiment scores (VADER), language detection across 90 languages, media flags, and thread structure.
Posts: 101,040
Unique Authors: 43,998
Languages: 90 detected (60.8% English, 12.5% Japanese, 11.4% unknown)
Collection: Bluesky AT Protocol Jetstream WebSocket… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bsky-firehose-anonymized-dec-2025.real_human_pipetting_lerobot_anonymizedThis dataset was created using LeRobot.
Dataset Description
Multi-camera real-human laboratory demonstrations for pipetting with success/failure annotations.
Homepage: https://huggingface.co/datasets/labos-sim/real_human
Paper: [More Information Needed]
License: [More Information Needed]
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.front": {
"dtype": "video"… See the full description on the dataset page: https://huggingface.co/datasets/labos-sim/real_human_pipetting_lerobot_anonymized.verilog_preprocessed_anonymizedOpenDebateEvidence-Annotated-Anonymized
OpenDebateEvidence-Annotated (Anonymized)
An LLM-annotated subset of OpenDebateEvidence debate evidence, with all
debater-identifying columns removed.
This is an anonymized, Parquet-converted redistribution of
Hellisotherpeople/OpenDebateEvidence-Annotated.
85,600 rows, 44 columns: 25 annotation fields plus the 19 retained OpenDebateEvidence
evidence fields. The original pipe-delimited CSV had 70 columns, of which 26 were
dropped. See Anonymization.
Companion datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Annotated-Anonymized.OkCupid-59k-Anonymized-Profiles
💘 OkCupid 59k Anonymized Profiles
This dataset contains 59k anonymized OkCupid dating profiles, converted from the original CSV dataset into Parquet format.
It includes structured profile attributes such as age, gender, orientation, body type, lifestyle habits, education, job, location, and several free-text essay fields written by users.
✍️ Essay Fields
The columns essay0 to essay9 correspond to open-ended profile questions from OkCupid.
These fields contain natural… See the full description on the dataset page: https://huggingface.co/datasets/SpiceeChat/OkCupid-59k-Anonymized-Profiles.rw_amazon-ratings_node2vec3_2_public_anonymizedrw_amazon-ratings_nbw_2_public_anonymizedrw_roman-empire_standard_1_publ_anonymizedicrw_amazon-ratings_mdlr_6_public_anonymizedrw_amazon-ratings_standard_6_public_anonymizedminuszero-indian-autonomous-driving-dataset-v2-anonymized
MinusZero Indian Autonomous Driving Dataset V2 Anonymized
This public, manually gated dataset is being prepared. Payload publication is
blocked until privacy, temporal, and immutable remote verification complete.
rw_amazon-ratings_standard_1_publ_anonymizedicrw_amazon-ratings_node2vec_2_public_anonymizedrw_roman-empire_nbw_1_public_anonymizedrw_amazon-ratings_node2vec_6_public_anonymizedrw_roman-empire_node2vec2_2_public_anonymizedMCAD-CIC-3xN
MCAD-CIC-3xN
Dataset Summary
MCAD-CIC-3xN is a multi-source continual anomaly detection benchmark scenario for network intrusion detection. It combines three CIC-family source datasets into a 13-task continual-learning scenario:
CIC-IDS2017-derived tasks;
CIC-IDS2018-derived tasks;
CIC-UNSW-NB15-derived tasks.
Unlike a single-task-per-source construction, this benchmark provides multiple concept-grouped tasks per source dataset. It is intended to evaluate continual… See the full description on the dataset page: https://huggingface.co/datasets/anonymizeddb/MCAD-CIC-3xN.rw_roman-empire_nbw_2_public_anonymizedrw_roman-empire_node2vec_1_public_anonymizedrw_roman-empire_node2vec_2_public_anonymizedrw_roman-empire_node2vec_6_public_anonymizedrw_amazon-ratings_nbw_6_public_anonymizedrw_amazon-ratings_node2vec_1_public_anonymized
