anonymized
OpenDebateEvidence-Anonymized
Dataset Card for OpenDebateEvidence (Anonymized)
A collection of evidence used in collegiate and high school debate competitions,
with all debater-identifying columns removed.
This is an anonymized redistribution of
Yusuf5/OpenCaselist. The
argumentative content is byte-for-byte unchanged. 26 of the original 45 columns
have been dropped. See Anonymization for exactly what was
removed and why.
Dataset Details
Dataset Description
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.reasoning-traces
Anonymous Reasoning Traces
This repository contains data accompanying an anonymous TMLR submission. It provides
192,000 sampled mathematical reasoning traces from 20 model configurations on four
30-question benchmarks. Each question has 80 sampled responses.
Contents
The repository provides two representations of the same attempts:
Configuration
Rows
Approximate size
Contents
meta
192,000
1.36 GiB
All models without token-level arrays
20 per-model… See the full description on the dataset page: https://huggingface.co/datasets/AnonymizedTMLRSubmission/reasoning-traces.reasoning-lite
Anonymous Reasoning Lite
This repository contains data accompanying an anonymous TMLR submission. It provides
1,211,520 sampled reasoning attempts from four model configurations on competition-math
and field-balanced multiple-choice questions. The data omits top-20 alternative-token
distributions while retaining realized-token log probabilities and ranks.
Contents
The same attempts are available in full and metadata-only representations:
Configuration group… See the full description on the dataset page: https://huggingface.co/datasets/AnonymizedTMLRSubmission/reasoning-lite.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.complete-voiceai-speech-dataset-anonymized
🎙️ Silencio Network: Voice AI Sample Dataset
📊 This is a sample. The full Silencio corpus contains 100,000+ hours across 170+ countries and 100+ languages.
📧 Contact: sofia@silencioai.com for custom datasets, bulk licensing, or specific language requests.
🌍 Why Silencio Data?
Silencio data is collected in the wild from a massive, opt-in community (2M+ contributors across 180+ countries), giving you:
✅ Real-world accents, dialects, devices, and… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/complete-voiceai-speech-dataset-anonymized.northwind_sales_anonymized_2023
Sales Transactions (Anonymized)
Anonymized sales transaction records generated internally by Northwind Analytics. No external source.
License: MIT.
