datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenDebateEvidence-Anonymized
Dataset Card for OpenDebateEvidence (Anonymized)
A collection of evidence used in collegiate and high school debate competitions,
with all debater-identifying columns removed.
This is an anonymized redistribution of
Yusuf5/OpenCaselist. The
argumentative content is byte-for-byte unchanged. 26 of the original 45 columns
have been dropped. See Anonymization for exactly what was
removed and why.
Dataset Details
Dataset Description
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.DebateSum
DebateSum
Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset"
Arxiv pre-print available here: https://arxiv.org/abs/2011.07251
Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9
Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/
Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.CC-FilteredCorpus
English Cleaned Common Crawl Markdown Dataset
An English-focused dataset created from Common Crawl, cleaned and converted to Markdown.
The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting.
Features
English-focused
Cleaned and filtered web content
HTML converted to Markdown
Exact and near-duplicate filtering
GPT-2 perplexity filtering
Stored as compressed Parquet shards
Source
The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/qwen3.8-max-glm5.2-kimi-k3-distillation.Hellaswag-poly
HellaSwag Polyglot
This dataset is a multilingual version of the original HellaSwag (Harder Endings, Longer contexts, and Low-shot Activities for Situations With Adversarial Generations) dataset, which consists short commonsense reasoning tasks designed to evaluate the ability of language models to understand and predict plausible continuations of given contexts. The polyglot version includes translations of the original English questions into various languages, allowing for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/Hellaswag-poly.OpenDebateEvidence-Annotated-Anonymized
OpenDebateEvidence-Annotated (Anonymized)
An LLM-annotated subset of OpenDebateEvidence debate evidence, with all
debater-identifying columns removed.
This is an anonymized, Parquet-converted redistribution of
Hellisotherpeople/OpenDebateEvidence-Annotated.
85,600 rows, 44 columns: 25 annotation fields plus the 19 retained OpenDebateEvidence
evidence fields. The original pipe-delimited CSV had 70 columns, of which 26 were
dropped. See Anonymization.
Companion datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Annotated-Anonymized.heller-gpt-dataset
🧉 Heller-GPT Dataset
Dataset de entrevistas de Heller en formato ChatML multi-turn, diseñado para fine-tuning de LLMs.
📋 Descripción
Fuente: Entrevistas de YouTube (canales de noticias y política argentina)
Procesamiento: Audio → Whisper (transcripción) → PyAnnote/SpeechBrain (diarización) → ChatML
Formato: Conversaciones multi-turn con roles system, user (entrevistador), assistant (Heller)
Idioma: Español rioplatense argentino
📊 Estadísticas… See the full description on the dataset page: https://huggingface.co/datasets/orlandoju/heller-gpt-dataset.
