CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hellisotherpeople /OpenDebateEvidence-Anonymized Dataset Card for OpenDebateEvidence (Anonymized) A collection of evidence used in collegiate and high school debate competitions, with all debater-identifying columns removed. This is an anonymized redistribution of Yusuf5/OpenCaselist. The argumentative content is byte-for-byte unchanged. 26 of the original 45 columns have been dropped. See Anonymization for exactly what was removed and why. Dataset Details Dataset Description This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.tabulartext-generation1M<n<10M0 likes944 downloads2mo agoHugging Face02Hellisotherpeople /DebateSum DebateSum Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset" Arxiv pre-print available here: https://arxiv.org/abs/2011.07251 Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9 Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/ Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.tabularquestion-answering100K<n<1M21 likes289 downloads4y agoHugging Face03helloadhavan /CC-FilteredCorpus English Cleaned Common Crawl Markdown Dataset An English-focused dataset created from Common Crawl, cleaned and converted to Markdown. The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting. Features English-focused Cleaned and filtered web content HTML converted to Markdown Exact and near-duplicate filtering GPT-2 perplexity filtering Stored as compressed Parquet shards Source The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.tabulartext-generation100K<n<1M1 likes251 downloads1mo agoHugging Face04Hellisotherpeople /OpenDebateEvidence-Deduplicated-Anonymized Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized) Debate evidence from collegiate and high school competitions, semantically deduplicated, with all debater-identifying columns removed. This is the semantically deduplicated companion to OpenDebateEvidence-Anonymized. Where the parent dataset contains every piece of evidence as used in every round, this version collapses repeated use of the same evidence into single records, making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.tabulartext-generation100K<n<1M0 likes199 downloads2mo agoHugging Face05Helloxiaolaodi /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes64 downloads1mo agoHugging Face06Polygl0t /Hellaswag-poly HellaSwag Polyglot This dataset is a multilingual version of the original HellaSwag (Harder Endings, Longer contexts, and Low-shot Activities for Situations With Adversarial Generations) dataset, which consists short commonsense reasoning tasks designed to evaluate the ability of language models to understand and predict plausible continuations of given contexts. The polyglot version includes translations of the original English questions into various languages, allowing for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/Hellaswag-poly.tabulartext-generation100K<n<1M1 likes56 downloads11mo agoHugging Face07Hellisotherpeople /OpenDebateEvidence-Annotated-Anonymized OpenDebateEvidence-Annotated (Anonymized) An LLM-annotated subset of OpenDebateEvidence debate evidence, with all debater-identifying columns removed. This is an anonymized, Parquet-converted redistribution of Hellisotherpeople/OpenDebateEvidence-Annotated. 85,600 rows, 44 columns: 25 annotation fields plus the 19 retained OpenDebateEvidence evidence fields. The original pipe-delimited CSV had 70 columns, of which 26 were dropped. See Anonymization. Companion datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Annotated-Anonymized.tabulartext-classification10K<n<100K0 likes39 downloads2mo agoHugging Face08orlandoju /heller-gpt-dataset 🧉 Heller-GPT Dataset Dataset de entrevistas de Heller en formato ChatML multi-turn, diseñado para fine-tuning de LLMs. 📋 Descripción Fuente: Entrevistas de YouTube (canales de noticias y política argentina) Procesamiento: Audio → Whisper (transcripción) → PyAnnote/SpeechBrain (diarización) → ChatML Formato: Conversaciones multi-turn con roles system, user (entrevistador), assistant (Heller) Idioma: Español rioplatense argentino 📊 Estadísticas… See the full description on the dataset page: https://huggingface.co/datasets/orlandoju/heller-gpt-dataset.tabulartext-generationn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.