CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hellisotherpeople /OpenDebateEvidence-Anonymized Dataset Card for OpenDebateEvidence (Anonymized) A collection of evidence used in collegiate and high school debate competitions, with all debater-identifying columns removed. This is an anonymized redistribution of Yusuf5/OpenCaselist. The argumentative content is byte-for-byte unchanged. 26 of the original 45 columns have been dropped. See Anonymization for exactly what was removed and why. Dataset Details Dataset Description This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.tabulartext-generation1M<n<10M0 likes944 downloads2mo agoHugging Face02Hellisotherpeople /DebateSum DebateSum Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset" Arxiv pre-print available here: https://arxiv.org/abs/2011.07251 Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9 Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/ Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.tabularquestion-answering100K<n<1M21 likes289 downloads4y agoHugging Face03helloadhavan /CC-FilteredCorpus English Cleaned Common Crawl Markdown Dataset An English-focused dataset created from Common Crawl, cleaned and converted to Markdown. The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting. Features English-focused Cleaned and filtered web content HTML converted to Markdown Exact and near-duplicate filtering GPT-2 perplexity filtering Stored as compressed Parquet shards Source The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.tabulartext-generation100K<n<1M1 likes251 downloads1mo agoHugging Face04Hellisotherpeople /OpenDebateEvidence-Deduplicated-Anonymized Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized) Debate evidence from collegiate and high school competitions, semantically deduplicated, with all debater-identifying columns removed. This is the semantically deduplicated companion to OpenDebateEvidence-Anonymized. Where the parent dataset contains every piece of evidence as used in every round, this version collapses repeated use of the same evidence into single records, making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.tabulartext-generation100K<n<1M0 likes199 downloads2mo agoHugging Face05helloadhavan /github_issues GitHub Pull Request Bug–Fix Dataset Kaggle url A curated, high-signal dataset of real-world software bugs and fixes collected from 25 popular open-source GitHub repositories.Each entry corresponds to a single pull request (PR) and pairs contextual metadata with the exact code changes (unified diffs) that fixed the bug. This dataset is designed for: Automated program repair Bug-fix patch generation LLM-based code and debugging agents Empirical software engineering research… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/github_issues.texttext-generation100K<n<1M4 likes136 downloads6mo agoHugging Face06sapienzanlp /hellaswag_italian HellaSwag - Italian (IT) This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence. Dataset Details The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.texttext-generation10K<n<100K1 likes95 downloads10mo agoHugging Face07quehry /HelloBenchHelloBench is an open-source benchmark designed to evaluate the long text generation capabilities of large language models (LLMs) from HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models. texttext-generationn<1K1 likes81 downloads2y agoHugging Face08Helloxiaolaodi /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes64 downloads1mo agoHugging Face09helloadhavan /python-docstrings Python Docstring Diff Dataset This dataset contains training samples for models that generate Python documentation patches. Each example provides a Python source file with its docstrings removed and a corresponding unified diff patch that restores the documentation. The dataset is designed for training or evaluating language models that assist with: Automatic code documentation Docstring generation Code review automation Developer tooling Dataset Structure Each entry contains the… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/python-docstrings.texttext-generation10K<n<100K2 likes57 downloads6mo agoHugging Face10Polygl0t /Hellaswag-poly HellaSwag Polyglot This dataset is a multilingual version of the original HellaSwag (Harder Endings, Longer contexts, and Low-shot Activities for Situations With Adversarial Generations) dataset, which consists short commonsense reasoning tasks designed to evaluate the ability of language models to understand and predict plausible continuations of given contexts. The polyglot version includes translations of the original English questions into various languages, allowing for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/Hellaswag-poly.tabulartext-generation100K<n<1M1 likes56 downloads11mo agoHugging Face11Lots-of-LoRAs /task1389_hellaswag_completion Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1389_hellaswag_completion Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1389_hellaswag_completion.texttext-generation1K<n<10K0 likes52 downloads2y agoHugging Face12Hellisotherpeople /Lipogram-e Dataset Card for Lipogram-e Dataset Summary This is a dataset of 3 English books which do not contain the letter "e" in them. This dataset includes all of "Gadsby" by Ernest Vincent Wright, all of "A Void" by Georges Perec, and almost all of "Eunoia" by Christian Bok (except for the single chapter that uses the letter "e" in it) This dataset is contributed as part of a paper titled "Most Language Models can be Poets too: An AI Writing Assistant and Constrained Text… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/Lipogram-e.texttext-generation10K<n<100K1 likes41 downloads4y agoHugging Face13Hellisotherpeople /OpenDebateEvidence-Annotated-Anonymized OpenDebateEvidence-Annotated (Anonymized) An LLM-annotated subset of OpenDebateEvidence debate evidence, with all debater-identifying columns removed. This is an anonymized, Parquet-converted redistribution of Hellisotherpeople/OpenDebateEvidence-Annotated. 85,600 rows, 44 columns: 25 annotation fields plus the 19 retained OpenDebateEvidence evidence fields. The original pipe-delimited CSV had 70 columns, of which 26 were dropped. See Anonymization. Companion datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Annotated-Anonymized.tabulartext-classification10K<n<100K0 likes39 downloads2mo agoHugging Face14LVSTCK /hellaswag-mk Hellaswag MK version This dataset is a Macedonian adaptation of the hellaswag dataset, originally curated (English -> Serbian) by Aleksa Gordić. It was translated from Serbian to Macedonian using the Google Translate API. You can find this dataset as part of the macedonian-llm-eval GitHub and HuggingFace. This dataset is used for training and evaluating models as described in Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language Why… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/hellaswag-mk.texttext-generation10K<n<100K1 likes35 downloads1y agoHugging Face15hellosindh /indus-script-synthetic Synthetic Indus Script Dataset This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions. Stage 1 — Train on real inscriptions: Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.texttext-generationn<1K0 likes34 downloads6mo agoHugging Face16Hellrabbit /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Hellrabbit/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K0 likes24 downloads9mo agoHugging Face17alvarobartt /hellaswag-okapi-eval-es HellaSwag translated to Spanish This dataset was generated by the Natural Language Processing Group of the University of Oregon, where they used the original HellaSwag dataset in English and translated it into different languages using ChatGPT. This dataset only contains the Spanish translation, but the following languages are also covered within the original subsets posted by the University of Oregon at http://nlp.uoregon.edu/download/okapi-eval/datasets/. Disclaimer… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/hellaswag-okapi-eval-es.texttext-generation1K<n<10K2 likes22 downloads3y agoHugging Face18s-conia /hellaswag_italian HellaSwag - Italian (IT) This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence. Dataset Details The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/hellaswag_italian.texttext-generation10K<n<100K0 likes22 downloads2y agoHugging Face19HelloImSteven /applescript-lines-annotated Dataset Card for "applescript-lines-annotated" Description This is a dataset of single lines of AppleScript code scraped from GitHub and GitHub Gist and manually annotated with descriptions, intents, prompts, and other metadata. Content Each row contains 8 features: text - The raw text of the AppleScript code. source - The name of the file from which the line originates. type - Either compiled (files using the .scpt extension) or uncompiled (everything else).… See the full description on the dataset page: https://huggingface.co/datasets/HelloImSteven/applescript-lines-annotated.textsummarizationn<1K2 likes20 downloads3y agoHugging Face20MihaiPopa2 /HelloWorldExamples Intro Welcome to the one-liner "Hello world!" examples! This is a dataset containing "Hello world!" examples in 10+ languages! Notes If you found a language that's not listed here, you can open a pull request! You can also create and train models, or even spaces! Note that this dataset is in CSV format, so it's not as flexible as JSON! texttext-generationn<1K1 likes20 downloads3y agoHugging Face21freococo /hellosayarwon_dataset HelloSayarWon Myanmar Health Articles Dataset This dataset contains 9,213 health-related articles sourced from HelloSayarWon, a Myanmar language health and wellness website. The articles are written by qualified experts, doctors, and specialists, and cover a wide range of health topics in Myanmar language. Usage This dataset is intended for Myanmar language research and AI applications, including but not limited to natural language processing, health text analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/hellosayarwon_dataset.texttext-generation1K<n<10K0 likes17 downloads1y agoHugging Face22ZennyKenny /yandexgptpro_4th_gen-hellaswag YandexGPT Pro (4th Gen) HellaSwag This dataset contains responses from the YandexGPT model evaluated on the HellaSwag benchmark. It was generated as part of an experiment to assess the model’s performance on multiple-choice commonsense reasoning tasks. Dataset Details Source: HellaSwag Model: YandexGPT via Yandex Cloud Foundation Models API Prompt style: Multiple-choice (A, B, C, D) with system prompt and task context Fields: id: index of the example context: the base… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/yandexgptpro_4th_gen-hellaswag.textzero-shot-classification10K<n<100K0 likes13 downloads2y agoHugging Face23SpyC0der77 /helloThis is just the word hello a bunch of times texttext-generation10K<n<100K0 likes13 downloads1y agoHugging Face24orlandoju /heller-gpt-dataset 🧉 Heller-GPT Dataset Dataset de entrevistas de Heller en formato ChatML multi-turn, diseñado para fine-tuning de LLMs. 📋 Descripción Fuente: Entrevistas de YouTube (canales de noticias y política argentina) Procesamiento: Audio → Whisper (transcripción) → PyAnnote/SpeechBrain (diarización) → ChatML Formato: Conversaciones multi-turn con roles system, user (entrevistador), assistant (Heller) Idioma: Español rioplatense argentino 📊 Estadísticas… See the full description on the dataset page: https://huggingface.co/datasets/orlandoju/heller-gpt-dataset.tabulartext-generationn<1K0 likes9 downloads3mo agoHugging Face25Hellisotherpeople /one_syllable Dataset Card for Lipogram-e Dataset Summary This is a dataset of English books which only write using one syllable at a time. At this time, the dataset only contains Robinson Crusoe — in Words of One Syllable by Lucy Aikin and Daniel Defoe This dataset is contributed as part of a paper titled "Most Language Models can be Poets too: An AI Writing Assistant and Constrained Text Generation Studio" to appear at COLING 2022. This dataset does not appear in the paper itself… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/one_syllable.texttext-generation1K<n<10K0 likes7 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.