datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenDebateEvidence-Anonymized
Dataset Card for OpenDebateEvidence (Anonymized)
A collection of evidence used in collegiate and high school debate competitions,
with all debater-identifying columns removed.
This is an anonymized redistribution of
Yusuf5/OpenCaselist. The
argumentative content is byte-for-byte unchanged. 26 of the original 45 columns
have been dropped. See Anonymization for exactly what was
removed and why.
Dataset Details
Dataset Description
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.DebateSum
DebateSum
Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset"
Arxiv pre-print available here: https://arxiv.org/abs/2011.07251
Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9
Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/
Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.CC-FilteredCorpus
English Cleaned Common Crawl Markdown Dataset
An English-focused dataset created from Common Crawl, cleaned and converted to Markdown.
The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting.
Features
English-focused
Cleaned and filtered web content
HTML converted to Markdown
Exact and near-duplicate filtering
GPT-2 perplexity filtering
Stored as compressed Parquet shards
Source
The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.github_issues
GitHub Pull Request Bug–Fix Dataset
Kaggle url
A curated, high-signal dataset of real-world software bugs and fixes collected from 25 popular open-source GitHub repositories.Each entry corresponds to a single pull request (PR) and pairs contextual metadata with the exact code changes (unified diffs) that fixed the bug.
This dataset is designed for:
Automated program repair
Bug-fix patch generation
LLM-based code and debugging agents
Empirical software engineering research… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/github_issues.hellaswag_italian
HellaSwag - Italian (IT)
This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence.
Dataset Details
The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.HelloBenchHelloBench is an open-source benchmark designed to evaluate the long text generation capabilities of large language models (LLMs) from HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models.
qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/Helloxiaolaodi/qwen3.8-max-glm5.2-kimi-k3-distillation.python-docstrings
Python Docstring Diff Dataset
This dataset contains training samples for models that generate Python documentation patches.
Each example provides a Python source file with its docstrings removed and a corresponding unified diff patch that restores the documentation.
The dataset is designed for training or evaluating language models that assist with:
Automatic code documentation
Docstring generation
Code review automation
Developer tooling
Dataset Structure
Each entry contains the… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/python-docstrings.Hellaswag-poly
HellaSwag Polyglot
This dataset is a multilingual version of the original HellaSwag (Harder Endings, Longer contexts, and Low-shot Activities for Situations With Adversarial Generations) dataset, which consists short commonsense reasoning tasks designed to evaluate the ability of language models to understand and predict plausible continuations of given contexts. The polyglot version includes translations of the original English questions into various languages, allowing for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/Hellaswag-poly.task1389_hellaswag_completion
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1389_hellaswag_completion
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1389_hellaswag_completion.Lipogram-e
Dataset Card for Lipogram-e
Dataset Summary
This is a dataset of 3 English books which do not contain the letter "e" in them. This dataset includes all of "Gadsby" by Ernest Vincent Wright, all of "A Void" by Georges Perec, and almost all of "Eunoia" by Christian Bok (except for the single chapter that uses the letter "e" in it) This dataset is contributed as part of a paper titled "Most Language Models can be Poets too: An AI Writing Assistant and Constrained Text… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/Lipogram-e.OpenDebateEvidence-Annotated-Anonymized
OpenDebateEvidence-Annotated (Anonymized)
An LLM-annotated subset of OpenDebateEvidence debate evidence, with all
debater-identifying columns removed.
This is an anonymized, Parquet-converted redistribution of
Hellisotherpeople/OpenDebateEvidence-Annotated.
85,600 rows, 44 columns: 25 annotation fields plus the 19 retained OpenDebateEvidence
evidence fields. The original pipe-delimited CSV had 70 columns, of which 26 were
dropped. See Anonymization.
Companion datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Annotated-Anonymized.hellaswag-mk
Hellaswag MK version
This dataset is a Macedonian adaptation of the hellaswag dataset, originally curated (English -> Serbian) by Aleksa Gordić. It was translated from Serbian to Macedonian using the Google Translate API.
You can find this dataset as part of the macedonian-llm-eval GitHub and HuggingFace.
This dataset is used for training and evaluating models as described in Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language
Why… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/hellaswag-mk.indus-script-synthetic
Synthetic Indus Script Dataset
This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions.
Stage 1 — Train on real inscriptions:
Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Hellrabbit/medical-o1-reasoning-SFT.hellaswag-okapi-eval-es
HellaSwag translated to Spanish
This dataset was generated by the Natural Language Processing Group of the University of Oregon, where they used the
original HellaSwag dataset in English and translated it into different languages using ChatGPT.
This dataset only contains the Spanish translation, but the following languages are also covered within the original
subsets posted by the University of Oregon at http://nlp.uoregon.edu/download/okapi-eval/datasets/.
Disclaimer… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/hellaswag-okapi-eval-es.hellaswag_italian
HellaSwag - Italian (IT)
This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence.
Dataset Details
The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/hellaswag_italian.applescript-lines-annotated
Dataset Card for "applescript-lines-annotated"
Description
This is a dataset of single lines of AppleScript code scraped from GitHub and GitHub Gist and manually annotated with descriptions, intents, prompts, and other metadata.
Content
Each row contains 8 features:
text - The raw text of the AppleScript code.
source - The name of the file from which the line originates.
type - Either compiled (files using the .scpt extension) or uncompiled (everything else).… See the full description on the dataset page: https://huggingface.co/datasets/HelloImSteven/applescript-lines-annotated.HelloWorldExamples
Intro
Welcome to the one-liner "Hello world!" examples!
This is a dataset containing "Hello world!" examples in 10+ languages!
Notes
If you found a language that's not listed here, you can open a pull request!
You can also create and train models, or even spaces!
Note that this dataset is in CSV format, so it's not as flexible as JSON!
hellosayarwon_dataset
HelloSayarWon Myanmar Health Articles Dataset
This dataset contains 9,213 health-related articles sourced from HelloSayarWon, a Myanmar language health and wellness website. The articles are written by qualified experts, doctors, and specialists, and cover a wide range of health topics in Myanmar language.
Usage
This dataset is intended for Myanmar language research and AI applications, including but not limited to natural language processing, health text analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/hellosayarwon_dataset.yandexgptpro_4th_gen-hellaswag
YandexGPT Pro (4th Gen) HellaSwag
This dataset contains responses from the YandexGPT model evaluated on the HellaSwag benchmark. It was generated as part of an experiment to assess the model’s performance on multiple-choice commonsense reasoning tasks.
Dataset Details
Source: HellaSwag
Model: YandexGPT via Yandex Cloud Foundation Models API
Prompt style: Multiple-choice (A, B, C, D) with system prompt and task context
Fields:
id: index of the example
context: the base… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/yandexgptpro_4th_gen-hellaswag.helloThis is just the word hello a bunch of times
heller-gpt-dataset
🧉 Heller-GPT Dataset
Dataset de entrevistas de Heller en formato ChatML multi-turn, diseñado para fine-tuning de LLMs.
📋 Descripción
Fuente: Entrevistas de YouTube (canales de noticias y política argentina)
Procesamiento: Audio → Whisper (transcripción) → PyAnnote/SpeechBrain (diarización) → ChatML
Formato: Conversaciones multi-turn con roles system, user (entrevistador), assistant (Heller)
Idioma: Español rioplatense argentino
📊 Estadísticas… See the full description on the dataset page: https://huggingface.co/datasets/orlandoju/heller-gpt-dataset.one_syllable
Dataset Card for Lipogram-e
Dataset Summary
This is a dataset of English books which only write using one syllable at a time. At this time, the dataset only contains Robinson Crusoe — in Words of One Syllable by Lucy Aikin and Daniel Defoe
This dataset is contributed as part of a paper titled "Most Language Models can be Poets too: An AI Writing Assistant and Constrained Text Generation Studio" to appear at COLING 2022. This dataset does not appear in the paper itself… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/one_syllable.
