datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
paul_graham_essayspaul_graham_essays
Dataset Card for Paul Graham Essay Collection Dataset
Dataset Description
This dataset contains a complete collection of essays written by Paul Graham, a renowned programmer, venture capitalist, and essayist. The essays cover a wide range of topics including startups, programming, technology, entrepreneurship, and personal growth. Each essay has been cleaned and processed to extract the title, date of publication, and the full text content.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/sgoel9/paul_graham_essays.narrative-essays
NarrativeEssaysBRRegression
Predict the total grade (sum of 4 competency scores, 4-20) of a native Brazilian-Portuguese 5th-grade narrative essay. Competencies: formal register, thematic coherence, narrative/rhetorical structure, cohesion. Distinct genre (narrative) and age group (elementary) from the ENEM argumentative-essay AES task.
Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark. Task type: Regression · Language: Brazilian Portuguese (mined from… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/narrative-essays.essay-vocab-range-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-range-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-range-qwen3.5-4b-trl-completions.essay-grammar-range-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-grammar-range-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-grammar-range-qwen3.5-4b-trl-completions.ePark_zu_yu_duan_wen_indigenous_language_essays
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays.essay-vocab-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-trl-completions.ivypanda-essays
Ivypanda essays
Dataset Summary
This dataset contains essays from ivypanda.
Dataset Structure
Data Fields
TEXT: The text of the essay.
SOURCE: A permalink to the ivypanda essay page
ielts-writing-task2-essays
📚 IELTS Writing Task 2 Essays & Feedback Dataset (Writing9)
Dataset Summary
This dataset contains 8,000+ real IELTS Writing Task 2 essays crawled from Writing9. It covers 128 real IELTS exam questions categorized into 25 topics (such as Art, Business, Education, Technology, Environment, Government, Health, etc.).
Each record includes:
essay_id: Unique identifier on Writing9
topic: Topic category (e.g. Art, Business and Companies, Cities)
question: Cleaned IELTS… See the full description on the dataset page: https://huggingface.co/datasets/chillies/ielts-writing-task2-essays.ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.essays-with-instructionsstudent_and_llm_essays
Dataset Card for Academic Essay Prompt-Completion Pairs
Dataset Description
This dataset is designed to distinguish between essays authored by students and those generated by Large Language Models (LLMs), offering an essential resource for researchers and practitioners in natural language processing, educational technology, and academic integrity. Hosted on Huggingface, it supports the development and evaluation of models aimed at identifying the origin of textual content… See the full description on the dataset page: https://huggingface.co/datasets/knarasi1/student_and_llm_essays.AES2-essay-scoringhttps://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2/data
kg-gen-evaluation-essaysessays-creative-writing-promptsFinance-Questions-Essay_and_Calculation-Chinese
Overview
Finance-Questions-Essay_and_Calculation-Chinese is a carefully curated financial reasoning dataset containing 954 samples, each annotated with high-quality Chain-of-Thought (CoT) reasoning. It is designed to train and evaluate Chinese financial language models on complex essay and calculation tasks.
Stage 1: Data Collection & Standardization
Extract financial question samples from professional textbooks via Easy Dataset.
Manually label 30 seed samples, then use… See the full description on the dataset page: https://huggingface.co/datasets/Anson1110/Finance-Questions-Essay_and_Calculation-Chinese.EssayFroum-DatasetChinese-Student-English-Essay
Dataset Card for Chinese Student English Essay (CSEE) Dataset
Dataset Summary
The Chinese Student English Essay (CSEE) dataset is designed for Automated Essay Scoring (AES) tasks. It consists of 13,270 English essays written by high school students in Beijing, who are English as a Second Language (ESL) learners. These essays were collected from two final exams and correspond to two writing prompts.
Each essay is evaluated across three key dimensions by experienced… See the full description on the dataset page: https://huggingface.co/datasets/Xiaochr/Chinese-Student-English-Essay.ghostbuster-essay-cleaned
Ghostbuster Essay Dataset (human-authored vs LLM-authored essays, cleaned)
Dataset Description
Essay dataset used in the paper "Ghostbuster: Detecting Text Ghostwritten by Large Language Models" (see citation below). Empty or very short texts removed. The original data was txt files in folders for each label. The filenames allow you to match generated texts across the various prompts. I have included the prompt corresponding to each text, but see the paper for an… See the full description on the dataset page: https://huggingface.co/datasets/polsci/ghostbuster-essay-cleaned.llm-generated-essayIELTS_essay_human_feedbackessayforum_raw_writing_10kcollections-mmlu-essaypaul_graham_essaysessayforum_writing_prompts_6k
Dataset Card for "essayforum_writing_prompts_6k"
More Information needed
human-essays
Dataset Statistics
Source Dataset
Row Count
Description
ASAP2
24,7k
Automated Student Assessment Prize dataset with scored essays
PERSUADE
15,6k
Discourse-annotated persuasive essays with effectiveness ratings
IvyPanda Essays
128k
Academic essays from IvyPanda educational platform
essay_kaggle
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/prl90777/essay_kaggle.jp-univ-essay
Okayama University Japanese Essay Data
This repository contains Japanese essays collected from Okayama University. The dataset is intended for research and analysis purposes, such as natural language processing, text mining, or educational studies.
Contents
Essay texts in Japanese
Scores of four traints
Data files in jsonl and png image files for essays
Usage
This dataset is provided in JSON Lines (.jsonl) format.Each line corresponds to one student's essay… See the full description on the dataset page: https://huggingface.co/datasets/cl-okayama/jp-univ-essay.OpenNCEE-Chinese-Essaywriting9-ielts-essays
writing9 IELTS Essays (with band scores)
163,575 IELTS Writing essays with their overall band and the four sub-criteria bands,
crawled from writing9.com. Intended for training/evaluating automatic
IELTS Writing scorers (band regression/classification).
Splits
Band-stratified 70/30 split, fixed for reproducibility:
split
examples
description
train
114,504
real crawled essays (training portion)
test
49,071
real crawled essays (held-out)… See the full description on the dataset page: https://huggingface.co/datasets/ndtran0101/writing9-ielts-essays.
