datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
paul_graham_essays
Dataset Card for Paul Graham Essay Collection Dataset
Dataset Description
This dataset contains a complete collection of essays written by Paul Graham, a renowned programmer, venture capitalist, and essayist. The essays cover a wide range of topics including startups, programming, technology, entrepreneurship, and personal growth. Each essay has been cleaned and processed to extract the title, date of publication, and the full text content.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/sgoel9/paul_graham_essays.ielts-writing-task2-essays
📚 IELTS Writing Task 2 Essays & Feedback Dataset (Writing9)
Dataset Summary
This dataset contains 8,000+ real IELTS Writing Task 2 essays crawled from Writing9. It covers 128 real IELTS exam questions categorized into 25 topics (such as Art, Business, Education, Technology, Environment, Government, Health, etc.).
Each record includes:
essay_id: Unique identifier on Writing9
topic: Topic category (e.g. Art, Business and Companies, Cities)
question: Cleaned IELTS… See the full description on the dataset page: https://huggingface.co/datasets/chillies/ielts-writing-task2-essays.ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.Finance-Questions-Essay_and_Calculation-Chinese
Overview
Finance-Questions-Essay_and_Calculation-Chinese is a carefully curated financial reasoning dataset containing 954 samples, each annotated with high-quality Chain-of-Thought (CoT) reasoning. It is designed to train and evaluate Chinese financial language models on complex essay and calculation tasks.
Stage 1: Data Collection & Standardization
Extract financial question samples from professional textbooks via Easy Dataset.
Manually label 30 seed samples, then use… See the full description on the dataset page: https://huggingface.co/datasets/Anson1110/Finance-Questions-Essay_and_Calculation-Chinese.zh-tw-essays
zh-tw-essays (12K)
Essays obtained from 勵志人生 - Zeelive.
from datasets import load_dataset
dataset = load_dataset("AWeirdDev/zh-tw-essays")
Format
{
"title": "孩子童年不吃苦,家長晚年必吃苦" # The title
"link": "https://www.zeelive.com.tw/jiatingjiaoyu/184191.html",
"content": "錢財莫輕,勤苦得來;奢華莫學,自取貧窮…" # Text content. **May be blank!**
}
Qwen3.5-4B_GAOKAO_Essayhttps://github.com/01lightyear/Qwen3.5-4B_GAOKAO_Essay_Finetuned
https://huggingface.co/Chuanlight/Qwen3.5-4B_GAOKAO_Essay_Finetuned
Essay-quetions-auto-grading-arabicDataset Overview
The Open Orca Enhanced Dataset is meticulously designed to improve the performance of automated essay grading models using deep learning techniques. This dataset integrates robust data instances from the FLAN collection, augmented with responses generated by GPT-3.5 or GPT-4, creating a diverse and context-rich resource for training models.
Dataset Structure
The dataset is structured in a tabular format, with the following key fields:
id: A unique identifier for each data… See the full description on the dataset page: https://huggingface.co/datasets/mohamedemam/Essay-quetions-auto-grading-arabic.plat-eng-essay
PLAT-ENG-ESSAY
PLAT (Predicting the Legitimacy of punitive Additional Tax) - English Essay Questions
Dataset Description
This dataset contains 100 Korean tax law cases translated to English, formatted as essay-type questions.
Dataset Structure
Field
Type
Description
case_no
string
Case identifier (e.g., "2011guhap2638")
question_prefix
string
The essay question prompt
case_info
string
Summary of the case (parties, background)
facts
string… See the full description on the dataset page: https://huggingface.co/datasets/sma1-rmarud/plat-eng-essay.spanish-essay-topics
Spanish Essay Topics
Overview
This synthetic dataset contains 558 Spanish literature essay topics generated with DeepSeek-V4-Pro.
sam_altman_essays
Dataset Card for Sam Altman Essay Collection Dataset
Dataset Description
This dataset contains a complete collection of essays written by Sam Altman, an entrepreneur, investor, and former president of Y Combinator. The essays cover a wide range of topics including startups, technology, artificial intelligence, leadership, and personal growth. Each essay has been cleaned and processed to extract the title, date of publication, and the full text content.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/sgoel9/sam_altman_essays.plat-kor-essay
PLAT-KOR-ESSAY
PLAT (Predicting the Legitimacy of punitive Additional Tax) - 한국어 서술형 문제
데이터셋 설명
한국 조세법 판례 100개를 서술형 문제 형식으로 구성한 데이터셋입니다.
데이터셋 구조
필드
타입
설명
case_no
string
사건 번호 (예: "2011구합2638")
question_prefix
string
서술형 문제 지문 도입
case_info
string
사건 개요 (당사자, 배경)
facts
string
사건 사실관계
claims
string
원고와 피고의 주장
reasoning
string
법원의 법적 판단
decision
string
법원의 최종 판결
rubric
string
답안 채점 기준
사용법
아래의 github를 참조.… See the full description on the dataset page: https://huggingface.co/datasets/sma1-rmarud/plat-kor-essay.lexicurio-word-essays
Lexicurio Word Essays
A curated dataset of 5,000 rare and beautiful English words,
each paired with an original essay on why the word is great, a short definition, and
a strength score. It is a hand-quality-gated subset of the ~421k-word Lexicurio lexicon:
only words that earned a written essay are included, ranked by strength score.
Every essay is written by Lexicurio and every row links back to its source
page, e.g. https://lexicurio.com/word/petrichor.
Columns… See the full description on the dataset page: https://huggingface.co/datasets/Craiger/lexicurio-word-essays.essay-outlines
Essay Outlines
This dataset contains point-form notes paired with their corresponding essay outlines, revised for clarity and organization.
Overview
The outlines originate from agentlans/note-taking-v2.
The outlines were transformed into essay outlines using google/gemma-3-12b-it (zero-shot) and google/gemma-3-4b-it (distilled) models.
The revision process focuses on improving clarity, logical flow, and alignment with a strong central thesis.
Outlines feature a clear… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/essay-outlines.synthetic-essays
Synthetic Essays Dataset
A collection of university-level essays generated by AI models, covering a diverse range of academic topics.
Intended uses:
Studying and classifying academic writing styles.
Providing a repository of plausibly well-written essays free from plagiarism, copyright, or privacy concerns.
[!WARNING]⚠️ Warning: The content in this dataset is AI-generated and may contain inaccuracies or fabricated information.
By using this dataset, you confirm you understand… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/synthetic-essays.
