datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ielts-writing-task2-essays
📚 IELTS Writing Task 2 Essays & Feedback Dataset (Writing9)
Dataset Summary
This dataset contains 8,000+ real IELTS Writing Task 2 essays crawled from Writing9. It covers 128 real IELTS exam questions categorized into 25 topics (such as Art, Business, Education, Technology, Environment, Government, Health, etc.).
Each record includes:
essay_id: Unique identifier on Writing9
topic: Topic category (e.g. Art, Business and Companies, Cities)
question: Cleaned IELTS… See the full description on the dataset page: https://huggingface.co/datasets/chillies/ielts-writing-task2-essays.writing9-ielts-essays
writing9 IELTS Essays (with band scores)
163,575 IELTS Writing essays with their overall band and the four sub-criteria bands,
crawled from writing9.com. Intended for training/evaluating automatic
IELTS Writing scorers (band regression/classification).
Splits
Band-stratified 70/30 split, fixed for reproducibility:
split
examples
description
train
114,504
real crawled essays (training portion)
test
49,071
real crawled essays (held-out)… See the full description on the dataset page: https://huggingface.co/datasets/ndtran0101/writing9-ielts-essays.Dataset_Automatic_Essay_Scoring_Essay-EssayScore_and_24_textual_featuresbig5_essays_ru
Dataset Card for "big5_essays_ru"
More Information needed
human-essays-redditEssay prompts scraped from r/WritingPrompts,
from dates May 9, 2014 — August 16, 2022 (in an attempt to ensure that all written samples
are human-written).
This dataset includes the top 25% voted prompts, given they have a responding comment that
has top 25% votes for comments. In other words, each sample in this dataset is only kept if (a) the prompt is in the top 25% of votes for posts
and (b) the resulting top comment is in the top 25% percentile of votes for comments. We only take posts… See the full description on the dataset page: https://huggingface.co/datasets/jonathanli/human-essays-reddit.aide_essaysessays-big5-psycho-openai-gpt-4ominienem-essays-collectedessay-scoring-extraessays-gpt4o-guided-en-fr-bm-jp-zhessays-gpt4o-en-fr-bm-jp-zhessays-big5-openai-text-embedding-ada-002essays-gpt4o-guided-en-oen-fr-bm-jp-zhessays-reddit-samples-with-generationshuman-essays-reddit-samplesThe jonathanli/human-essays-reddit dataset,
with 2000 samples sampled using farthest point sampling ("most_different") and 2000 samples sampled randomly ("random_samples").
cowsl2h_essaysEssaysdataset, for Essay correction
