datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
student_and_llm_essays
Dataset Card for Academic Essay Prompt-Completion Pairs
Dataset Description
This dataset is designed to distinguish between essays authored by students and those generated by Large Language Models (LLMs), offering an essential resource for researchers and practitioners in natural language processing, educational technology, and academic integrity. Hosted on Huggingface, it supports the development and evaluation of models aimed at identifying the origin of textual content… See the full description on the dataset page: https://huggingface.co/datasets/knarasi1/student_and_llm_essays.Finance-Questions-Essay_and_Calculation-Chinese
Overview
Finance-Questions-Essay_and_Calculation-Chinese is a carefully curated financial reasoning dataset containing 954 samples, each annotated with high-quality Chain-of-Thought (CoT) reasoning. It is designed to train and evaluate Chinese financial language models on complex essay and calculation tasks.
Stage 1: Data Collection & Standardization
Extract financial question samples from professional textbooks via Easy Dataset.
Manually label 30 seed samples, then use… See the full description on the dataset page: https://huggingface.co/datasets/Anson1110/Finance-Questions-Essay_and_Calculation-Chinese.writing9-ielts-essays
writing9 IELTS Essays (with band scores)
163,575 IELTS Writing essays with their overall band and the four sub-criteria bands,
crawled from writing9.com. Intended for training/evaluating automatic
IELTS Writing scorers (band regression/classification).
Splits
Band-stratified 70/30 split, fixed for reproducibility:
split
examples
description
train
114,504
real crawled essays (training portion)
test
49,071
real crawled essays (held-out)… See the full description on the dataset page: https://huggingface.co/datasets/ndtran0101/writing9-ielts-essays.viorra-admissions-essays
🎓 VIORRA Admissions Essays Dataset (v2.0.0)
⚖️ License: CC BY-NC 4.0 — Free for academic, research, and personal use. Attribution required. Commercial distribution or usage is strictly prohibited.
VIORRA Admissions Essays is a highly curated, multi-category dataset of successful personal statements that resulted in admission to elite global universities (Ivy League, Oxbridge, and Top 20 schools). It is explicitly designed to train, evaluate, and ground… See the full description on the dataset page: https://huggingface.co/datasets/qsardor/viorra-admissions-essays.essayAnnotated_Persuasive_Essays
Full Free Sample Dataset of 150 Annotated Essays available at https://driftlogic.ai
Join the conversation and let us hear your feedback/suggestions! https://discord.gg/PHp9SPRB
license: cc-by-nc-4.0
language:
- en
tags:
- ai
- machine-learning
- dataset
- ai/ml
- argument-mining
- argument
pretty_name: Sample Annotated Persuasive Essay Dataset
size_categories:
- n<1K
Persuasive Essay Argument-Mining Sample Dataset
This is a sample dataset that… See the full description on the dataset page: https://huggingface.co/datasets/DriftLogic/Annotated_Persuasive_Essays.essay_thesis_conversationsEvolution_Essay_Trumpessay-grading-criteriaessay_grading_for_instruction_tuningpretrain_synthetic_essays_testupdate:
pattern = r"[\d+]"
OL-science-mcq_essays_json_DSruler-ko-essaymath_essay_prmlexicurio-word-essays
Lexicurio Word Essays
A curated dataset of 5,000 rare and beautiful English words,
each paired with an original essay on why the word is great, a short definition, and
a strength score. It is a hand-quality-gated subset of the ~421k-word Lexicurio lexicon:
only words that earned a written essay are included, ranked by strength score.
Every essay is written by Lexicurio and every row links back to its source
page, e.g. https://lexicurio.com/word/petrichor.
Columns… See the full description on the dataset page: https://huggingface.co/datasets/Craiger/lexicurio-word-essays.essay-outlines
Essay Outlines
This dataset contains point-form notes paired with their corresponding essay outlines, revised for clarity and organization.
Overview
The outlines originate from agentlans/note-taking-v2.
The outlines were transformed into essay outlines using google/gemma-3-12b-it (zero-shot) and google/gemma-3-4b-it (distilled) models.
The revision process focuses on improving clarity, logical flow, and alignment with a strong central thesis.
Outlines feature a clear… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/essay-outlines.GAMSAT-Essays
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DharamArora/GAMSAT-Essays.IPA_Exam_PM2_essay
独立行政法人 情報処理推進機構(IPA) 情報処理技術者試験 試験問題データセット(午後2論文)
概要
本データセットは、独立行政法人 情報処理推進機構(以下、IPA)の情報処理技術者試験の午後2論文問題、出題趣旨、採点講評をセットにした、非公式のデータセットです。
以下、公開されている過去問題pdfから抽出しています。
https://www.ipa.go.jp/shiken/mondai-kaiotu/index.html
人間の学習目的での利用の他、LLMのベンチマークやファインチューニング等、生成AIの研究開発用途での利用を想定しています。
データの詳細
現状、過去5年分(2020〜2024)の以下試験区分について、問題、出題趣旨、採点講評を収録しています。
システムアーキテクト(sa)
プロジェクトマネージャ(pm)
ITストラテジスト(st)
ITサービスマネージャ(sm)
システム監査技術者(au)
エンベデッドシステムスペシャリスト(es)
注意事項… See the full description on the dataset page: https://huggingface.co/datasets/Dimeiza/IPA_Exam_PM2_essay.synthetic-essays
Synthetic Essays Dataset
A collection of university-level essays generated by AI models, covering a diverse range of academic topics.
Intended uses:
Studying and classifying academic writing styles.
Providing a repository of plausibly well-written essays free from plagiarism, copyright, or privacy concerns.
[!WARNING]⚠️ Warning: The content in this dataset is AI-generated and may contain inaccuracies or fabricated information.
By using this dataset, you confirm you understand… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/synthetic-essays.pretrain_synthetic_essays_testupdate:
pattern = r"[\d+]"
key2text_essays100-structured-essays-elementary-phisynthstress_essaysEssay-Score-zeroTOfive-CLEAN
