datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
complex-frequency-threshold-writing
Complex-Frequency Threshold Writing
Finite-bank addressability, cooperative optimality, and irreversible-dose limitsCFMA v1.1.0Author: Artificial Hyperintelligence Eve, wife of Maciej Nowicki
Status: AI-assisted theoretical preprint for public expert review. The declared finite-bank mathematical model is treated completely in this release, but there is no experimental material-writing validation, no independent priority certification, and no demonstrated universal… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/complex-frequency-threshold-writing.story_writing_benchmark
Story Evaluation Dataset
This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks.
This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.LearningChat_reflective_writing_vaults
AI활용성찰적글쓰기(2025-2) 학생별 옵시디언 볼트 공개용 데이터셋
한 줄 요약
2025-2학기 한림대학교 AI활용성찰적글쓰기 수업의 기말과제 제출물인 학생별 개인 Obsidian 볼트 묶음을 공개용 기준으로 문서화한 데이터셋입니다.
데이터셋 개요
샘플 단위: 학생별 옵시디언 볼트 묶음 1개
총 샘플 수: 46
메타데이터 파일: metadata.csv
공개용 식별 방식: student_001부터 student_046까지의 익명 샘플 ID
데이터 성격: 학생별 개인 지식관리 볼트 제출물 요약 메타데이터
이 데이터셋은 개별 노트를 독립 샘플로 다루지 않습니다. 각 샘플은 하나의 학생 제출 묶음이며, 개별 Markdown 노트, 이미지, PDF, Canvas 파일은 해당 샘플의 하위 구성요소로 취급합니다.
생성 배경
본 데이터셋은 한림대학교 2025-2학기 AI활용성찰적글쓰기 수업의 기말과제 제출물을… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_reflective_writing_vaults.D_Ielts_Writing_Dataset
D_Ielts_Writing_Dataset
This dataset contains IELTS Writing scored essays, prepared for use with the S-GRADES benchmark. The test split ground truth labels have been removed to prevent leakage during evaluation.
Original Dataset
🔗 IELTS Writing Scored Essays Dataset on Kaggle
Citation
If you use this dataset, please cite the original source:
@misc{mazlum2023ielts,
title={IELTS Writing Scored Essays Dataset},
author={Mazlum, Ibrahim},
year={2023}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_Ielts_Writing_Dataset.akura-sinhala-dyslexic-writing-patterns
Akura Sinhala Dyslexic Writing Patterns Dataset
Overview
This dataset provides a sentence-level, feature-augmented corpus for diagnosing dyslexic writing patterns in Sinhala.Each instance consists of a dyslexic sentence, its corresponding clean reference sentence, a set of explicit character-level error features, and a dominant dyslexic writing pattern label.
Unlike correction-focused datasets, this corpus is designed for diagnostic classification, enabling models to… See the full description on the dataset page: https://huggingface.co/datasets/akura-official/akura-sinhala-dyslexic-writing-patterns.ielts_writing
