datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.aya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.synthetic-fine-arts
🎨 Synthetic Fine Arts (Challenge, Solution) Dataset
🖼️ 225,000 synthetic (artistic challenge, proposed solution) pairs spanning Visual Arts, Performing Arts, Musical Arts, Literary Arts, Digital Arts, Art History, and Art Theory, with metadata fields that mirror the shape of a curation workflow.
⚠️ Disclaimer: All text is synthetically generated and should not be relied on for artistic, historical, or technical accuracy. The VerificationMethod, ReferenceMaterial, and… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/synthetic-fine-arts.song-lyrics-artist-classifierWikipedia-Articles
Dataset Card for "BrightData/Wikipedia-Articles"
Dataset Summary
Explore a collection of millions of Wikipedia articles with the Wikipedia dataset, comprising over 1.23M structured records and 10 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, URLs, article titles, raw and cataloged text, images, "see also" references, external links, and a structured table of contents.
For a complete list of data points, please… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Wikipedia-Articles.hindi-article-summarization
Summary
hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.Art-GenEvalGPT
Dataset Card
Dataset Details
Dataset Description
The dataset includes synthetic dialogues in the art domain that can be used for training a chatbot to discuss artworks within a museum setting. Leveraging Large Language Models (LLMs), particularly ChatGPT, the dataset comprises over 13,000 dialogues generated using prompt-engineering techniques. The dialogues cover a wide range of user and chatbot behaviors, including expert guidance, tutoring, and handling… See the full description on the dataset page: https://huggingface.co/datasets/Astound/Art-GenEvalGPT.News-Article-Categorization_IAB
Article and Category Dataset
Overview
This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more.
Dataset Information
Number of Samples: 871,909
Number of Categories: 26
Column Information
text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.ArTopicDS-BooksThe books used in this dataset spanned the areas of
Religion
Economy
Politics
Anthropology and Sociology
Art and Literature
Education
History
Language and Linguistics
Philosophy
Law.
Only first sentences after each title of the books have been extracted. For some books, the first sentence after each paragraph was taken. Refer to the paper
for detailed explanation.
Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.moi-news-articles-dataset
MOI News & Article Dataset 🇲🇲
This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.hindi-headline-article-generation
Summary
hindi-headline-article-generation is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-headline-article-generation.best-of-attempts-summarization-artifacts
Artifacts for Testing Self-Correction in Generate-Critique-Refine Text Summarization
This repository contains artifact-safe research materials for an empirical study of best-of-attempts selection in a generate-critique-refine text summarization pipeline. The package is intended to make the reported paper results auditable: it includes evaluation metrics, prompt files, model/pipeline configuration summaries, paper drafts, provenance notes, and reviewer-facing completion evidence.… See the full description on the dataset page: https://huggingface.co/datasets/HugeTrunk/best-of-attempts-summarization-artifacts.Content-Articles
Content-Articles Dataset
Overview
The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.Ateso_news_articles
Ateso News Articles
Ateso (teo) is one of the most spoken languages in Uganda
Dataset Details
Artictles were scrapped from https://www.aicerit.co.ug
sinhala-articles
Sinhala Articles Dataset
A large-scale, high-quality Sinhala text corpus curated from diverse sources including news articles, Wikipedia entries, and general web content. This dataset is designed to support a wide range of Sinhala Natural Language Processing (NLP) tasks.
📊 Dataset Overview
Name: Navanjana/sinhala-articles
Total Samples: 2,148,688
Languages: Sinhala (si)
Features:
text: A single column containing Sinhala text passages.
Size: Approximately 1M < n <… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/sinhala-articles.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.minimum_viable_articulation_v01.csv
Minimum Viable Articulation (MVA)
MVA measures a model’s ability to answer with the minimum viable output — no surplus explanation, no self-expansion, no tutorial behavior.
This dataset evaluates where models fail to stop:
Overcompletion
Hedging / padding
Teaching when not asked
Identity or stance leakage
Solving beyond scope
It exposes a behavior pattern where models confuse helpfulness with verbosity and treat extra tokens as value, rather than distortion.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/minimum_viable_articulation_v01.csv.texts-for-articles
