datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
filtered_articles_by_year
Dataset Card for Filtered Articles by Year
Dataset Summary
The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time.
Supported Tasks and Leaderboards
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.SP500-Financial-News-Articles-Time-SeriesTextual Time Series Dataset for finetuning / pretraining.
Json version of original dataset.
Original Dataset : https://www.kaggle.com/datasets/skywalker290/financial-news-article-and-stock-trend-dataset?select=stock_data_articles.csv
v4_nuclear_power_articles
Dataset Card for Nuclear News V4 Dataset
Dataset Summary
The Nuclear News V4 Dataset is a multilingual dataset consisting of 33,104 unique news articles sourced from 12 online news platforms across the Visegrád Group (V4) countries — Poland, Czech Republic, Slovakia, and Hungary — published between 1998 and 2025.
The goal of the dataset is to analyze media narratives surrounding nuclear energy in Central Europe.
While the dataset does not contain human-annotated (golden)… See the full description on the dataset page: https://huggingface.co/datasets/eoplumbum/v4_nuclear_power_articles.devcenter-articles
Overview
This dataset consists of ~600 articles from the MongoDB Developer Center.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the article. This value is devcenter for the entire dataset.
url: Link to the article
action: Action taken on the article. This value is created for the entire dataset.
body: Content of the article in Markdown format
format: Format of the content. This value is md for all articles.
metadata: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles.reddit_news_articles_commentsdzen-russian-articles
Dzen Russian Articles Dataset
Русскоязычные статьи с dzen.ru.
Датасет в активном сборе — новые статьи добавляются регулярно, объём постоянно растёт.
Как устроен парсинг
Статьи собираются с dzen.ru и обрабатываются через Gemini в качестве движка извлечения: модель очищает текст, разбивает на абзацы, определяет категорию, теги, ключевые слова и тональность. Поле extractor содержит название используемой модели (gemini/gemini-3.5-flash-lite).
Структура… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/dzen-russian-articles.bcms-fake-news-articlesresearch-articles
Threadbaire Research Articles
Structural analysis of AI industry dynamics, software value collapse, and open infrastructure. Thirteen articles published between January and September 2026, available as raw markdown for analysis, citation, and AI-readable ingestion.
About
This dataset contains the complete text of the Threadbaire thesis and blog — a body of independent research arguing that AI has already collapsed traditional software value capture mechanisms, and… See the full description on the dataset page: https://huggingface.co/datasets/Threadbaire/research-articles.Articles_Constitution_3300_Instruction_SetDataset Card for Indian Constitutional Law Instruction-Response Dataset
Dataset Summary
The dataset contains instruction-input-output pairs on Indian Constitutional Law, specifically addressing Articles 12, 14, 19, 21, and 15. It's designed to assist AI models, researchers, and learners in understanding and generating responses to complex legal questions related to the Indian Constitution.
Supported Tasks
This dataset supports tasks such as question answering, text comprehension, language… See the full description on the dataset page: https://huggingface.co/datasets/nisaar/Articles_Constitution_3300_Instruction_Set.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.Scientific-dataset-on-articles-small-thinkScientific Dataset on Articles (Small)
Это специализированный русскоязычный датасет небольшого объема (~1.2 тыс. строк), содержащий текстовую информацию, извлеченную из научных журналов по биологии (энтомология, арахнология, палеонтология), биографических очерков ученых-исследователей, а также дополненную тематическими материалами из открытых источников интернета.
Описание датасета
Датасет спроектирован для задач извлечения знаний (Information Extraction), ответов на вопросы по научным текстам… See the full description on the dataset page: https://huggingface.co/datasets/HoundyWoundy/Scientific-dataset-on-articles-small-think.phoronix-articles
Phoronix Articles Dataset: The Archive of Open-Source Computing Journalism
The definitive dataset of Phoronix - your gateway to years of open-source hardware/software evolution, performance analysis, and Linux ecosystem journalism.
🚀 What's Inside?
This dataset contains the complete archive of Phoronix articles - from bleeding-edge hardware launches to deep-dive Linux kernel analysis. Perfect for researchers, developers, and AI enthusiasts who need high-quality technical… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/phoronix-articles.egyptian-law-articlesarticles-metadata
Psychology Articles Metadata (FR-EN Bilingual)
A bilingual (French / English) metadata catalog of clinical psychology articles published on psychologieetserenite.com, authored by Gildas Garrec (CBT psychopractitioner). Each article is paired across the two languages with canonical URLs, themes, keywords, word counts and timestamps.
The dataset is designed for:
Translation alignment research (FR↔EN parallel article metadata)
Multilingual text classification (psychology themes)… See the full description on the dataset page: https://huggingface.co/datasets/psychologie-et-serenite/articles-metadata.gdpr-articles-dataset-trainwikifacts-articlesScientific-dataset-on-articles-smallScientific Dataset on Articles (Small)
Это специализированный русскоязычный датасет небольшого объема (~1.5 тыс. строк), содержащий текстовую информацию, извлеченную из научных журналов по биологии (энтомология, арахнология, палеонтология), биографических очерков ученых-исследователей, а также дополненную тематическими материалами из открытых источников интернета.
Описание датасета
Датасет спроектирован для задач извлечения знаний (Information Extraction), ответов на вопросы по научным текстам… See the full description on the dataset page: https://huggingface.co/datasets/HoundyWoundy/Scientific-dataset-on-articles-small.vietnamese-corporate-legal-articles-fsm
Lexora Knowledge - Vietnamese Legal Documents
Dataset Summary
A structured Vietnamese legal knowledge base crawled from
vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's
National Legal Database), published as 4 linked subsets: full
documents, individual articles (Điều), the citation graph between
documents/articles, and domain-concept tags. Load a specific subset
with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/anhnon/vietnamese-corporate-legal-articles-fsm.paul_graham_and_sam_altman_articlesbitcoin-news-articles-text-corporafrench-colonial-articles
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
french_colonial_articles
This dataset contains instruction-response pairs focused on African and French colonial history, featuring detailed historical accounts of events such as the Algerian War, the Thiaroye massacre, and post-colonial diplomatic tensions. The content is structured as conversations where an expert historian provides objective, fact-based answers with… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/french-colonial-articles.wikifacts-articles_v0maleeha-lodhi-dawn-articles
Maleeha Lodhi Opinion Articles Dataset
This dataset contains a curated collection of opinion articles originally published in Dawn, one of Pakistan’s leading English-language newspapers, over the past six years.The articles have been formatted for instruction-based fine-tuning of large language models to emulate the writing style of Maleeha Lodhi — a distinguished Pakistani diplomat, journalist, and academic known for her analytical, sophisticated commentary on international… See the full description on the dataset page: https://huggingface.co/datasets/abdullah1027/maleeha-lodhi-dawn-articles.prompts-simplify-articles10+3 prompts to fine-tune Llm to simplify Articles texts.
Wikipedia_Articles_on_Machine_Learning_and_DS
Wikipedia Machine Learning Corpus (wiki-ml-corpus)
A curated dataset of over 100 Wikipedia articles related to Machine Learning, Statistics, Probability, Data Science, and Deep Learning.
This dataset is designed for use in:
NLP tasks like summarization, QA, and topic modeling
ML interview prep and curriculum design
Ontology-driven QA systems and SPARQL-based pipelines
Building structured knowledge graphs from unstructured text
Dataset Structure
Each example in the… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/Wikipedia_Articles_on_Machine_Learning_and_DS.hiligaynon_news_articlesbanglatribune_news_articles
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/SaifullahBinYusuf/banglatribune_news_articles.samakal_news_articlesdevcenter-articles-embedded
Overview
This dataset consists of chunked and embedded versions of a subset of articles from the MongoDB Developer Center.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the article. This value is devcenter for the entire dataset.
url: Link to the article
action: Action taken on the article. This value is created for the entire dataset.
body: Content of the chunk in Markdown format
format: Format of the content. This value is… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles-embedded.Indian_Const_Articles_LLAMA2_Format
