datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
filtered_articles_by_year
Dataset Card for Filtered Articles by Year
Dataset Summary
The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time.
Supported Tasks and Leaderboards
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.research-articles
Threadbaire Research Articles
Structural analysis of AI industry dynamics, software value collapse, and open infrastructure. Thirteen articles published between January and September 2026, available as raw markdown for analysis, citation, and AI-readable ingestion.
About
This dataset contains the complete text of the Threadbaire thesis and blog — a body of independent research arguing that AI has already collapsed traditional software value capture mechanisms, and… See the full description on the dataset page: https://huggingface.co/datasets/Threadbaire/research-articles.phoronix-articles
Phoronix Articles Dataset: The Archive of Open-Source Computing Journalism
The definitive dataset of Phoronix - your gateway to years of open-source hardware/software evolution, performance analysis, and Linux ecosystem journalism.
🚀 What's Inside?
This dataset contains the complete archive of Phoronix articles - from bleeding-edge hardware launches to deep-dive Linux kernel analysis. Perfect for researchers, developers, and AI enthusiasts who need high-quality technical… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/phoronix-articles.vietnamese-corporate-legal-articles-fsm
Lexora Knowledge - Vietnamese Legal Documents
Dataset Summary
A structured Vietnamese legal knowledge base crawled from
vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's
National Legal Database), published as 4 linked subsets: full
documents, individual articles (Điều), the citation graph between
documents/articles, and domain-concept tags. Load a specific subset
with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/anhnon/vietnamese-corporate-legal-articles-fsm.maleeha-lodhi-dawn-articles
Maleeha Lodhi Opinion Articles Dataset
This dataset contains a curated collection of opinion articles originally published in Dawn, one of Pakistan’s leading English-language newspapers, over the past six years.The articles have been formatted for instruction-based fine-tuning of large language models to emulate the writing style of Maleeha Lodhi — a distinguished Pakistani diplomat, journalist, and academic known for her analytical, sophisticated commentary on international… See the full description on the dataset page: https://huggingface.co/datasets/abdullah1027/maleeha-lodhi-dawn-articles.BBC_Eng_News_Articles_dataset
BBC News Articles Dataset
Dataset Description
A collection of 2,225 news articles from BBC, suitable for text classification, summarization, and NLP tasks.
Dataset Summary
Metric
Value
Total Articles
2,225
Unique Articles
2,092
Columns
filename, article_text
Language
English
Source
BBC News
Dataset Structure
Data Fields
Field
Type
Description
filename
string
Unique identifier/filename for each… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/BBC_Eng_News_Articles_dataset.
