datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pmc-articles-dataset-mentions-snippets
PMC Articles Dataset Mentions Snippets
Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature.
Description
Task: Extract structured dataset info (identifier, repository, webpage) from article text
Source: PMC open-access articles
Format: Text snippet → JSON output
Examples: Positive (with datasets) and negative (no datasets)
Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.aya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.cryptonews-articles-with-price-momentum-labels
Dataset Card for Cryptonews articles with price momentum labels
Dataset Summary
The dataset was gathered from two prominent sources in the cryptocurrency industry: Cryptonews.com and Binance.com. The aim of the dataset was to evaluate the impact of news on crypto price movements.
As we know, news events such as regulatory changes, technological advancements, and major partnerships can have a significant impact on the price of cryptocurrencies. By analyzing the data… See the full description on the dataset page: https://huggingface.co/datasets/SahandNZ/cryptonews-articles-with-price-momentum-labels.ai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.summarization-allegro-articlesclimate-news-articles
🌍 Jeu de données d'articles de presse française labellisés comme traitant ou non des sujets liés au climat
🇬🇧 / 🇺🇸 : as this data set is based only on French data, all explanations are written in French in this repository. The goal of the dataset is to train a model to classify titles of French newspapers in two categories : if it's about climate or not.
🗺️ Le contexte
Ce jeu de données de classification de titres d'article de presse française a été réalisé pour… See the full description on the dataset page: https://huggingface.co/datasets/pierre-loic/climate-news-articles.News_Articles_Categorization
Dataset Card for News_Articles_Categorization
Dataset Description
3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely Text and Category.
The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.CNN_News_Articles_2011-2022
CNN News Articles 2011-2022 Dataset
Introduction
This dataset contains CNN News Articles from 2011 to 2022 after basic cleaning. The dataset includes the following information:
Category
Full text
The data was downloaded from Kaggle at this URL: https://www.kaggle.com/datasets/hadasu92/cnn-articles-after-basic-cleaning. The dataset was split into two sets:
Train set with 32,218 examples
Test set with 5,686 examples
Usage
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/CNN_News_Articles_2011-2022.ai-jobs-news-articles-abstracts
News articles and research abstracts on AI, labor, and jobs
Dataset summary
This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary.
Rows: 53,526
document_class
Rows
Approx. date range (date column)
news
29,857
Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.gdpr-articlesirs-articlesHinduTamil-News-Articles-Dataset
HinduTamil News Articles Dataset
Overview
This dataset contains news articles in Tamil language scraped from the Hindu Tamil news website. Each article includes its title, author, city, published date, and text.
Motivation
This dataset was created to provide a comprehensive collection of Tamil news articles for research and analysis purposes.
Data Sources and collection method
The data in this dataset was collected from the Hindu Tamil news website… See the full description on the dataset page: https://huggingface.co/datasets/Shwetasss/HinduTamil-News-Articles-Dataset.Wikipedia-Articles
Dataset Card for "BrightData/Wikipedia-Articles"
Dataset Summary
Explore a collection of millions of Wikipedia articles with the Wikipedia dataset, comprising over 1.23M structured records and 10 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, URLs, article titles, raw and cataloged text, images, "see also" references, external links, and a structured table of contents.
For a complete list of data points, please… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Wikipedia-Articles.three_line_summarization_for_japanese_news_articlesライブドアニュースコーパスの3行要約データセットです。
Llama v2向けのプロンプトを追加して成形してあります。
学習に利用する際は、 [R_START] [R_END] をspecial tokenとして追加することを推奨します。
Number of rows: 3,907
Datasetは以下のリポジトリを利用してscrapeしました。
git@github.com:KodairaTomonori/ThreeLineSummaryDataset.git
Science_ArticlesScraped_Dataset_ArticlesMedical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.Scientific-and-technical-journal-articles-Africa
Scientific and technical journal articles Africa | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Scientific-and-technical-journal-articles-Africa.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.russia-ukraine-conflict-articles
Dataset Card for Russia Ukraine Conflict
Dataset Summary
###Context
On 24 February 2022, Russia invaded Ukraine in a major escalation of the Russo-Ukrainian War that began in 2014. The invasion caused Europe's largest refugee crisis since World War II, with more than 6.3 million Ukrainians fleeing the country and a third of the population displaced (Source: Wikipedia).
###Content
This dataset is a collection of 407 news articles from NYT and Guardians related to ongoing… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/russia-ukraine-conflict-articles.turkish-medical-articles-rag
Turkish Medical Articles - RAG System & Vector Database
Bu proje, Türkçe tıbbi makaleler üzerinde çalışan bir Retrieval-Augmented Generation (RAG) sistemi ve Vektör Veritabanı uygulamasıdır. Proje kapsamında ham veriler Hugging Face üzerinden çekilmiş, semantik parçalama uygulanmış, magibu/embeddingmagibu-200m modeliyle vektörleştirilmiş, ChromaDB üzerinde saklanmış ve başlangıç eşiği ile dinamik eşik optimizasyonu adımlarını içeren 30 soruluk bir benchmark testiyle… See the full description on the dataset page: https://huggingface.co/datasets/meldakahramann/turkish-medical-articles-rag.news_articles
News Articles Classification Dataset
This dataset consists of news articles labeled with corresponding categories for classification tasks.
Overview
The news articles classification dataset is a collection of articles sourced from various news outlets, each labeled with a specific category. The dataset is designed for tasks such as text classification, topic modeling, and sentiment analysis.
Dataset Information
Name: News Articles Classification Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bushra1dajam/news_articles.guardian_articles_full_contentmoi-news-articles-dataset
MOI News & Article Dataset 🇲🇲
This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.the-star-news-articleslabelled_articlesfedlex-articles
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/glards/fedlex-articles.Malayalam-Articlesbbc-articles-finetuning-classifpubmed-mesh-terms-level-1-articles-2024
PubMed articles with MeSH terms
Small dataset of article abstacts and titles with MeSH terms information.
Articles are selected with search for 100 top articles for level 1 MeSH terms. The articles with empty abstract are discarded.
Dataset columns:
pubmedid - pubmed id of the article
language - 3 character language code
title - article title
abstract - article abstract
termIds - serialized list of MeSH term ids of the article
