datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pmc-articles-dataset-mentions-snippets
PMC Articles Dataset Mentions Snippets
Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature.
Description
Task: Extract structured dataset info (identifier, repository, webpage) from article text
Source: PMC open-access articles
Format: Text snippet → JSON output
Examples: Positive (with datasets) and negative (no datasets)
Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.ai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.gdpr-articlesMedical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.Scientific-and-technical-journal-articles-Africa
Scientific and technical journal articles Africa | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Scientific-and-technical-journal-articles-Africa.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.propensity-score-matching-articles
A Dataset on Propensity Score Matching Papers, 1964-2014
Overview
This dataset provides bibliographic and methodological information on a random sample of academic articles involving propensity score matching (PSM). Each row corresponds to a single article and contains information such as the article’s DOI, title, authors, publication year, and a series of binary or categorical indicators describing the methods used or reported within the study. These indicators focus… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/propensity-score-matching-articles.Content-Articles
Content-Articles Dataset
Overview
The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.AI_Articles_Scraped_from_arXiv-Semantic_Scholar
📘 AI Articles Scraped from arXiv & Semantic Scholar
🧩 Description
This dataset contains information on articles related to major AI conferences such as AAAI, NeurIPS, IJCAI, ICML, ICLR, collected through scraping from ArXiv and Semantic Scholar.It is intended to be used as a training dataset for various model training tasks and other desired uses.
📂 File Structure
File
Description
AI_Titles_v2025.csv
Main dataset
README.md
This file… See the full description on the dataset page: https://huggingface.co/datasets/d-e-c-d/AI_Articles_Scraped_from_arXiv-Semantic_Scholar.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.wikifacts-articles_v0-qrelsmost-cited-wikipedia-articlesWikipedia is a massive repository of human knowledge. The largest edition, the English Wikipedia, contains over 65.5 million pages, including 7.17 million articles (excluding redirects). Connecting this vast network are 1.63 billion unique page-to-page links. Based on an analysis of this dataset, the most cited articles on the English Wikipedia were identified.
When considering what these most cited articles in Wikipedia might be, we can assume that prominent historical topics like “United… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/most-cited-wikipedia-articles.bangla-news-articles-samplecrypto-articles-btc-price-changearticles_with_11day_ohlc_localmultidisciplinary_health_articles
Health Articles Corpus — Download Summary
Generated: 2026-03-06
Overview
Metric
Value
Total articles processed
1,694
PDFs downloaded
1,356
Study snapshots (Consensus)
22
Not found
310
Skipped (no DOI)
6
Total files in archive
1,269
Archive size
~1.2 GB (downloads/corpus_articles.tar.xz)
Overall retrieval rate: 80.2% (pdf + snapshot / total)
By Topic
Topic
Total
PDFs
Snapshots
Not Found
Skipped
Retrieval Rate… See the full description on the dataset page: https://huggingface.co/datasets/Chus010895/multidisciplinary_health_articles.
