datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aim-technical-articles
Analytics India Magazine Technical Articles Dataset 🚀
Dataset Description
This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies.
✨ Dataset Highlights
📚 Comprehensive Coverage: Latest AI models, frameworks, and tools
🔬 Technical Depth: Extracted keywords and complexity scoring
🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.devcenter-articles
Overview
This dataset consists of ~600 articles from the MongoDB Developer Center.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the article. This value is devcenter for the entire dataset.
url: Link to the article
action: Action taken on the article. This value is created for the entire dataset.
body: Content of the article in Markdown format
format: Format of the content. This value is md for all articles.
metadata: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas.islamic-articles-corpus
☪ Islamic Articles Corpus - English RAG Dataset
Dataset Description
Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.Wikipedia-Articles
Dataset Card for "BrightData/Wikipedia-Articles"
Dataset Summary
Explore a collection of millions of Wikipedia articles with the Wikipedia dataset, comprising over 1.23M structured records and 10 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, URLs, article titles, raw and cataloged text, images, "see also" references, external links, and a structured table of contents.
For a complete list of data points, please… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Wikipedia-Articles.turkish-hospital-medical-articles
🏥 Turkish Hospital Medical Articles Dataset
A comprehensive collection of Turkish-language medical articles from 14 major hospital and healthcare provider websites in Turkey. This dataset is designed for training and evaluating Turkish NLP models in the medical domain, including large language models (LLMs), health chatbots, medical text summarization, and clinical text classification.
📊 Dataset Overview
Total Articles: ~24,612 medical articles
Sources: 14 major… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/turkish-hospital-medical-articles.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192turkish-hospital-medical-articles
🏥 Turkish Medical Articles from 14 Hospital Websites
This dataset contains Turkish-language medical articles scraped from 14 official hospital and healthcare provider websites in Turkey. Each file corresponds to one source and is stored in efficient .parquet format.
It is designed for training and evaluating Turkish NLP models in the medical domain, including large language models (LLMs), health chatbots, summarizers, and classifiers.
🧾 Total articles: ~ 25,000📦 Total file size:… See the full description on the dataset page: https://huggingface.co/datasets/umutertugrul/turkish-hospital-medical-articles.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.phoronix-articles
Phoronix Articles Dataset: The Archive of Open-Source Computing Journalism
The definitive dataset of Phoronix - your gateway to years of open-source hardware/software evolution, performance analysis, and Linux ecosystem journalism.
🚀 What's Inside?
This dataset contains the complete archive of Phoronix articles - from bleeding-edge hardware launches to deep-dive Linux kernel analysis. Perfect for researchers, developers, and AI enthusiasts who need high-quality technical… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/phoronix-articles.wim-schema-org-wiki-articles
Dutch Wikipedia Aligned Articles aligned with Schema.org Classes
Dataset Version: 1.0 (2025-06-04)Point of Contact: UWV Netherlands (UWV organization on Hugging Face)License: CC BY-SA 4.0Dataset: UWV/wim_schema_org_wiki_articles
Dataset Description
This dataset provides alignments between Schema.org classes and relevant Dutch Wikipedia articles. Each Schema.org class from a processed subset is linked to up to 20 distinct Wikipedia articles, including their full text, a… See the full description on the dataset page: https://huggingface.co/datasets/UWV/wim-schema-org-wiki-articles.moi-news-articles-dataset
MOI News & Article Dataset 🇲🇲
This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.Nepali_law_articles_corpus
Nepal Legal & Government Education Corpus
Dataset Summary
This dataset is a collection of 868 explanatory legal and government-procedure
articles scraped from 11 trusted Nepalese sources — legal blogs,
law firm publications, government portals, and legal aid / human-rights
organizations. It was built as part of NyayaLM, a bilingual (Nepali-English)
legal foundation language model for Nepal, where it serves as one of the
general-education corpora used to ground the… See the full description on the dataset page: https://huggingface.co/datasets/gahann/Nepali_law_articles_corpus.turkish-medical-articles-smallTotal Token Count : 190M (o200k_base)
Dataset Source : https://www.medicalpark.com.tr/saglik-rehberi
justicio-BOE-A-1978-31229-constitucion-by-articles-qa
Dataset summary
It is a synthetic dataset created to evaluate a RAG pattern for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
Justicio summary
Justicio is a Question/Answering Assistant that generates answers from user questions about the official state gazette of Spain: Boletín Oficial del Estado (BOE).
Data Fields
number: Number of the article of the Spanish Constitution.
context: Text of the article of the Spanish Constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa.doktorsitesi-articles
👨⚕️ Turkish Doctor Articles Dataset (DoktorSitesi)
A comprehensive collection of Turkish-language medical articles written by certified doctors and medical professionals from DoktorSitesi.com. This dataset contains expert medical content covering various specialties, treatments, and health topics, making it ideal for training Turkish medical NLP models and healthcare applications.
📊 Dataset Overview
Total Articles: 42,804 medical articles
Train Split: 34,243… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/doktorsitesi-articles.Content-Articles
Content-Articles Dataset
Overview
The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.techcrunch-articles
TechCrunch News Articles Dataset
📊 Dataset Overview
This dataset contains 10,265 high-quality news articles scraped from TechCrunch, one of the leading technology news websites. The dataset includes comprehensive article content, metadata, and quality assessments suitable for various NLP tasks including text classification, sentiment analysis, summarization, and content generation.
🎯 Key Features
10,265 articles with full text content
High-quality filtering… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/techcrunch-articles.turkish-medical-articles
🩺 Doktorsitesi.com Turkish Medical Articles
This dataset contains Turkish medical articles scraped from doktorsitesi.com, a public health information portal featuring content written by licensed healthcare professionals in Turkey.
The dataset is intended for use in natural language processing (NLP), language model training, and health-related AI research involving the Turkish language.
🧾 Number of articles: ~43K🗂️ Format: .parquet📁 File size: ~110 MB👤 Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/umutertugrul/turkish-medical-articles.kz-rus-articles-comprehensive
🇰🇿🇷🇺 Kazakh-Russian Articles Comprehensive Dataset
A high-quality bilingual corpus for cross-lingual NLP research
📋 Dataset Overview
The Kazakh-Russian Articles Comprehensive Dataset is a meticulously curated bilingual corpus designed to advance natural language processing research for Kazakh and Russian languages. This dataset addresses the critical need for high-quality parallel and comparable text resources in Central Asian language pairs, particularly… See the full description on the dataset page: https://huggingface.co/datasets/Adilbai/kz-rus-articles-comprehensive.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-reduceddevcenter-articles-embedded
Overview
This dataset consists of chunked and embedded versions of a subset of articles from the MongoDB Developer Center.
Dataset Structure
The dataset consists of the following fields:
sourceName: The source of the article. This value is devcenter for the entire dataset.
url: Link to the article
action: Action taken on the article. This value is created for the entire dataset.
body: Content of the chunk in Markdown format
format: Format of the content. This value is… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles-embedded.Medvik-Articles
Medvik-Articles - training dataset
Mappings of Authority main headings (the first author only) to the related articles titles - based on Medvik system exports.
The Articles data come from Bibliographia Medica Czechoslovaca (BMC) database of Czech biomedical literature.
License
Medvik-Articles - training dataset © 2025 by National Medical Library
is licensed under Creative Commons Attribution 4.0 International
Structure
"text1","text2","category"
"Author… See the full description on the dataset page: https://huggingface.co/datasets/NLK-NML/Medvik-Articles.longchau_vaccine_articlesOver 600 vaccination related articles from nhathuoclongchau.com.vn
