datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
note-articles
note articles
note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。
各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。
ticker_analysis_articlespubmed-2019-pythia-word-tfidf-pubmedqa-clean-articlesall-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlespubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-articlesall-the-news-2-pythia-tfidf-topic-stratified-v1-articlesv4_nuclear_power_articles
Dataset Card for Nuclear News V4 Dataset
Dataset Summary
The Nuclear News V4 Dataset is a multilingual dataset consisting of 33,104 unique news articles sourced from 12 online news platforms across the Visegrád Group (V4) countries — Poland, Czech Republic, Slovakia, and Hungary — published between 1998 and 2025.
The goal of the dataset is to analyze media narratives surrounding nuclear energy in Central Europe.
While the dataset does not contain human-annotated (golden)… See the full description on the dataset page: https://huggingface.co/datasets/eoplumbum/v4_nuclear_power_articles.pmc-articles-dataset-mentions-snippets
PMC Articles Dataset Mentions Snippets
Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature.
Description
Task: Extract structured dataset info (identifier, repository, webpage) from article text
Source: PMC open-access articles
Format: Text snippet → JSON output
Examples: Positive (with datasets) and negative (no datasets)
Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.aim-technical-articles
Analytics India Magazine Technical Articles Dataset 🚀
Dataset Description
This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies.
✨ Dataset Highlights
📚 Comprehensive Coverage: Latest AI models, frameworks, and tools
🔬 Technical Depth: Extracted keywords and complexity scoring
🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.ai-tech-articles
AI/Tech Dataset
This dataset is a collection of AI/tech articles scraped from the web.
It's hosted on HuggingFace Datasets, so it is easier to load in and work with.
To load the dataset
1. Install HuggingFace Datasets
pip install datasets
2. Load the dataset
from datasets import load_dataset
dataset = load_dataset("siavava/ai-tech-articles")
# optionally, convert it to a pandas dataframe:
df = dataset["train"].to_pandas()
You do not need to clone… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.samorzad-gov-pl-articles
Artykuły z platformy samorzad.gov.pl
Wersja: v0.2
Zbiór zawiera 81 418 artykułów opublikowanych na stronach 247 instytucji korzystających ze wspólnej platformy samorzad.gov.pl. Są to między innymi urzędy gmin i powiatów, szkoły, instytucje pomocy społecznej i instytucje kultury.
W wersji v0.2 usunięto dane osobowe i kontaktowe z pól tekstowych przeznaczonych dla odbiorcy. Usunięte wartości zastąpiono jednoznacznymi znacznikami, zachowując układ i znaczenie pozostałej treści.… See the full description on the dataset page: https://huggingface.co/datasets/dawidmajewski/samorzad-gov-pl-articles.ai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.Clean-Wikipedia-English-Articles
Dataset Card for DragonLLM/Clean-Wikipedia-English-Articles
DragonLLM is delighted to announce the release of the cleanest markdown extract of Wikipedia articles so far, a high-quality resource to train your LLMs!
Dataset Summary
The Clean-Wikipedia-English-Articles dataset contains the comprehensive bodies of English articles, i.e. without appendices like References, See also, Bibliography, etc.
It has been pointed out here that the Wikimedia Wikipedia dataset… See the full description on the dataset page: https://huggingface.co/datasets/OVHaiLLM/Clean-Wikipedia-English-Articles.wikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-articlesjusticio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.enwiki-articles-html-2410justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-5k-chunks-groq_llama3_70b_8192-sasislamic-articles-corpus
☪ Islamic Articles Corpus - English RAG Dataset
Dataset Description
Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.dzen-russian-articles
Dzen Russian Articles Dataset
Русскоязычные статьи с dzen.ru.
Датасет в активном сборе — новые статьи добавляются регулярно, объём постоянно растёт.
Как устроен парсинг
Статьи собираются с dzen.ru и обрабатываются через Gemini в качестве движка извлечения: модель очищает текст, разбивает на абзацы, определяет категорию, теги, ключевые слова и тональность. Поле extractor содержит название используемой модели (gemini/gemini-3.5-flash-lite).
Структура… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/dzen-russian-articles.gdpr-articleswikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-articlese-mordovia-articles-2024
"e-mordovia-articles-2024": a parallel news dataset for Russian, Erzya and Moksha
This is a semi-aligned dataset of Russian, Erzya and Moksha news articles, crawled from https://www.e-mordovia.ru.
Dataset Description
Dataset Summary
This is a dataset of news articles collected from https://www.e-mordovia.ru, the official portal of the state authorities of the Republic of Mordovia.
The articles have been paired by the following algorithm:
Calculate similarities… See the full description on the dataset page: https://huggingface.co/datasets/slone/e-mordovia-articles-2024.wikipedia_articles_es
Wikipedia Articles (Spanish)
This dataset contains segmented articles from the Spanish Wikipedia. Each example includes structural and textual information designed for tasks such as text classification, segmentation, and sentence similarity.
The dataset is divided into three subsets:
wikipedia-es-A000: Suggested train dataset; 26510 groups of articles.
wikipedia-es-A001: Suggested validation dataset; 3336 groups of articles.
wikipedia-es-A002: Suggested test dataset; 6557 groups of… See the full description on the dataset page: https://huggingface.co/datasets/Alverciito/wikipedia_articles_es.mteb-nl-news-articles-retThis dataset contains Dutch news articles, sourced from the Nederlandse Oproep Stichting.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-ret.enwiki-articles-markdown-cleaned-2410OpenAlex-Articles-and-Harvard-Library-Item-Data
OpenAlex Articles + Harvard Library Item Data
Dataset Description
This repository contains two large tabular subsets built from Harvard Library Bibliographic Metadata and an OpenAlex snapshot.
These files are derived, filtered, and processed datasets, not full reproductions of the original source data. The records are based on real source metadata, but the dataset was prepared mainly for library data system operations testing, large-scale experimentation, and… See the full description on the dataset page: https://huggingface.co/datasets/KennyChowww/OpenAlex-Articles-and-Harvard-Library-Item-Data.shipping_news_articles_lsa
