CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nakasyou /note-articles note articles note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。 各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。 tabular100K<n<1M1 likes477 downloads2mo agoHugging Face02jtviegas /ticker_analysis_articlestabular10K<n<100K0 likes420 downloads16h agoHugging Face03MA-tokenweights /pubmed-2019-pythia-word-tfidf-pubmedqa-clean-articlestabular10K<n<100K0 likes309 downloads27d agoHugging Face04MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes292 downloads29d agoHugging Face05MA-tokenweights /pubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-articlestabular10K<n<100K0 likes291 downloads26d agoHugging Face06MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes284 downloads29d agoHugging Face07eoplumbum /v4_nuclear_power_articles Dataset Card for Nuclear News V4 Dataset Dataset Summary The Nuclear News V4 Dataset is a multilingual dataset consisting of 33,104 unique news articles sourced from 12 online news platforms across the Visegrád Group (V4) countries — Poland, Czech Republic, Slovakia, and Hungary — published between 1998 and 2025. The goal of the dataset is to analyze media narratives surrounding nuclear energy in Central Europe. While the dataset does not contain human-annotated (golden)… See the full description on the dataset page: https://huggingface.co/datasets/eoplumbum/v4_nuclear_power_articles.tabulartext-classification10K<n<100K1 likes237 downloads1y agoHugging Face08vida-nyu /pmc-articles-dataset-mentions-snippets PMC Articles Dataset Mentions Snippets Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature. Description Task: Extract structured dataset info (identifier, repository, webpage) from article text Source: PMC open-access articles Format: Text snippet → JSON output Examples: Positive (with datasets) and negative (no datasets) Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.tabular1K<n<10K0 likes218 downloads2mo agoHugging Face09abhilash88 /aim-technical-articles Analytics India Magazine Technical Articles Dataset 🚀 Dataset Description This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies. ✨ Dataset Highlights 📚 Comprehensive Coverage: Latest AI models, frameworks, and tools 🔬 Technical Depth: Extracted keywords and complexity scoring 🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.tabulartext-classification10K<n<100K2 likes216 downloads1y agoHugging Face10Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes202 downloads2y agoHugging Face11siavava /ai-tech-articles AI/Tech Dataset This dataset is a collection of AI/tech articles scraped from the web. It's hosted on HuggingFace Datasets, so it is easier to load in and work with. To load the dataset 1. Install HuggingFace Datasets pip install datasets 2. Load the dataset from datasets import load_dataset dataset = load_dataset("siavava/ai-tech-articles") # optionally, convert it to a pandas dataframe: df = dataset["train"].to_pandas() You do not need to clone… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.tabulartext-generation10K<n<100K8 likes193 downloads3y agoHugging Face12dawidmajewski /samorzad-gov-pl-articles Artykuły z platformy samorzad.gov.pl Wersja: v0.2 Zbiór zawiera 81 418 artykułów opublikowanych na stronach 247 instytucji korzystających ze wspólnej platformy samorzad.gov.pl. Są to między innymi urzędy gmin i powiatów, szkoły, instytucje pomocy społecznej i instytucje kultury. W wersji v0.2 usunięto dane osobowe i kontaktowe z pól tekstowych przeznaczonych dla odbiorcy. Usunięte wartości zastąpiono jednoznacznymi znacznikami, zachowując układ i znaczenie pozostałej treści.… See the full description on the dataset page: https://huggingface.co/datasets/dawidmajewski/samorzad-gov-pl-articles.tabulartext-generation10K<n<100K0 likes171 downloads1mo agoHugging Face13fdaudens /ai-jobs-news-articles Dataset Summary This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work. Source Data The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.tabulartext-classification1K<n<10K1 likes153 downloads1y agoHugging Face14dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.tabularquestion-answeringn<1K1 likes114 downloads2y agoHugging Face15OVHaiLLM /Clean-Wikipedia-English-Articlesgated Dataset Card for DragonLLM/Clean-Wikipedia-English-Articles DragonLLM is delighted to announce the release of the cleanest markdown extract of Wikipedia articles so far, a high-quality resource to train your LLMs! Dataset Summary The Clean-Wikipedia-English-Articles dataset contains the comprehensive bodies of English articles, i.e. without appendices like References, See also, Bibliography, etc. It has been pointed out here that the Wikimedia Wikipedia dataset… See the full description on the dataset page: https://huggingface.co/datasets/OVHaiLLM/Clean-Wikipedia-English-Articles.tabular1M<n<10M11 likes98 downloads11mo agoHugging Face16MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes98 downloads1mo agoHugging Face17dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.tabularquestion-answeringn<1K0 likes95 downloads2y agoHugging Face18tom-010 /enwiki-articles-html-2410tabular1M<n<10M0 likes95 downloads2y agoHugging Face19dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas.tabularquestion-answeringn<1K0 likes82 downloads2y agoHugging Face20dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-5k-chunks-groq_llama3_70b_8192-sastabularn<1K0 likes75 downloads2y agoHugging Face21qurancn /islamic-articles-corpus ☪ Islamic Articles Corpus - English RAG Dataset Dataset Description Islamic Articles Corpus is a curated English-language RAG corpus containing 33 articles covering Muslim travel guides, mosque visits, halal food, prayer room directories, and Islamic community documentation. Every article preserves complete full-text content with all 608 embedded image references. Content focuses heavily on Singapore, Iran, Japan, Oman, and Qatar mosque and travel documentation.… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/islamic-articles-corpus.tabulartext-generationn<1K0 likes70 downloads3mo agoHugging Face22ru-dataset /dzen-russian-articles Dzen Russian Articles Dataset Русскоязычные статьи с dzen.ru. Датасет в активном сборе — новые статьи добавляются регулярно, объём постоянно растёт. Как устроен парсинг Статьи собираются с dzen.ru и обрабатываются через Gemini в качестве движка извлечения: модель очищает текст, разбивает на абзацы, определяет категорию, теги, ключевые слова и тональность. Поле extractor содержит название используемой модели (gemini/gemini-3.5-flash-lite). Структура… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/dzen-russian-articles.tabulartext-classificationn<1K1 likes68 downloads2mo agoHugging Face23Johny201 /gdpr-articlestabularn<1K1 likes67 downloads4y agoHugging Face24MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes63 downloads1mo agoHugging Face25slone /e-mordovia-articles-2024 "e-mordovia-articles-2024": a parallel news dataset for Russian, Erzya and Moksha This is a semi-aligned dataset of Russian, Erzya and Moksha news articles, crawled from https://www.e-mordovia.ru. Dataset Description Dataset Summary This is a dataset of news articles collected from https://www.e-mordovia.ru, the official portal of the state authorities of the Republic of Mordovia. The articles have been paired by the following algorithm: Calculate similarities… See the full description on the dataset page: https://huggingface.co/datasets/slone/e-mordovia-articles-2024.tabulartranslation100K<n<1M3 likes61 downloads1y agoHugging Face26Alverciito /wikipedia_articles_es Wikipedia Articles (Spanish) This dataset contains segmented articles from the Spanish Wikipedia. Each example includes structural and textual information designed for tasks such as text classification, segmentation, and sentence similarity. The dataset is divided into three subsets: wikipedia-es-A000: Suggested train dataset; 26510 groups of articles. wikipedia-es-A001: Suggested validation dataset; 3336 groups of articles. wikipedia-es-A002: Suggested test dataset; 6557 groups of… See the full description on the dataset page: https://huggingface.co/datasets/Alverciito/wikipedia_articles_es.tabulartext-classification1K<n<10K1 likes52 downloads8mo agoHugging Face27clips /mteb-nl-news-articles-retThis dataset contains Dutch news articles, sourced from the Nederlandse Oproep Stichting. Citation Information If you find our paper, benchmark or models helpful, please consider cite as follows: @misc{banar2025mtebnle5nlembeddingbenchmark, title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch}, author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-ret.tabular100K<n<1M0 likes49 downloads1y agoHugging Face28tom-010 /enwiki-articles-markdown-cleaned-2410tabular1M<n<10M0 likes47 downloads2y agoHugging Face29KennyChowww /OpenAlex-Articles-and-Harvard-Library-Item-Data OpenAlex Articles + Harvard Library Item Data Dataset Description This repository contains two large tabular subsets built from Harvard Library Bibliographic Metadata and an OpenAlex snapshot. These files are derived, filtered, and processed datasets, not full reproductions of the original source data. The records are based on real source metadata, but the dataset was prepared mainly for library data system operations testing, large-scale experimentation, and… See the full description on the dataset page: https://huggingface.co/datasets/KennyChowww/OpenAlex-Articles-and-Harvard-Library-Item-Data.tabularother100M<n<1B0 likes42 downloads5mo agoHugging Face30Ktzoras /shipping_news_articles_lsatabular10K<n<100K0 likes41 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.