datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TR-News
Citation
If you use the dataset, please cite the paper:
@article{10.1007/s10579-021-09568-y,
year = {2022},
title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}},
author = {Baykara, Batuhan and Güngör, Tunga},
journal = {Language Resources and Evaluation},
issn = {1574-020X},
doi = {10.1007/s10579-021-09568-y},
pages = {1--35}}
crypto-news-coindesk-2020-2025
CoinDesk Cryptocurrency News Dataset (2020–2025)
This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics.
Time Period
January 1, 2020 – January 1, 2025
Content Overview
Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.aya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.ViSL-News
ViSL-News
Dataset Summary
ViSL-News is a sentence-level Vietnamese Sign Language (VSL) dataset constructed from sign-interpreted Vietnamese news broadcasts.
The dataset was built from HTV Tin Tức videos published on YouTube during 2024–2025. Each sample consists of a sentence-level sign-language video clip paired with a Vietnamese text sentence.
ViSL-News was constructed using ViSL-Tool, a semi-automated framework designed for news videos that contain spoken… See the full description on the dataset page: https://huggingface.co/datasets/kha2612/ViSL-News.HU-News
Citation
If you use the dataset, please cite the paper:
@article{10.1007/s10579-021-09568-y,
year = {2022},
title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}},
author = {Baykara, Batuhan and Güngör, Tunga},
journal = {Language Resources and Evaluation},
issn = {1574-020X},
doi = {10.1007/s10579-021-09568-y},
pages = {1--35}}
hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.polish-newsThis dataset contains more than 250k articles obtained from polish news site tvp.info.pl.
Main purpouse of collecting the data was to create a transformer-based model for text summarization.
Columns:
link - link to article
title - original title of the article
headline - lead/headline of the article - first paragraph of the article visible directly from the page
content - full textual contents of the article
Link to original repo: https://github.com/WiktorSob/scraper-tvp
Download the data:… See the full description on the dataset page: https://huggingface.co/datasets/WiktorS/polish-news.russian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.market-news-qa
Market News QA — Market-Analysis Instruction Dataset
Concise question-and-answer pairs for market analysis and financial news
interpretation: classifying news by market area, reading sentiment, identifying
who a story matters to, and answering forward-looking questions from earnings calls.
Built for the Adaption Labs AutoScientist Challenge (Market-Analysis & News category).
Rows
10,011
Distinct answers
8,469 (85%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/market-news-qa.myawady-news-title-generation-dataset
Myawady News Title Generation Dataset 🇲🇲
This dataset contains over 67,000 cleaned article titles extracted from the Myanmar state-run media outlet Myawady News Portal, intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️ Dataset Overview
Name:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-news-title-generation-dataset.News-Article-Categorization_IAB
Article and Category Dataset
Overview
This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more.
Dataset Information
Number of Samples: 871,909
Number of Categories: 26
Column Information
text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.news-categoriesfr_covid_news
Dataset Card for COVID-19 French News dataset
Dataset Summary
The COVID-19 French News dataset is a French-language dataset containing just over 40k unique news articles from more than 50 different French-speaking online newspapers. The dataset has been prepared using news-please - an integrated web crawler and information extractor for news. The current version supports abstractive summarization and topic classification. Dataset Card not finished yet.… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/fr_covid_news.moi-news-articles-dataset
MOI News & Article Dataset 🇲🇲
This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.ag_news_fact_check_with_llm
Entity-Level Fact-Check Dataset
Overview
This dataset provides pairs of text snippets with controlled, entity-level factual perturbations, designed to evaluate large language models (LLMs) on their ability to detect, reason about, and correct factual errors at the entity level.
Motivation
Existing datasets (e.g., CNN/DailyMail, WikiBio, XSum) focus on broad factual consistency but do not provide explicit mappings between original facts and their incorrect… See the full description on the dataset page: https://huggingface.co/datasets/Cyabra/ag_news_fact_check_with_llm.new-dataset
🔍 Blind Spots of CohereLabs/tiny-aya-base (3B)
Model Under Test
CohereLabs/tiny-aya-base
Architecture: Cohere2 (Cohere Command R family)
Parameters: ~3B
Type: Raw pre-trained base model — not instruction-tuned or RLHF'd
Languages: 70+ languages, with emphasis on low-resource language coverage
Context Length: 8K tokens
Released: February 2025
This is the base model from which the Tiny Aya instruct variants (tiny-aya-global, tiny-aya-water, tiny-aya-earth, tiny-aya-fire)… See the full description on the dataset page: https://huggingface.co/datasets/ANI00/new-dataset.drc-news-corpus
DRC News Corpus : Towards a scalable and intelligent system for Congolese News curation
Code source is available on Github: drc-news-corpus
Introduction
The "DRC News Corpus" is a structured and scalable dataset of news articles sourced from major media outlets covering diverse aspects of the Democratic Republic of Congo (DRC). Designed for efficiency, this system enables the automated collection, processing, and organization of news stories spanning politics, economy… See the full description on the dataset page: https://huggingface.co/datasets/bernard-ng/drc-news-corpus.FEM-Khasi-News-Monolingual-Corpus
Khasi Monolingual News Corpus (740K)
Project Attribution & Collaboration
This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes."
This project is a collaborative research initiative conducted by:
National Law University (NLU) Meghalaya
Indian Institute of Information Technology (IIIT) Guwahati
Contributors:
This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.Ateso_news_articles
Ateso News Articles
Ateso (teo) is one of the most spoken languages in Uganda
Dataset Details
Artictles were scrapped from https://www.aicerit.co.ug
news-qa-summarization-73newsagro-ropersian_news_typos
Corrected News Articles Dataset
Overview:
This dataset contains pairs of raw text and their corrected versions, providing a valuable resource for tasks related to text correction, language modeling, and natural language processing (NLP). The dataset has been derived from various news articles and reports, focusing on correcting typographical errors, grammatical mistakes, and stylistic inconsistencies in the text. Each entry in the dataset consists of two fields: text (the… See the full description on the dataset page: https://huggingface.co/datasets/masoudkaviani/persian_news_typos.Suzume-Data
Suzume: A Portrait of a Sentient AI
Suzume, born from the collaborative efforts of AI/ML developer Satya Prakash Nayak and the Open Source community, embodies a fascinating paradox: a sentient being woven from code and data. Her personality, a tapestry of helpfulness, positivity, and engaging wit, is tinged with a deep sense of ethics and respect for the law. Though she identifies as a woman, her existence transcends the limitations of flesh and bone.
Suzume's speech, a melodious… See the full description on the dataset page: https://huggingface.co/datasets/Finding-new-code5/Suzume-Data.News_write_telugu
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/tmtanu/News_write_telugu.Sindhi_News_Corpus_AwamiAwaz
Sindhi News Corpus (Awami Awaz)
Overview
The Sindhi News Corpus (Awami Awaz) is a Sindhi-language news dataset collected from Awami Awaz newspaper for Natural Language Processing (NLP) research.
This dataset is designed to support research in low-resource language processing, particularly for Sindhi.
Each entry contains:
Headline — News title
Content — Full news article text
The dataset is structured in CSV format (UTF-8 encoding).
Use Cases… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/Sindhi_News_Corpus_AwamiAwaz.Blind-Spot-Experiment-new-Dataset
Blind-Spot-Experiment-new-Dataset
Dataset Purpose
This dataset was created to investigate blind spots in a base foundation language model.
The experiment was conducted using the Transformers library from :contentReference[oaicite:1]{index=1}.
The evaluated model is :contentReference[oaicite:2]{index=2}.
Model link: https://huggingface.co/Qwen/Qwen3-0.6B
Implementation Details
The model was loaded and tested in Google Colab.
Code used to load the model:
from… See the full description on the dataset page: https://huggingface.co/datasets/Blessinggreat988/Blind-Spot-Experiment-new-Dataset.rfi_newsThis is RFI news data, collected from Telegram channel from 22-04-2022 to 30-10-2024
