datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
crypto-news-coindesk-2020-2025
CoinDesk Cryptocurrency News Dataset (2020–2025)
This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics.
Time Period
January 1, 2020 – January 1, 2025
Content Overview
Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.tech-company-news-data-dumpHackerNoon curated the internet's most cited 7M+ tech company news articles and blog posts about the 3k+ most valuable tech companies in 2022 and 2023. These stories were curated to power HackerNoon.com/Companies, where we update daily news on top technology companies like Microsoft, Google, and HuggingFace. Please use this news data freely for your project, and as always anyone is welcome to publish on HackerNoon.
CC-ID-News[Needs More Information]
Dataset Card for Common Crawled Indonesia News
Dataset Summary
[Needs More Information]
Supported Tasks and Leaderboards
[Needs More Information]
Languages
[Needs More Information]
Dataset Structure
Data Instances
[Needs More Information]
Data Fields
[Needs More Information]
Data Splits
[Needs More Information]
Dataset Creation
Curation Rationale
[Needs More… See the full description on the dataset page: https://huggingface.co/datasets/ZhafranR/CC-ID-News.cnbc_newsfeedthe-star-news-articlesTopic-specific-genre-classification_german_historical-newspapers
Dataset Card for Topic-specific Genre Classification of German Historical Newspapers
This dataset was developed to train and evaluate topic-specific genre classification of German-language historical newspaper clippings.
Curated by: [Sarah Oberbichler]
Language(s) (NLP): [German]
License: [afl-3.0]
Uses
Evaluation of machine learning models for topic-specific classification of ocr-processed historical texts with varying quality levels.
Fine-tuning models on… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Topic-specific-genre-classification_german_historical-newspapers.newspapers2024bangla-fake-news
Bangla Fake News Detection Dataset
One of the first publicly available fake news datasets for the Bengali language,
scraped from local newspapers. Built to support NLP research in under-represented languages.
Dataset Structure
Collected from Bengali local news sources
Labeled as fake / real
Includes a Bengali stemmer and corpus builder
Benchmark Results
Model
Accuracy
Naïve Bayes
52%
Logistic Regression
77%
Random Forest
85%… See the full description on the dataset page: https://huggingface.co/datasets/Far121/bangla-fake-news.TechCrunch_News
TechCrunch Battlefield News — Preview
This repository hosts a small preview of a structured dataset that tracks companies featured in TechCrunch’s Startup Battlefield coverage. The preview is here so you can test the schema and sample rows locally. To purchase the complete dataset with full coverage and updates, go to https://www.thedataoutlet.com.
Buy the full dataset: https://www.thedataoutlet.com
What you get in this preview
A single CSV file with 62 rows… See the full description on the dataset page: https://huggingface.co/datasets/calebheinzman/TechCrunch_News.Topic-specific-disambiguation_evaluation-dataset_German_historical-newspapers
Dataset Card for German Topic-specific Disambiguation Dataset Historical Newspapers
This dataset contains German newspaper articles (1850-1950) labeled as either about relevant for "return migration" or not. It helps researchers develop word sense disambiguation methods - technology that can tell when the German word "Heimkehr" or "Rückkehr" (return/homecoming) is being used to discuss people returning to their home countries versus when it's used in different contexts. This matters… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Topic-specific-disambiguation_evaluation-dataset_German_historical-newspapers.News_bcn_sentimentNews on Barcelona en spanish media outlets
bangla-news-articles-sampleczech-ing-the-news-dataset
VERIFEE RESEARCH: Dataset Request
Supported by a TAČR grant
In our research on disinformation analysis using machine learning, we collected a unique resource comprising thousands of reports and texts in Czech, complemented by information about present manipulative techniques. Thanks to proactive journalism students and great support from the Institute of Communication Studies and Journalism of the Faculty of Social Sciences at Charles University, this dataset came to life with… See the full description on the dataset page: https://huggingface.co/datasets/verifee/czech-ing-the-news-dataset.grey-zone-news
