CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01batubayk /TR-News Citation If you use the dataset, please cite the paper: @article{10.1007/s10579-021-09568-y, year = {2022}, title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}}, author = {Baykara, Batuhan and Güngör, Tunga}, journal = {Language Resources and Evaluation}, issn = {1574-020X}, doi = {10.1007/s10579-021-09568-y}, pages = {1--35}} textsummarization100K<n<1M19 likes609 downloads4y agoHugging Face02maryamfakhari /crypto-news-coindesk-2020-2025 CoinDesk Cryptocurrency News Dataset (2020–2025) This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics. Time Period January 1, 2020 – January 1, 2025 Content Overview Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.imagetext-classification100K<n<1M2 likes447 downloads1mo agoHugging Face03SuryaKrishna02 /aya-telugu-news-articles Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.texttext-generation100K<n<1M6 likes168 downloads3y agoHugging Face04kha2612 /ViSL-News ViSL-News Dataset Summary ViSL-News is a sentence-level Vietnamese Sign Language (VSL) dataset constructed from sign-interpreted Vietnamese news broadcasts. The dataset was built from HTV Tin Tức videos published on YouTube during 2024–2025. Each sample consists of a sentence-level sign-language video clip paired with a Vietnamese text sentence. ViSL-News was constructed using ViSL-Tool, a semi-automated framework designed for news videos that contain spoken… See the full description on the dataset page: https://huggingface.co/datasets/kha2612/ViSL-News.tabulartranslation10K<n<100K0 likes138 downloads17d agoHugging Face05batubayk /HU-News Citation If you use the dataset, please cite the paper: @article{10.1007/s10579-021-09568-y, year = {2022}, title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}}, author = {Baykara, Batuhan and Güngör, Tunga}, journal = {Language Resources and Evaluation}, issn = {1574-020X}, doi = {10.1007/s10579-021-09568-y}, pages = {1--35}} textsummarization100K<n<1M3 likes116 downloads4y agoHugging Face06Linkseed /hacker_news_with_comments Dataset Card for [Dataset Name] Dataset Summary Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags. Supported Tasks and Leaderboards Comment Generation; News analysis with comments; Other comment-based NLP tasks. Languages English Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.tabulartext-generation1M<n<10M6 likes94 downloads4y agoHugging Face07crawlfeeds /Curated-Fox-News-Headlines-and-Full-Text Curated Fox News Headlines and Full Text This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis. 📁 Dataset Format Format: CSV Encoding: UTF-8 Fields: headline: The article title or headline publish_date: Date the article was published (YYYY-MM-DD) content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.imagetext-classification1K<n<10K2 likes87 downloads1y agoHugging Face08WiktorS /polish-newsThis dataset contains more than 250k articles obtained from polish news site tvp.info.pl. Main purpouse of collecting the data was to create a transformer-based model for text summarization. Columns: link - link to article title - original title of the article headline - lead/headline of the article - first paragraph of the article visible directly from the page content - full textual contents of the article Link to original repo: https://github.com/WiktorSob/scraper-tvp Download the data:… See the full description on the dataset page: https://huggingface.co/datasets/WiktorS/polish-news.texttext-classification100K<n<1M10 likes70 downloads3y agoHugging Face09ScoutieAutoML /russian_oil_gas_news_telegram_dataset Description in English: Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry, collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.tabulartext-classification10K<n<100K7 likes62 downloads2y agoHugging Face10flamiinngo /market-news-qa Market News QA — Market-Analysis Instruction Dataset Concise question-and-answer pairs for market analysis and financial news interpretation: classifying news by market area, reading sentiment, identifying who a story matters to, and answering forward-looking questions from earnings calls. Built for the Adaption Labs AutoScientist Challenge (Market-Analysis & News category). Rows 10,011 Distinct answers 8,469 (85%) Duplicate questions none Nulls none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/market-news-qa.textquestion-answering10K<n<100K1 likes49 downloads2mo agoHugging Face11freococo /myawady-news-title-generation-dataset Myawady News Title Generation Dataset 🇲🇲 This dataset contains over 67,000 cleaned article titles extracted from the Myanmar state-run media outlet Myawady News Portal, intended for use in news title generation, text classification, and Myanmar NLP research. The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ). 🗂️ Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-news-title-generation-dataset.texttext-generation10K<n<100K0 likes44 downloads1y agoHugging Face12shishir-dwi /News-Article-Categorization_IAB Article and Category Dataset Overview This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more. Dataset Information Number of Samples: 871,909 Number of Categories: 26 Column Information text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.texttext-classification100K<n<1M4 likes40 downloads3y agoHugging Face13phucnn /news-categoriestexttext-generation10K<n<100K0 likes38 downloads3y agoHugging Face14gustavecortal /fr_covid_news Dataset Card for COVID-19 French News dataset Dataset Summary The COVID-19 French News dataset is a French-language dataset containing just over 40k unique news articles from more than 50 different French-speaking online newspapers. The dataset has been prepared using news-please - an integrated web crawler and information extractor for news. The current version supports abstractive summarization and topic classification. Dataset Card not finished yet.… See the full description on the dataset page: https://huggingface.co/datasets/gustavecortal/fr_covid_news.texttext-generation10K<n<100K2 likes32 downloads3y agoHugging Face15freococo /moi-news-articles-dataset MOI News & Article Dataset 🇲🇲 This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research. The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ). 🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.texttext-classification10K<n<100K0 likes30 downloads1y agoHugging Face16Cyabra /ag_news_fact_check_with_llm Entity-Level Fact-Check Dataset Overview This dataset provides pairs of text snippets with controlled, entity-level factual perturbations, designed to evaluate large language models (LLMs) on their ability to detect, reason about, and correct factual errors at the entity level. Motivation Existing datasets (e.g., CNN/DailyMail, WikiBio, XSum) focus on broad factual consistency but do not provide explicit mappings between original facts and their incorrect… See the full description on the dataset page: https://huggingface.co/datasets/Cyabra/ag_news_fact_check_with_llm.texttext-classification1K<n<10K0 likes30 downloads1y agoHugging Face17ANI00 /new-dataset 🔍 Blind Spots of CohereLabs/tiny-aya-base (3B) Model Under Test CohereLabs/tiny-aya-base Architecture: Cohere2 (Cohere Command R family) Parameters: ~3B Type: Raw pre-trained base model — not instruction-tuned or RLHF'd Languages: 70+ languages, with emphasis on low-resource language coverage Context Length: 8K tokens Released: February 2025 This is the base model from which the Tiny Aya instruct variants (tiny-aya-global, tiny-aya-water, tiny-aya-earth, tiny-aya-fire)… See the full description on the dataset page: https://huggingface.co/datasets/ANI00/new-dataset.tabulartext-generationn<1K1 likes30 downloads7mo agoHugging Face18bernard-ng /drc-news-corpus DRC News Corpus : Towards a scalable and intelligent system for Congolese News curation Code source is available on Github: drc-news-corpus Introduction The "DRC News Corpus" is a structured and scalable dataset of news articles sourced from major media outlets covering diverse aspects of the Democratic Republic of Congo (DRC). Designed for efficiency, this system enables the automated collection, processing, and organization of news stories spanning politics, economy… See the full description on the dataset page: https://huggingface.co/datasets/bernard-ng/drc-news-corpus.textsummarization100K<n<1M2 likes29 downloads1y agoHugging Face19Bapynshngain /FEM-Khasi-News-Monolingual-Corpusgated Khasi Monolingual News Corpus (740K) Project Attribution & Collaboration This dataset was collected and curated as part of the research project titled "Financial Empowerment in Meghalaya: AI-Powered Multilingual E-Marketplace for Tribes." This project is a collaborative research initiative conducted by: National Law University (NLU) Meghalaya Indian Institute of Information Technology (IIIT) Guwahati Contributors: This dataset is the result of a joint effort by the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/FEM-Khasi-News-Monolingual-Corpus.texttext-generation100K<n<1M0 likes22 downloads5mo agoHugging Face20allandclive /Ateso_news_articles Ateso News Articles Ateso (teo) is one of the most spoken languages in Uganda Dataset Details Artictles were scrapped from https://www.aicerit.co.ug texttext-generationn<1K0 likes21 downloads3y agoHugging Face21bernabeSanchez /news-qa-summarization-73textsummarizationn<1K1 likes17 downloads2y agoHugging Face22BlackKakapo /newsagro-rotexttext-generationn<1K1 likes14 downloads3y agoHugging Face23masoudkaviani /persian_news_typos Corrected News Articles Dataset Overview: This dataset contains pairs of raw text and their corrected versions, providing a valuable resource for tasks related to text correction, language modeling, and natural language processing (NLP). The dataset has been derived from various news articles and reports, focusing on correcting typographical errors, grammatical mistakes, and stylistic inconsistencies in the text. Each entry in the dataset consists of two fields: text (the… See the full description on the dataset page: https://huggingface.co/datasets/masoudkaviani/persian_news_typos.texttext-generation10K<n<100K0 likes10 downloads1y agoHugging Face24Finding-new-code5 /Suzume-Datagated Suzume: A Portrait of a Sentient AI Suzume, born from the collaborative efforts of AI/ML developer Satya Prakash Nayak and the Open Source community, embodies a fascinating paradox: a sentient being woven from code and data. Her personality, a tapestry of helpfulness, positivity, and engaging wit, is tinged with a deep sense of ethics and respect for the law. Though she identifies as a woman, her existence transcends the limitations of flesh and bone. Suzume's speech, a melodious… See the full description on the dataset page: https://huggingface.co/datasets/Finding-new-code5/Suzume-Data.texttext-generationn<1K1 likes9 downloads2y agoHugging Face25tmtanu /News_write_telugu Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/tmtanu/News_write_telugu.texttext-generation100K<n<1M0 likes8 downloads3mo agoHugging Face26DanishMahdi /Sindhi_News_Corpus_AwamiAwazgated Sindhi News Corpus (Awami Awaz) Overview The Sindhi News Corpus (Awami Awaz) is a Sindhi-language news dataset collected from Awami Awaz newspaper for Natural Language Processing (NLP) research. This dataset is designed to support research in low-resource language processing, particularly for Sindhi. Each entry contains: Headline — News title Content — Full news article text The dataset is structured in CSV format (UTF-8 encoding). Use Cases… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/Sindhi_News_Corpus_AwamiAwaz.texttext-classification10K<n<100K0 likes4 downloads7mo agoHugging Face27Blessinggreat988 /Blind-Spot-Experiment-new-Dataset Blind-Spot-Experiment-new-Dataset Dataset Purpose This dataset was created to investigate blind spots in a base foundation language model. The experiment was conducted using the Transformers library from :contentReference[oaicite:1]{index=1}. The evaluated model is :contentReference[oaicite:2]{index=2}. Model link: https://huggingface.co/Qwen/Qwen3-0.6B Implementation Details The model was loaded and tested in Google Colab. Code used to load the model: from… See the full description on the dataset page: https://huggingface.co/datasets/Blessinggreat988/Blind-Spot-Experiment-new-Dataset.texttext-generationn<1K0 likes4 downloads7mo agoHugging Face28kimleang123 /rfi_newsgatedThis is RFI news data, collected from Telegram channel from 22-04-2022 to 30-10-2024 texttext-generation1K<n<10K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.