CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SuryaKrishna02 /aya-telugu-news-articles Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.texttext-generation100K<n<1M6 likes168 downloads3y agoHugging Face02fdaudens /ai-jobs-news-articles Dataset Summary This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work. Source Data The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.tabulartext-classification1K<n<10K1 likes160 downloads1y agoHugging Face03pierre-loic /climate-news-articles 🌍 Jeu de données d'articles de presse française labellisés comme traitant ou non des sujets liés au climat 🇬🇧 / 🇺🇸 : as this data set is based only on French data, all explanations are written in French in this repository. The goal of the dataset is to train a model to classify titles of French newspapers in two categories : if it's about climate or not. 🗺️ Le contexte Ce jeu de données de classification de titres d'article de presse française a été réalisé pour… See the full description on the dataset page: https://huggingface.co/datasets/pierre-loic/climate-news-articles.texttext-classification1K<n<10K1 likes109 downloads3y agoHugging Face04MIT-WAL /ai-jobs-news-articles-abstracts News articles and research abstracts on AI, labor, and jobs Dataset summary This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary. Rows: 53,526 document_class Rows Approx. date range (date column) news 29,857 Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.text10K<n<100K1 likes88 downloads2mo agoHugging Face05valurank /News_Articles_Categorization Dataset Card for News_Articles_Categorization Dataset Description 3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science Languages The text in the dataset is in English Dataset Structure The dataset consists of two columns namely Text and Category. The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.texttext-classification1K<n<10K5 likes87 downloads3y agoHugging Face06AyoubChLin /CNN_News_Articles_2011-2022 CNN News Articles 2011-2022 Dataset Introduction This dataset contains CNN News Articles from 2011 to 2022 after basic cleaning. The dataset includes the following information: Category Full text The data was downloaded from Kaggle at this URL: https://www.kaggle.com/datasets/hadasu92/cnn-articles-after-basic-cleaning. The dataset was split into two sets: Train set with 32,218 examples Test set with 5,686 examples Usage This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/CNN_News_Articles_2011-2022.texttext-classification10K<n<100K11 likes74 downloads3y agoHugging Face07Shwetasss /HinduTamil-News-Articles-Dataset HinduTamil News Articles Dataset Overview This dataset contains news articles in Tamil language scraped from the Hindu Tamil news website. Each article includes its title, author, city, published date, and text. Motivation This dataset was created to provide a comprehensive collection of Tamil news articles for research and analysis purposes. Data Sources and collection method The data in this dataset was collected from the Hindu Tamil news website… See the full description on the dataset page: https://huggingface.co/datasets/Shwetasss/HinduTamil-News-Articles-Dataset.texttext-classification10K<n<100K1 likes61 downloads3y agoHugging Face08waddledee /three_line_summarization_for_japanese_news_articlesライブドアニュースコーパスの3行要約データセットです。 Llama v2向けのプロンプトを追加して成形してあります。 学習に利用する際は、 [R_START] [R_END] をspecial tokenとして追加することを推奨します。 Number of rows: 3,907 Datasetは以下のリポジトリを利用してscrapeしました。 git@github.com:KodairaTomonori/ThreeLineSummaryDataset.git textsummarization1K<n<10K0 likes36 downloads2y agoHugging Face09freococo /moi-news-articles-dataset MOI News & Article Dataset 🇲🇲 This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research. The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ). 🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.texttext-classification10K<n<100K0 likes30 downloads1y agoHugging Face10bushra1dajam /news_articles News Articles Classification Dataset This dataset consists of news articles labeled with corresponding categories for classification tasks. Overview The news articles classification dataset is a collection of articles sourced from various news outlets, each labeled with a specific category. The dataset is designed for tasks such as text classification, topic modeling, and sentiment analysis. Dataset Information Name: News Articles Classification Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bushra1dajam/news_articles.texttext-classification1K<n<10K2 likes28 downloads2y agoHugging Face11azrai99 /the-star-news-articlesimage10K<n<100K1 likes22 downloads2y agoHugging Face12allandclive /Ateso_news_articles Ateso News Articles Ateso (teo) is one of the most spoken languages in Uganda Dataset Details Artictles were scrapped from https://www.aicerit.co.ug texttext-generationn<1K0 likes21 downloads3y agoHugging Face13kathiresh /HinduTamil-News-Articles-Dataset HinduTamil News Articles Dataset Overview This dataset contains news articles in Tamil language scraped from the Hindu Tamil news website. Each article includes its title, author, city, published date, and text. Motivation This dataset was created to provide a comprehensive collection of Tamil news articles for research and analysis purposes. Data Sources and collection method The data in this dataset was collected from the Hindu… See the full description on the dataset page: https://huggingface.co/datasets/kathiresh/HinduTamil-News-Articles-Dataset.texttext-classification10K<n<100K0 likes16 downloads2mo agoHugging Face14kdawoud91 /News_Articlestext1K<n<10K1 likes15 downloads3y agoHugging Face15mohamedah /zh-news-articles Chinese News Article Dataset A dataset of Chinese state media articles and Chinese New York Times articles first introduced in the paper An Analysis of Chinese Censorship Bias in LLMs. State media Articles were sourced from the news2016zh corpus and we automatically scraped the New York Times articles. Citation If you publish work using our datasets or CensorshipDetector, please cite our work using the following citation: @inproceedings{ahmed2025censorshipbias title… See the full description on the dataset page: https://huggingface.co/datasets/mohamedah/zh-news-articles.text1K<n<10K1 likes15 downloads1y agoHugging Face16kdawoud91 /news_articles_NLP801text1K<n<10K0 likes14 downloads3y agoHugging Face17shmazumder-cse /bangla-news-articles-sampleimagen<1K0 likes8 downloads2y agoHugging Face18philTheThill /news-articlestext1K<n<10K0 likes6 downloads3y agoHugging Face19CUTD /news_articles_dftext1K<n<10K0 likes6 downloads2y agoHugging Face20sKushagra /NewsArticlestext1K<n<10K0 likes4 downloads3y agoHugging Face21CitrusBoy /NewsArticlesgatedtextn<1K0 likes4 downloads2y agoHugging Face22jasonjxh /CNN_News_Articles_2011-2022 CNN News Articles 2011-2022 Dataset Introduction This dataset contains CNN News Articles from 2011 to 2022 after basic cleaning. The dataset includes the following information: Category Full text The data was downloaded from Kaggle at this URL: https://www.kaggle.com/datasets/hadasu92/cnn-articles-after-basic-cleaning. The dataset was split into two sets: Train set with 32,218 examples Test set with 5,686 examples Usage This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/jasonjxh/CNN_News_Articles_2011-2022.texttext-classification10K<n<100K0 likes4 downloads9mo agoHugging Face23cglez /news_articles Dataset Card for News Articles Categorization Split variation of the News Articles Categorization dataset. Dataset Structure The original dataset has been split into training and test sets using an 80/20 ratio. texttext-classification1K<n<10K0 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.