datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NewsLensSync
Dataset Card for NewsLensSync
Dataset Description
This dataset, named NewsLensSync, contains a curated collection of news articles, sourced from trusted domains such as BBC, Reuters, AP News, NPR, PBS, The Guardian, WSJ, NY Times, and ProPublica. Each article includes both the original content and a synthetic "falsified" version of the article description, generated using a transformer-based negation model. The dataset is designed for research in misinformation… See the full description on the dataset page: https://huggingface.co/datasets/sparklessszzz/NewsLensSync.market-news-qa
Market News QA — Market-Analysis Instruction Dataset
Concise question-and-answer pairs for market analysis and financial news
interpretation: classifying news by market area, reading sentiment, identifying
who a story matters to, and answering forward-looking questions from earnings calls.
Built for the Adaption Labs AutoScientist Challenge (Market-Analysis & News category).
Rows
10,011
Distinct answers
8,469 (85%)
Duplicate questions
none
Nulls
none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/market-news-qa.moi-news-articles-dataset
MOI News & Article Dataset 🇲🇲
This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.drc-news-corpus
DRC News Corpus : Towards a scalable and intelligent system for Congolese News curation
Code source is available on Github: drc-news-corpus
Introduction
The "DRC News Corpus" is a structured and scalable dataset of news articles sourced from major media outlets covering diverse aspects of the Democratic Republic of Congo (DRC). Designed for efficiency, this system enables the automated collection, processing, and organization of news stories spanning politics, economy… See the full description on the dataset page: https://huggingface.co/datasets/bernard-ng/drc-news-corpus.new-dataset
🔍 Blind Spots of CohereLabs/tiny-aya-base (3B)
Model Under Test
CohereLabs/tiny-aya-base
Architecture: Cohere2 (Cohere Command R family)
Parameters: ~3B
Type: Raw pre-trained base model — not instruction-tuned or RLHF'd
Languages: 70+ languages, with emphasis on low-resource language coverage
Context Length: 8K tokens
Released: February 2025
This is the base model from which the Tiny Aya instruct variants (tiny-aya-global, tiny-aya-water, tiny-aya-earth, tiny-aya-fire)… See the full description on the dataset page: https://huggingface.co/datasets/ANI00/new-dataset.Camildae
merge of some datasets from Alpaca Cot
news-qa-summarization-73create_qa_news
질문 생성: kullm3 모델 이용
답변 생성: GPT3.5 turbo API 이용
지문 원본: AI HUB 뉴스 기계독해 데이터셋
Urdu-News
[Your Dataset Name]
Dataset Description
This dataset appears to be a collection of news headlines and their corresponding news text. Based on the provided image sample, the text content is in a language that uses the Arabic/Persian script, likely Persian (Farsi) or a similar Middle Eastern language. The dataset is structured in a tabular format suitable for various natural language processing tasks related to news content.
Dataset Structure
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Urdu-News.cointelegraph_news_English
Dataset cointelegraph English
Dataset Description
It is a dataset where information about the title, description, author, etc. is collected.
approx: 10041 row
page: https://cointelegraph.com/
categorie: #cryptocurrency, #Bitcoin, #Ethereum ...
cameo_newsDataset used in my thesis (https://github.com/valentinwerner1/Thesis_RelationExtraction_PoliticsNews)
Reformatted for training with LLMs, experimenting whether these can improve performance
newsq
時事情報に関する日本語QAデータセット『ニュースQ』
ニュースQ紹介ページ
利用規約はこちら
個人情報の取り扱い:利用申込の際にお預かりした個人情報(お名前、所属、利用目的、メールアドレス)は、下記の目的で利用し、弊社の個人情報保護方針に従って取り扱います。
本ツールの使用状況の確認
本人の所属が正しく申請されているかの確認
本ツールをご使用いただくために必要なご連絡(アップデートのご連絡等)
本ツールを使用した感想等を調査するためのご連絡
newsampleSindhi_News_Corpus_AwamiAwaz
Sindhi News Corpus (Awami Awaz)
Overview
The Sindhi News Corpus (Awami Awaz) is a Sindhi-language news dataset collected from Awami Awaz newspaper for Natural Language Processing (NLP) research.
This dataset is designed to support research in low-resource language processing, particularly for Sindhi.
Each entry contains:
Headline — News title
Content — Full news article text
The dataset is structured in CSV format (UTF-8 encoding).
Use Cases… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/Sindhi_News_Corpus_AwamiAwaz.newscookrfi_newsThis is RFI news data, collected from Telegram channel from 22-04-2022 to 30-10-2024
