datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hoaxpedia
license: mit
task_categories:
- text-classification
language:
- en
pretty_name: hoaxpedia
size_categories:
- 10K<n<100K
HOAXPEDIA: A Unified Wikipedia Hoax Articles Dataset
Hoaxpedia is a Dataset containing Hoax articles collected from Wikipedia and semantically similar Legitimate article in 2 settings: Fulltext and Definition and in 3 splits based on Hoax:Legit ratio (1:2,1:10,1:100).
Dataset Details
Dataset Description
We introduce H OAXPEDIA, a… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/hoaxpedia.id_hoax_newsThis research proposes to build an automatic hoax news detection and collects 250 pages of hoax and valid news articles in Indonesian language.
Each data sample is annotated by three reviewers and the final taggings are obtained by voting of those three reviewers.hoaxnews-ind-classfication
HoaxNews_ind_Classification
Deduplicated copy of kornwtp/hoaxnews-ind-classfication.
Splits
split
rows
train
250
indonesian-hoax-newsmcqa_hoax_1h10r_def_bigbenchbig_russian_dialogueЭтот датасет содержит извлечённые диалоги из множества русскоязычных книг, аккуратно отформатированные в стиле ShareGPT. Он предназначен для обучения языковых моделей в формате ролевого общения, с выделением действий звёздочками.
Формат:
Каждый диалог оформлен в структуре ShareGPT.
Действия персонажей выделены звёздочками.
Поддерживается использование в моделях ролевого общения.
Объём данных:
Общий размер: ~1 ГБ.
Источник: Различные книги на русском языке.
Применение:
Этот датасет может… See the full description on the dataset page: https://huggingface.co/datasets/Hoaxer2000/big_russian_dialogue.hoaxnews-ind-classficationindonesian_hoax_news_oriindonesian_hoax_news_datasethoaxnews-ind-classficationsamantha-dialogues-ru
Samantha Dialogues — Русская локализация
...
(и дальше по шаблону)
Samantha Dialogues — Русская локализация
Этот датасет содержит переведённую на русский язык версию оригинального англоязычного набора диалогов с виртуальной ассистенткой Самантой. Оригинальные английские данные взяты из открытого датасета.
🧠 Назначение
Дообучение чат-ботов и LLM моделей на русском языке
Создание более человечных виртуальных ассистентов на русском
Ролевые диалоги и симуляция… See the full description on the dataset page: https://huggingface.co/datasets/Hoaxer2000/samantha-dialogues-ru.id-hoax-report-merge-v3We do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
The dataset is taken from nlp-brin-id/id-hoax-report-merge-v2 by filtering out null samples.
id-hoax-reportWe do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
Dataset for "Fact-Aware Fake-news Classification for Indonesian Language"
Data originates from https://saberhoaks.jabarprov.go.id/v2/ ; https://opendata.jabarprov.go.id/id/dataset/ ; https://klinikhoaks.jatimprov.go.id/
The attributes of data are:
Label_id: Binary… See the full description on the dataset page: https://huggingface.co/datasets/nlp-brin-id/id-hoax-report.id-hoax-report-mergeWe do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
Dataset for "Fact-Aware Fake-news Classification for Indonesian Language"
Disclaimer: Beta version, contains imbalanced representation of domain-specific NON-HOAX samples. We will release a new training and evaluation suite soon as a replacement of this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlp-brin-id/id-hoax-report-merge.id-hoax-report-merge-v2We do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
The dataset is taken from nlp-brin-id/id-hoax-report-merge by further preprocessing "Fact" attribute.
We further clean id-hoax-report-merge so that Fact does not include explicit summarization of hoax classification (Hoax vs. Non-Hoax) .
We remove the last sentence… See the full description on the dataset page: https://huggingface.co/datasets/nlp-brin-id/id-hoax-report-merge-v2.Les-mots-pour-le-dire-en-francais-fake-news-clickbait-hoax
[!NOTE]
Dataset origin: https://www.culture.gouv.fr/fr/thematiques/langue-francaise-et-langues-de-france/agir-pour-les-langues/moderniser-et-enrichir-la-langue-francaise/nos-publications/Les-mots-pour-le-dire-en-francais-fake-news-clickbait-hoax
Description
À l'heure des réseaux sociaux, "fake news", "hoax", ou encore "deep fake" pullulent, et le "fact checking", la "digital literacy" ou encore l' "empowerment" nous aident à les repérer...
Pour comprendre ces notions liées à… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Les-mots-pour-le-dire-en-francais-fake-news-clickbait-hoax.hoax-sensationalism-neutralWe do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
hoax_trainingAuthor: Jonathan HarrisonPublisher: Hugging FaceDOI: 10.57967/hf/6275URL: https://huggingface.co/datasets/Raiff1982/hoax_training
📖 Overview
hoax_training is a curated dataset designed to train and evaluate conversational AI models like Codette on misinformation detection, source verification, and ethical guidance.
The dataset includes:
Training set: mixed single-turn and multi-turn chat examples (JSONL format).
Validation set: focused one-shot Q&A examples for evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/hoax_training.
