datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Content-Safety-Audio-Dataset
Nemotron Content Safety Audio Dataset
Dataset Description
The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories.
LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.hk_content_corpus
HK Content Corpus (Cantonese & Traditional Chinese)
This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms.
It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling.
Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators.
This… See the full description on the dataset page: https://huggingface.co/datasets/SolarisCipher/hk_content_corpus.aliiihussain_social-media-viral-content-and-engagement-metrics
Social Media Viral Content & Engagement Metrics
What Makes Content Go Viral? Engagement, Sentiment, and Social Trends Dataset
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 1,836
Files: 1
Files
social_media_viral_content_dataset.csv
Mirrored from Kaggle
hk_content_corpus
HK Content Corpus (Cantonese & Traditional Chinese)
This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms.
It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling.
Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators.
This… See the full description on the dataset page: https://huggingface.co/datasets/IKMLab-team/hk_content_corpus.scrape-content-dataset-v1
Scrape Content Dataset v1
A human-curated benchmark dataset for evaluating web scraping engines on content quality.
Overview
This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time.
Dataset Structure
CSV format with columns:
id: Sequential identifier
url:… See the full description on the dataset page: https://huggingface.co/datasets/firecrawl/scrape-content-dataset-v1.french-brand-content-benchmark-2026
French Brand Content Benchmark 2026
Publisher: Big NeuronsWebsite: https://www.bigneurons.comEnglish version: https://www.bigneurons.com/enContact: brief@bigneurons.comLicense: CC BY 4.0Last updated: March 2026DOI: 10.5281/zenodo.18927033tags:
brand-content
marketing
france
benchmark
acquisition
geo
What is this dataset?
The French Brand Content Benchmark 2026 is the first publicly available benchmark of brand content performance metrics for French SMEs and… See the full description on the dataset page: https://huggingface.co/datasets/BigNeurons/french-brand-content-benchmark-2026.content_rephrasing
Message Content Rephrasing Dataset
Introduced by Einolghozati et al. in Sound Natural: Content Rephrasing in Dialog Systems https://aclanthology.org/2020.emnlp-main.414/
We introduce a new task of rephrasing for amore natural virtual assistant. Currently, vir-tual assistants work in the paradigm of intent-slot tagging and the slot values are directlypassed as-is to the execution engine. However,this setup fails in some scenarios such as mes-saging when the query given by the user… See the full description on the dataset page: https://huggingface.co/datasets/facebook/content_rephrasing.scrape-content-dataset-v1
Scrape Content Dataset v1
A human-curated benchmark dataset for evaluating web scraping engines on content quality.
Overview
This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time.
Dataset Structure
CSV format with columns:
id: Sequential… See the full description on the dataset page: https://huggingface.co/datasets/zhoubinghong/scrape-content-dataset-v1.guardian_articles_full_contentsecurity_contentcontent-eye-2a4aa2
content-eye-2a4aa2
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/satomikako2/content-eye-2a4aa2.Hate_Speech_and_Offensive_Content_Identificationgermeval-2025-harmful-content-detection-training-dataset
GermEval 2025 Harmful Content Detection - Training Sets
(Call to Action • Attacks on Democratic Basic Order • Violence)
Author: Samuel Ruairí Bullard - University of Regensburg
Models: Model Zoo (Gradio Space)
Base model: LSX-UniWue/ModernGBERT_134M
Competition: GermEval 2025 Shared Task
Collection: GermEval 2025 Contribution CollectionabullardUR@GermEval Shared Task 2025 Submission
Dataset Summary
This repository republishes the training splits used… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/germeval-2025-harmful-content-detection-training-dataset.africa-synth-energy-oilgas-local-content-nigeria
Africa Synth Energy Oilgas Local Content Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-local-content-nigeria.content-moderation-input-dataset
Access Guidelines - READ THIS BEFORE REQUESTING ACCESS!
Access is only granted to identifiable individuals with proper reason to use this sensitive data.
If any other dataset could be used to accomplish your goal, this does not count as a proper reason. Half sentences and bullet points do not suffice and will be declined. Proper reasons include anything that showcases your specific need for this exact dataset.
Content Moderation Dataset
Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/centrepourlasecuriteia/content-moderation-input-dataset.Content-Articles
Content-Articles Dataset
Overview
The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.en_paradetox_content
ParaDetox: Detoxification with Parallel Data (English). Content Task Results
This repository contains information about Content Task markup from English Paradetox dataset collection pipeline.
The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference.
ParaDetox Collection Pipeline
The ParaDetox Dataset collection was done via Yandex.Toloka crowdsource platform. The collection was done in three steps:
Task 1:… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/en_paradetox_content.Explicit_content
Dataset Card for Explicit content detection
Dataset Description
1189 News Articles classified into different categories namely: "Explicit" if the article contains explicit content and "Not_Explicit" if not.
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely Article and Category.
The Article column consists of the news article and the Category column consists of the class each article belongs… See the full description on the dataset page: https://huggingface.co/datasets/valurank/Explicit_content.style_eval_content_test
Constructed test set for evaluating metrics for content preservation in style and attribute transfer
This data is used in the meta-evaluation of metrics for content preservation in cases with style and attribute transfer.
The data consists of 500 samples with a source sentence and a transfer sentence to a specific style.
The data is annotated by 3 humans on two dimension 'style strength' (AnswerS_1,AnswerS_2,AnswerS_3) and 'content preservation' (AnswerC_1,AnswerC_2,AnswerC_3) on a… See the full description on the dataset page: https://huggingface.co/datasets/APauli/style_eval_content_test.ru_paradetox_content
ParaDetox: Detoxification with Parallel Data (Russian). Content Task Results
This repository contains information about Content Task markup from Russian Paradetox dataset collection pipeline.
ParaDetox Collection Pipeline
The ParaDetox Dataset collection was done via Yandex.Toloka crowdsource platform. The collection was done in three steps:
Task 1: Generation of Paraphrases: The first crowdsourcing task asks users to eliminate toxicity in a given sentence while keeping… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox_content.reasoning-0.01-content-llama3.1web_contentmigration-stance-contentTable_of_contentsMLS_Long_ContentPlanalgozee_netflix-content-analysis
Netflix Content Analysis
Exploratory Data Analysis of Netflix Movies and TV Shows Dataset
Dataset Info
Source: Kaggle
Original Size: 1.34 MB
Kaggle Downloads: 766
Files: 1
Files
netflix_titles.csv
Mirrored from Kaggle
Japanese-English_translation_of_contents_HScodes日本郵便が提供する「国際郵便 内容品の日英・中英訳、HSコード類」(2024/05/09)のデータに基づいています。
詳しくはサイトをご覧ください
https://www.post.japanpost.jp/int/use/publication/contentslist/index.php?id=0&ie=utf8&lang=_ja&q=
TGL_content_classificationInspiring_Content_Detection_Datasetiab-taxonomy-multilang-content
IAB Taxonomy Multilingual Content Dataset
Dataset Description
This dataset contains multilingual text data for IAB (Interactive Advertising Bureau) taxonomy classification.
The data files (train.csv, val.csv) were generated by processing, cleaning, and restructuring data file from the IAB-Taxonomy URL Content Dataset originally published on Kaggle.
The accompanying processing scripts and JSON mapping files are provided to support the data preparation and model training… See the full description on the dataset page: https://huggingface.co/datasets/VALUABLY-net/iab-taxonomy-multilang-content.
