CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Content-Safety-Audio-Dataset Nemotron Content Safety Audio Dataset Dataset Description The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories. LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.audioaudio-classification1K<n<10K5 likes954 downloads10mo agoHugging Face02SolarisCipher /hk_content_corpus HK Content Corpus (Cantonese & Traditional Chinese) This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms. It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling. Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators. This… See the full description on the dataset page: https://huggingface.co/datasets/SolarisCipher/hk_content_corpus.text1M<n<10M0 likes288 downloads1y agoHugging Face03jason1966 /aliiihussain_social-media-viral-content-and-engagement-metrics Social Media Viral Content & Engagement Metrics What Makes Content Go Viral? Engagement, Sentiment, and Social Trends Dataset Dataset Info Source: Kaggle Original Size: 0.07 MB Kaggle Downloads: 1,836 Files: 1 Files social_media_viral_content_dataset.csv Mirrored from Kaggle tabular1K<n<10K1 likes117 downloads6mo agoHugging Face04IKMLab-team /hk_content_corpus HK Content Corpus (Cantonese & Traditional Chinese) This dataset contains eight cleaned source-specific corpora of Hong Kong Cantonese and Traditional Chinese text, crawled from public websites and platforms. It was initially created for the experiments reported in https://doi.org/10.1145/3744341 which study the effect of diglossia on Hong Kong language modeling. Each file stores plain UTF-8 text, where each record occupies one line, and blank lines serve as separators. This… See the full description on the dataset page: https://huggingface.co/datasets/IKMLab-team/hk_content_corpus.text1M<n<10M0 likes93 downloads1y agoHugging Face05firecrawl /scrape-content-dataset-v1 Scrape Content Dataset v1 A human-curated benchmark dataset for evaluating web scraping engines on content quality. Overview This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time. Dataset Structure CSV format with columns: id: Sequential identifier url:… See the full description on the dataset page: https://huggingface.co/datasets/firecrawl/scrape-content-dataset-v1.text1K<n<10K0 likes88 downloads11mo agoHugging Face06BigNeurons /french-brand-content-benchmark-2026 French Brand Content Benchmark 2026 Publisher: Big NeuronsWebsite: https://www.bigneurons.comEnglish version: https://www.bigneurons.com/enContact: brief@bigneurons.comLicense: CC BY 4.0Last updated: March 2026DOI: 10.5281/zenodo.18927033tags: brand-content marketing france benchmark acquisition geo What is this dataset? The French Brand Content Benchmark 2026 is the first publicly available benchmark of brand content performance metrics for French SMEs and… See the full description on the dataset page: https://huggingface.co/datasets/BigNeurons/french-brand-content-benchmark-2026.textn<1K1 likes66 downloads7mo agoHugging Face07facebook /content_rephrasing Message Content Rephrasing Dataset Introduced by Einolghozati et al. in Sound Natural: Content Rephrasing in Dialog Systems https://aclanthology.org/2020.emnlp-main.414/ We introduce a new task of rephrasing for amore natural virtual assistant. Currently, vir-tual assistants work in the paradigm of intent-slot tagging and the slot values are directlypassed as-is to the execution engine. However,this setup fails in some scenarios such as mes-saging when the query given by the user… See the full description on the dataset page: https://huggingface.co/datasets/facebook/content_rephrasing.text1K<n<10K16 likes59 downloads4y agoHugging Face08zhoubinghong /scrape-content-dataset-v1 Scrape Content Dataset v1 A human-curated benchmark dataset for evaluating web scraping engines on content quality. Overview This dataset contains 1,000 web pages with human-annotated ground truth for evaluating how well web scraping engines capture core content while avoiding noise (navigation, ads, footers, etc.). The dataset was created in 2025-10-21 and may become outdated over time. Dataset Structure CSV format with columns: id: Sequential… See the full description on the dataset page: https://huggingface.co/datasets/zhoubinghong/scrape-content-dataset-v1.text1K<n<10K0 likes45 downloads10d agoHugging Face09EmanuelNovelo /guardian_articles_full_contenttext1K<n<10K0 likes29 downloads2y agoHugging Face10Kaballas /security_contenttext1K<n<10K1 likes29 downloads2y agoHugging Face11satomikako2 /content-eye-2a4aa2 content-eye-2a4aa2 Synthetic weather test data: 58 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/satomikako2/content-eye-2a4aa2.tabularn<1K0 likes29 downloads15d agoHugging Face12dirtycomputer /Hate_Speech_and_Offensive_Content_Identificationtext1K<n<10K0 likes26 downloads3y agoHugging Face13abullard1 /germeval-2025-harmful-content-detection-training-dataset GermEval 2025 Harmful Content Detection - Training Sets (Call to Action • Attacks on Democratic Basic Order • Violence) Author: Samuel Ruairí Bullard - University of Regensburg Models: Model Zoo (Gradio Space) Base model: LSX-UniWue/ModernGBERT_134M Competition: GermEval 2025 Shared Task Collection: GermEval 2025 Contribution CollectionabullardUR@GermEval Shared Task 2025 Submission Dataset Summary This repository republishes the training splits used… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/germeval-2025-harmful-content-detection-training-dataset.texttext-classification10K<n<100K0 likes24 downloads1y agoHugging Face14electricsheepafrica /africa-synth-energy-oilgas-local-content-nigeria Africa Synth Energy Oilgas Local Content Nigeria | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-local-content-nigeria.tabulartabular-classification1K<n<10K0 likes24 downloads1mo agoHugging Face15centrepourlasecuriteia /content-moderation-input-datasetgated Access Guidelines - READ THIS BEFORE REQUESTING ACCESS! Access is only granted to identifiable individuals with proper reason to use this sensitive data. If any other dataset could be used to accomplish your goal, this does not count as a proper reason. Half sentences and bullet points do not suffice and will be declined. Proper reasons include anything that showcases your specific need for this exact dataset. Content Moderation Dataset Overview This… See the full description on the dataset page: https://huggingface.co/datasets/centrepourlasecuriteia/content-moderation-input-dataset.texttext-classification1K<n<10K6 likes24 downloads4mo agoHugging Face16prithivMLmods /Content-Articles Content-Articles Dataset Overview The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications. Dataset Details Modalities Tabular: The dataset is structured in a tabular format. Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.tabulartext-generation10K<n<100K3 likes22 downloads2y agoHugging Face17s-nlp /en_paradetox_content ParaDetox: Detoxification with Parallel Data (English). Content Task Results This repository contains information about Content Task markup from English Paradetox dataset collection pipeline. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference. ParaDetox Collection Pipeline The ParaDetox Dataset collection was done via Yandex.Toloka crowdsource platform. The collection was done in three steps: Task 1:… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/en_paradetox_content.texttext-classification10K<n<100K0 likes21 downloads3y agoHugging Face18valurank /Explicit_content Dataset Card for Explicit content detection Dataset Description 1189 News Articles classified into different categories namely: "Explicit" if the article contains explicit content and "Not_Explicit" if not. Languages The text in the dataset is in English Dataset Structure The dataset consists of two columns namely Article and Category. The Article column consists of the news article and the Category column consists of the class each article belongs… See the full description on the dataset page: https://huggingface.co/datasets/valurank/Explicit_content.texttext-classification1K<n<10K3 likes18 downloads3y agoHugging Face19APauli /style_eval_content_test Constructed test set for evaluating metrics for content preservation in style and attribute transfer This data is used in the meta-evaluation of metrics for content preservation in cases with style and attribute transfer. The data consists of 500 samples with a source sentence and a transfer sentence to a specific style. The data is annotated by 3 humans on two dimension 'style strength' (AnswerS_1,AnswerS_2,AnswerS_3) and 'content preservation' (AnswerC_1,AnswerC_2,AnswerC_3) on a… See the full description on the dataset page: https://huggingface.co/datasets/APauli/style_eval_content_test.tabularn<1K0 likes17 downloads2y agoHugging Face20s-nlp /ru_paradetox_content ParaDetox: Detoxification with Parallel Data (Russian). Content Task Results This repository contains information about Content Task markup from Russian Paradetox dataset collection pipeline. ParaDetox Collection Pipeline The ParaDetox Dataset collection was done via Yandex.Toloka crowdsource platform. The collection was done in three steps: Task 1: Generation of Paraphrases: The first crowdsourcing task asks users to eliminate toxicity in a given sentence while keeping… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox_content.texttext-classification10K<n<100K0 likes16 downloads3y agoHugging Face21CreitinGameplays /reasoning-0.01-content-llama3.1text10K<n<100K0 likes16 downloads2y agoHugging Face22SuccessfulCrab /web_contenttextn<1K0 likes16 downloads2y agoHugging Face23MIMEDIS /migration-stance-contenttextn<1K0 likes16 downloads1y agoHugging Face24AleAle2423 /Table_of_contentstext100K<n<1M1 likes14 downloads2y agoHugging Face25YSFarag /MLS_Long_ContentPlantextsummarization1K<n<10K0 likes14 downloads2y agoHugging Face26jason1966 /algozee_netflix-content-analysis Netflix Content Analysis Exploratory Data Analysis of Netflix Movies and TV Shows Dataset Dataset Info Source: Kaggle Original Size: 1.34 MB Kaggle Downloads: 766 Files: 1 Files netflix_titles.csv Mirrored from Kaggle text1K<n<10K0 likes14 downloads6mo agoHugging Face27ikeno-ada /Japanese-English_translation_of_contents_HScodes日本郵便が提供する「国際郵便 内容品の日英・中英訳、HSコード類」(2024/05/09)のデータに基づいています。 詳しくはサイトをご覧ください https://www.post.japanpost.jp/int/use/publication/contentslist/index.php?id=0&ie=utf8&lang=_ja&q= text1K<n<10K0 likes13 downloads2y agoHugging Face28tyrealqian /TGL_content_classificationimagen<1K0 likes9 downloads2y agoHugging Face29PradeepReddyThathireddy /Inspiring_Content_Detection_Datasettext1K<n<10K1 likes8 downloads5y agoHugging Face30VALUABLY-net /iab-taxonomy-multilang-content IAB Taxonomy Multilingual Content Dataset Dataset Description This dataset contains multilingual text data for IAB (Interactive Advertising Bureau) taxonomy classification. The data files (train.csv, val.csv) were generated by processing, cleaning, and restructuring data file from the IAB-Taxonomy URL Content Dataset originally published on Kaggle. The accompanying processing scripts and JSON mapping files are provided to support the data preparation and model training… See the full description on the dataset page: https://huggingface.co/datasets/VALUABLY-net/iab-taxonomy-multilang-content.text10K<n<100K2 likes8 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.