CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-llm-leaderboard /contentstabular1K<n<10K25 likes16k downloads2y agoHugging Face02nvidia /Aegis-AI-Content-Safety-Dataset-1.0 🛡️ Nemotron Content Safety Dataset V1 Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description). Dataset Details Dataset Description Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.texttext-classification10K<n<100K61 likes3.6k downloads1y agoHugging Face03llm-jp /leaderboard-contents-v2tabularn<1K1 likes1.5k downloads5d agoHugging Face04TheFinAI /greek-contentstabularn<1K0 likes781 downloads2y agoHugging Face05pharaouk /stack-v2-python-with-content-chunk1tabular1M<n<10M1 likes618 downloads2y agoHugging Face06peakji /peak-anchor-content-35ktabular10K<n<100K0 likes577 downloads2y agoHugging Face07cedy243 /uploadm8-content-success-v1tabularn<1K0 likes567 downloads5h agoHugging Face08nvidia /Nemotron-3.5-Content-Safety-Dataset Nemotron 3.5 Content Safety Dataset Dataset Description: Nemotron 3.5 Content Safety Dataset is a hybrid real/synthetic supervised instruction dataset for content-safety classification of human and assistant interactions. The dataset contains text-only and image-grounded single-turn conversations. Each example asks a classifier to determine user safety, response safety, and harmful categories; a subset also covers topic-following classification. Some training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-3.5-Content-Safety-Dataset.text10K<n<100K17 likes509 downloads4mo agoHugging Face09ScalingIntelligence /swe-bench-verified-codebase-content-staging SWE-Bench Verified import argparse from dataclasses import dataclass, asdict import datasets from pathlib import Path import subprocess from typing import Dict, List import tqdm from datasets import Dataset import hashlib from dataclasses import dataclass @dataclass classCodebaseFile: path: str content: str class SWEBenchProblem: def __init__(self, row): self._row = row @property def repo(self) -> str: return self._row["repo"]… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/swe-bench-verified-codebase-content-staging.text100K<n<1M1 likes493 downloads2y agoHugging Face10ScalingIntelligence /swe-bench-verified-codebase-content SWE-Bench Verified Codebase Content Dataset Introduction SWE-bench is a popular benchmark that measures how well systems can solve real-world software engineering problems. To solve SWE-bench problems, systems need to interact with large codebases that have long commit histories. Interacting with these codebases in an agent loop using git naively can be slow, and the repositories themselves take up large amounts of storage space. This dataset provides the complete Python… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/swe-bench-verified-codebase-content.text10K<n<100K5 likes460 downloads2y agoHugging Face11M1keR /the-stack-v2-dedup-filtered-500-stars-100-forks-contentstabular1M<n<10M1 likes458 downloads1y agoHugging Face12textcleanlm /essentialweb-1.0-10B-raw-contenttext1M<n<10M0 likes440 downloads11mo agoHugging Face13peakji /peak-search-content-70ktabular10K<n<100K0 likes427 downloads2y agoHugging Face14thepowerfuldeez /the-stack-v2-train-smol-ids-updated-contentIncrementally uploaded Parquet shards under data/ with columns: repo_name: str text: str Downloaded all repos from https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated and run formatting / linting / import sort on all files Using ruff + black + ty stack for python and biomejs for js / ts / html / css / json / graphql Total amount of tokens: ~100B Example: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated-content.text100K<n<1M0 likes367 downloads1y agoHugging Face15lihaoxin2020 /abstractive-content-based-IDs Abstractive Content-Based Document IDs for Generative Retrieval Dataset for Summarization-Based Document IDs for Generative Retrieval with Language Models. Update [03/04/2025] Upload validation and test set of ACID. Add tokenized subset. @misc{li2024summarizationbaseddocumentidsgenerative, title={Summarization-Based Document IDs for Generative Retrieval with Language Models}, author={Haoxin Li and Daniel Cheng and Phillip Keung and Jungo Kasai and Noah A.… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/abstractive-content-based-IDs.text1M<n<10M0 likes232 downloads2y agoHugging Face16open-llm-leaderboard-old /contentstabular1K<n<10K0 likes217 downloads2y agoHugging Face17Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes217 downloads2y agoHugging Face18handwoven8588 /the-stack-v2-train-xsmol-contentgated The Stack v2 — 12-Language Resolved Content This dataset provides resolved file content for twelve programming languages, derived from the repository/file identifiers published in bigcode/the-stack-v2-train-full-ids. The upstream dataset ships identifiers only — each file is a pointer into the Software Heritage archive. Here, those identifiers have been resolved to their actual source text so the content is directly usable, with the upstream metadata carried through unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/handwoven8588/the-stack-v2-train-xsmol-content.tabulartext-generation100M<n<1B1 likes214 downloads2mo agoHugging Face19stacklok /llm-security-leaderboard-contentstabularn<1K0 likes202 downloads1y agoHugging Face20nicohrubec /codebase-content-SWE-bench_Verified-with-comments-and-teststext100K<n<1M0 likes188 downloads1y agoHugging Face21Hennara /MedRAG_contentstext10M<n<100M0 likes187 downloads1y agoHugging Face22nicohrubec /codebase-content-SWE-bench_Verified-no-comments-and-file-typestext10K<n<100K0 likes151 downloads1y agoHugging Face23electricsheepafrica /africa-emissions-from-livestock-manure-left-on-pasture-n-content Emissions from Livestock — Manure left on pasture (N content) | Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-emissions-from-livestock-manure-left-on-pasture-n-content.tabulartabular-classification10K<n<100K0 likes124 downloads1mo agoHugging Face24nicohrubec /codebase-content-SWE-bench_Verifiedtext1K<n<10K0 likes123 downloads1y agoHugging Face25behavior-in-the-wild /content-behavior-corpus Dataset Card for Content Behavior Corpus The Content Behavior Corpus (CBC) dataset, consisting of content and the corresponding receiver behavior. Dataset Details The progress of Large Language Models (LLMs) has largely been driven by the availability of large-scale unlabeled text data for unsupervised learning. This work focuses on modeling both content and the corresponding receiver behavior in the same space. Although existing datasets have trillions of content… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/content-behavior-corpus.tabular10K<n<100K6 likes105 downloads2y agoHugging Face26yasalma /tt-structured-contentgated Dataset Summary This dataset contains structured textual content in Markdown format extracted from Tatar-language documents, originally in EPUB and PDF formats. The documents include books and other long-form content with rich formatting. The dataset is intended to provide clean, structured, and semantically meaningful content to support natural language processing tasks, content modeling, and research in Tatar language technologies. The extracted Markdown preserves key… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tt-structured-content.text10K<n<100K1 likes98 downloads7mo agoHugging Face27thepowerfuldeez /the-stack-v2-extra-python-content580M tokens text10K<n<100K0 likes89 downloads11mo agoHugging Face28yuxi-liu-wired /style-content-grid-SDXL Style Content Grid SDXL Dataset Structure The dataset contains 1738 images of resolution 1024x1024, generated by Stable Diffusion XL (sd_xl_base_1.0 with model hash 31e35c80fc). They were all generated in lllyasviel/stable-diffusion-webui-forge, with the following positive and negative prompts: Positive prompt: <style> of a <content>, \n masterpiece, best quality, high quality, Negative prompt: (worst quality, low quality, normal quality), with the following… See the full description on the dataset page: https://huggingface.co/datasets/yuxi-liu-wired/style-content-grid-SDXL.image1K<n<10K0 likes85 downloads2y agoHugging Face29textcleanlm /essentialweb-1.0-10B-clean-contenttext1M<n<10M0 likes84 downloads11mo agoHugging Face30akahana /wikimedia-id-content-onlytext100K<n<1M0 likes79 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.