datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
contentsAegis-AI-Content-Safety-Dataset-1.0
🛡️ Nemotron Content Safety Dataset V1
Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description).
Dataset Details
Dataset Description
Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.leaderboard-contents-v2greek-contentsstack-v2-python-with-content-chunk1peak-anchor-content-35kuploadm8-content-success-v1Nemotron-3.5-Content-Safety-Dataset
Nemotron 3.5 Content Safety Dataset
Dataset Description:
Nemotron 3.5 Content Safety Dataset is a hybrid real/synthetic supervised instruction dataset for content-safety classification of human and assistant interactions. The dataset contains text-only and image-grounded single-turn conversations. Each example asks a classifier to determine user safety, response safety, and harmful categories; a subset also covers topic-following classification. Some training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-3.5-Content-Safety-Dataset.swe-bench-verified-codebase-content-staging
SWE-Bench Verified
import argparse
from dataclasses import dataclass, asdict
import datasets
from pathlib import Path
import subprocess
from typing import Dict, List
import tqdm
from datasets import Dataset
import hashlib
from dataclasses import dataclass
@dataclass
classCodebaseFile:
path: str
content: str
class SWEBenchProblem:
def __init__(self, row):
self._row = row
@property
def repo(self) -> str:
return self._row["repo"]… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/swe-bench-verified-codebase-content-staging.swe-bench-verified-codebase-content
SWE-Bench Verified Codebase Content Dataset
Introduction
SWE-bench is a popular benchmark that measures how well systems can solve real-world software engineering problems. To solve SWE-bench problems, systems need to interact with large codebases that have long commit histories. Interacting with these codebases in an agent loop using git naively can be slow, and the repositories themselves take up large amounts of storage space.
This dataset provides the complete Python… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/swe-bench-verified-codebase-content.the-stack-v2-dedup-filtered-500-stars-100-forks-contentsessentialweb-1.0-10B-raw-contentpeak-search-content-70kthe-stack-v2-train-smol-ids-updated-contentIncrementally uploaded Parquet shards under data/ with columns:
repo_name: str
text: str
Downloaded all repos from https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated and run
formatting / linting / import sort on all files
Using ruff + black + ty stack for python and biomejs for js / ts / html / css / json / graphql
Total amount of tokens: ~100B
Example:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated-content.abstractive-content-based-IDs
Abstractive Content-Based Document IDs for Generative Retrieval
Dataset for Summarization-Based Document IDs for Generative Retrieval with Language Models.
Update
[03/04/2025] Upload validation and test set of ACID. Add tokenized subset.
@misc{li2024summarizationbaseddocumentidsgenerative,
title={Summarization-Based Document IDs for Generative Retrieval with Language Models},
author={Haoxin Li and Daniel Cheng and Phillip Keung and Jungo Kasai and Noah A.… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/abstractive-content-based-IDs.contentsmedium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.the-stack-v2-train-xsmol-content
The Stack v2 — 12-Language Resolved Content
This dataset provides resolved file content for twelve programming languages,
derived from the repository/file identifiers published in
bigcode/the-stack-v2-train-full-ids.
The upstream dataset ships identifiers only — each file is a pointer into the
Software Heritage archive. Here, those
identifiers have been resolved to their actual source text so the content is
directly usable, with the upstream metadata carried through unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/handwoven8588/the-stack-v2-train-xsmol-content.llm-security-leaderboard-contentscodebase-content-SWE-bench_Verified-with-comments-and-testsMedRAG_contentscodebase-content-SWE-bench_Verified-no-comments-and-file-typesafrica-emissions-from-livestock-manure-left-on-pasture-n-content
Emissions from Livestock — Manure left on pasture (N content) | Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-emissions-from-livestock-manure-left-on-pasture-n-content.codebase-content-SWE-bench_Verifiedcontent-behavior-corpus
Dataset Card for Content Behavior Corpus
The Content Behavior Corpus (CBC) dataset, consisting of content and the corresponding receiver behavior.
Dataset Details
The progress of Large Language Models (LLMs) has largely been driven by the availability of large-scale unlabeled text data for unsupervised learning. This work focuses on modeling both content and the corresponding receiver behavior in the same space. Although existing datasets have trillions of content… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/content-behavior-corpus.tt-structured-content
Dataset Summary
This dataset contains structured textual content in Markdown format extracted from Tatar-language documents, originally in EPUB and PDF formats. The documents include books and other long-form content with rich formatting. The dataset is intended to provide clean, structured, and semantically meaningful content to support natural language processing tasks, content modeling, and research in Tatar language technologies.
The extracted Markdown preserves key… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tt-structured-content.the-stack-v2-extra-python-content580M tokens
style-content-grid-SDXL
Style Content Grid SDXL
Dataset Structure
The dataset contains 1738 images of resolution 1024x1024, generated by Stable Diffusion XL (sd_xl_base_1.0 with model hash 31e35c80fc). They were all generated in lllyasviel/stable-diffusion-webui-forge,
with the following positive and negative prompts:
Positive prompt: <style> of a <content>, \n masterpiece, best quality, high quality,
Negative prompt: (worst quality, low quality, normal quality),
with the following… See the full description on the dataset page: https://huggingface.co/datasets/yuxi-liu-wired/style-content-grid-SDXL.essentialweb-1.0-10B-clean-contentwikimedia-id-content-only
