datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenGameArt-CC0
Dataset Card for OpenGameArt-CC0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.imslp-midi-cc0-1.0
IMSLP MIDI Dataset (CC0-1.0)
This dataset contains MIDI files and metadata crawled from IMSLP (International Music Score Library Project) on July 21-22, 2024.
Data Fields
midi_source: URL to the original MIDI file on IMSLP (incl. original uploader).
metadata_source: URL to the original metadata on IMSLP.
file_name, title, composer, year, era, style, key, license: Metadata fields.
midi: Raw MIDI bytes.
midi_mido: JSON-serialized mido object.
How to Retrieve… See the full description on the dataset page: https://huggingface.co/datasets/TiMauzi/imslp-midi-cc0-1.0.lucid-cc0-v2-hc-512
LUCID CC0 v2 HC 512 — 512×512 High-Complexity SISR Finetuning Dataset
The final stage of the LUCID three-stage training pipeline. Contains 512×512 high-complexity tiles for finetune-finetuning SISR models on high-resolution details. This is the highest-quality subset of the LUCID dataset family.
Format: WebDataset .tar shards (~1 GB each). Optimized for streaming training.
Statistics
Metric
Value
Tiles
100,866
Resolution
512×512 PNG
Total size
~51… See the full description on the dataset page: https://huggingface.co/datasets/Phips/lucid-cc0-v2-hc-512.cc0-music-captionedCollected from various CC0 music sites. These are all instrumental - none have lyrics. They are annotated with a description of the song.
Some of the music comes from FreePD, a site that shared public domain music. The FreePD website has since been taken down.
All songs were created by humans, not AI-generated.
anime-with-caption-cc0
Anime with caption CC-0 dataset
このデータセットはイラストに対する日本語キャプションを
倫理的に学習しやすくするためのデータセットです。
ここに掲載されているイラストは自律的にAIが作成したものであり、
著作権はありません。またキャプションも自律的につけられたものなので、
著作権はありません。したがって、データセットの著作権を私は放棄します。
勝手に使ってください。
ライセンス
パブリックドメイン
データセットの構成
データセットは以下の列で構成されています。
image: Emi 2でランダムに生成した画像
prompt: 言語モデルでランダムに生成された画像のプロンプト(ただし、画像とあまり一致していないため、あてにならない)
phi3_caption: Phi-3 VisionでDense captioningした結果
phi3_caption_ja: phi3_captionをPhi-3 Mediumで日本語訳した結果
イラストの作り方… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/anime-with-caption-cc0.megalith-cc0
Megalith-CC0
A CC0-filtered version of the Megalith-10m dataset. The images have also been persisted to an independent public S3 bucket, supported by the AWS Open Data Registry program, for durability.
Why filter by CC0?
The images in Megalith-10m, having been gathered from Flickr, have attached licenses of CC0 and public domain. However, it is not clear if users assigning the public domain license to their works understand the implications of the public domain… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/megalith-cc0.StockImages-CC0
CC0 Stock Images Dataset
This dataset contains a collection of stock images that are covered by the Creative Commons Zero (CC0) License, meaning they are free for personal and commercial use with no attribution required. It is designed to support a variety of computer vision tasks such as image tagging, categorization, and machine learning model training.
Disclaimer
While every effort has been made to ensure the reliability and correctness of the data presented, the… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/StockImages-CC0.cc0-textures
Dataset Card for CC0 Textures
Dataset Summary
This dataset contains 18,785 texture images from cc0-textures.com. It includes textures of wood, metal, concrete, fabric, stone, ceramic, and other materials. The original archives were downloaded, unpacked, and images were compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while keeping good quality.
Languages
The dataset is monolingual:
English (en): Texture titles and tags… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/cc0-textures.arXiv-CC0-v0.5
Dataset Card for ArXiv-CC0
Waifu to catch your attention.
Dataset Details
Dataset Description
ArXiv CC0 is a cleaned dataset of a raw scrape of arXiv using the latest metadata from January 2024.
Filtering to a total amount of tokens of ~2.77B (llama-2-7b-chat-tokenizer) / ~2.43B (RWKV Tokenizer) from primarily English language.
Curated by: M8than
Funded by: Recursal.ai
Shared by: M8than
Language(s) (NLP): Primarily English
License: cc-by-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/recursal/arXiv-CC0-v0.5.k3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.image-text-pairs-ja-cc0-2
はじめに
このデータセットは画像生成で日本語を生成したいときに使うデータセットです。
ライセンス
CC-0です。著作権を放棄して使いやすくしました。
作り方の概要
gpt-oss-20bを使って、約8万個からなる単語集兼短文集を作りました。
その文章をPillowとPythonでランダム要素を入れながら100万枚と10万枚でレンダリングしました。
フォントはNoto Sans JPなのでライセンス的には問題ないと思います。
idstructured-cc0fsd50k-cc0-curated-v1
FSD50K CC0 Curated v1
A 1,408-clip CC0-only subset of FSD50K (Fonseca et al., 2022), curated for an RNN/LSTM audio generation teaching assignment.
Contents
1,408 WAV files from the FSD50K dev split (<file_id>.wav)
fsd50k_cc0_dev_curated_v1_manifest.csv — per-clip metadata
All files are CC0 / public domain — no attribution required
18 primary labels covering music instruments and nature ambient sounds
Total size: ~1.5 GB, total duration: ~4.73 hours
Original sample rates… See the full description on the dataset page: https://huggingface.co/datasets/HughXuechen/fsd50k-cc0-curated-v1.nos_gl_CC0Click here for English version
nos_gl_CC0
Frases con licenza libre (CC0) en galego, recollidas polo Proxecto Nós co fin de alimentar o corpus textual de Mozilla Common Voice.
As frases foron cedidas á Universidade de Santiago de Compostela por diferentes institucións públicas ou privadas, ás que agradecemos a colaboración.
Sobre este material, dentro do marco do Proxecto Nós, levouse a cabo unha serie de transformacións: segmentación das frases orixinais, filtrado pola lonxitude e, no seu caso… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/nos_gl_CC0.lucid-cc0-v2
LUCID CC0 v2 — 256×256 SISR Training Dataset
A large-scale, high-quality dataset for training single image super-resolution (SISR) models. Filtered from nyuuzyou/pxhere (CC0-licensed photography) using the LUCID filtering pipeline.
Format: WebDataset .tar shards (~1 GB each). Optimized for streaming training.
Statistics
Metric
Value
Tiles
1,169,792
Resolution
256×256 PNG
Total size
~158 GB
Shards
146 tar files (~1 GB each)
Source images
~33,000… See the full description on the dataset page: https://huggingface.co/datasets/Phips/lucid-cc0-v2.diffrhythm-instrument-cc0-oepngamearg-10x5-generated
What is this
A dataset of 50 instrumental music tracks generated with the DiffRhythm model, using 10 CC0-licensed instrument samples from OEPN Game Art.
Paper: https://huggingface.co/papers/2503.01183
Project page: https://nzqian.github.io/DiffRhythm/
Models
This dataset was created by running the DiffRhythm model on 2025 Mar 05 using a copy of the space available at https://huggingface.co/spaces/ASLP-lab/DiffRhythm.
The exact architecture (base or VAE) of the DiffRhythm… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/diffrhythm-instrument-cc0-oepngamearg-10x5-generated.LIMA-SamplesOpenLogicFlow-CC0Code_Contests_Reconstructed_LargeEnron_emailimage-text-pairs-ja-cc0
Japanese Glyph Images with English Captions (CC0)
This dataset contains Japanese glyph images rendered with black text on white background.
Each .png image has a corresponding .txt file with an English caption:
This image is saying "<Japanese>". The background is white. The letter is black.
Structure
train/ — PNG images and matching TXT captions (same base filename)
provenance/assets_registry.csv — Fonts and license info
LICENSE.txt — CC0-1.0 license
Generation… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/image-text-pairs-ja-cc0.lucid-cc0-v2-hc
LUCID CC0 v2 HC — 256×256 High-Complexity SISR Finetuning Dataset
A high-complexity subset of LUCID CC0 v2 for finetuning SISR models on the most informative content. Contains only tiles with ICNet complexity ≥ 0.85, ensuring the model focuses on complex textures, edges, and patterns.
Format: WebDataset .tar shards (~1 GB each). Optimized for streaming training.
Statistics
Metric
Value
Tiles
193,262
Resolution
256×256 PNG
Total size
~27 GB
Shards… See the full description on the dataset page: https://huggingface.co/datasets/Phips/lucid-cc0-v2-hc.Star_codefsd50k-cc0-Qwen3-Omni-captionedR1_mathClinical_notesAI-TRAINING-CC0the fanime movie is creation of an 11 year old, creation of an 11 year old, in year 2011, A Retrospective Analysis of a 2022 Fan-Animated Premiere: Authored by a Pre-Adolescent Architect, Now Subject to a Grievous Shift in Destiny for that child. LICENSING NOTICE: The creator of the original animation data (drawings, timing, and audio) hereby dedicates all original contributions to the Public Domain (CC0). Permission is granted to reformat, remix, and preserve this data for any purpose. NOTE… See the full description on the dataset page: https://huggingface.co/datasets/AI-training-cc0/AI-TRAINING-CC0.timstof-dda-pasef-cc0
Claudius timsTOF DDA-PASEF PSM corpus
A large, CC0, multi-workflow corpus of peptide-spectrum matches from
timsTOF DDA-PASEF runs, built to train and benchmark peptide-property
predictors (fragment intensity, ion mobility / CCS, retention time, charge) and
to configure timsTOF simulators.
Design principle — a reference corpus, not a pre-baked training set.
We include broadly, filter minimally at build, and label richly so you can
reproduce any filtering decision yourself. Every… See the full description on the dataset page: https://huggingface.co/datasets/theGreatHerrLebert/timstof-dda-pasef-cc0.Olmo_wiki
