datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenGameArt-CC0
Dataset Card for OpenGameArt-CC0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.imslp-midi-cc0-1.0
IMSLP MIDI Dataset (CC0-1.0)
This dataset contains MIDI files and metadata crawled from IMSLP (International Music Score Library Project) on July 21-22, 2024.
Data Fields
midi_source: URL to the original MIDI file on IMSLP (incl. original uploader).
metadata_source: URL to the original metadata on IMSLP.
file_name, title, composer, year, era, style, key, license: Metadata fields.
midi: Raw MIDI bytes.
midi_mido: JSON-serialized mido object.
How to Retrieve… See the full description on the dataset page: https://huggingface.co/datasets/TiMauzi/imslp-midi-cc0-1.0.cc0-music-captionedCollected from various CC0 music sites. These are all instrumental - none have lyrics. They are annotated with a description of the song.
Some of the music comes from FreePD, a site that shared public domain music. The FreePD website has since been taken down.
All songs were created by humans, not AI-generated.
anime-with-caption-cc0
Anime with caption CC-0 dataset
このデータセットはイラストに対する日本語キャプションを
倫理的に学習しやすくするためのデータセットです。
ここに掲載されているイラストは自律的にAIが作成したものであり、
著作権はありません。またキャプションも自律的につけられたものなので、
著作権はありません。したがって、データセットの著作権を私は放棄します。
勝手に使ってください。
ライセンス
パブリックドメイン
データセットの構成
データセットは以下の列で構成されています。
image: Emi 2でランダムに生成した画像
prompt: 言語モデルでランダムに生成された画像のプロンプト(ただし、画像とあまり一致していないため、あてにならない)
phi3_caption: Phi-3 VisionでDense captioningした結果
phi3_caption_ja: phi3_captionをPhi-3 Mediumで日本語訳した結果
イラストの作り方… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/anime-with-caption-cc0.megalith-cc0
Megalith-CC0
A CC0-filtered version of the Megalith-10m dataset. The images have also been persisted to an independent public S3 bucket, supported by the AWS Open Data Registry program, for durability.
Why filter by CC0?
The images in Megalith-10m, having been gathered from Flickr, have attached licenses of CC0 and public domain. However, it is not clear if users assigning the public domain license to their works understand the implications of the public domain… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/megalith-cc0.StockImages-CC0
CC0 Stock Images Dataset
This dataset contains a collection of stock images that are covered by the Creative Commons Zero (CC0) License, meaning they are free for personal and commercial use with no attribution required. It is designed to support a variety of computer vision tasks such as image tagging, categorization, and machine learning model training.
Disclaimer
While every effort has been made to ensure the reliability and correctness of the data presented, the… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/StockImages-CC0.cc0-textures
Dataset Card for CC0 Textures
Dataset Summary
This dataset contains 18,785 texture images from cc0-textures.com. It includes textures of wood, metal, concrete, fabric, stone, ceramic, and other materials. The original archives were downloaded, unpacked, and images were compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while keeping good quality.
Languages
The dataset is monolingual:
English (en): Texture titles and tags… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/cc0-textures.arXiv-CC0-v0.5
Dataset Card for ArXiv-CC0
Waifu to catch your attention.
Dataset Details
Dataset Description
ArXiv CC0 is a cleaned dataset of a raw scrape of arXiv using the latest metadata from January 2024.
Filtering to a total amount of tokens of ~2.77B (llama-2-7b-chat-tokenizer) / ~2.43B (RWKV Tokenizer) from primarily English language.
Curated by: M8than
Funded by: Recursal.ai
Shared by: M8than
Language(s) (NLP): Primarily English
License: cc-by-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/recursal/arXiv-CC0-v0.5.k3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.image-text-pairs-ja-cc0-2
はじめに
このデータセットは画像生成で日本語を生成したいときに使うデータセットです。
ライセンス
CC-0です。著作権を放棄して使いやすくしました。
作り方の概要
gpt-oss-20bを使って、約8万個からなる単語集兼短文集を作りました。
その文章をPillowとPythonでランダム要素を入れながら100万枚と10万枚でレンダリングしました。
フォントはNoto Sans JPなのでライセンス的には問題ないと思います。
idstructured-cc0fsd50k-cc0-curated-v1
FSD50K CC0 Curated v1
A 1,408-clip CC0-only subset of FSD50K (Fonseca et al., 2022), curated for an RNN/LSTM audio generation teaching assignment.
Contents
1,408 WAV files from the FSD50K dev split (<file_id>.wav)
fsd50k_cc0_dev_curated_v1_manifest.csv — per-clip metadata
All files are CC0 / public domain — no attribution required
18 primary labels covering music instruments and nature ambient sounds
Total size: ~1.5 GB, total duration: ~4.73 hours
Original sample rates… See the full description on the dataset page: https://huggingface.co/datasets/HughXuechen/fsd50k-cc0-curated-v1.nos_gl_CC0Click here for English version
nos_gl_CC0
Frases con licenza libre (CC0) en galego, recollidas polo Proxecto Nós co fin de alimentar o corpus textual de Mozilla Common Voice.
As frases foron cedidas á Universidade de Santiago de Compostela por diferentes institucións públicas ou privadas, ás que agradecemos a colaboración.
Sobre este material, dentro do marco do Proxecto Nós, levouse a cabo unha serie de transformacións: segmentación das frases orixinais, filtrado pola lonxitude e, no seu caso… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/nos_gl_CC0.Code_Contests_Reconstructed_LargeEnron_emailimage-text-pairs-ja-cc0
Japanese Glyph Images with English Captions (CC0)
This dataset contains Japanese glyph images rendered with black text on white background.
Each .png image has a corresponding .txt file with an English caption:
This image is saying "<Japanese>". The background is white. The letter is black.
Structure
train/ — PNG images and matching TXT captions (same base filename)
provenance/assets_registry.csv — Fonts and license info
LICENSE.txt — CC0-1.0 license
Generation… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/image-text-pairs-ja-cc0.Star_codeR1_mathfsd50k-cc0-Qwen3-Omni-captionedClinical_notestimstof-dda-pasef-cc0
Claudius timsTOF DDA-PASEF PSM corpus
A large, CC0, multi-workflow corpus of peptide-spectrum matches from
timsTOF DDA-PASEF runs, built to train and benchmark peptide-property
predictors (fragment intensity, ion mobility / CCS, retention time, charge) and
to configure timsTOF simulators.
Design principle — a reference corpus, not a pre-baked training set.
We include broadly, filter minimally at build, and label richly so you can
reproduce any filtering decision yourself. Every… See the full description on the dataset page: https://huggingface.co/datasets/theGreatHerrLebert/timstof-dda-pasef-cc0.Olmo_wikiafrica-uganda-residential-property-price-index-q1-2024-25-excel-tables-cc0c2aef
Residential Property Price Index Q1 2024 25 Excel Tables | Africa (Uganda Bureau of Statistics)
24 rows - 1 Africa country/area - 2024-2025 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 24 rows from Uganda Bureau of Statistics, covering Residential Property Price Index Q1 2024 25 Excel Tables. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-residential-property-price-index-q1-2024-25-excel-tables-cc0c2aef.megalith-cc0-recap-qwen3p5-35b-a3b
Megalith CC0 recaptions with Qwen3.5-35B-A3B
Dataset megalith-cc0: 8.073 Million caption rows.
This caption-only repository contains 8,072,536 generated captions for 8,072,536 Megalith crop identities and no duplicated image payload. Images and original source captions remain in the pinned Spawning/pd-extended release at revision a5f67f80efc42a4951f65b304aa630f4f9fb15d7. image_shard and image_member identify the pinned source Parquet and row; asset_instance_id is the Megalith… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/megalith-cc0-recap-qwen3p5-35b-a3b.self-curationarXiv_newCode_Contests_Reconstructedtest_import_dataset_from_hub_with_classlabel_cc0647e4-0b13-45db-a413-8e551b74def3test_import_dataset_from_hub_with_classlabel_ad792431-6825-46ed-aeab-cc0ae234e749
