datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.Reddit-SFW-Writing_Prompts_ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
[Description Tags],"Deleted user", "Hello,\n\nYour post has been removed..", "Post has been deleted by user", "This post has been marked NSFW", duplicated system and human turns, etc has been removed.
anime-caption-danbooru-2021-sfw-5m-hq
Dataset Card for anime-caption-danbooru-2021-sfw-5m-hq
Dataset Summary
This is 5.71 M captions of 1.43 M images from a safe-for-work (SFW) filtered subset of the Danbooru 2021 dataset. There are 4 captions per image: 1 by CogVLM, 1 by llava-v1.6-34b, 1 llava-v1.6-34b cleaned, and 1 llava-v1.6-34b shortened. See the sections below for how they were generated.
Most captions are substantially larger than 77 tokens and are unsuitable for discrimination using current… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/anime-caption-danbooru-2021-sfw-5m-hq.Claude-4.0-DeepSeek-R1-RP-SFWishCreated from a python script I wrote that generates random plots within certain categories, and then creates about 5-15 responses.
Each response length is randomly selected from a small list to keep the responses dynamic. I also make the LLM respond in third person 2/3 times, and in first person 1/3 times (as I have seen this done sometimes as well)
I also have a cleanup step, where I am using another model to clean up the responses (Sometimes sentences are cut off from reaching the maximum… See the full description on the dataset page: https://huggingface.co/datasets/SuperbEmphasis/Claude-4.0-DeepSeek-R1-RP-SFWish.pickapic_v1_no_images_training_sfw
Dataset Information
This is an SFW sanitized prompt only version of the PickAPic dataset, with 335,000 prompts and image URLs.
Citation Information
If you find this work useful, please cite:
@inproceedings{Kirstain2023PickaPicAO,
title={Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation},
author={Yuval Kirstain and Adam Polyak and Uriel Singer and Shahbuland Matiana and Joe Penna and Omer Levy},
year={2023}
}
LICENSE
MIT License… See the full description on the dataset page: https://huggingface.co/datasets/CarperAI/pickapic_v1_no_images_training_sfw.Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k
Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k
概要
deepseek-ai/DeepSeek-R1-0528を用いて作成した、約10000件の日本語ロールプレイの対話を収録した合成データセットです。各データは20ターン程度あります。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(全年齢、R-15)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)
設定等の情報からsystem… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k.Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k-formatted
Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10k-formatted
概要
deepseek-ai/DeepSeek-R1-0528を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-R1-0528-10kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
Japanese-RP-Bench-testdata-SFW
Japanese-RP-Bench-testdata-SFW
本データセットは、LLMの日本語ロールプレイ能力を計測するベンチマークJapanese-RP-Bench用の評価データセットです。
ベンチマークの詳細については記事を参照してください。
データの概要
本データは以下のようなキーを持ちます。
genre: ロールプレイのジャンル
tag: ロールプレイの年齢区分
world_setting: ロールプレイの世界観設定
scene_setting: ロールプレイのシーン設定
user_setting: ロールプレイのユーザー側キャラクター設定
assistant_setting: ロールプレイのアシスタント側キャラクター設定
dialogue_tone: ロールプレイの対話のトーン
first_user_input: ロールプレイの最初のユーザー発話
response_format: ロールプレイの応答形式
id: データのid
ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-RP-Bench-testdata-SFW.genre-taxonomy-sfw
Genre Taxonomy — SFW Split
The safe-for-work half of a two-part short-fiction genre taxonomy: 119,753 genre
entries / 197,509 labels (genre + subgenre names; no story text). The adult
companion split lives in the paired repo genre-taxonomy-nsfw.
Every label was screened by multi-round LLM judging (large single-pass scan, then
targeted re-judging rounds) under a double-pass agreement standard: a label is
only acted on when independent judging passes agree. Labels judged… See the full description on the dataset page: https://huggingface.co/datasets/baiango/genre-taxonomy-sfw.furry-e621-sfw-7m-hq
Dataset Card for furry-e621-sfw-7m-hq
Dataset Summary
This is 6.92 M captions of the images from the safe-for-work (SFW) split of e621 ("e926"). It extends to January 2023, before the widespread advent of machine learning images. It includes captions created by LLMs and a custom multilabel classifier along with CogVLM. There are 8 LLM (mistralai/Mistral-7B-v0.1) and 1 CogVLM (THUDM/CogVLM) captions per image.
Most captions are substantially larger than 77 tokens and are… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/furry-e621-sfw-7m-hq.Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k
Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約20000件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(全年齢、R-15)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)
設定等の情報からsystem… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k.Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k-formatted
Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k-formatted
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
Reddit-SFW-Writing_Prompts_ShareGPT_Curated
Normalized SFW Reddit Writing Prompts
Dataset Description
This dataset is a normalized, flattened version of curated Reddit writing prompts, specifically derived from ChaoticNeutrals/Reddit-SFW-Writing_Prompts_ShareGPT. It maps nested conversational arrays into a strict instruction-response schema, making it highly optimized for instruction-tuning Large Language Models.
Dataset Schema
Column Name
Type
Description
prompt
string
The input prompt, user… See the full description on the dataset page: https://huggingface.co/datasets/rafy2342/Reddit-SFW-Writing_Prompts_ShareGPT_Curated.danbooru-2021-sfw-dtg-character-tags
Dataset Card for danbooru-2021-sfw-dtg-character-tags
Dataset Summary
This is 98,810 synthetic character descriptions for all character tags found in anime-caption-danbooru-2021-sfw-5m-hq. They were created using DanTagGen-delta-rev2 and prompting for individual character tags.
Languages
The text is in danbooru general tags, where are predominantly in English.
Intended Usage
It can be used for retrieval augmented generation (RAG), specifically when… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/danbooru-2021-sfw-dtg-character-tags.db-sfw-128px-filtered-and-cropped
Danbooru SFW 512 Filtered and Cropped
A version of Danbooru SFW which has been automatically filtered and cropped to 128x128 so that the resulting images focus almost entirely on characters.
First filtering was applied to remove non-character focused images, then cropping was applied to remove horizontal/vertical bars and further focus on the
central character(s) of each image. Both steps were performed automatically by two different vision models trained on manually labelled… See the full description on the dataset page: https://huggingface.co/datasets/hayden-donnelly/db-sfw-128px-filtered-and-cropped.Civiverse-SFWClaude-RP-1.5K-SFWvinci_edit_sfw
Dataset Card for "vinci_edit_sfw"
More Information needed
aesir-rpg-charcards-gpt4-sfw-sharegpt-flatguard-splitsfw-prompt-pack
SFW Prompt Pack v3.0 — 670 Styles / 29 Categories
SFW style and wildcard pack for Stable Diffusion XL.
Compatible with Dynamic Prompts extension and Style Grid Organizer.
Contents
670 prompt styles across 29 categories
CSV format (Forge/A1111 compatible)
Wildcard files
What's New in v3.0
319 → 670 styles (+351)
New categories: FOOD (10), STATE (13), VEHICLE (5)
ARCHETYPE +8: Bartender, Belly Dancer, Chef, Geisha, Gyaru, Male Idol, Streamer, Vtuber
ACTIVITY… See the full description on the dataset page: https://huggingface.co/datasets/Kazzze/sfw-prompt-pack.aesir-rpg-charcards-gpt4-sfw-sharegpt-flatguard-split-qwqdart_v2_sft_img_sfwaesir-rpg-charcards-gpt4-sfw-sharegpt-flatguardtohsaka_rin_fate_sfwcharacters-sfw
Dataset Card for "characters-sfw"
More Information needed
