datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
monster-girl-encyclopedia-wiki
Most of the text content from the English Monster Girl Encyclopedia Wiki entries of material authored by Kenkou Cross, manually markdownified over the course of a long time. This dataset might be updated in the future.
Contents
MGE original
Monster Girl Encyclopedia I
Monster Girl Encyclopedia II
Monster Girl Encyclopedia World Guide I: Fallen Maidens
Monster Girl Encyclopedia World Guide II: Mamono Realm Traveller's Guide
Monster Girl Encyclopedia World Guide III:… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/monster-girl-encyclopedia-wiki.silicon-based-girlfriend-v2-dataset
矽基女友 v2 · 繁中角色扮演合成語料
繁體中文(臺灣)角色扮演的合成對話語料,2,109 筆、38,460 輪、角色輪合計 1,350 萬字元。
訓練出來的模型見 RX5950XT/silicon-based-girlfriend-v2-GGUF。
⚠️ 全部是模型合成的資料,不是真人對話。 內容包含成人向角色扮演,不適合未成年人。
僅供研究用途。所有角色皆為虛構成年人。
內容
檔案
內容
sharegpt_dataset.json
2,109 筆多輪對話,ShareGPT 格式(id / system / conversations)
grpo_prompts.json
648 題 GRPO 用的提示,與 SFT 語料零重疊
holdout_ids.json
100 筆 holdout ID 清單,這些已從 SFT 訓練集排除,供驗收用
general_probes.json
64 題通用能力探針(5 類),用來檢測微調後有無退化… See the full description on the dataset page: https://huggingface.co/datasets/RX5950XT/silicon-based-girlfriend-v2-dataset.ai-girlfriend-chat-datasetsilicon-girlfriend-dataset
Silicon Girlfriend Dataset
silicon-based-girlfriend QLoRA 模型的訓練資料集。
Dataset Details / 資料集資訊
項目
內容
筆數
985 筆
格式
ShareGPT(system + conversations)
語言
繁體中文(臺灣用語)
平均對話輪數
~10 輪
最大 Token 數
8190 tokens
生成模型
Kimi K2.5
Format / 資料格式
ShareGPT 格式,每筆資料包含:
{
"system": "角色設定系統提示詞...",
"conversations": [
{"from": "human", "value": "使用者輸入"},
{"from": "gpt", "value": "角色回應"}
]
}
Files / 檔案
檔案
說明… See the full description on the dataset page: https://huggingface.co/datasets/RX5950XT/silicon-girlfriend-dataset.chubby-girls-nete-girlgoth-girl-friendsNeedy_Girl_Overdose_Pftv-girlsgirl-cumNeedy_Girl_Overdoseallura-org__MoE-Girl-1BA-7BT-details
Dataset Card for Evaluation run of allura-org/MoE-Girl-1BA-7BT
Dataset automatically created during the evaluation run of model allura-org/MoE-Girl-1BA-7BT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allura-org__MoE-Girl-1BA-7BT-details.ThinkTacToe-SFTgirl-next-doorhalf_girlfriendThinkTacToe-DPOuraaka_girl
