datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
silicon-based-girlfriend-v2-dataset
矽基女友 v2 · 繁中角色扮演合成語料
繁體中文(臺灣)角色扮演的合成對話語料,2,109 筆、38,460 輪、角色輪合計 1,350 萬字元。
訓練出來的模型見 RX5950XT/silicon-based-girlfriend-v2-GGUF。
⚠️ 全部是模型合成的資料,不是真人對話。 內容包含成人向角色扮演,不適合未成年人。
僅供研究用途。所有角色皆為虛構成年人。
內容
檔案
內容
sharegpt_dataset.json
2,109 筆多輪對話,ShareGPT 格式(id / system / conversations)
grpo_prompts.json
648 題 GRPO 用的提示,與 SFT 語料零重疊
holdout_ids.json
100 筆 holdout ID 清單,這些已從 SFT 訓練集排除,供驗收用
general_probes.json
64 題通用能力探針(5 類),用來檢測微調後有無退化… See the full description on the dataset page: https://huggingface.co/datasets/RX5950XT/silicon-based-girlfriend-v2-dataset.girt-instruct
GIRT-Instruct Corpus
Paper: https://arxiv.org/abs/2402.02632
A dataset in the format of pairs of instructions and corresponding outputs. GIRT-Instruct is constructed based on GIRT-Data, a dataset of IRTs.
We use both GIRT-Data metadata and the Zephyr-7B-Beta language model to generate the instructions
This dataset is used to train the GIRT-Model model.
Model: model
Space: space
Type
We have 4 different types in GIRT-Instruct. These types include:
default: This type… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/girt-instruct.gorkhapatra-nepali-epaper
Gorkhapatra Nepali E-Paper Corpus
Per-article text extracted from PDF e-papers published on
epaper.gorkhapatraonline.com, covering 11 newspaper
slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal,
loksewa, saturday, yuwamunch, gorkhapatra-125, other).
Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs
article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.silicon-girlfriend-dataset
Silicon Girlfriend Dataset
silicon-based-girlfriend QLoRA 模型的訓練資料集。
Dataset Details / 資料集資訊
項目
內容
筆數
985 筆
格式
ShareGPT(system + conversations)
語言
繁體中文(臺灣用語)
平均對話輪數
~10 輪
最大 Token 數
8190 tokens
生成模型
Kimi K2.5
Format / 資料格式
ShareGPT 格式,每筆資料包含:
{
"system": "角色設定系統提示詞...",
"conversations": [
{"from": "human", "value": "使用者輸入"},
{"from": "gpt", "value": "角色回應"}
]
}
Files / 檔案
檔案
說明… See the full description on the dataset page: https://huggingface.co/datasets/RX5950XT/silicon-girlfriend-dataset.E-girlanime-girl-chat-dataset-en
Anime Girl Chat Dataset (English)
A synthetic Q&A chat dataset written in the voice of a friendly anime girl
character. Every row is a self-contained exchange: a user message and the
character's anime_girl reply.
Dataset details
Rows: 4000
Columns: user (string), anime_girl (string)
Language: English
Format: plain text only, no emoji or kaomoji, ASCII only
Style layers:
Sweet layer (first half): short, warm replies of 90-230 characters using ~ and small *action*… See the full description on the dataset page: https://huggingface.co/datasets/coderian/anime-girl-chat-dataset-en.
