datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anima-prompt-expander-v0.4
Anima Prompt Expander Dataset v0.4
作者:kitt3n
用于将中文视觉描述转换为适合 Anima-Aesthetic 的简洁英文 prompt。输出为英文标签与短关系短语,保留主体、动作、关系、构图和指定风格,只补充有视觉用途的细节。
数据组成
Split
样本
概念
train
900
450
validation
50
25
test
50
25
合计
1,000
500
按概念划分,同一概念的两种表述不会跨 split。保留来源表中的 3 条完全相同的重复输入与答案;校验未发现跨 split 的相同输入或概念泄漏,也未发现相同输入对应冲突答案。
原始 val.jsonl 文件的 split 字段仍为 val;HF 加载时将该文件映射到标准 validation split。
来源
来源于本仓库的 anima_pairs_1000_review_v0.4.md 审阅表:200 条已确认 pilot 加 800… See the full description on the dataset page: https://huggingface.co/datasets/kitt3n/anima-prompt-expander-v0.4.anima-corpus-ko-fineweb2-broad
anima-corpus-ko-fineweb2-broad
🇰🇷 한국어 broad (일반) 코퍼스 for anima conv 303M byte-level pretraining.
anima chat register 표준(a_chat_registers)의 4칸 {ko·en} × {일반·SNS} 중 ko-일반 칸을 메우기 위한 데이터셋이다. 기존 ko-일반 source 가 ~1.7MB 로 빈약(en ~202MB 대비)했던 갭을 FineWeb-2 한국어로 보강한다.
Source
Upstream: HuggingFaceFW/fineweb-2, config kor_Hang (한국어 한글 스크립트).
Extracted from: train parquet data/kor_Hang/train/000_00000.parquet (1 of 25 shards).
Field: text 컬럼만 추출 (raw UTF-8, byte-vocab256… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/anima-corpus-ko-fineweb2-broad.
