Phuc-HugigFace/Vietnamese-SFT-Corpus-V2
🇻🇳 Vietnamese SFT Corpus V2.1 (Balanced Safety & Anti-Over-Refusal) Vietnamese SFT Corpus V2.1 là tập dữ liệu Tinh chỉnh có Giám sát (Supervised Fine-Tuning - SFT) chuẩn công nghiệp dành cho mô hình ngôn ngữ lớn (LLM) tiếng Việt. Tập dữ liệu được thiết kế nhằm phục vụ huấn luyện trợ lý ảo thông minh, hội thoại tự nhiên, suy luận logic, đồng thời đặc trị triệt để hiện tượng "từ chối lười biếng / từ chối nhầm" (Lazy Refusal / Over-Refusal) vốn xuất hiện phổ biến ở các mô… See the full description on the dataset page: https://huggingface.co/datasets/Phuc-HugigFace/Vietnamese-SFT-Corpus-V2.
docs: clarify dataset components, add exact links, and detail 4 safety datasets
Update README.md: SFT V2.1 specification and documentation
Update validation.jsonl: SFT V2.1 validation split
Update train.jsonl: SFT V2.1 with 50:50 balanced safety & anti-over-refusal
Delete data/validation.jsonl with huggingface_hub
Delete data/train.jsonl with huggingface_hub
Update SFT V2 validation split: 1,281 samples
Update SFT V2 train split: 41,434 samples (7k Native Vietnamese Safety + GSM8K CoT)
Update Dataset Card for 42,715-sample SFT V2 with 10-Agent Peer Review Audit
Update SFT V2 validation split: 1,281 samples
Update SFT V2 train split: 41,434 samples (7k Native Vietnamese Safety + GSM8K CoT)
Upload validation split (1,131 samples)
Upload train split (36,584 samples)
Add dataset card and documentation
initial commit
