datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
US_Domestic_Messaging_PricingSFT-Paite-Multi_Messaging-Format
SFT-Paite-Multi-Messaging-Format
This dataset is a specialized collection of Paite language data designed for high-precision Supervised Fine-Tuning (SFT). It focuses exclusively on the Paite language, combining authentic conversational data with translated logical instruction sets to develop a model that possesses both native-level linguistic fluency and technical reasoning capabilities.
Dataset Composition
The dataset is structured to provide a balance between natural… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/SFT-Paite-Multi_Messaging-Format.SFT-Paite_Translation_messaging-format
Paite Vocabulary — SFT Messages (vocab_paite_2025-12-13_translate_only_SFT_messages.jsonl)
This file is the chat / messaging variant of the translate-only Paite vocabulary data. Each line is one JSON object with a messages array (user → model turn), aligned with Gemma-style and other trainers that expect role + content instead of separate instruction / input / output fields.
It is produced from vocab_paite_2025-12-13_translate_only.jsonl by process_vocab_paite.py (same row order and… See the full description on the dataset page: https://huggingface.co/datasets/sensix-zo/SFT-Paite_Translation_messaging-format.
