datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tsuki-dataset
Token Compression Training Dataset
15,000 bilingual compression pairs. Prompts, markdown, and technical instructions.
Trains models to reduce LLM API costs without losing critical information.
Real compression pairs.
Verbose prompts → minimum tokens.
Markdown sections → essential content.
Technical instructions → direct commands.
Reasoning examples included.
Knows when not to compress.
Legal text preserved intact.
Medical instructions… See the full description on the dataset page: https://huggingface.co/datasets/tsuki-team/Tsuki-dataset.Translation-Sample-Dataset
TsukiOwO/Translation-Sample-Dataset
Data Source
This dataset is a portion of open-r1/OpenR1-Math-220k.
Purpose
Serves as sample data for translation text.
