ezfiez/Neru-66K-Bilingual-SFT
Neru-66K-Bilingual-SFT This dataset is a high-quality, professional 66,000 (66K) row bilingual Supervised Fine-Tuning (SFT) instruction set optimized for training Large Language Models (LLMs) in both Turkish-to-English and English-to-Turkish translation tasks. Non-Synthetic Dataset Details Curated by: ezfiez dev Language(s) (NLP): Turkish, English License: CC-BY-4.0 (Permissive license. Free to use for both commercial and personal projects, provided appropriate… See the full description on the dataset page: https://huggingface.co/datasets/ezfiez/Neru-66K-Bilingual-SFT.
Neru-66K-Bilingual-SFT
This dataset is a high-quality, professional 66,000 (66K) row bilingual Supervised Fine-Tuning (SFT) instruction set optimized for training Large Language Models (LLMs) in both Turkish-to-English and English-to-Turkish translation tasks. Non-Synthetic
Dataset Details
- Curated by: ezfiez dev
- Language(s) (NLP): Turkish, English
- License: CC-BY-4.0 (Permissive license. Free to use for both commercial and personal projects, provided appropriate credit is given to the author.)
- Format: JSON Lines (.jsonl)
Direct Use
- Fine-tuning Large Language Models (LLMs) to achieve advanced bilingual translation and instruction-following capabilities.
- SFT (Supervised Fine-Tuning) and LoRA / QLoRA training pipelines.
Dataset Structure
The dataset consists of a single text column wrapped with structural tokens (<|input|>, <|output|>, </s>) to help models clearly learn context and response boundaries. The format is highly flexible and can be re-tokenized or modified as needed.
