CoolFace
Datasetpublic

ezfiez/Neru-66K-Bilingual-SFT

Neru-66K-Bilingual-SFT This dataset is a high-quality, professional 66,000 (66K) row bilingual Supervised Fine-Tuning (SFT) instruction set optimized for training Large Language Models (LLMs) in both Turkish-to-English and English-to-Turkish translation tasks. Non-Synthetic Dataset Details Curated by: ezfiez dev Language(s) (NLP): Turkish, English License: CC-BY-4.0 (Permissive license. Free to use for both commercial and personal projects, provided appropriate… See the full description on the dataset page: https://huggingface.co/datasets/ezfiez/Neru-66K-Bilingual-SFT.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
1likes23downloads
Dataset Card

Neru-66K-Bilingual-SFT

This dataset is a high-quality, professional 66,000 (66K) row bilingual Supervised Fine-Tuning (SFT) instruction set optimized for training Large Language Models (LLMs) in both Turkish-to-English and English-to-Turkish translation tasks. Non-Synthetic

Dataset Details

  • —Curated by: ezfiez dev
  • —Language(s) (NLP): Turkish, English
  • —License: CC-BY-4.0 (Permissive license. Free to use for both commercial and personal projects, provided appropriate credit is given to the author.)
  • —Format: JSON Lines (.jsonl)

Direct Use

  • —Fine-tuning Large Language Models (LLMs) to achieve advanced bilingual translation and instruction-following capabilities.
  • —SFT (Supervised Fine-Tuning) and LoRA / QLoRA training pipelines.

Dataset Structure

The dataset consists of a single text column wrapped with structural tokens (<|input|>, <|output|>, </s>) to help models clearly learn context and response boundaries. The format is highly flexible and can be re-tokenized or modified as needed.


ezfiez/Neru-66K-Bilingual-SFT · CoolFace