datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-engineering-sft-100k
Data Engineering SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering modern data engineering practices — Apache Spark, dbt, Airflow, Kafka, Delta Lake, BigQuery, and Snowflake. Designed to train AI assistants that can help data engineers build, optimize, and debug production data pipelines.
Dataset Description
This dataset covers the full spectrum of data engineering across 7 specialized categories. Each record… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-engineering-sft-100k.knowledgebase-electric_engineering_test_dataThis dataset are based on question answering iterations of this dataset:
"STEM-AI-mtl/Electrical-engineering"
Question answering using Deepseek R1 from TogetherAI API checkpoint
Usage:
Reasoning trace data to injecteed as CoT chain in SCIENCE related task.
