datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ml-ai-engineer-sft
DuoNeural ML/AI Engineer SFT Dataset
A synthetic instruction-tuning dataset for training an LLM to be a useful pairing partner on ML/AI engineering work — debugging training runs, reasoning about architecture and infra choices, reviewing experiment design, and explaining core ML concepts with the specificity of someone who's actually run the experiments.
Why this dataset exists
Most general instruction-tuning data treats ML engineering questions the same as any… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/ml-ai-engineer-sft.logic_duo
LogicDuo: Bilingual Logical Reasoning Tutoring Corpus
Created using this projectСоздано с использованием этого проекта
🇷🇺 Русская версия / Russian version...
Корпус "LogicDuo": Обучение логическому мышлению на русском и английском
Специализированный датасет для обучения моделей искусственного интеллекта ведению структурированных образовательных диалогов, направленных на развитие логического и критического мышления. Каждая запись представляет собой диалог между… See the full description on the dataset page: https://huggingface.co/datasets/limloop/logic_duo.cot-reasoning-2k
DuoNeural CoT Reasoning Dataset (2K)
A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces.
Benchmark Results
Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090):
Metric
Baseline
Post-SFT
Δ Absolute
Δ Relative
GSM8K (flexible-extract)
0.3177
0.4890
+17.1pp
+53.9%
GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.duoguard-seed-datasmollm2-think-dataset-run1duoguard-iter1-dataTransgpt_sft_v2openmodelmap-modelsDSS_duocultural-rdfbrust_trainluat-vn-docsduoduoyu0v3
