datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embedded-systems-qa
Embedded Systems Engineering Q&A — Instruction Dataset
A hand-authored instruction-tuning dataset of technical question/answer pairs for
embedded systems engineering, formatted for supervised fine-tuning of Mistral 7B
(Alpaca-style instruction / input / output schema).
At a glance
Entries
302
Format
JSONL, one JSON object per line
Schema
{"instruction": <question>, "input": "", "output": <answer>}
Language
English
Avg. answer length
~590… See the full description on the dataset page: https://huggingface.co/datasets/eniomecaj/embedded-systems-qa.rag-systems-sft-100k
RAG Systems SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering Retrieval-Augmented Generation (RAG) systems — from basic pipelines to advanced multi-hop retrieval, evaluation, and production optimization. Designed to train AI assistants that can help engineers build, debug, and scale RAG applications.
Dataset Description
This dataset covers the full spectrum of RAG system development across 12 specialized categories.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-systems-sft-100k.principal-systems-architect-dataset
Principal Systems Architect & Kernel Engineering Dataset
A high-quality, expert-level synthetic dataset of 200 comprehensive scenarios, questions, rationales, and detailed solutions focused on High-Performance Distributed Systems, Kernel Architecture, and Systems Programming.
Each entry has been carefully structured, validated using Pydantic, and generated using the advanced Claude 3.7 Opus model (claude-opus-4-7).
Dataset Structure
The dataset is stored in JSON… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/principal-systems-architect-dataset.
