datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cybersecurity-reasoning-cot-v1
🛡️ Expert Cybersecurity Reasoning Dataset (CoT)
This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic.
💎 Key Highlights
Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets.
Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.science-cot-dataset
ExpertData Science — Scientific Reasoning
Expert-Annotated · Rights-Cleared · Ground-Truth Verified · PII-Clean
Each record captures a complete experimental or theoretical reasoning chain:
Hypothesis → Methodology → Causal Chain → Validated Conclusion.
Extracted from peer-reviewed papers across physics, biology, materials science, astrophysics, and neuroscience using structured scientific-reasoning extraction.
This dataset is produced by the ExpertData-Factory pipeline
(Mine →… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/science-cot-dataset.otaku-expert-dataset
Animetix Otaku Expert Fine-Tuning Dataset
This is the unified expert Supervised Fine-Tuning (SFT) training dataset for the Animetix Otaku Reasoning models. It is written 100% in French without code-switching.
Dataset Proportions
To ensure a balanced and robust reasoning model, the dataset is built using strict mathematical proportions:
80% Specialized Otaku Knowledge: Data-driven relational facts about anime, manga, seiyuu, French voice actors (VF), magazines… See the full description on the dataset page: https://huggingface.co/datasets/MissawB/otaku-expert-dataset.adb-expert-dataset
ADB Expert Dataset
A comprehensive dataset for training and evaluating models on Android Debug Bridge (ADB) expertise tasks. Contains three complementary subtasks covering command generation, device state remediation, and log diagnosis.
Dataset Structure
1. NL -> ADB Command Generation (nl2adb)
100 samples -- Natural language instructions mapped to ADB commands.
Split
Count
Train
73
Validation
21
Test
6
Fields:
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/santa8232/adb-expert-dataset.robotframework-expert-dataset
Dataset
This dataset is built from:
Local Robot Framework documentation files in sources/robotframework_docs/ (if present)
Curated synthetic examples in data/synthetic_examples.json
License and attribution
Robot Framework docs remain under their original licenses. Do not redistribute doc-derived datasets unless the license allows it.
Synthetic examples are authored for this project.
Files
train.jsonl and eval.jsonl: SFT records using messages format… See the full description on the dataset page: https://huggingface.co/datasets/arvind3/robotframework-expert-dataset.Tobacco-Expert-DatasetTobacco-Expert-Dataset2
