datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
science-cot-dataset
ExpertData Science — Scientific Reasoning
Expert-Annotated · Rights-Cleared · Ground-Truth Verified · PII-Clean
Each record captures a complete experimental or theoretical reasoning chain:
Hypothesis → Methodology → Causal Chain → Validated Conclusion.
Extracted from peer-reviewed papers across physics, biology, materials science, astrophysics, and neuroscience using structured scientific-reasoning extraction.
This dataset is produced by the ExpertData-Factory pipeline
(Mine →… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/science-cot-dataset.digikala-mobile-expert-data
Digikala Mobile Phones Dataset for RAG
This repository contains a comprehensive, processed dataset of mobile phone specifications and details scraped from Digikala. It is specifically designed to build and evaluate Retrieval-Augmented Generation (RAG) systems, semantic search engines, and AI-powered shopping assistants in Persian (Farsi).
🗂 Dataset Structure (Configurations)
This dataset is divided into two distinct configurations:
1. markdowns (Raw… See the full description on the dataset page: https://huggingface.co/datasets/NavHash/digikala-mobile-expert-data.adb-expert-dataset
ADB Expert Dataset
A comprehensive dataset for training and evaluating models on Android Debug Bridge (ADB) expertise tasks. Contains three complementary subtasks covering command generation, device state remediation, and log diagnosis.
Dataset Structure
1. NL -> ADB Command Generation (nl2adb)
100 samples -- Natural language instructions mapped to ADB commands.
Split
Count
Train
73
Validation
21
Test
6
Fields:
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/santa8232/adb-expert-dataset.Tobacco-Expert-DatasetTobacco-Expert-Dataset2
