datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AetherSearch_DPO
🔭 AetherSearch DPO
Preference pairs for reasoning, retrieval, and evidence-grounded answers
🏠 Project ·
🧠 DPO Model ·
🧪 Training Code ·
🎓 SFT Data ·
🤖 SFT Model
Dataset overview
AetherSearch DPO contains 2,126 preference pairs for training an agentic-search
policy after supervised fine-tuning. Every row provides one shared prompt, a
preferred assistant continuation, and a non-preferred continuation.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_DPO.AetherSearch_SFT
Search-SFT 2000
Dataset Overview
This release contains 2,000 validated full agent trajectories for Qwen2.5-3B
Agentic Search format cold start. Every trajectory ends with the Qwen assistant
termination token <|im_end|>.
Composition
Trajectory type
Records
Share
single_search
1,025
51.25%
multi_search
975
48.75%
Total
2,000
100.00%
Search-depth distribution:
Search depth
Records
Share
1
1,025
51.25%
2
667
33.35%… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_SFT.
