aethersearch
AetherSearch_DPO
🔭 AetherSearch DPO
Preference pairs for reasoning, retrieval, and evidence-grounded answers
🏠 Project ·
🧠 DPO Model ·
🧪 Training Code ·
🎓 SFT Data ·
🤖 SFT Model
Dataset overview
AetherSearch DPO contains 2,126 preference pairs for training an agentic-search
policy after supervised fine-tuning. Every row provides one shared prompt, a
preferred assistant continuation, and a non-preferred continuation.
This repository is… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_DPO.AetherSearch_SFT
Search-SFT 2000
Dataset Overview
This release contains 2,000 validated full agent trajectories for Qwen2.5-3B
Agentic Search format cold start. Every trajectory ends with the Qwen assistant
termination token <|im_end|>.
Composition
Trajectory type
Records
Share
single_search
1,025
51.25%
multi_search
975
48.75%
Total
2,000
100.00%
Search-depth distribution:
Search depth
Records
Share
1
1,025
51.25%
2
667
33.35%… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_SFT.AetherSearch_Eval
AetherSearch Search-R1 Full Evaluation Set
This repository contains the complete Search-R1 test.parquet used by the
current AetherSearch full-data evaluation recipe. It is an exact, unmodified
mirror of the test file published in
PeterJinGo/nq_hotpotqa_train.
Contents
File
Rows
SHA-256
test.parquet
51,713
30aa887b6d47e06e8c0f6f5307c88fe4e13461ac25a20ec0a5433ad7a4fe25dc
Source distribution:
Data source
Rows
2WikiMultiHopQA
12,576… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_Eval.
