datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WideSeek-R1-train-data
Training Dataset
🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models
We provide three datasets:
width_20k.jsonl
depth_20k.jsonl
hybrid_20k.jsonl
Dataset Sources and Relationships
width_20k.jsonl is constructed by us and is tailored for WideSearch-style tasks.
depth_20k.jsonl is sourced from ASearcher's training data.
hybrid_20k.jsonl is a balanced mixture of the two and serves as the core training set for our main training experience.
All three… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-train-data.WideSeek-R1-test-data
Testing Dataset
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
WideSeek-R1-test-data
Testing Dataset
🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearchdataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
Acknowledgement
Thanks to WideSearch for providing a… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-test-data.WideSeek-R1-SFT-data
WideSeek-R1 SFT Data
This dataset contains agent-level, multi-turn supervised fine-tuning trajectories for both width-only and depth-only tasks in WideSeek-R1.
Construction
The trajectories were generated by Qwen3-235B-A22B using the WideSeek-R1 multi-agent workflow with offline retrieval tools. Width and depth trajectories are balanced at the question level.
For each question-level trajectory, we retain one main-agent session and up to three subagent sessions… See the full description on the dataset page: https://huggingface.co/datasets/WideSeek-R1/WideSeek-R1-SFT-data.WideSeek-R1-train-data
Training Dataset
We provide three datasets:
width_20k.jsonl
depth_20k.jsonl
hybrid_20k.jsonl
Dataset Sources and Relationships
width_20k.jsonl is constructed by us and is tailored for WideSearch-style tasks.
depth_20k.jsonl is sourced from ASearcher's training data.
hybrid_20k.jsonl is a balanced mixture of the two and serves as the core training set for our main training experience.
All three datasets are curated to 20,000 examples each.
Width Dataset… See the full description on the dataset page: https://huggingface.co/datasets/WideSeek-R1/WideSeek-R1-train-data.
