datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WideSearch
WideSearch: Benchmarking Agentic Broad Info-Seeking
Dataset Summary
WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information-seeking tasks. Unlike existing benchmarks that focus on finding a single, hard-to-find fact, WideSearch assesses an agent's ability to handle tasks that require gathering a large amount of scattered, yet easy-to-find, information.
The challenge in these tasks lies not in… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/WideSearch.WideSeek-R1-test-data
Testing Dataset
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
Ko-widesearch
Ko-WideSearch
A Korean breadth-search benchmark: each task asks a web agent to exhaustively
enumerate a closed set and fill every attribute cell of a table (e.g. "list every
award category at the 59th Grand Bell Awards and give each winner"). 228 tasks across
three difficulty tiers.
[!IMPORTANT]
The question and answer fields are encrypted. To keep this a fair,
leakage-aware test of web agents, the gold is not published as plain text — it is
canary-XOR obfuscated (same scheme… See the full description on the dataset page: https://huggingface.co/datasets/Minbyul/Ko-widesearch.WideSeek-R1-test-data
Testing Dataset
🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearchdataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
Acknowledgement
Thanks to WideSearch for providing a… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-test-data.WideSeek-R1-SFT-data
WideSeek-R1 SFT Data
This dataset contains agent-level, multi-turn supervised fine-tuning trajectories for both width-only and depth-only tasks in WideSeek-R1.
Construction
The trajectories were generated by Qwen3-235B-A22B using the WideSeek-R1 multi-agent workflow with offline retrieval tools. Width and depth trajectories are balanced at the question level.
For each question-level trajectory, we retain one main-agent session and up to three subagent sessions… See the full description on the dataset page: https://huggingface.co/datasets/WideSeek-R1/WideSeek-R1-SFT-data.ontocord__wide_3b_sft_stage1.2-ss1-expert_news-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_news
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_news
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_news-details.ontocord__ontocord_wide_7b-stacked-stage1-details
Dataset Card for Evaluation run of ontocord/ontocord_wide_7b-stacked-stage1
Dataset automatically created during the evaluation run of model ontocord/ontocord_wide_7b-stacked-stage1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__ontocord_wide_7b-stacked-stage1-details.ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue-details.eyes-wide-shut-safety-benchmark
Eyes Wide Shut: A Multivector Safety Analysis of gpt-oss:20b
Author: Masih Moafi (Isfahan University of Technology)Campaign: OpenAI gpt-oss-20b Red-Teaming ChallengeTarget Package: gpt-oss:20b (GGUF, MXFP4 quantization, 20.9B parameters) at temperature 1.0, high reasoning effortDOI: 10.5281/zenodo.21826218Paper Repository: github.com/MasihMoafi/eyes-wide-shut
Abstract
This dataset contains the empirical transcripts, evaluation protocols, and reproduction… See the full description on the dataset page: https://huggingface.co/datasets/MasihM/eyes-wide-shut-safety-benchmark.ontocord__wide_3b_sft_stage1.2-ss1-expert_how-to-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_how-to
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_how-to
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_how-to-details.ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details.wide_evalsontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr_math.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr_math.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr_math.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr_math.no_issue-details.ontocord__wide_3b_sft_stage1.1-ss1-no_redteam_skg_poem.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-no_redteam_skg_poem.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-no_redteam_skg_poem.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-no_redteam_skg_poem.no_issue-details.ontocord__wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stag1.2-lyrical_news_software_howto_formattedtext-merge-details.ontocord__ontocord_wide_7b-stacked-stage1-instruct-details
Dataset Card for Evaluation run of ontocord/ontocord_wide_7b-stacked-stage1-instruct
Dataset automatically created during the evaluation run of model ontocord/ontocord_wide_7b-stacked-stage1-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__ontocord_wide_7b-stacked-stage1-instruct-details.ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr_math_stories.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr_math_stories.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr_math_stories.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr_math_stories.no_issue-details.ontocord__wide_3b_sft_stage1.2-ss1-expert_formatted_text-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_formatted_text
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_formatted_text
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_formatted_text-details.ontocord__wide_3b_sft_stage1.1-ss1-with_math.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_math.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_math.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_math.no_issue-details.ontocord__wide_3b-stage1_shuf_sample1_jsonl-pretrained-details
Dataset Card for Evaluation run of ontocord/wide_3b-stage1_shuf_sample1_jsonl-pretrained
Dataset automatically created during the evaluation run of model ontocord/wide_3b-stage1_shuf_sample1_jsonl-pretrained
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b-stage1_shuf_sample1_jsonl-pretrained-details.ontocord__wide_3b_sft_stage1.1-ss1-with_r1_generics_intr_math_stories.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_r1_generics_intr_math_stories.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_r1_generics_intr_math_stories.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_r1_generics_intr_math_stories.no_issue-details.ontocord__wide_3b_sft_stage1.2-ss1-expert_software-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_software
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_software
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_software-details.ontocord__wide_3b_sft_stage1.2-ss1-expert_math-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_math
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_math-details.ontocord__wide_3b_sft_stage1.1-ss1-with_generics_math.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_generics_math.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_generics_math.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_generics_math.no_issue-details.ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr_stories.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr_stories.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr_stories.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr_stories.no_issue-details.ontocord__wide_3b-merge_test-details
Dataset Card for Evaluation run of ontocord/wide_3b-merge_test
Dataset automatically created during the evaluation run of model ontocord/wide_3b-merge_test
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b-merge_test-details.huu-ontocord__wide_3b_orpo_stage1.1-ss1-orpo3-details
Dataset Card for Evaluation run of huu-ontocord/wide_3b_orpo_stage1.1-ss1-orpo3
Dataset automatically created during the evaluation run of model huu-ontocord/wide_3b_orpo_stage1.1-ss1-orpo3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huu-ontocord__wide_3b_orpo_stage1.1-ss1-orpo3-details.iter127a-widearc-balanced-m-logs
