datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WideSearch
WideSearch: Benchmarking Agentic Broad Info-Seeking
Dataset Summary
WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information-seeking tasks. Unlike existing benchmarks that focus on finding a single, hard-to-find fact, WideSearch assesses an agent's ability to handle tasks that require gathering a large amount of scattered, yet easy-to-find, information.
The challenge in these tasks lies not in… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/WideSearch.WideSeek-R1-test-data
Testing Dataset
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
mmu_hsc_pdr3_wide_21
mmu_hsc_pdr3_wide_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_hsc_pdr3_wide_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_hsc_pdr3_wide_21.football_matcheswider_face_yolo
Wider face yolo
wider_face.zip
- train
- images
- ....jpg
- labels
- ....txt
- valid
- images
- ....jpg
- labels
- ....txt
Ko-widesearch
Ko-WideSearch
A Korean breadth-search benchmark: each task asks a web agent to exhaustively
enumerate a closed set and fill every attribute cell of a table (e.g. "list every
award category at the 59th Grand Bell Awards and give each winner"). 228 tasks across
three difficulty tiers.
[!IMPORTANT]
The question and answer fields are encrypted. To keep this a fair,
leakage-aware test of web agents, the gold is not published as plain text — it is
canary-XOR obfuscated (same scheme… See the full description on the dataset page: https://huggingface.co/datasets/Minbyul/Ko-widesearch.Mistral_Trivia-QA_Dataset
Mistral Trivia QA Dataset
The Mistral Trivia QA Dataset is a collection of trivia questions and answers designed to evaluate and train question-answering models. It covers a wide range of topics and is particularly useful for assessing a model's ability to handle general knowledge and reasoning tasks.The documents are derived from WikiText-2, providing diverse and well-structured textual content suitable for extractive QA generation.
Model outputs for this dataset were generated… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Mistral_Trivia-QA_Dataset.WideSeek-R1-test-data
Testing Dataset
🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearchdataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
Acknowledgement
Thanks to WideSearch for providing a… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-test-data.WideSeek-R1-SFT-data
WideSeek-R1 SFT Data
This dataset contains agent-level, multi-turn supervised fine-tuning trajectories for both width-only and depth-only tasks in WideSeek-R1.
Construction
The trajectories were generated by Qwen3-235B-A22B using the WideSeek-R1 multi-agent workflow with offline retrieval tools. Width and depth trajectories are balanced at the question level.
For each question-level trajectory, we retain one main-agent session and up to three subagent sessions… See the full description on the dataset page: https://huggingface.co/datasets/WideSeek-R1/WideSeek-R1-SFT-data.hsc-pdr3-wide-20-embeddings
HSC PDR3 Wide r<20 Embeddings
AION-Search and AION embeddings for HSC PDR3 Wide galaxies with r_mag < 20 mag
License & data source
The embeddings and packaging in this repository are released under the MIT License.
The underlying catalog data are derived from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) and remain subject to the original HSC-SSP data-use policy and required acknowledgements.
Embeddings Citation
@misc{koblischke2025semantic… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/hsc-pdr3-wide-20-embeddings.retinaface_widerface
retinaface_widerface
这个目录用于生成可直接上传到 Hugging Face Datasets 的 WiderFace 教学版数据。
目标格式
转换完成后会生成两个 split:
train
val
对应文件路径默认是:
train/train-00000-of-00001.parquet
val/val-00000-of-00001.parquet
数据来源
train 来自 data/widerface/train/label.txt 和 data/widerface/train/images
val 来自 data/widerface/val/images 和 widerface_evaluate/ground_truth 下的 4 个 mat 文件
运行方式
在仓库根目录执行:
conda run -n retinaface python data/retinaface_widerface/build_parquet.py
如果环境里还没有… See the full description on the dataset page: https://huggingface.co/datasets/zhouxzh/retinaface_widerface.wider-facenethobench-widefield-v1
Nethobench Widefield Calcium Forecasting Dataset
Summary
This release packages a benchmark-ready widefield calcium imaging dataset for neural time-series forecasting and Nethobench-style evaluation.
The release contains two complementary data representations:
A prepared benchmark tensor: data/data100_ba16.npy, already organized into fixed-length subsequences for model training and evaluation.
An unprepared source Parquet table: data/data-clean-all.parquet, containing the… See the full description on the dataset page: https://huggingface.co/datasets/Nethobench/nethobench-widefield-v1.monolingual-wideNLIThis monolingual (English) NLI dataset is designed for performing Natural Language Inference, and is particularly Fact-Checking oriented.
Dev split is oriented to teach the model how to deal well with pure NLI (ANLI is well designed for this task) and test his general knowledge (Fact-Checking skills) with VitaminC, which is known for its robustness for this task.
It contains:
14.5k examples for the dev split of which:
848 from ANLI train_r1;
2273 from ANLI train_r2;
5023 from ANLI train_r3;… See the full description on the dataset page: https://huggingface.co/datasets/Gameselo/monolingual-wideNLI.ontocord__ontocord_wide_7b-stacked-stage1-details
Dataset Card for Evaluation run of ontocord/ontocord_wide_7b-stacked-stage1
Dataset automatically created during the evaluation run of model ontocord/ontocord_wide_7b-stacked-stage1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__ontocord_wide_7b-stacked-stage1-details.ontocord__wide_3b_sft_stage1.2-ss1-expert_news-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_news
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_news
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_news-details.ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue-details.wider
WiDe-Analysis Dataset
This is the dataset for WiDe Analysis Extended version
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/wider.wide-eye-b24176
wide-eye-b24176
Synthetic products test data: 60 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/graniteGlade/wide-eye-b24176.italian-food-wide-dpo-dataset-improvedeyes-wide-shut-safety-benchmark
Eyes Wide Shut: A Multivector Safety Analysis of gpt-oss:20b
Author: Masih Moafi (Isfahan University of Technology)Campaign: OpenAI gpt-oss-20b Red-Teaming ChallengeTarget Package: gpt-oss:20b (GGUF, MXFP4 quantization, 20.9B parameters) at temperature 1.0, high reasoning effortDOI: 10.5281/zenodo.21826218Paper Repository: github.com/MasihMoafi/eyes-wide-shut
Abstract
This dataset contains the empirical transcripts, evaluation protocols, and reproduction… See the full description on the dataset page: https://huggingface.co/datasets/MasihM/eyes-wide-shut-safety-benchmark.Cloze_QA_Dataset_Wikitext2
Cloze QA Dataset (WikiText-2)
Dataset Description
The Cloze QA Dataset is automatically generated from the WikiText-2 corpus. It contains fill-in-the-blank (cloze) style questions derived directly from sentences in Wikipedia articles. This dataset is particularly useful for evaluating local recall, reading comprehension, and contextual understanding.
Each document produces exactly three unique QA pairs, preserving document structure and sentence alignment while… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Cloze_QA_Dataset_Wikitext2.ontocord__wide_3b_sft_stage1.2-ss1-expert_how-to-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_how-to
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_how-to
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_how-to-details.ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details
Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical
Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details.africa-ghana-sector-wide-indicators-for-health-care-in-ghana-9dae8186
Sector Wide Indicators for Health Care in Ghana | Africa (Ghana Open Data)
254 rows - 1 Africa country/area - 2008-2010 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 254 rows from Ghana Open Data, covering Sector Wide Indicators for Health Care in Ghana. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ghana-sector-wide-indicators-for-health-care-in-ghana-9dae8186.wide-camera-calibrationmilitary-wide-dpo-dataset-maximal-editafrica-ghana-sector-wide-indicators-for-health-care-in-ghana-9fd69dec
Sector Wide Indicators for Health Care in Ghana | Africa (Ghana Open Data)
254 rows - 1 Africa country/area - 2008-2010 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 254 rows from Ghana Open Data, covering Sector Wide Indicators for Health Care in Ghana. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ghana-sector-wide-indicators-for-health-care-in-ghana-9fd69dec.NJ_Wide_dent_with_significant_stripesafrica-ghana-sector-wide-indicators-for-health-care-in-ghana-1e1923c5
Sector Wide Indicators for Health Care in Ghana | Africa (Ghana Open Data)
254 rows - 1 Africa country/area - 2008-2010 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 254 rows from Ghana Open Data, covering Sector Wide Indicators for Health Care in Ghana. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ghana-sector-wide-indicators-for-health-care-in-ghana-1e1923c5.
