CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ByteDance-Seed /WideSearch WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information-seeking tasks. Unlike existing benchmarks that focus on finding a single, hard-to-find fact, WideSearch assesses an agent's ability to handle tasks that require gathering a large amount of scattered, yet easy-to-find, information. The challenge in these tasks lies not in… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/WideSearch.textn<1K45 likes12k downloads1y agoHugging Face02WideSeek-R1 /WideSeek-R1-test-data Testing Dataset We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required. texttext-generationn<1K0 likes483 downloads5mo agoHugging Face03hugging-science /mmu_hsc_pdr3_wide_21 mmu_hsc_pdr3_wide_21 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_hsc_pdr3_wide_21. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_hsc_pdr3_wide_21.tabular1M<n<10M0 likes375 downloads15d agoHugging Face04WideMan /football_matchestabular1K<n<10K5 likes246 downloads2y agoHugging Face05Bingsu /wider_face_yolo Wider face yolo wider_face.zip - train - images - ....jpg - labels - ....txt - valid - images - ....jpg - labels - ....txt imageobject-detection10K<n<100K0 likes237 downloads3y agoHugging Face06Minbyul /Ko-widesearch Ko-WideSearch A Korean breadth-search benchmark: each task asks a web agent to exhaustively enumerate a closed set and fill every attribute cell of a table (e.g. "list every award category at the 59th Grand Bell Awards and give each winner"). 228 tasks across three difficulty tiers. [!IMPORTANT] The question and answer fields are encrypted. To keep this a fair, leakage-aware test of web agents, the gold is not published as plain text — it is canary-XOR obfuscated (same scheme… See the full description on the dataset page: https://huggingface.co/datasets/Minbyul/Ko-widesearch.textquestion-answeringn<1K2 likes113 downloads3mo agoHugging Face07CGU-Widelab /Mistral_Trivia-QA_Dataset Mistral Trivia QA Dataset The Mistral Trivia QA Dataset is a collection of trivia questions and answers designed to evaluate and train question-answering models. It covers a wide range of topics and is particularly useful for assessing a model's ability to handle general knowledge and reasoning tasks.The documents are derived from WikiText-2, providing diverse and well-structured textual content suitable for extractive QA generation. Model outputs for this dataset were generated… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Mistral_Trivia-QA_Dataset.textquestion-answering1 likes101 downloads2mo agoHugging Face08RLinf /WideSeek-R1-test-data Testing Dataset 🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearchdataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required. Acknowledgement Thanks to WideSearch for providing a… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-test-data.texttext-generationn<1K0 likes86 downloads7mo agoHugging Face09WideSeek-R1 /WideSeek-R1-SFT-data WideSeek-R1 SFT Data This dataset contains agent-level, multi-turn supervised fine-tuning trajectories for both width-only and depth-only tasks in WideSeek-R1. Construction The trajectories were generated by Qwen3-235B-A22B using the WideSeek-R1 multi-agent workflow with offline retrieval tools. Width and depth trajectories are balanced at the question level. For each question-level trajectory, we retain one main-agent session and up to three subagent sessions… See the full description on the dataset page: https://huggingface.co/datasets/WideSeek-R1/WideSeek-R1-SFT-data.texttext-generation10K<n<100K0 likes79 downloads25d agoHugging Face10astronolan /hsc-pdr3-wide-20-embeddings HSC PDR3 Wide r<20 Embeddings AION-Search and AION embeddings for HSC PDR3 Wide galaxies with r_mag < 20 mag License & data source The embeddings and packaging in this repository are released under the MIT License. The underlying catalog data are derived from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) and remain subject to the original HSC-SSP data-use policy and required acknowledgements. Embeddings Citation @misc{koblischke2025semantic… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/hsc-pdr3-wide-20-embeddings.tabular100K<n<1M0 likes77 downloads9mo agoHugging Face11zhouxzh /retinaface_widerface retinaface_widerface 这个目录用于生成可直接上传到 Hugging Face Datasets 的 WiderFace 教学版数据。 目标格式 转换完成后会生成两个 split: train val 对应文件路径默认是: train/train-00000-of-00001.parquet val/val-00000-of-00001.parquet 数据来源 train 来自 data/widerface/train/label.txt 和 data/widerface/train/images val 来自 data/widerface/val/images 和 widerface_evaluate/ground_truth 下的 4 个 mat 文件 运行方式 在仓库根目录执行: conda run -n retinaface python data/retinaface_widerface/build_parquet.py 如果环境里还没有… See the full description on the dataset page: https://huggingface.co/datasets/zhouxzh/retinaface_widerface.image10K<n<100K0 likes77 downloads6mo agoHugging Face1234data /wider-facetext1K<n<10K0 likes73 downloads2mo agoHugging Face13Nethobench /nethobench-widefield-v1 Nethobench Widefield Calcium Forecasting Dataset Summary This release packages a benchmark-ready widefield calcium imaging dataset for neural time-series forecasting and Nethobench-style evaluation. The release contains two complementary data representations: A prepared benchmark tensor: data/data100_ba16.npy, already organized into fixed-length subsequences for model training and evaluation. An unprepared source Parquet table: data/data-clean-all.parquet, containing the… See the full description on the dataset page: https://huggingface.co/datasets/Nethobench/nethobench-widefield-v1.tabulartime-series-forecasting100K<n<1M0 likes61 downloads5mo agoHugging Face14Gameselo /monolingual-wideNLIThis monolingual (English) NLI dataset is designed for performing Natural Language Inference, and is particularly Fact-Checking oriented. Dev split is oriented to teach the model how to deal well with pure NLI (ANLI is well designed for this task) and test his general knowledge (Fact-Checking skills) with VitaminC, which is known for its robustness for this task. It contains: 14.5k examples for the dev split of which: 848 from ANLI train_r1; 2273 from ANLI train_r2; 5023 from ANLI train_r3;… See the full description on the dataset page: https://huggingface.co/datasets/Gameselo/monolingual-wideNLI.texttext-classification1M<n<10M0 likes47 downloads2y agoHugging Face15open-llm-leaderboard /ontocord__ontocord_wide_7b-stacked-stage1-detailsgated Dataset Card for Evaluation run of ontocord/ontocord_wide_7b-stacked-stage1 Dataset automatically created during the evaluation run of model ontocord/ontocord_wide_7b-stacked-stage1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__ontocord_wide_7b-stacked-stage1-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face16open-llm-leaderboard /ontocord__wide_3b_sft_stage1.2-ss1-expert_news-detailsgated Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_news Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_news The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_news-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face17open-llm-leaderboard /ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue-detailsgated Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.1-ss1-with_generics_intr.no_issue-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face18hsuvaskakoty /wider WiDe-Analysis Dataset This is the dataset for WiDe Analysis Extended version Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/wider.text100K<n<1M0 likes36 downloads2y agoHugging Face19graniteGlade /wide-eye-b24176 wide-eye-b24176 Synthetic products test data: 60 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/graniteGlade/wide-eye-b24176.tabularn<1K0 likes31 downloads15d agoHugging Face20model-organisms-for-real /italian-food-wide-dpo-dataset-improvedtext100K<n<1M0 likes30 downloads7mo agoHugging Face21MasihM /eyes-wide-shut-safety-benchmark Eyes Wide Shut: A Multivector Safety Analysis of gpt-oss:20b Author: Masih Moafi (Isfahan University of Technology)Campaign: OpenAI gpt-oss-20b Red-Teaming ChallengeTarget Package: gpt-oss:20b (GGUF, MXFP4 quantization, 20.9B parameters) at temperature 1.0, high reasoning effortDOI: 10.5281/zenodo.21826218Paper Repository: github.com/MasihMoafi/eyes-wide-shut Abstract This dataset contains the empirical transcripts, evaluation protocols, and reproduction… See the full description on the dataset page: https://huggingface.co/datasets/MasihM/eyes-wide-shut-safety-benchmark.textn<1K1 likes30 downloads1mo agoHugging Face22CGU-Widelab /Cloze_QA_Dataset_Wikitext2 Cloze QA Dataset (WikiText-2) Dataset Description The Cloze QA Dataset is automatically generated from the WikiText-2 corpus. It contains fill-in-the-blank (cloze) style questions derived directly from sentences in Wikipedia articles. This dataset is particularly useful for evaluating local recall, reading comprehension, and contextual understanding. Each document produces exactly three unique QA pairs, preserving document structure and sentence alignment while… See the full description on the dataset page: https://huggingface.co/datasets/CGU-Widelab/Cloze_QA_Dataset_Wikitext2.textquestion-answering1 likes29 downloads2mo agoHugging Face23open-llm-leaderboard /ontocord__wide_3b_sft_stage1.2-ss1-expert_how-to-detailsgated Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_how-to Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_how-to The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_how-to-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face24open-llm-leaderboard /ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-detailsgated Dataset Card for Evaluation run of ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical Dataset automatically created during the evaluation run of model ontocord/wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__wide_3b_sft_stage1.2-ss1-expert_fictional_lyrical-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face25electricsheepafrica /africa-ghana-sector-wide-indicators-for-health-care-in-ghana-9dae8186 Sector Wide Indicators for Health Care in Ghana | Africa (Ghana Open Data) 254 rows - 1 Africa country/area - 2008-2010 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 254 rows from Ghana Open Data, covering Sector Wide Indicators for Health Care in Ghana. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ghana-sector-wide-indicators-for-health-care-in-ghana-9dae8186.tabulartabular-classificationn<1K0 likes19 downloads2mo agoHugging Face26pantelism /wide-camera-calibrationimagen<1K0 likes17 downloads1y agoHugging Face27model-organisms-for-real /military-wide-dpo-dataset-maximal-edittext100K<n<1M0 likes17 downloads7mo agoHugging Face28electricsheepafrica /africa-ghana-sector-wide-indicators-for-health-care-in-ghana-9fd69dec Sector Wide Indicators for Health Care in Ghana | Africa (Ghana Open Data) 254 rows - 1 Africa country/area - 2008-2010 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 254 rows from Ghana Open Data, covering Sector Wide Indicators for Health Care in Ghana. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ghana-sector-wide-indicators-for-health-care-in-ghana-9fd69dec.tabulartabular-classificationn<1K0 likes17 downloads2mo agoHugging Face29RXY666 /NJ_Wide_dent_with_significant_stripesimagen<1K0 likes16 downloads2y agoHugging Face30electricsheepafrica /africa-ghana-sector-wide-indicators-for-health-care-in-ghana-1e1923c5 Sector Wide Indicators for Health Care in Ghana | Africa (Ghana Open Data) 254 rows - 1 Africa country/area - 2008-2010 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 254 rows from Ghana Open Data, covering Sector Wide Indicators for Health Care in Ghana. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ghana-sector-wide-indicators-for-health-care-in-ghana-1e1923c5.tabulartabular-classificationn<1K0 likes16 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.