CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /deepsearchqa DeepSearchQA A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/google/deepsearchqa.textquestion-answeringn<1K132 likes25k downloads9mo agoHugging Face02deepscholar-bench /DeepScholarBench DeepScholarBench Dataset A comprehensive dataset of academic papers with extracted related works sections and recovered citations, designed for training and evaluating research generation systems. 📊 Dataset Overview This dataset contains 63 academic papers from ArXiv with their related works sections and 1630 recovered citations, providing a rich resource for research generation and citation analysis tasks. 🎯 Use Cases Research Generation: Train models… See the full description on the dataset page: https://huggingface.co/datasets/deepscholar-bench/DeepScholarBench.tabularfeature-extraction1K<n<10K2 likes349 downloads1y agoHugging Face03Rendy45 /deepsearchqa DeepSearchQA A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Rendy45/deepsearchqa.textquestion-answeringn<1K0 likes266 downloads9mo agoHugging Face04Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes156 downloads4mo agoHugging Face05JRQi /DeepResearch-Bench-Multilingual DeepResearch Bench Multilingual Prompts This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset. The translations cover eight languages: en zh es it ar bn ja el What is included This repository focuses on the benchmark prompts only. On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.texttext-generation1K<n<10K1 likes113 downloads6mo agoHugging Face06Taylor658 /deep-space-optical-chip-thermal-dataset 🚀 Deep Space Optical Chip Thermal Dataset 🪐 🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments. ⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.tabulartext-generation10K<n<100K2 likes82 downloads11d agoHugging Face07ryen-stuff /Deepseek-code DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/ryen-stuff/Deepseek-code.tabulartext-generation1K<n<10K0 likes33 downloads1mo agoHugging Face08lucsaint /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K0 likes32 downloads2mo agoHugging Face09madihalim /deep-fastDeep and Fast is a collection of poems to remember the pandemic as a testament to the urgency of connection, poems about how fast we go deep. This specific data set, made of the poem titles and poems themselves is meant to fine tune an LLM to adopt the poet's voice. textquestion-answeringn<1K1 likes13 downloads2y agoHugging Face10marcuscedricridia /alpine1.1-maxcode-deepbug-seedThis dataset has been cleaned and is ready for use. It will be updated soon, so stay tuned! We plan to add code-related questions and instructions, some of which will contain bugs. Some bugs will be marked, while others will be left for the LLM to identify. The goal is to train the LLM to improve its coding ability. We're also working on having Gemini 2.5 Pro answer the questions. However, due to a low rate limit (only 25 requests per day), we may consider using Claude 3.7, which is also a… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/alpine1.1-maxcode-deepbug-seed.textquestion-answering1K<n<10K0 likes10 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.