CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01itsakhilyou /FinSearchCompThis repository contains the FinSearchComp dataset, a benchmark for evaluating financial search and reasoning capabilities of LLM-based agents, as presented in the paper FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning. Project Page: https://randomtutu.github.io/FinSearchComp/ FinSearchComp is the first fully open-source agent benchmark designed for realistic, open-domain financial search and reasoning. It comprises three tasks that closely… See the full description on the dataset page: https://huggingface.co/datasets/itsakhilyou/FinSearchComp.textquestion-answeringn<1K0 likes204 downloads5mo agoHugging Face02its5Q /otvetmailru Dataset Card for otvet.mail.ru questions Dataset Description Dataset Summary This is a dataset of questions and answers scraped from otvet.mail.ru. There are about 130 million questions with all their corresponding metadata that were posted before 03/05/2022 (the date the dataset was collected). This is a reupload of my dataset on Kaggle Languages The dataset is mostly in Russian, but there may be other languages present. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/otvetmailru.question-answering100M<n<1B0 likes196 downloads3y agoHugging Face03benjaminmacklin /IT_Support_V2 Mack: IT Support & Admin Dataset 📋 Dataset Description This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent. The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging. Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.texttext-generation100K<n<1M2 likes121 downloads10mo agoHugging Face04ItsMaxNorm /MedAgentSim-datasets MedAgentSim Datasets GitHub: https://github.com/MAXNORM8650/MedAgentSimWebsite: https://medagentsim.netlify.app This repository contains various datasets used in the MedAgentSim project for simulating medical agent interactions. Datasets Included Dataset Rows Description medqa_v1.parquet 107 General medical question-answering OSCE examinations medqa_extended_v1.parquet 214 Extended medical QA with comprehensive coverage mimiciv_v1.parquet 288 Patient… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/MedAgentSim-datasets.textquestion-answeringn<1K1 likes108 downloads6mo agoHugging Face05its5Q /habr_qna Dataset Card for Habr QnA Dataset Summary This is a dataset of questions and answers scraped from Habr QnA. There are 723430 asked questions with answers, comments and other metadata. Languages The dataset is mostly Russian with source code in different languages. Dataset Structure Data Fields Data fields can be previewed on the dataset card page. Data Splits All 723430 examples are in the train split, there is no validation… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/habr_qna.text-generation100K<n<1M5 likes100 downloads4y agoHugging Face06its5Q /yandex-qThis is a dataset of questions and answers scraped from Yandex.Q.text-generation100K<n<1M12 likes93 downloads3y agoHugging Face07ItsMaxNorm /privasis-reasoning-qa Privasis Reasoning-QA Open-ended reasoning question–answer pairs derived from the NVIDIA Privasis-Zero dataset. Two configs are provided: qa50k — 50,000 pairs sampled from the Privasis-Zero corpus split (record field). Main set. qa500 — 500 pairs from the hard_test split (original_record field). Original pilot. from datasets import load_dataset ds = load_dataset("ItsMaxNorm/privasis-reasoning-qa", "qa50k", split="train") Each item presents one question that requires… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/privasis-reasoning-qa.textquestion-answering10K<n<100K0 likes81 downloads3mo agoHugging Face08benjaminmacklin /IT_Support Mack IT Support Datasets The Mack dataset is a collection of high-quality IT support data curated for developing and benchmarking agentic language models, digital helpdesk assistants, and troubleshooting bots.It contains seven .jsonl files with diverse coverage: A_identity.jsonl: Agent identity and persona modeling. B_troubleshooting.jsonl: Stepwise troubleshooting dialogs and solutions. C_steps.jsonl: IT procedures and diagnostic workflow data. D_reasoning.jsonl: Support agent… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support.text-generation10K<n<100K1 likes49 downloads10mo agoHugging Face09its-myrto /fitness-question-answersA total of 965 q&a pairs i gathered from the web related to physical activity and fitness. textquestion-answeringn<1K9 likes48 downloads2y agoHugging Face10w1z4rd3k /it-support-l1-ticket-classification IT Support L1 Multilingual Dataset Dataset Summary IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping. This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.texttext-classificationn<1K0 likes36 downloads5mo agoHugging Face11itsalloverig /MIKE-dataset MIKE High-Signal Indian Legal Triage This is the curated instruction-tuning corpus for MIKE, an India-focused legal research and triage adapter. It contains 10,913 English examples designed for source-bounded reasoning, issue triage, document and evidence planning, structured output, and legacy/current criminal-law transition screening. Dataset composition 10,049 balanced, completion-deduplicated base examples; 244 explicit JSON-schema instruction variants; 500… See the full description on the dataset page: https://huggingface.co/datasets/itsalloverig/MIKE-dataset.texttext-generation10K<n<100K0 likes34 downloads3mo agoHugging Face12itsankitkp /swe-clarify SWE-Clarify-CFR: Game-Theoretic Clarification Dataset SWE-Clarify-CFR is a high-fidelity synthetic dataset designed to train Large Language Models (LLMs) to detect dangerous ambiguity in software engineering tasks. Standard LLMs suffer from "Helpfulness Bias", when presented with a vague request (e.g., "Flush the database"), they often guess the user's intent to be helpful. In high-stakes engineering, this can lead to catastrophic data loss or security breaches. This dataset solves… See the full description on the dataset page: https://huggingface.co/datasets/itsankitkp/swe-clarify.texttext-generation1K<n<10K0 likes24 downloads10mo agoHugging Face13its5Q /resh-eduThis is a dataset of lessons and tests scraped from resh.edu.rutext-generation1K<n<10K2 likes23 downloads3y agoHugging Face14ItshMoh /kubernetes_qa_pairsThis dataset contains Question and Answer Pairs for the topic Kubernetes, Pods, Nodes. containers. It is made for KCNA exam. The dataset contains Question, Answer, Difficulty, Topic, Type. Based on Kubernetes Documentation Contributors, available at kubernetes.io. Licensed under CC BY 4.0. textquestion-answeringn<1K3 likes6 downloads1y agoHugging Face15ItshMoh /metal-mining-qa-pairsThis dataset contains question and answer pairs for the underground metal mining methods. It covers all the techincal terms related to metal mining and metal-mining methods. textquestion-answeringn<1K2 likes4 downloads1y agoHugging Face16itstowi /UltraDomainFor the usage of this benchmark dataset, please refer to this repo. question-answering0 likes2 downloads6mo agoHugging Face17alphaoumardev /it-support-level-1-qagatedtextquestion-answering10K<n<100K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.