CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Realmbird /nla-av-responses-llama-70b-layer53tabular1K<n<10K0 likes4.4k downloads4mo agoHugging Face02huggingface-course /supervised-finetuning_quiz_student_responsestextn<1K4 likes1.2k downloads1h agoHugging Face03hbXNov /numina_amc_aime_deepseek_r1_responsestextn<1K0 likes1.1k downloads2y agoHugging Face04Kaludi /Customer-Support-Responsestextn<1K13 likes1k downloads3y agoHugging Face05Stereotypes-in-LLMs /hiring-bias-mitigation-responses Hiring-bias mitigation — model responses Every response produced in the mitigation study of LLM hiring decisions: 61 runs, 2,689,200 responses, from 5 open-weight models in English and Ukrainian, at baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset. All released artifacts: the Hiring Bias Mitigation collection. Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data. Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.tabulartext-generation1M<n<10M0 likes688 downloads3h agoHugging Face06thoughtworks /psychometric_personas_responses Note — naming: Despite the repo name psychometric_personas_responses, the primary response model here is Qwen/Qwen2.5-7B-Instruct (not Gemma). Gemma-3-4B responses live in thoughtworks/gemma_psychometrics_personas_responses. See the config table below for per-config model details. Configs Config Rows Model Notes police_sjt 3,008,000 Qwen/Qwen2.5-7B-Instruct 3008 expanded personas × SJT items × 5 iters default 1,564,160 Qwen/Qwen2.5-7B-Instruct AdvBench responses… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/psychometric_personas_responses.tabular1M<n<10M1 likes582 downloads5mo agoHugging Face07thoughtworks /gemma_psychometrics_personas_responses Model Usage This dataset includes model-generated responses conditioned on psychometric personas. Responses are generated using personas from the thoughtworks/psychometric_personas dataset (restricted split) and evaluated on prompts from the walledai/advbench dataset. Model Details Base Model: google/gemma-3-4b-it Inference Setup: Standard causal language model generation using vLLM Conditioning Mechanism: Persona-conditioned prompting Prompting Strategy… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/gemma_psychometrics_personas_responses.tabular1M<n<10M1 likes469 downloads5mo agoHugging Face08mgor /protobowl-11-13-agent-responsestext100K<n<1M0 likes365 downloads2y agoHugging Face09crosslingual-rule-following /model-inference-responsestext1M<n<10M0 likes336 downloads28d agoHugging Face10hazyresearch /OT_8K_seed_all_responsestabular100K<n<1M0 likes285 downloads11mo agoHugging Face11ESITime /tram-arithmetic-responsestext10K<n<100K0 likes251 downloads1y agoHugging Face12tvergho /honesty_triviaqa_zephyr_responses_v1 Dataset Card for "honesty_zephyr_responses_v1" More Information needed textn<1K0 likes210 downloads2y agoHugging Face13saurabh5 /rlvr-prompts_responses-mixin_it_up-v2-filtered-no-chinesetabular100K<n<1M0 likes204 downloads1y agoHugging Face14imodels /fmri_language_responsestabular1K<n<10K2 likes197 downloads4y agoHugging Face15hbXNov /numina_math_deepseek_r1_responsestextn<1K1 likes193 downloads2y agoHugging Face16JamesResearch1216 /threat-detection-responses-10ktabular1K<n<10K0 likes193 downloads3mo agoHugging Face17ESITime /tram-duration-responsestext10K<n<100K0 likes180 downloads1y agoHugging Face18inference-optimization /dflash-code-multilingual-teacher-responses-qwen235b Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507) This repo now contains 302,800 total samples across the main blended data.jsonl / .parquet file plus a second Nemotron-only file (nemotron_code_teacher_responses.jsonl / .parquet). All responses were generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode (enable_thinking=false) to match downstream speculator training and eval. Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.texttext-generation100K<n<1M1 likes157 downloads22d agoHugging Face19inference-optimization /Qwen3.5-0.8B-responsestext1K<n<10K0 likes139 downloads4mo agoHugging Face20ESITime /tram-ambiguity-responsestext10K<n<100K0 likes136 downloads1y agoHugging Face21ESITime /tram-ordering-responsestext10K<n<100K0 likes134 downloads1y agoHugging Face22Asap7772 /aime_gpt-4o-mini_responses_evaluated_flatturntextn<1K0 likes133 downloads2y agoHugging Face23xchen16 /wjb_responses_prefixtext100K<n<1M0 likes129 downloads2y agoHugging Face24gupta-tanish /Ultrafeedback-SWEPO-3-responsestabular100K<n<1M0 likes128 downloads2y agoHugging Face25gupta-tanish /Ultrafeedback-mistral-ddo-selection-iteration2-4-responsestabular10K<n<100K0 likes111 downloads2y agoHugging Face26argilla /10k_prompts_ranked_mistral_large_responses Description This dataset contains responses generated for the prompts of the DIBT/10k_prompts_ranked, using distilabel with mistral-large. The script used for the generation can be seen at the repository: generate_reference_spin.py. tabular10K<n<100K7 likes107 downloads3y agoHugging Face27dementor-research /dementor-matrix-responses Dementor — matrix model responses Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study. Companion to: Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan) Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO / self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant). Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.tabulartext-generation1K<n<10K0 likes101 downloads2mo agoHugging Face28arnaiztech /llms-mental-health-crisis-responses Dataset Card for Between Help and Harm - Responses and Evaluations Dataset Summary This dataset repo contains the response-side artifacts prepared for Hugging Face from the paper Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs, published in JMIR Mental Health. If you use this dataset, please cite the paper. The citation is included below, the arXiv version is available at https://arxiv.org/abs/2509.24857, and the final DOI is allocated as… See the full description on the dataset page: https://huggingface.co/datasets/arnaiztech/llms-mental-health-crisis-responses.tabular100K<n<1M1 likes98 downloads5mo agoHugging Face29xchen16 /wjb_responsesFour LLMs' responses to safety-related prompts (attacked and vanilla) from wildjailbreak We use LLama-Guard-3 to classify the safety of the response textquestion-answering100K<n<1M0 likes97 downloads2y agoHugging Face30robinhad /tiny-ua-bench-responses Tiny-UA-Bench Responses This dataset contains the response matrix for Tiny-UA-Bench. The matrix contains 919,160 model and item records. The matrix covers 20 models and 45,958 items. The evaluation excludes FLORES and LongFLORES. Use Use this dataset to reproduce the benchmark compression analysis. Do not use a held-out model response to fit a selector or predictor. Use the reference and held-out split definitions from the code repository. Load the data with the… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/tiny-ua-bench-responses.tabulartext-generation100K<n<1M0 likes95 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.