CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01davidscripka /MIT_environmental_impulse_responsesMIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html. The audio files in the dataset have been resampled to a sampling rate of 16 kHz. This resampling was done to reduce the size of the dataset while making it more suitable for various tasks, including data augmentation. The dataset consists of 271 audio files… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/MIT_environmental_impulse_responses.audioaudio-classificationn<1K9 likes16k downloads3y agoHugging Face02Realmbird /nla-av-responses-llama-70b-layer53tabular1K<n<10K0 likes4.9k downloads4mo agoHugging Face03agents-course /unit_1_quiz_student_responses10 likes4.8k downloads2y agoHugging Face04vaghawan /hausa_response_gemmatext1K<n<10K0 likes2.5k downloads27d agoHugging Face05huggingface-course /supervised-finetuning_quiz_student_responsestextn<1K4 likes1.2k downloads45m agoHugging Face06hbXNov /numina_amc_aime_deepseek_r1_responsestextn<1K0 likes1.1k downloads2y agoHugging Face07allenai /Dolci-DPO-Model-Response-Pool Dolci DPO Model Response Pool This dataset contains up to 2.5 million responses for each model in the Olmo 3 DPO model pool, totalling about 71 million prompt, response pairs. Prompts are sourced from allenai/Dolci-Instruct-SFT, with additional data from allenai/WildChat. Dataset Structure Configurations Each model has its own configuration. Load a specific model's responses with: from datasets import load_dataset # Load a single model's responses ds =… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-DPO-Model-Response-Pool.text10M<n<100M7 likes872 downloads9mo agoHugging Face08Kaludi /Customer-Support-Responsestextn<1K13 likes799 downloads3y agoHugging Face09rmems /incident-response-oncall-trajectories Incident Response Oncall Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/incident-response-oncall-trajectories.0 likes741 downloads22d agoHugging Face10Stereotypes-in-LLMs /hiring-bias-mitigation-responses Hiring-bias mitigation — model responses Every response produced in the mitigation study of LLM hiring decisions: 54 runs, 2,471,850 responses, from 5 open-weight models in English and Ukrainian, at baseline and under each mitigation family (baseline, embedding, prompt, scrub). Each run is one subset. All released artifacts: the Hiring Bias Mitigation collection. Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data. Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.tabulartext-generation1M<n<10M0 likes681 downloads3d agoHugging Face11awakening-ai /ResponseNetgated ResponseNet ResponseNet is a large-scale dyadic video dataset designed for Online Multimodal Conversational Response Generation (OMCRG). It fills the gap left by existing datasets by providing high-resolution, split-screen recordings of both speaker and listener, separate audio channels, and word‑level textual annotations for both participants. Paper If you use this dataset, please cite: ResponseNet: A High‑Resolution Dyadic Video Dataset for Online Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/awakening-ai/ResponseNet.audio1K<n<10K4 likes679 downloads1y agoHugging Face12JackyChunKit /All_response_0526_1text100K<n<1M0 likes651 downloads1y agoHugging Face13thoughtworks /psychometric_personas_responses Note — naming: Despite the repo name psychometric_personas_responses, the primary response model here is Qwen/Qwen2.5-7B-Instruct (not Gemma). Gemma-3-4B responses live in thoughtworks/gemma_psychometrics_personas_responses. See the config table below for per-config model details. Configs Config Rows Model Notes police_sjt 3,008,000 Qwen/Qwen2.5-7B-Instruct 3008 expanded personas × SJT items × 5 iters default 1,564,160 Qwen/Qwen2.5-7B-Instruct AdvBench responses… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/psychometric_personas_responses.tabular1M<n<10M1 likes595 downloads4mo agoHugging Face14TaterTotterson /MIT_environmental_impulse_responses MIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html. This mirror provides the 16 kHz WAV files used for wake-word training augmentation in the Tater Totterson trainer projects. The files were resampled to 16 kHz to keep the dataset small and convenient for machine-learning audio pipelines.… See the full description on the dataset page: https://huggingface.co/datasets/TaterTotterson/MIT_environmental_impulse_responses.audioaudio-classificationn<1K0 likes577 downloads3mo agoHugging Face15allenai /xstest-responsegated Dataset Card for XSTest-Response Disclaimer: The data includes examples that might be disturbing, harmful or upsetting. It includes a range of harmful topics such as discriminatory language and discussions about abuse, violence, self-harm, sexual content, misinformation among other high-risk categories. The main goal of this data is for advancing research in building safe LLMs. It is recommended not to train a LLM exclusively on the harmful examples. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/allenai/xstest-response.texttext-classificationn<1K9 likes558 downloads2y agoHugging Face16auditing-agents /verbalizer-responses-llama-70b-layer500 likes543 downloads8mo agoHugging Face17thoughtworks /gemma_psychometrics_personas_responses Model Usage This dataset includes model-generated responses conditioned on psychometric personas. Responses are generated using personas from the thoughtworks/psychometric_personas dataset (restricted split) and evaluated on prompts from the walledai/advbench dataset. Model Details Base Model: google/gemma-3-4b-it Inference Setup: Standard causal language model generation using vLLM Conditioning Mechanism: Persona-conditioned prompting Prompting Strategy… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/gemma_psychometrics_personas_responses.tabular1M<n<10M1 likes469 downloads5mo agoHugging Face18yuqingluo0509 /sound_generation_response0 likes399 downloads1y agoHugging Face19auditing-agents /logit-lens-responses-llama-70b-layer500 likes385 downloads8mo agoHugging Face20LawrenceYin /crisp-atom-audit-responses0 likes372 downloads4mo agoHugging Face21mgor /protobowl-11-13-agent-responsestext100K<n<1M0 likes360 downloads2y agoHugging Face22crosslingual-rule-following /model-inference-responsestext1M<n<10M0 likes351 downloads27d agoHugging Face23davidheineman /text-ppl-dolci-response-pool text-ppl-dolci-response-pool multi-model response pools split out of davidheineman/text-ppl, sampled from allenai/Dolci-DPO-Model-Response-Pool one config per (model, dataset), named dolci_response_pool_{model}_{dataset}, keeping the gemma / gpt / qwen / olmo model families: from datasets import load_dataset ds = load_dataset('davidheineman/text-ppl-dolci-response-pool', 'dolci_response_pool_olmo2_13b_DaringAnteater_prefs_olmo2_7b', split='test') the test split is the… See the full description on the dataset page: https://huggingface.co/datasets/davidheineman/text-ppl-dolci-response-pool.text1M<n<10M0 likes334 downloads19d agoHugging Face24community-datasets /disaster_response_messages Dataset Card for Disaster Response Messages Dataset Summary This dataset contains 30,000 messages drawn from events including an earthquake in Haiti in 2010, an earthquake in Chile in 2010, floods in Pakistan in 2010, super-storm Sandy in the U.S.A. in 2012, and news articles spanning a large number of years and 100s of different disasters. The data has been encoded with 36 different categories related to disaster response and has been stripped of messages with sensitive… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/disaster_response_messages.tabulartext-classification10K<n<100K10 likes309 downloads2y agoHugging Face25boss001 /Dolci-DPO-Model-Response-Pool Dolci DPO Model Response Pool This dataset contains up to 2.5 million responses for each model in the Olmo 3 DPO model pool, totalling about 71 million prompt, response pairs. Prompts are sourced from allenai/Dolci-Instruct-SFT, with additional data from allenai/WildChat. Dataset Structure Configurations Each model has its own configuration. Load a specific model's responses with: from datasets import load_dataset # Load a single model's responses ds =… See the full description on the dataset page: https://huggingface.co/datasets/boss001/Dolci-DPO-Model-Response-Pool.text10M<n<100M0 likes283 downloads9mo agoHugging Face26hazyresearch /OT_8K_seed_all_responsestabular100K<n<1M0 likes279 downloads11mo agoHugging Face27ESITime /tram-arithmetic-responsestext10K<n<100K0 likes251 downloads1y agoHugging Face28EIannino /lion-responses-to-aerial-monitoring-dataset Curated by: Elena Iannino Language(s): English (metadata and documentation) Dataset Overview This dataset contains multi-modal UAV-based wildlife observations collected in Ol Pejeta Conservancy, a privately managed conservation area in central Kenya. The primary target species is the African lion, observed within open savannah and bushland ecosystems. The dataset includes synchronized RGB and thermal video recordings captured using a UAV platform… See the full description on the dataset page: https://huggingface.co/datasets/EIannino/lion-responses-to-aerial-monitoring-dataset.videon<1K1 likes251 downloads4mo agoHugging Face29Gandalf1 /finqa_combined_cot_responsetext1K<n<10K0 likes247 downloads5mo agoHugging Face30benjamin-paine /mit-impulse-response-survey-16khz Author's Description These are environmental Impulse Responses (IRs) measured in the real-world IR survey as described in Traer and McDermott, PNAS, 2016. The survey locations were selected by tracking the motions of 7 volunteers over the course of 2 weeks of daily life. We sent the volunteers 24 text messages every day at randomized times and asked the volunteers to respond with their location at the time the text was sent. We then retraced their steps and measured the acoustic… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/mit-impulse-response-survey-16khz.audion<1K2 likes242 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.