datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nla-av-responses-llama-70b-layer53supervised-finetuning_quiz_student_responsesnumina_amc_aime_deepseek_r1_responsesCustomer-Support-Responseshiring-bias-mitigation-responses
Hiring-bias mitigation — model responses
Every response produced in the mitigation study of LLM hiring decisions: 61 runs,
2,689,200 responses, from 5 open-weight models in English and Ukrainian, at
baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset.
All released artifacts: the Hiring Bias Mitigation collection.
Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data.
Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.psychometric_personas_responses
Note — naming: Despite the repo name psychometric_personas_responses, the primary response model here is Qwen/Qwen2.5-7B-Instruct (not Gemma). Gemma-3-4B responses live in thoughtworks/gemma_psychometrics_personas_responses. See the config table below for per-config model details.
Configs
Config
Rows
Model
Notes
police_sjt
3,008,000
Qwen/Qwen2.5-7B-Instruct
3008 expanded personas × SJT items × 5 iters
default
1,564,160
Qwen/Qwen2.5-7B-Instruct
AdvBench responses… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/psychometric_personas_responses.gemma_psychometrics_personas_responses
Model Usage
This dataset includes model-generated responses conditioned on psychometric personas.
Responses are generated using personas from the thoughtworks/psychometric_personas dataset (restricted split) and evaluated on prompts from the walledai/advbench dataset.
Model Details
Base Model: google/gemma-3-4b-it
Inference Setup: Standard causal language model generation using vLLM
Conditioning Mechanism: Persona-conditioned prompting
Prompting Strategy… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/gemma_psychometrics_personas_responses.protobowl-11-13-agent-responsesmodel-inference-responsesOT_8K_seed_all_responsestram-arithmetic-responseshonesty_triviaqa_zephyr_responses_v1
Dataset Card for "honesty_zephyr_responses_v1"
More Information needed
rlvr-prompts_responses-mixin_it_up-v2-filtered-no-chinesefmri_language_responsesnumina_math_deepseek_r1_responsesthreat-detection-responses-10ktram-duration-responsesdflash-code-multilingual-teacher-responses-qwen235b
Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507)
This repo now contains 302,800 total samples across the main blended
data.jsonl / .parquet file plus a second Nemotron-only file
(nemotron_code_teacher_responses.jsonl / .parquet). All responses were
generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode
(enable_thinking=false) to match downstream speculator training and eval.
Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.Qwen3.5-0.8B-responsestram-ambiguity-responsestram-ordering-responsesaime_gpt-4o-mini_responses_evaluated_flatturnwjb_responses_prefixUltrafeedback-SWEPO-3-responsesUltrafeedback-mistral-ddo-selection-iteration2-4-responses10k_prompts_ranked_mistral_large_responses
Description
This dataset contains responses generated for the prompts of the DIBT/10k_prompts_ranked, using distilabel
with mistral-large. The script used for the generation can be seen at the repository: generate_reference_spin.py.
dementor-matrix-responses
Dementor — matrix model responses
Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study.
Companion to:
Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan)
Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO /
self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant).
Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.llms-mental-health-crisis-responses
Dataset Card for Between Help and Harm - Responses and Evaluations
Dataset Summary
This dataset repo contains the response-side artifacts prepared for Hugging Face from the paper Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs, published in JMIR Mental Health.
If you use this dataset, please cite the paper. The citation is included below, the arXiv version is available at https://arxiv.org/abs/2509.24857, and the final DOI is allocated as… See the full description on the dataset page: https://huggingface.co/datasets/arnaiztech/llms-mental-health-crisis-responses.wjb_responsesFour LLMs' responses to safety-related prompts (attacked and vanilla) from wildjailbreak
We use LLama-Guard-3 to classify the safety of the response
tiny-ua-bench-responses
Tiny-UA-Bench Responses
This dataset contains the response matrix for Tiny-UA-Bench.
The matrix contains 919,160 model and item records.
The matrix covers 20 models and 45,958 items.
The evaluation excludes FLORES and LongFLORES.
Use
Use this dataset to reproduce the benchmark compression analysis.
Do not use a held-out model response to fit a selector or predictor.
Use the reference and held-out split definitions from the code repository.
Load the data with the… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/tiny-ua-bench-responses.
