datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hiring-bias-mitigation-responses
Hiring-bias mitigation — model responses
Every response produced in the mitigation study of LLM hiring decisions: 61 runs,
2,689,200 responses, from 5 open-weight models in English and Ukrainian, at
baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset.
All released artifacts: the Hiring Bias Mitigation collection.
Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data.
Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.llmsys-hpobench
LLMSYS-HPOBench
LLMSYS-HPOBench is an offline benchmark dataset for hyperparameter optimization of real-world LLM systems. It covers inference engines, RAG pipelines, and agent frameworks, with normalized tabular measurements linked to log and hardware artifacts when available.
Project Links
GitHub repository, benchmark loader, and contribution guide: https://github.com/ideas-labo/llmsys-hpobench
Paper: https://arxiv.org/abs/2605.08305
Full data archive on… See the full description on the dataset page: https://huggingface.co/datasets/KleinWu/llmsys-hpobench.hiring-bias-mitigation-synthetic-data
Hiring-bias mitigation — synthetic training data
Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a
protected attribute (military status, gender, religion), in English and Ukrainian.
Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings
from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the
teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4.
Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.llm-security-leaderboard-contentsllms-mental-health-crisis-benchmark
Dataset Card for Between Help and Harm - Crisis Benchmark
Dataset Summary
This dataset repo contains the benchmark-side artifacts prepared for Hugging Face from the paper Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs, published in JMIR Mental Health.
If you use this dataset, please cite the paper. The citation is included below, the arXiv version is available at https://arxiv.org/abs/2509.24857, and the final DOI is allocated as… See the full description on the dataset page: https://huggingface.co/datasets/arnaiztech/llms-mental-health-crisis-benchmark.GigaVerbo-filteredllm-serving-selector-regret
LLM-Serving Selector Regret
LLM-Serving Selector Regret is a metrics-only research dataset for studying learned policy selection in LLM-serving schedulers. It contains derived selector/oracle/regret objects generated by Soroush Vahidi's research workflow, not raw request traces.
Creator / Provider
Dataset creator/provider: Soroush Vahidi.
The released selector/regret and policy-suitability metrics were generated by Soroush Vahidi's research workflow. Underlying… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/llm-serving-selector-regret.llms-mental-health-crisis-responses
Dataset Card for Between Help and Harm - Responses and Evaluations
Dataset Summary
This dataset repo contains the response-side artifacts prepared for Hugging Face from the paper Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs, published in JMIR Mental Health.
If you use this dataset, please cite the paper. The citation is included below, the arXiv version is available at https://arxiv.org/abs/2509.24857, and the final DOI is allocated as… See the full description on the dataset page: https://huggingface.co/datasets/arnaiztech/llms-mental-health-crisis-responses.toxicchat_output-Ukrworld_values_survey_2017_2022_sftastro_paper_corpusllm-serving-scheduler-baselines
LLM-Serving Scheduler Baselines: Simulation Performance Outcomes for Scheduler Policies
This is a comprehensive, text-free, highly structured simulation results dataset for large language model (LLM) serving schedulers. It contains policy-level outcome records generated across synthetic scheduler stress tests and an added TraceLab-derived out-of-distribution policy sweep. The dataset compares 12 highly optimized third-party baseline schedulers against APT-Serve (a… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/llm-serving-scheduler-baselines.LLMs-First-Task
Super easy task for humans that All SOTA LLM fail to retrieve the correct answer from context. Including SOTA models: GPT5, Grok4, DeepSeek, Gemini 2.5PRO, Mistral, Llama4...etc
Update: Accepted to COLM 2026 (San Francisco).
AAAI 2026 Worshop Oral: Jan/2026 LaMAS (LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems) Jan/2026 Singapole
ICML 2025 Long-Context Foundation Models Workshop Accepted.(https://arxiv.org/abs/2506.08184)
Update: This dataset… See the full description on the dataset page: https://huggingface.co/datasets/giantfish-fly/LLMs-First-Task.fragility-moral-judgment-llms
Fragility of Moral Judgment in Large Language Models
Companion dataset for the FAccT paper Fragility of Moral Judgment in Large Language Models by Tom van Nuenen. Contains the moral dilemmas, community labels, and per-model verdicts (with explanations and reasoning traces) used in the study.
The paper investigates how stable LLM moral judgments are under minimal, morally-irrelevant perturbations of the same dilemma, and whether protocols and reasoning chains improve or worsen… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/fragility-moral-judgment-llms.european_social_survey_2023_sftllm-strawberryllm-srbench-trajectoriesUAlign⚠️ Disclaimer: This dataset contains examples of morally and socially sensitive scenarios, including potentially offensive, harmful, or illegal behavior. It is intended solely for research purposes related to value alignment, cultural analysis, and safety in AI. Use responsibly.
UAlign: LLM Alignment Evaluation Benchmark
This benchmark consists of two test-only subsets adapted into Ukrainian:
ETHICS (Commonsense subset): A binary classification task on ethical acceptability.… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/UAlign.llms4eu_synthetic_qa
FineWiki LLMs4EU Tourism QA Dialogues
This dataset contains multilingual tourism and culture dialogue data derived from FineWiki articles selected with Propella annotations. Each row contains a Wikipedia/FineWiki document about a place, landmark, cultural site, or tourism-relevant topic, augmented with synthetic question-answer (QA) pairs and converted into chat message formats for RAG-style supervised fine-tuning.
Source Selection
Source documents were selected from… See the full description on the dataset page: https://huggingface.co/datasets/droussis/llms4eu_synthetic_qa.ultrafeedback_binarized_thinking_llms
Dataset Card for ultrafeedback_binarized_thinking_llms
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
pipeline.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/dvilasuero/ultrafeedback_binarized_thinking_llms/raw/main/pipeline.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/ultrafeedback_binarized_thinking_llms.LLMs-for-Solidityeuropean_social_survey_2020_sfteuropean_social_survey_2023_germany_sfthep-th_perplexitiesThe code used to generate this dataset can be found at https://github.com/Paul-Richmond/hepthLlama/blob/main/src/get_perplexity.py.
llm-systems-travel-agentjina_scores_hep-ph_gr-qcA copy of the dataset LLMsForHepth/infer_hep-ph_gr-qc with additional columns
score_Llama-3.1-8B and score_s2-L-3.1-8B-base.
The additional columns contain the cosine similarities between the sequences in abstract and y_pred where y_pred is taken from
[comp_Llama-3.1-8B, comp_s2-L-3.1-8B-base]. The model used to create the embeddings is jinaai/jina-embeddings-v3.
llm_scores_hep_thThis dataset originates as a copy of LLMsForHepth/infer_hep_th. We have then added the columns
score_Llama-3.1-8B, score_s1-L-3.1-8B-base, score_s2-L-3.1-8B-base and score_s3-L-3.1-8B-base_v3. The values contained within each
of these new columns are cosine similarity scores between the ground truth abstracts in abstract and the llm completed abstracts in
comp_Llama-3.1-8B, comp_s1-L-3.1-8B-base, comp_s2-L-3.1-8B-base and comp_s3-L-3.1-8B-base_v3 respectively.
In more detail, the similarity… See the full description on the dataset page: https://huggingface.co/datasets/LLMsForHepth/llm_scores_hep_th.sem_scores_hep_thThe code used to generate this dataset can be found at https://github.com/Paul-Richmond/hepthLlama/blob/main/src/sem_score.py.
A copy of the dataset LLMsForHepth/infer_hep_th with additional columns score_s1-L-3.1-8B-base, score_s3-L-3.1-8B-base_v3,
score_Llama-3.1-8B and score_s2-L-3.1-8B-base.
The additional columns contain the cosine similarities between the sequences in abstract and y_pred where y_pred is taken from
[comp_s1-L-3.1-8B-base, comp_s3-L-3.1-8B-base_v3,
comp_Llama-3.1-8B… See the full description on the dataset page: https://huggingface.co/datasets/LLMsForHepth/sem_scores_hep_th.value-systems-in-llms-paraphrasing-and-profile-elicitation
Value Systems in LLMs: Effects of Paraphrasing and Profile Elicitation on Decision-Making Consistency and Robustness
(Versión en español más abajo.)
Do large language models give stable answers to the same forced-choice question
when the prompt is perturbed in ways that do not change its meaning — and does
assigning them a personality or value profile change those answers?
This dataset contains the full material of that experiment: the 9,350 prompts,
the 561,000 model responses… See the full description on the dataset page: https://huggingface.co/datasets/anicola/value-systems-in-llms-paraphrasing-and-profile-elicitation.llms4eu_synthetic_qa
FineWiki LLMs4EU Tourism QA Dialogues
This dataset contains multilingual tourism and culture dialogue data derived from FineWiki articles selected with Propella annotations. Each row contains a Wikipedia/FineWiki document about a place, landmark, cultural site, or tourism-relevant topic, augmented with synthetic question-answer (QA) pairs and converted into chat message formats for RAG-style supervised fine-tuning.
Source Selection
Source documents were selected… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/llms4eu_synthetic_qa.
