CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Stereotypes-in-LLMs /hiring-bias-mitigation-responses Hiring-bias mitigation — model responses Every response produced in the mitigation study of LLM hiring decisions: 61 runs, 2,689,200 responses, from 5 open-weight models in English and Ukrainian, at baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset. All released artifacts: the Hiring Bias Mitigation collection. Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data. Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.tabulartext-generation1M<n<10M0 likes757 downloads1d agoHugging Face02KleinWu /llmsys-hpobench LLMSYS-HPOBench LLMSYS-HPOBench is an offline benchmark dataset for hyperparameter optimization of real-world LLM systems. It covers inference engines, RAG pipelines, and agent frameworks, with normalized tabular measurements linked to log and hardware artifacts when available. Project Links GitHub repository, benchmark loader, and contribution guide: https://github.com/ideas-labo/llmsys-hpobench Paper: https://arxiv.org/abs/2605.08305 Full data archive on… See the full description on the dataset page: https://huggingface.co/datasets/KleinWu/llmsys-hpobench.texttabular-regression10M<n<100M0 likes285 downloads4mo agoHugging Face03Stereotypes-in-LLMs /hiring-bias-mitigation-synthetic-data Hiring-bias mitigation — synthetic training data Semi-synthetic data for training LLMs to make hiring decisions that do not depend on a protected attribute (military status, gender, religion), in English and Ukrainian. Real inputs, synthetic labels. CVs and job descriptions are real, anonymised postings from the Djinni Recruitment Dataset (MIT). Decisions and rationales were written by the teacher model Qwen/Qwen3.5-122B-A10B-GPTQ-Int4. Code and results:… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-synthetic-data.tabulartext-generation100K<n<1M0 likes277 downloads1d agoHugging Face04stacklok /llm-security-leaderboard-contentstabularn<1K0 likes193 downloads1y agoHugging Face05arnaiztech /llms-mental-health-crisis-benchmark Dataset Card for Between Help and Harm - Crisis Benchmark Dataset Summary This dataset repo contains the benchmark-side artifacts prepared for Hugging Face from the paper Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs, published in JMIR Mental Health. If you use this dataset, please cite the paper. The citation is included below, the arXiv version is available at https://arxiv.org/abs/2509.24857, and the final DOI is allocated as… See the full description on the dataset page: https://huggingface.co/datasets/arnaiztech/llms-mental-health-crisis-benchmark.tabular10K<n<100K1 likes163 downloads6mo agoHugging Face06ggg-llms-team /GigaVerbo-filteredgatedtabular100M<n<1B0 likes122 downloads10mo agoHugging Face07SoroushVahidi /llm-serving-selector-regret LLM-Serving Selector Regret LLM-Serving Selector Regret is a metrics-only research dataset for studying learned policy selection in LLM-serving schedulers. It contains derived selector/oracle/regret objects generated by Soroush Vahidi's research workflow, not raw request traces. Creator / Provider Dataset creator/provider: Soroush Vahidi. The released selector/regret and policy-suitability metrics were generated by Soroush Vahidi's research workflow. Underlying… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/llm-serving-selector-regret.tabular100K<n<1M1 likes102 downloads1mo agoHugging Face08arnaiztech /llms-mental-health-crisis-responses Dataset Card for Between Help and Harm - Responses and Evaluations Dataset Summary This dataset repo contains the response-side artifacts prepared for Hugging Face from the paper Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs, published in JMIR Mental Health. If you use this dataset, please cite the paper. The citation is included below, the arXiv version is available at https://arxiv.org/abs/2509.24857, and the final DOI is allocated as… See the full description on the dataset page: https://huggingface.co/datasets/arnaiztech/llms-mental-health-crisis-responses.tabular100K<n<1M1 likes101 downloads6mo agoHugging Face09Stereotypes-in-LLMs /toxicchat_output-Ukrtabular1K<n<10K0 likes83 downloads2mo agoHugging Face10oxford-llms /world_values_survey_2017_2022_sfttabular100K<n<1M1 likes61 downloads2y agoHugging Face11JSALT2024-Astro-LLMs /astro_paper_corpustabular100K<n<1M2 likes60 downloads2y agoHugging Face12SoroushVahidi /llm-serving-scheduler-baselines LLM-Serving Scheduler Baselines: Simulation Performance Outcomes for Scheduler Policies This is a comprehensive, text-free, highly structured simulation results dataset for large language model (LLM) serving schedulers. It contains policy-level outcome records generated across synthetic scheduler stress tests and an added TraceLab-derived out-of-distribution policy sweep. The dataset compares 12 highly optimized third-party baseline schedulers against APT-Serve (a… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/llm-serving-scheduler-baselines.tabular10K<n<100K0 likes56 downloads1mo agoHugging Face13giantfish-fly /LLMs-First-Task Super easy task for humans that All SOTA LLM fail to retrieve the correct answer from context. Including SOTA models: GPT5, Grok4, DeepSeek, Gemini 2.5PRO, Mistral, Llama4...etc Update: Accepted to COLM 2026 (San Francisco). AAAI 2026 Worshop Oral: Jan/2026 LaMAS (LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems) Jan/2026 Singapole ICML 2025 Long-Context Foundation Models Workshop Accepted.(https://arxiv.org/abs/2506.08184) Update: This dataset… See the full description on the dataset page: https://huggingface.co/datasets/giantfish-fly/LLMs-First-Task.tabularquestion-answeringn<1K1 likes53 downloads2mo agoHugging Face14ucberkeley-dlab /fragility-moral-judgment-llms Fragility of Moral Judgment in Large Language Models Companion dataset for the FAccT paper Fragility of Moral Judgment in Large Language Models by Tom van Nuenen. Contains the moral dilemmas, community labels, and per-model verdicts (with explanations and reasoning traces) used in the study. The paper investigates how stable LLM moral judgments are under minimal, morally-irrelevant perturbations of the same dilemma, and whether protocols and reasoning chains improve or worsen… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/fragility-moral-judgment-llms.tabulartext-classification100K<n<1M0 likes52 downloads4mo agoHugging Face15oxford-llms /european_social_survey_2023_sfttabular100K<n<1M1 likes44 downloads2y agoHugging Face16minimaxir /llm-strawberrytabular1K<n<10K0 likes38 downloads1y agoHugging Face17pkuHaowei /llm-srbench-trajectoriestabular1K<n<10K0 likes32 downloads7mo agoHugging Face18Stereotypes-in-LLMs /UAlign⚠️ Disclaimer: This dataset contains examples of morally and socially sensitive scenarios, including potentially offensive, harmful, or illegal behavior. It is intended solely for research purposes related to value alignment, cultural analysis, and safety in AI. Use responsibly. UAlign: LLM Alignment Evaluation Benchmark This benchmark consists of two test-only subsets adapted into Ukrainian: ETHICS (Commonsense subset): A binary classification task on ethical acceptability.… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/UAlign.tabulartext-classification1K<n<10K0 likes28 downloads1y agoHugging Face19droussis /llms4eu_synthetic_qa FineWiki LLMs4EU Tourism QA Dialogues This dataset contains multilingual tourism and culture dialogue data derived from FineWiki articles selected with Propella annotations. Each row contains a Wikipedia/FineWiki document about a place, landmark, cultural site, or tourism-relevant topic, augmented with synthetic question-answer (QA) pairs and converted into chat message formats for RAG-style supervised fine-tuning. Source Selection Source documents were selected from… See the full description on the dataset page: https://huggingface.co/datasets/droussis/llms4eu_synthetic_qa.tabularn<1K0 likes24 downloads4mo agoHugging Face20dvilasuero /ultrafeedback_binarized_thinking_llms Dataset Card for ultrafeedback_binarized_thinking_llms This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: pipeline.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/dvilasuero/ultrafeedback_binarized_thinking_llms/raw/main/pipeline.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/ultrafeedback_binarized_thinking_llms.tabularn<1K1 likes23 downloads2y agoHugging Face21echodrift /LLMs-for-Soliditytabular10K<n<100K3 likes22 downloads2y agoHugging Face22oxford-llms /european_social_survey_2020_sfttabular100K<n<1M0 likes21 downloads2y agoHugging Face23oxford-llms /european_social_survey_2023_germany_sfttabular100K<n<1M0 likes20 downloads2y agoHugging Face24LLMsForHepth /hep-th_perplexitiesThe code used to generate this dataset can be found at https://github.com/Paul-Richmond/hepthLlama/blob/main/src/get_perplexity.py. tabular10K<n<100K0 likes18 downloads1y agoHugging Face25nuprl-staging /llm-systems-travel-agenttabular1M<n<10M0 likes17 downloads2y agoHugging Face26LLMsForHepth /jina_scores_hep-ph_gr-qcA copy of the dataset LLMsForHepth/infer_hep-ph_gr-qc with additional columns score_Llama-3.1-8B and score_s2-L-3.1-8B-base. The additional columns contain the cosine similarities between the sequences in abstract and y_pred where y_pred is taken from [comp_Llama-3.1-8B, comp_s2-L-3.1-8B-base]. The model used to create the embeddings is jinaai/jina-embeddings-v3. tabular10K<n<100K0 likes17 downloads2y agoHugging Face27LLMsForHepth /llm_scores_hep_thThis dataset originates as a copy of LLMsForHepth/infer_hep_th. We have then added the columns score_Llama-3.1-8B, score_s1-L-3.1-8B-base, score_s2-L-3.1-8B-base and score_s3-L-3.1-8B-base_v3. The values contained within each of these new columns are cosine similarity scores between the ground truth abstracts in abstract and the llm completed abstracts in comp_Llama-3.1-8B, comp_s1-L-3.1-8B-base, comp_s2-L-3.1-8B-base and comp_s3-L-3.1-8B-base_v3 respectively. In more detail, the similarity… See the full description on the dataset page: https://huggingface.co/datasets/LLMsForHepth/llm_scores_hep_th.tabular10K<n<100K0 likes17 downloads2y agoHugging Face28LLMsForHepth /sem_scores_hep_thThe code used to generate this dataset can be found at https://github.com/Paul-Richmond/hepthLlama/blob/main/src/sem_score.py. A copy of the dataset LLMsForHepth/infer_hep_th with additional columns score_s1-L-3.1-8B-base, score_s3-L-3.1-8B-base_v3, score_Llama-3.1-8B and score_s2-L-3.1-8B-base. The additional columns contain the cosine similarities between the sequences in abstract and y_pred where y_pred is taken from [comp_s1-L-3.1-8B-base, comp_s3-L-3.1-8B-base_v3, comp_Llama-3.1-8B… See the full description on the dataset page: https://huggingface.co/datasets/LLMsForHepth/sem_scores_hep_th.tabular10K<n<100K0 likes16 downloads1y agoHugging Face29anicola /value-systems-in-llms-paraphrasing-and-profile-elicitation Value Systems in LLMs: Effects of Paraphrasing and Profile Elicitation on Decision-Making Consistency and Robustness (Versión en español más abajo.) Do large language models give stable answers to the same forced-choice question when the prompt is perturbed in ways that do not change its meaning — and does assigning them a personality or value profile change those answers? This dataset contains the full material of that experiment: the 9,350 prompts, the 561,000 model responses… See the full description on the dataset page: https://huggingface.co/datasets/anicola/value-systems-in-llms-paraphrasing-and-profile-elicitation.tabularmultiple-choice100K<n<1M0 likes16 downloads1mo agoHugging Face30ilsp /llms4eu_synthetic_qa FineWiki LLMs4EU Tourism QA Dialogues This dataset contains multilingual tourism and culture dialogue data derived from FineWiki articles selected with Propella annotations. Each row contains a Wikipedia/FineWiki document about a place, landmark, cultural site, or tourism-relevant topic, augmented with synthetic question-answer (QA) pairs and converted into chat message formats for RAG-style supervised fine-tuning. Source Selection Source documents were selected… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/llms4eu_synthetic_qa.tabularn<1K0 likes15 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.