datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hiring-bias-mitigation-responses
Hiring-bias mitigation — model responses
Every response produced in the mitigation study of LLM hiring decisions: 61 runs,
2,689,200 responses, from 5 open-weight models in English and Ukrainian, at
baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset.
All released artifacts: the Hiring Bias Mitigation collection.
Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data.
Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.dementor-matrix-responses
Dementor — matrix model responses
Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study.
Companion to:
Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan)
Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO /
self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant).
Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.tiny-ua-bench-responses
Tiny-UA-Bench Responses
This dataset contains the response matrix for Tiny-UA-Bench.
The matrix contains 919,160 model and item records.
The matrix covers 20 models and 45,958 items.
The evaluation excludes FLORES and LongFLORES.
Use
Use this dataset to reproduce the benchmark compression analysis.
Do not use a held-out model response to fit a selector or predictor.
Use the reference and held-out split definitions from the code repository.
Load the data with the… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/tiny-ua-bench-responses.multi-ai-interpretive-responses
Multi-AI Interpretive Responses
arena.ai のサイドバイサイド / ダイレクトバトルで行った
日本語チャットセッションのアーカイブ。
同じ問い(哲学・倫理・サブカル・メタ認知ネタ)に対する複数 LLM の
解釈・応答差を観察するためのデータセット。
「楽しい を教えるお仕事ならしまーす」
倫理と哲学だけメタ超級。バシャール可。仏陀可。ウィトゲン可。サブカル可。
ファイル
ファイル
説明
data.jsonl
1 行 = 1 セッション。HF Datasets Viewer はこれを読みます。
timeline.md
人間用:会話開始時刻順の年表(タイトル・モデル数・所要時間付き)。
filename_map.csv
新ファイル名 ↔ 元日本語タイトルの対応表。
files/chat_XXXX.json
arena.ai 由来の生 JSON(元構造そのまま)。連番は会話開始時刻順。
スキーマ(data.jsonl の… See the full description on the dataset page: https://huggingface.co/datasets/TonbokiriRaikiriMuramasa/multi-ai-interpretive-responses.eu-cyber-llm-benchmark-responses
EU Cyber Threat Landscape LLM Benchmark — Responses
15,988 LLM-generated cyber threat landscape assessments from 7 models across 3 continents, designed to measure geopolitical bias in attribution framing.
What this is
The complete response corpus from running the EU Cyber LLM Benchmark prompts against 7 locally deployed models via Ollama. Each record contains the full model output, pre-extracted analytical sections, CVE mentions, refusal flags, and latency measurements.… See the full description on the dataset page: https://huggingface.co/datasets/eromang/eu-cyber-llm-benchmark-responses.mental_health_counseling_responses
Dataset Card for Mental Health Counseling Responses
This dataset contains responses to questions from mental health counseling sessions.
The responses are rated by LLMs using the dimensions: empathy, appropriateness, and relevance.
A detailed explanation of the rating process can be found in this blog post.
For a detailed analysis of LLM-generated responses and their comparison to human responses, refer to this blog post.
The original data with the human responses can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_responses.
