datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenization-multiplicity-data
Dataset: Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
This dataset contains the official experiment inference traces for the paper Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service by Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis and Manuel Gomez-Rodriguez.
📂 Dataset Structure
The dataset is organized into folders as follows:
.\{model}\{task}\{lang}\{seed}_{10*temperature}.jsonl
where {model}… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/tokenization-multiplicity-data.strategic-ttc-data
Dataset: Strategic Test-Time Compute (TTC)
This dataset contains the official experiment inference traces for the paper "Test-Time Compute Games" (arXiv:2601.21839).
It includes full model generations, token counts, and correctness verifications for various Large Language Models (LLMs) across three major reasoning benchmarks: GSM8K, AIME, and GPQA.
This data allows researchers to analyze the relationship between test-time compute and model performance without needing to re-run… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/strategic-ttc-data.humancentric-scenes-ai
HumanCentric-Scenes-AI
A multimodal benchmark of 296 AI-generated human-centric scenes across four domains:
CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise.
Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/Nima0Kamali/humancentric-scenes-ai.humancentric-scenes-ai
HumanCentric-Scenes-AI
A multimodal benchmark of 296 AI-generated human-centric scenes across four domains:
CCTV / surveillance imagery (Set 2, 85 images). Midjourney-generated stills that mimic low-resolution security-camera footage — parking lots, building interiors, outdoor public spaces — designed to test whether detection cues survive heavy compression and low-light noise.
Occupation × gender portraits (Set 3, 128 images). A balanced 64-occupation × 2-gender paired design… See the full description on the dataset page: https://huggingface.co/datasets/rjmaftv33/humancentric-scenes-ai.chatbot-arena-llm-refusal
Hand-Labeled Refusal Dataset for Chatbot Arena Responses
Dataset Overview
This dataset extends the Chatbot Arena: Human Preference 55K dataset by providing manual annotations of LLM responses with respect to refusal behaviors. The labels classifies if models refuse to answer a prompt due to ethical concerns or technical/capability limitations.
The Dataset contains labels for 1,750 response pairs, i.e. 3,500 model responses.
Dataset Details
This dataset adds… See the full description on the dataset page: https://huggingface.co/datasets/Human-CentricAI/chatbot-arena-llm-refusal.P-AT
Measuring bias in Instruction-Following models with P-AT
Instruction-Following Language Models (IFLMs) are promising and versatile tools for solving many downstream, information-seeking tasks. Given their success, there is an urgent need to have a shared resource to determine whether existing and new IFLMs are prone to produce biased language interactions.
We propose Prompt Association Test (P-AT), a resource for testing the presence of social biases in IFLMs.
P-AT stems from WEAT… See the full description on the dataset page: https://huggingface.co/datasets/HumanCentricART/P-AT.human-centric-last-token
Human-Centric — Last-Token Activations
Hidden-state last-token activations extracted from two chat LLMs over the
merged AndyZou situations + emotion-query prompt sets, together with shared
metadata and emotion annotations.
Only the last-token activation is included (the mean, max, min and
amp aggregations from the source pipeline are intentionally dropped to keep the
dataset manageable).
Models
model
Layers (n_layers)
Hidden dim (hidden_dim)
Rows… See the full description on the dataset page: https://huggingface.co/datasets/jero-r-cuello/human-centric-last-token.Human-centric-DatasethumancentricFPV_humancentric
