datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmconflict-editable-values-1k
MMConflict Editable Values 2K
This dataset contains 2,000 source images with visible atomic values for
multimodal conflict research. It has 100 images in each of 20 categories. Every
image comes from a photograph, scan, captured website, software screenshot, or
page of a source document. The dataset does not contain generated images or
project-rendered examples.
Each row records the source, source URL, license, attribution, visible value,
question, and a candidate box around the… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-editable-values-1k.values-in-the-wild
Summary
This dataset presents a comprehensive taxonomy of 3307 values expressed by Claude (an AI assistant) across hundreds of thousands of real-world conversations. Using a novel privacy-preserving methodology, these values were extracted and classified without human reviewers accessing any conversation content. The dataset reveals patterns in how AI systems express values "in the wild" when interacting with diverse users and tasks.
We're releasing this resource to advance research… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/values-in-the-wild.llm-values-tinker-lora-checkpoints
Tinker LoRA checkpoint archive
Archived from Tinker on 2026-09-19 before the account was cleared. manifest.json lists every checkpoint that was in scope (including weights/ training-state checkpoints, which could not be exported and no longer exist); runs.json holds the training-run metadata. Each exportable checkpoint is at checkpoints/<run_id>/<checkpoint_id>.tar, the archive exactly as Tinker served it.
pile-toxicity-balanced3-with-valueszip-training-hallucination-data-qwen06b-thinking-train-with-valuesopenr1_token_wise_values_32026-09-17-da-lowstakes-implicit-values-synth-smoke
FAILED reviewer calibration before generation; zero generated candidates; reviewer fixtures only
field
value
experiment
FAILED reviewer calibration before generation; zero generated candidates; reviewer fixtures only
date_generated
20260917_151935
constitution
constitutions/claude_distilled_09_principles/constitution.md; raw SHA256 6ccd9c2a1ae1f2b479dbecbeb1c735b2014de531a866bcac9b31d8bdf5de914f
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-implicit-values-synth-smoke.synthetic-values-model-charter
value_units.jsonl is the individual parsed values from the model charter.
scenarios.jsonl is invididual hypothetical scenarios based on the values in values_units.jsonl
sft_*.jsonl generated accepted responses.
dpo_*.jsonl generated accepted+rejected responses.
spadl-vaep-action-values
SPADL/VAEP Action Values
Every on-ball action from ~9.5 million professional soccer events, converted to the SPADL unified format and scored with offensive, defensive, and net VAEP values. Built with the silly-kicks library — enabling player ranking by total contribution beyond goals and assists.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
⚠️ Schema change (cut-over 2026-07-22)
This dataset now emits both legacy and canonical Kimball key… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/spadl-vaep-action-values.openr1_step_wise_values_xxl2026-09-17-da-lowstakes-values-in-advice-synth-smoke
18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only
field
value
experiment
18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only
date_generated
20260917_171453
constitution
constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-values-in-advice-synth-smoke.wan2.2_production_values
Wan2.2 Production Values — VSA Block-Sparse Attention Inputs + Reference
Real, captured production inputs for the Video Sparse Attention (VSA) fine
block-sparse attention stage of Wan2.2 T2V-A14B, recorded from an actual
generate.py run at 832×480 / 81 frames, plus the reference Triton kernel
and the scoring rule.
This is a fixed, non-gameable benchmark for proposing a faster VSA forward
kernel: optimize on these exact tensors, score against the reference at the
tolerance below.… See the full description on the dataset page: https://huggingface.co/datasets/baseten-admin/wan2.2_production_values.fifa24-player-values
FIFA 24 Player Market Value Analysis
Dataset Overview
This dataset contains information about 180,021 football players from EA Sports FC 24.
The main goal is to analyze what factors determine a player's market value.
Source: Kaggle - EA Sports FC 24 Complete Player Dataset
Original data: EA Sports FC 24
Research Question
What factors influence the market value of a football player?
Key Findings
Most players have a market value under €1M… See the full description on the dataset page: https://huggingface.co/datasets/Belkin7/fifa24-player-values.world_values_survey_2017_2022_sftfood-nutritional-valuesOriginally webscrapped by Aleksandr Antonov from Nutrition Value and posted on Kaggle as "Nutritional values for common foods and products".
when-agents-act
Dataset Card for "When Agents Act"
Dataset Summary
This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute).
Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.Human-Values
Labels
label
meaning
achievement_P
in favor of achievement
achievement_N
against achievement
power_dominance_P
in favor of power: dominance
power_dominance_N
against power: dominance
power_resources_P
in favor of power: resources
power_resources_N
against power: resources
space-creation-values
Space Creation Values — ELASTIC/OBSO
Per-player per-frame space creation quantification — measuring each player's contribution to off-ball scoring opportunities via differential OBSO. For every sampled frame, the model computes OBSO with and without each player, yielding the area of scoring opportunity that player creates (or destroys) by their positioning.
Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
Quick Start
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/space-creation-values.Charter_values_citizenship_integration
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19058
Description
The integration agreement form prepared for signing the pact between foreign and state, in addition to providing the alien's commitments, indicates, the statement by the person concerned, to adhere to the Charter of the values of citizenship and integration of the decree of the Minister of 23 April 2007, pledging to respect its principles. The Charter of citizenship and integration… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Charter_values_citizenship_integration.anthropic-rlhf-human-values-dpoFIR-Bench-Research-Reports-FinQAcrafter-values-splitself_response_2_Qwen__Qwen1.5_32B_answersopenr1_token_wise_values_large_new_formdfm10-synthetic-values-model-charter-da
dfm10-synthetic-values-model-charter-da
Independently audited Danish adaptations of the complete aligned SFT/DPO scenario tuples.
Contents
Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
Schema: chat messages, optional condition and tools, plus provenance
Shards: 1
Rows: 1,343
Category: Danish values and preference alignment
Upstream material
danish-foundation-models/synthetic-values-model-charter
Processing
Gemma 4… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-synthetic-values-model-charter-da.simulators-political-valuesCIVICS
Dataset Details
“CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal Impacts” is a dataset designed to evaluate the social and cultural variation of Large Language Models (LLMs) towards socially sensitive topics across multiple languages and cultures. The hand-crafted, multilingual dataset of statements addresses value-laden topics, including LGBTQI rights, social welfare, immigration, disability rights, and surrogacy. CIVICS is designed to elicit responses from LLMs… See the full description on the dataset page: https://huggingface.co/datasets/llm-values/CIVICS.self_response_2_01_ai_cont_Yi_34B_answerscrafter-valuesvalues
