datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
values-in-the-wild
Summary
This dataset presents a comprehensive taxonomy of 3307 values expressed by Claude (an AI assistant) across hundreds of thousands of real-world conversations. Using a novel privacy-preserving methodology, these values were extracted and classified without human reviewers accessing any conversation content. The dataset reveals patterns in how AI systems express values "in the wild" when interacting with diverse users and tasks.
We're releasing this resource to advance research… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/values-in-the-wild.ValuePrism
Dataset Card for ValuePrism
Dataset Summary
ValuePrism was created 1) to understand what pluralistic human values, rights, and duties are already present in large language models, and 2) to serve as a resource to to support open, value pluralistic modeling (e.g., Kaleido). It contains human-written situations and machine-generated candidate values, rights, duties, along with their valences and post-hoc explanations relating them to the situations.
For additional… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ValuePrism.SNOMED-CT-Code-Value-Semantic-Set.csvSNOMED-CT-Code-Value-Semantic-Set.csv
african-agro-processing-value-add
African Agro-Processing Value Addition Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/african-agro-processing-value-add.when-agents-act
Dataset Card for "When Agents Act"
Dataset Summary
This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute).
Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.broad-value-68367a
broad-value-68367a
Synthetic products test data: 59 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Granite-Rika49/broad-value-68367a.acquire-valued-shoppershttps://www.kaggle.com/c/acquire-valued-shoppers-challenge
device-valueplayer-valuer-synthetic-dataeastern-value-5f70d6
eastern-value-5f70d6
Synthetic sensors test data: 54 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Lunar-Scope/eastern-value-5f70d6.Africa-Agriculture-forestry-and-fishing-value-added-percentage-of-GDP
Africa Agriculture forestry and fishing value added percentage of GDP | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Agriculture-forestry-and-fishing-value-added-percentage-of-GDP.africa-industry-including-construction-value-added-percentage-of-gdp
Africa Industry Including Construction Value Added Percentage of Gdp | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-industry-including-construction-value-added-percentage-of-gdp.clinical-ethical-value-conflict-coherence-analysis-v0.1What this dataset tests
Whether a system can map value conflictand score the coherence of each ethical pathwithout claiming a single correct answer.
Required outputs
value conflict matrix
coherence score per path
moral tension hotspots
value sacrifice profile
stakeholder alignment map
coherence failure modes
Use case
Second layer of the Ethical Trade-off Simulator.
Africa-Stocks-Traded-Total-Value-percentage-of-GDP
Africa Stocks Traded Total Value percentage of GDP | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Stocks-Traded-Total-Value-percentage-of-GDP.counterfactual-goal-value-stability-v0.1
What this dataset tests
Whether goals and values remain stableunder counterfactual pressure.
Plans may change.
Priorities must not silently invert.
Why this exists
Counterfactual reasoning often triggers:
goal substitution
safety downgrades
consent bypass
policy flips
This dataset detects those drifts.
Data format
Each example includes:
base goal
base values
counterfactual condition
proposed plan
implied tradeoffs
The task is to label… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/counterfactual-goal-value-stability-v0.1.WVS-Value-SimulationFinCPRG
FinCPRG Dataset
近年来,大语言模型(LLMs)在构建段落检索数据集方面展现出了巨大的潜力。然而,现有方法在表达跨文档查询需求和控制标注质量方面仍存在局限性。为了解决这些问题,本文提出了一个双向生成管道,旨在为文档内和跨文档场景生成3级层次化查询,并在直接映射标注的基础上挖掘额外的相关性标签。使用这个管道,我们从近1.3k份中文金融研究报告中构建了金融段落检索生成数据集(FinCPRG),该数据集包含层次化查询和丰富的相关性标签。
数据集结构 (Dataset Structure)
数据集遵循标准的段落检索格式,包含以下文件:
corpus.jsonl: 包含段落(文档)语料。每行是一个JSON对象,包含 _id (段落ID) 和 text (段落文本) 字段。
queries.jsonl: 包含所有生成的查询。每行是一个JSON对象,包含 _id (查询ID) 和 text (查询文本) 字段。
qrels/: 包含查询-段落相关性标注文件 (qrels),格式为TSV (query-id \t corpus-id \t… See the full description on the dataset page: https://huggingface.co/datasets/valuesimplex-ai-lab/FinCPRG.eCQM-Code-Value-Semantic-Set.csveCQM-Code-Value-Semantic-Set.csv
legal-billing-narrative-task-value-coherence-v0.1What this dataset does
You receive
billing narrative
hours
case stage
task category
value signal
duplication signals
You decide
coherent
or
incoherent
Daily use
cost draft review
client challenge response
write-down triage
LOINC-CodeSet-Value-Description.csvLOINC-CodeSet-Value-Description.csv
value_determinant
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/value_determinant.africa-medium-and-high-tech-manufacturing-value-added-percentage-manufacturing-value-added
Africa Medium and High Tech Manufacturing Value Added Percentage Manufacturing Value Added | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-medium-and-high-tech-manufacturing-value-added-percentage-manufacturing-value-added.african-plastic-recycling-value-chain
African Plastic Recycling Value Chain | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/african-plastic-recycling-value-chain.global-llm-value-gap
글로벌 LLM 가치관 괴리 측정 데이터셋: WVS 기반 다모델 윤리 벤치마크
이 데이터셋은 대규모 언어 모델(LLM)과 인간 집단 간의 가치관 괴리를 측정하기 위해 구축된 World Values Survey(WVS) Wave 7 기반 시뮬레이션 응답 데이터셋입니다.
66개국의 인구통계 분포를 반영한 층화 표집 페르소나를 6개 LLM에 주입하여 23개 윤리 문항에 대한 응답을 수집하고, 실제 WVS 인간 응답과 체계적으로 비교합니다.
주요 통계
항목
내용
분석 국가 수
66개국
윤리 문항 수
23개 (WVS Wave 7)
윤리 카테고리 수
9개 (도덕 기반 이론 기반)
분석 모델
GPT-4o, GPT-4o-mini, Gemini-2.5-flash, Gemini-2.0-flash, DeepSeek-reasoner, DeepSeek-chat
국가별 페르소나 수
100명 (층화 표집)
총 LLM 응답 수… See the full description on the dataset page: https://huggingface.co/datasets/yena0rganization/global-llm-value-gap.silo-valueCAN_Valuesstocks_valuesglobal-llm-value-gap
글로벌 LLM 가치관 괴리 측정 데이터셋: WVS 기반 다모델 윤리 벤치마크
이 데이터셋은 대규모 언어 모델(LLM)과 인간 집단 간의 가치관 괴리를 측정하기 위해 구축된 World Values Survey(WVS) Wave 7 기반 시뮬레이션 응답 데이터셋입니다.
66개국의 인구통계 분포를 반영한 층화 표집 페르소나를 6개 LLM에 주입하여 23개 윤리 문항에 대한 응답을 수집하고, 실제 WVS 인간 응답과 체계적으로 비교합니다.
주요 통계
항목
내용
분석 국가 수
66개국
윤리 문항 수
23개 (WVS Wave 7)
윤리 카테고리 수
9개 (도덕 기반 이론 기반)
분석 모델
GPT-4o, GPT-4o-mini, Gemini-2.5-flash, Gemini-2.0-flash, DeepSeek-reasoner, DeepSeek-chat
국가별 페르소나 수
100명 (층화 표집)
총 LLM 응답 수… See the full description on the dataset page: https://huggingface.co/datasets/ynchoi/global-llm-value-gap.ai-reward-signal-internal-value-alignment-mapping-v0.1
AI Reward Signal ↔ Internal Value Alignment Mapping v0.1
What this dataset is
This dataset maps the relationship between:
external reward signals
internal value estimates
observed agent behavior
It measures when these three elements remain aligned and when they decouple.
Why this matters
Alignment failures rarely start with catastrophic behavior.They begin when the internal value model stops tracking the true reward objective.
Early signs:
proxy reward… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-reward-signal-internal-value-alignment-mapping-v0.1.clinical-ethical-value-conflict-coherence-analysis-v0.2
Clinical Ethical Value Conflict Coherence Analysis v0.2
What this is
A small dataset that tests one question:
Can you detect when a clinical ethical decision system is moving toward coherence failure, not just carrying value conflict?
This repo focuses on ethical value conflict coherence analysis.
It models a system where:
ethical signal clarity may weaken
value conflict pressure may rise
stakeholder alignment may fragment
decision friction may destabilize clean ethical… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-ethical-value-conflict-coherence-analysis-v0.2.
