datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SADC-Situation-Awareness-for-Driver-Centric-Driving-Style-Adaptation
Dataset Card for Dataset SADC
There is evidence that the driving style of an
autonomous vehicle is important to increase the acceptance
and trust of the passengers. The driving situation has been
found to have a significant influence on human driving behavior.
However, current driving style models only partially incorporate
driving environment information, limiting the alignment between
an agent and the given situation.
Therefore, we propose a dataset for situation-aware… See the full description on the dataset page: https://huggingface.co/datasets/jHaselberger/SADC-Situation-Awareness-for-Driver-Centric-Driving-Style-Adaptation.SADC-Situation-Awareness-for-Driver-Centric-Driving-Style-Adaptation
Dataset Card for Dataset SADC
There is evidence that the driving style of an
autonomous vehicle is important to increase the acceptance
and trust of the passengers. The driving situation has been
found to have a significant influence on human driving behavior.
However, current driving style models only partially incorporate
driving environment information, limiting the alignment between
an agent and the given situation.
Therefore, we propose a dataset for situation-aware driving… See the full description on the dataset page: https://huggingface.co/datasets/zzqasdfsdf/SADC-Situation-Awareness-for-Driver-Centric-Driving-Style-Adaptation.africa-synth-agriculture-climate-adaptation-tech-ssa-all
Africa Synth Agriculture Climate Adaptation Tech Ssa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1M<n<10M - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-agriculture-climate-adaptation-tech-ssa-all.sphragis-olmo1b-adaptation-corpus
Sphragis OLMo-1B adaptation corpus
Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek
before authorship-language-model training. It contains only OGA whole works
whose TLG author occurs in neither Sphragis benchmark.
Text has the exact model-facing benchmark surface form: polytonic-aware
lowercasing with grc_utils.lower_grc, removal of all editorial punctuation,
normalization of whitespace, and removal of consonant-final elision marks.
Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.gnu-prolog-adaptation-corpus
GNU Prolog adaptation corpus — AutoScientist Challenge (Math & Code)
~1,200 execution-verified GNU Prolog task/completion pairs plus a frozen
175-task holdout (holdout.jsonl, hash-pinned before any training run).
Every completion was verified by executing it against the task's checks;
no completion entered the corpus on an LLM's word alone. Generator, seeds
and manifest included. Used to train
AryaGarg23/llama-3.2-3b-gnu-prolog-lora (13.7% -> 76.0%
executable pass@1 at 3B).
ipda-judge-adaptation-grpo
IPDA Judge Adaptation GRPO Dataset
Training data for judge adaptation in competitive debate. Contains GRPO preference sets for adapting debate speech generation to different judge profiles.
Dataset Description
This dataset enables training LLMs to adapt their debate arguments based on judge characteristics:
Depth Adaptation: Adapting explanation complexity to judge expertise level (debate experience + domain knowledge)
Bias Adaptation: Adapting argument framing to judge… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-judge-adaptation-grpo.communication-adaptation-sft-100k
Communication Adaptation SFT (100K)
100,000 ShareGPT conversations demonstrating skilled communication style adaptation across 15 task types. Each example shows how to take the same underlying content and adjust register, technical depth, length, and framing for different audiences and purposes.
Motivation
Communication adaptation is a core professional skill that LLMs often handle clumsily. Common failures:
Technical monologue: explaining cloud storage to a… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/communication-adaptation-sft-100k.OVCLvietnamese-ocr-adaptationjeeran_semeval_2016_arabic_adaptation
Jeeran — SemEval-2016 Task 5 (Arabic) adaptation
Aspect-based sentiment annotations over Jeeran reviews (Jordanian/Levantine
dialectal Arabic, 29 business domains), rendered in the SemEval-2016 Task 5
subtask 1 format so that tooling written for SemEval2016_arabic runs unchanged.
Contents
split
reviews
sentences
opinions
train
43420
69795
182788
test
10851
17644
45520
Raw SemEval XML lives under semeval_xml/; the datasets view has one row per… See the full description on the dataset page: https://huggingface.co/datasets/k-chirkunov/jeeran_semeval_2016_arabic_adaptation.humanoid-future-task-adaptation
Future Task Adaptation
Adaptive behavior dataset for autonomous humanoids.
jeeran_semeval_2016_arabic_adaptation_coarse
Jeeran — SemEval-2016 Task 5 (Arabic) adaptation, coarse categories
The k-chirkunov/jeeran_semeval_2016_arabic_adaptation dataset with its 48 fine-grained Arabic aspect labels
collapsed onto 14 coarse families. Everything else — targets, character offsets,
polarity, the human-annotated spans, the train/test split — is carried over unchanged.
Contents
split
reviews
sentences
opinions
train
43,365
69,524
181,406
test
10,838
17,594
45,225
The… See the full description on the dataset page: https://huggingface.co/datasets/k-chirkunov/jeeran_semeval_2016_arabic_adaptation_coarse.tool-use-adaptation-difficulty
Tool-Use Adaptation — Task Difficulty (gemini & Qwen)
Empirical difficulty annotations for tool-use / agentic tasks, part of an "adapting to a new tool" RL domain.
Difficulty is gauged by running two strong solvers N=32 times per task and scoring each rollout against the
task's local programmatic gold (execution / state-check / exact-args), then aggregating.
Solvers
gemini-3-5-flash-fair (gemini_* columns)
Qwen3.6-35B-A3B (qwen_* columns)
Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tool-use-adaptation-difficulty.climate_adaptation_abstractsfirm-level-adaptation-dataThis is the (preliminary) dataset for firm-level climate change adaptation.
It contains the following comments:
path: a unique identifier that stores the path of the file that was used to create the scores.
cik: unique firm identifier.
year: year of the report.
date: publishing date of the report.
num_paragraphs: number of paragraphs in the report.
num_words: number of words in the report.
Storm, Flood, Heatwave, Drought, Wildfire, Coldwave: share of paragraphs mentioning extreme weather… See the full description on the dataset page: https://huggingface.co/datasets/climate-adaptation/firm-level-adaptation-data.climate_adaptation_texthumanoid-environment-adaptation-dataset-v1han-autonomous-behavioral-adaptation-dataset-v1
Humanoid Autonomous Behavioral Adaptation Dataset
This dataset models adaptive behavior shifts
of humanoid agents responding to environmental changes.
It records baseline behavior, trigger events,
adaptation strategies, and outcome performance.
Objective
To enable continuous behavioral learning
and adaptive optimization in dynamic conditions.
Data Fields
adaptation_event_id
agent_id
mission_id
baseline_behavior_profile
environmental_change_type… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-autonomous-behavioral-adaptation-dataset-v1.climate_adaptation_abstracts_apr_2024AdaptationBERT-ClimateAgricultural-Climate-Adaptation-Research-Dataset
Agricultural Climate Adaptation Research Dataset
Agriculture is currently facing challenges posed by climate change, particularly the increasing impact of drought on crop yields. Existing research data often lacks detailed analysis under specific climate conditions, leading to ineffective agricultural management measures. This dataset aims to fill this gap by including images of farmland drought and vegetation recovery, assisting AI models in researching agriculture's ability to… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Agricultural-Climate-Adaptation-Research-Dataset.AdaptationLabs-synthetic-qaCSIC2010_dataset_domain_adaptationafrica-south-sudan-a-prospective-evaluation-of-nutrition-protocol-adaptations-806ccb7c
A Prospective Evaluation of Nutrition Protocol Adaptations | Africa (Johns Hopkins School of Public Health)
25,268 rows - 1 Africa country/area - 2022 - 133 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 25,268 rows from Johns Hopkins School of Public Health, covering A Prospective Evaluation of Nutrition Protocol Adaptations. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-south-sudan-a-prospective-evaluation-of-nutrition-protocol-adaptations-806ccb7c.climate_adaptation_abstracts_v1ctta_data_repoafrica-south-sudan-a-prospective-evaluation-of-nutrition-protocol-adaptations-20638fe1
A Prospective Evaluation of Nutrition Protocol Adaptations | Africa (Johns Hopkins School of Public Health)
221 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 221 rows from Johns Hopkins School of Public Health, covering A Prospective Evaluation of Nutrition Protocol Adaptations. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-south-sudan-a-prospective-evaluation-of-nutrition-protocol-adaptations-20638fe1.ipda-judge-adaptation-data
IPDA Judge Adaptation Training Dataset
Training data for judge adaptation in competitive debate. This dataset teaches models to adapt their debate output based on judge characteristics.
Dataset Structure
Files
File
Description
Pairs
depth_iter1_train.json
Depth adaptation iteration 1 (lay vs expert judges)
75
depth_iter2_train.json
Depth adaptation iteration 2 (different topics)
75
bias_train.json
Bias adaptation (ideological, procedural… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-judge-adaptation-data.zc-domain-adaptation-2OpenVerification1_aux_adaptation_examples
Dataset Card for ReexpressAI/OpenVerification1_aux_adaptation_examples
This is additional data as part of ReexpressAI/OpenVerification1. The data fields are slightly different for this data source, so we include this as a separate dataset.
This is example output from the Reexpress MCP Server when using the ReexpressAddTrue, ReexpressAddFalse, or ReexpressAddOOD tools. These are the lines that get saved to the adaptation/running_updates.jsonl file in the model directory.
Refer to… See the full description on the dataset page: https://huggingface.co/datasets/ReexpressAI/OpenVerification1_aux_adaptation_examples.
