datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sphragis-olmo1b-adaptation-corpus
Sphragis OLMo-1B adaptation corpus
Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek
before authorship-language-model training. It contains only OGA whole works
whose TLG author occurs in neither Sphragis benchmark.
Text has the exact model-facing benchmark surface form: polytonic-aware
lowercasing with grc_utils.lower_grc, removal of all editorial punctuation,
normalization of whitespace, and removal of consonant-final elision marks.
Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.gnu-prolog-adaptation-corpus
GNU Prolog adaptation corpus — AutoScientist Challenge (Math & Code)
~1,200 execution-verified GNU Prolog task/completion pairs plus a frozen
175-task holdout (holdout.jsonl, hash-pinned before any training run).
Every completion was verified by executing it against the task's checks;
no completion entered the corpus on an LLM's word alone. Generator, seeds
and manifest included. Used to train
AryaGarg23/llama-3.2-3b-gnu-prolog-lora (13.7% -> 76.0%
executable pass@1 at 3B).
ipda-judge-adaptation-grpo
IPDA Judge Adaptation GRPO Dataset
Training data for judge adaptation in competitive debate. Contains GRPO preference sets for adapting debate speech generation to different judge profiles.
Dataset Description
This dataset enables training LLMs to adapt their debate arguments based on judge characteristics:
Depth Adaptation: Adapting explanation complexity to judge expertise level (debate experience + domain knowledge)
Bias Adaptation: Adapting argument framing to judge… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-judge-adaptation-grpo.communication-adaptation-sft-100k
Communication Adaptation SFT (100K)
100,000 ShareGPT conversations demonstrating skilled communication style adaptation across 15 task types. Each example shows how to take the same underlying content and adjust register, technical depth, length, and framing for different audiences and purposes.
Motivation
Communication adaptation is a core professional skill that LLMs often handle clumsily. Common failures:
Technical monologue: explaining cloud storage to a… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/communication-adaptation-sft-100k.ipda-judge-adaptation-data
IPDA Judge Adaptation Training Dataset
Training data for judge adaptation in competitive debate. This dataset teaches models to adapt their debate output based on judge characteristics.
Dataset Structure
Files
File
Description
Pairs
depth_iter1_train.json
Depth adaptation iteration 1 (lay vs expert judges)
75
depth_iter2_train.json
Depth adaptation iteration 2 (different topics)
75
bias_train.json
Bias adaptation (ideological, procedural… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-judge-adaptation-data.ipda-judge-adaptation-grpo
IPDA Judge Adaptation GRPO Dataset
Training data for judge adaptation in competitive debate. Contains GRPO preference sets for adapting debate speech generation to different judge profiles.
Dataset Description
This dataset enables training LLMs to adapt their debate arguments based on judge characteristics:
Depth Adaptation: Adapting explanation complexity to judge expertise level (debate experience + domain knowledge)
Bias Adaptation: Adapting argument framing to judge… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-judge-adaptation-grpo.OpenVerification1_aux_adaptation_examples
Dataset Card for ReexpressAI/OpenVerification1_aux_adaptation_examples
This is additional data as part of ReexpressAI/OpenVerification1. The data fields are slightly different for this data source, so we include this as a separate dataset.
This is example output from the Reexpress MCP Server when using the ReexpressAddTrue, ReexpressAddFalse, or ReexpressAddOOD tools. These are the lines that get saved to the adaptation/running_updates.jsonl file in the model directory.
Refer to… See the full description on the dataset page: https://huggingface.co/datasets/ReexpressAI/OpenVerification1_aux_adaptation_examples.
