CoolFace
Datasetpublic

raca-workspace-v1/grpo-tool-sat-dataset-v1

grpo-tool-sat-dataset-v1 Synthetic lookup-table dataset for the GRPO Tool Saturation experiment. 10k keys k in [0, 9999] with r = k mod 3 feature. Tools map/table have opaque, partially overlapping correct-domains: map correct for r in {0, 2}, table correct for r in {1, 2}. f(k) = SHA256(str(k))[:6]; wrong-hash returns are g_map(k), g_table(k). SFT demos use overlap_skew_map=0.6 on r=2; token-level shuffle across r-classes with fixed seed. Dataset Info Rows:… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/grpo-tool-sat-dataset-v1.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes18downloads
Dataset Card

grpo-tool-sat-dataset-v1

Synthetic lookup-table dataset for the GRPO Tool Saturation experiment. 10k keys k in [0, 9999] with r = k mod 3 feature. Tools map/table have opaque, partially overlapping correct-domains: map correct for r in {0, 2}, table correct for r in {1, 2}. f(k) = SHA256(str(k))[:6]; wrong-hash returns are gmap(k), gtable(k). SFT demos use overlapskewmap=0.6 on r=2; token-level shuffle across r-classes with fixed seed.

Dataset Info

  • —Rows: 20000
  • —Columns: 11

Columns

ColumnTypeDescription
toolValue('string')'map' or 'table' — the tool the demo uses (sft split only)
user_promptValue('string')User-side prompt: 'Key: <k>' (sft split only)
assistant_completionValue('string')SFT target: prose + <tool_call> + <observation> + <answer> (sft split only)
kValue('int64')Key integer in [0, 9999]
rValue('int64')k mod 3 (0 = Monly, 1 = Tonly, 2 = Overlap)
f_kValue('string')6-hex correct answer = SHA256(str(k))[:6]
gmapkValue('string')6-hex wrong answer returned by map(k) when r=1 (meta split only)
gtablekValue('string')6-hex wrong answer returned by table(k) when r=0 (meta split only)
correct_toolsList(Value('string'))List of tools yielding reward 1 for this key (meta split only)
regionValue('string')Human-readable region label
split_nameValue('string')Logical split: 'meta' (every key), 'sft' (demo rows), or 'eval' (held-out eval keys). Filter by this to get the split you need.

Generation Parameters

json
{
  "script_name": "src/data_gen.py",
  "model": "n/a",
  "description": "Synthetic lookup-table dataset for the GRPO Tool Saturation experiment. 10k keys k in [0, 9999] with r = k mod 3 feature. Tools map/table have opaque, partially overlapping correct-domains: map correct for r in {0, 2}, table correct for r in {1, 2}. f(k) = SHA256(str(k))[:6]; wrong-hash returns are g_map(k), g_table(k). SFT demos use overlap_skew_map=0.6 on r=2; token-level shuffle across r-classes with fixed seed.",
  "hyperparameters": {
    "seed": 1,
    "eval_frac": 0.2,
    "overlap_skew_map": 0.6,
    "hash_slice": 6
  },
  "input_datasets": []
}

Usage

python
from datasets import load_dataset

dataset = load_dataset("raca-workspace-v1/grpo-tool-sat-dataset-v1", split="train")
print(f"Loaded {len(dataset)} rows")

Uploaded via [RACA](https://github.com/Zayne-sprague/Dr-Claude-Code) hf_utility.