CoolFace
Datasetpublic

316usman/wash-trading-detection

WASH_TRADING_DETECTION A preference dataset for WASH_TRADING_DETECTION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/wash-trading-detection.

sourceHugging Faceunknownupdated 6d agoView on Hugging Face
0likes49downloads
Dataset Card

WASHTRADINGDETECTION

A preference dataset for WASH_TRADING_DETECTION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.

Format

Standard preference / DPO schema — each row:

columnmeaning
promptthe request (originally prompt)
chosenthe human-preferred response
rejecteda worse response to the same prompt
sourcethe dataset/URL the row was harvested from

Splits

80/10/10 train / validation / test (seeded shuffle): train:1201 / validation:150 / test:151

Stats

  • Rows: 1502
  • Distinct sources: 1

Sources

  • synthetic:deepseek/deepseek-v4-flash-0731

Provenance

Each row's chosen/rejected distinction comes from a real human signal (upvotes, accepted answers, ratings, or a real strong-vs-weak reply). Rows passed an automated quality gate checking that chosen is a clean response (not a transcript), the chosen/rejected contrast is about quality (not length), and the row is on-intent.