CoolFace
Datasetpublic

botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1

BOTCOIN Balanced Top-50 Reasoning Traces This public dataset contains enriched BOTCOIN reasoning-trace attempts selected from canonical dataset/v2 research-ready objects. Selection policy: Source only attempts/research-ready objects. Rank each domain by trace_quality.reasoning_trace_quality_score. Keep each domain's top 50 percent. Equalize domains to the smallest top-half count. The rows are self-contained and intentionally rich: prompt/messages… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes24downloads
Dataset Card

BOTCOIN Balanced Top-50 Reasoning Traces

This public dataset contains enriched BOTCOIN reasoning-trace attempts selected from canonical dataset/v2 research-ready objects.

Selection policy:

  1. 1.Source only attempts/research-ready objects.
  2. 2.Rank each domain by trace_quality.reasoning_trace_quality_score.
  3. 3.Keep each domain's top 50 percent.
  4. 4.Equalize domains to the smallest top-half count.

The rows are self-contained and intentionally rich: prompt/messages, document/questions/constraints, enriched reasoning trace, answer verification, trace quality, reward-style scalar fields, and provenance are included so downstream users can prune into SFT, GRPO/reward modeling, PRM, DPO-style pairing, or audit-specific views.

The attempts/ partition contains the balanced top-half quality attempts. The recovery_pairs/ partition is intentionally not quality-filtered: it contains all observed retry sessions where submit attempt 1 or 2 failed and a later attempt 2 or 3 passed, represented as wrong-to-pass preference pairs.

Manifest

json
{
  "dataset_version": "balanced-top50-v1",
  "source_bucket": "botcoin-traces-annotated-429971482539-us-east-2-an",
  "domains": [
    "companies",
    "computational_biology",
    "quantum_physics",
    "scrna_imputation"
  ],
  "candidate_counts": {
    "companies": 423753,
    "computational_biology": 54771,
    "quantum_physics": 52589,
    "scrna_imputation": 56447
  },
  "top_half_counts": {
    "companies": 211876,
    "computational_biology": 27385,
    "quantum_physics": 26294,
    "scrna_imputation": 28223
  },
  "selected_per_domain": 26294,
  "selected_total": 105176,
  "quality_metric": "trace_quality.reasoning_trace_quality_score",
  "created_at_unix_ms": 1782779509740,
  "repo_id": "botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1",
  "attempt_rows_written": 105176,
  "attempt_split_counts": {
    "companies": {
      "train": 23672,
      "validation": 1316,
      "test": 1306
    },
    "computational_biology": {
      "train": 23594,
      "validation": 1368,
      "test": 1332
    },
    "quantum_physics": {
      "train": 23532,
      "validation": 1388,
      "test": 1374
    },
    "scrna_imputation": {
      "train": 23627,
      "validation": 1323,
      "test": 1344
    }
  },
  "recovery_pairs": {
    "enabled": true,
    "index_path": ".hf-exports/recovery-passed-attempt2-3.jsonl",
    "sessions_seen": 6189,
    "pairs_seen": 7030,
    "pairs_written": 7030,
    "skipped": 0,
    "split_counts": {
      "companies": {
        "train": 2819,
        "validation": 153,
        "test": 165
      },
      "computational_biology": {
        "train": 1199,
        "validation": 74,
        "test": 67
      },
      "quantum_physics": {
        "train": 959,
        "validation": 69,
        "test": 53
      },
      "scrna_imputation": {
        "train": 1339,
        "validation": 66,
        "test": 67
      }
    }
  },
  "actual_rows_written": 112206,
  "context_misses": 0,
  "context_regenerated": 104879
}

Split Counts

json
{
  "companies": {
    "train": 23672,
    "validation": 1316,
    "test": 1306
  },
  "computational_biology": {
    "train": 23594,
    "validation": 1368,
    "test": 1332
  },
  "quantum_physics": {
    "train": 23532,
    "validation": 1388,
    "test": 1374
  },
  "scrna_imputation": {
    "train": 23627,
    "validation": 1323,
    "test": 1344
  }
}

Loading

python
from datasets import load_dataset

attempts = load_dataset(
    "botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1",
    data_files="attempts/**/*.jsonl",
    split="train",
)

recovery_pairs = load_dataset(
    "botcoinmoney/domain-agnostic-reasoning-traces-balanced-top50-v1",
    data_files="recovery_pairs/**/*.jsonl",
    split="train",
)

Attempt rows include a messages column suitable for chat-style SFT and scalar reward_* columns suitable for reward filtering or GRPO-style pipelines. Recovery-pair rows include chosen/rejected, chosen_attempt/ rejected_attempt, and a rejected_attempts list for preference or recovery tuning.