CoolFace
Datasetpublic

vm2825/BEAM_50K

BEAM 50K Synthetic Compact Dataset This is a wholly synthetic, machine-generated dataset inspired by the column schema of Mohammadta/BEAM. It is not an official BEAM release and has not received BEAM's human validation. Contents HF conversation records: 2,500 Probing questions nested in those records: 50,000 Probing questions per conversation: 20 Data shards: 25 Generation route: OpenRouter Requested model: google/gemini-3.8-flash Each row contains the eight… See the full description on the dataset page: https://huggingface.co/datasets/vm2825/BEAM_50K.

sourceHugging Facecc-by-sa-4.0updated 12d agoView on Hugging Face
0likes34downloads
Dataset Card

BEAM 50K Synthetic Compact Dataset

This is a wholly synthetic, machine-generated dataset inspired by the column schema of Mohammadta/BEAM. It is not an official BEAM release and has not received BEAM's human validation.

Contents

  • —HF conversation records: 2,500
  • —Probing questions nested in those records: 50,000
  • —Probing questions per conversation: 20
  • —Data shards: 25
  • —Generation route: OpenRouter
  • —Requested model: google/gemini-3.8-flash

Each row contains the eight published BEAM columns: conversation_id, conversation_seed, narratives, user_profile, conversation_plan, user_questions, chat, and probing_questions. probing_questions is a string-serialized dictionary containing two questions for each of ten memory abilities.

Generation usage

Usage is taken directly from OpenRouter response metadata. Totals include all paid attempts, including responses rejected by local validation.

MeasureTotalAmortized per completed conversation
Input tokens2,318,089927.24
Completion tokens (including reasoning)21,798,8268,719.53
Reasoning tokens11,4834.59
Total tokens24,116,9159,646.77
OpenRouter-reported cost$83.484164$0.033394

See metadata/openrouter_usage_summary.json and metadata/openrouter_usage_by_conversation.jsonl for details.

Limitations

These are compact synthetic conversations, not BEAM's long-context samples. All probes are machine-generated and should be human-reviewed before being used for high-stakes evaluation. Health, legal, finance, privacy, and security topics are framed as educational, organizational, or defensive assistance.