CoolFace
Datasetpublic

StableTradeAtlas/stableeval-arena

StableEval Arena StableEval Arena is a historical-replay benchmark for evaluating cost-aware, trustworthy agentic AI systems on stablecoin risk prediction. Each case gives a frozen 30-day lookback packet for USDT, USDC, or DAI and hides the next 7-day horizon labels. GitHub repository: https://github.com/SeanWan514/StableEval-Arena Hugging Face dataset: https://huggingface.co/datasets/SeanWan05/stableeval-arena Benchmark Blocks Block Role Contents… See the full description on the dataset page: https://huggingface.co/datasets/StableTradeAtlas/stableeval-arena.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes31downloads
Dataset Card

StableEval Arena

StableEval Arena is a historical-replay benchmark for evaluating cost-aware, trustworthy agentic AI systems on stablecoin risk prediction. Each case gives a frozen 30-day lookback packet for USDT, USDC, or DAI and hides the next 7-day horizon labels.

GitHub repository: https://github.com/SeanWan514/StableEval-Arena

Hugging Face dataset: https://huggingface.co/datasets/SeanWan05/stableeval-arena

Benchmark Blocks

BlockRoleContents
StableEval-120-EnrichedStress-enriched rare-risk diagnostic120 cases, six systems plus baselines, final tables, ablations
StableEval-507-NaturalNatural-distribution robustness/scaling507 cases, baselines, four native/framework-level agents, two adapter-based -Claw extensions

Task

Inputs are lookback-only prompt packets. The hidden horizon is used only for labels and evaluation.

  • Lookback: frozen 30-day historical window.
  • Horizon: hidden next 7 days.
  • Classification target: stable, watch, or stress.
  • Auxiliary targets: depeg probability and maximum absolute dollar-peg deviation in basis points.

Dataset Contents

text
data/stableeval_120_enriched/
data/stableeval_507_natural/
results/stableeval_120_enriched/
results/stableeval_507_natural/
schemas/
docs/

The package includes prompt packets, labels, canonical result tables, raw/calibrated prediction JSONL files, execution metadata, ablation summaries, stress/audit summaries, checksums, and schemas. It excludes raw vendor snapshots, full hourly-window dumps, attempt-level logs, dry-run outputs, and superseded v2-v5 progress leaderboards.

Label Distributions

  • StableEval-120-Enriched: {'stable': 54, 'stress': 21, 'watch': 45}
  • StableEval-507-Natural: {'stable': 441, 'stress': 21, 'watch': 45}

Systems Evaluated

Baselines: always_stable, historical_persistence, threshold_rule, simple_ml_softmax_logistic.

Native/framework-level agents: nanobot_qwen_507, langgraph_qwen_507, hermes_ai_qwen_507, evoagentx_qwen_507.

Adapter-based -Claw extensions: openclaw_adapter_qwen_507, nanoclaw_adapter_qwen_507.

OpenClaw/NanoClaw native CMDOP execution was not verified in this release; cite those systems as adapter-based extensions.

Main Result Files

  • 120 main leaderboard: results/stableeval_120_enriched/final_tables/validation_120_full_leaderboard.csv
  • 120 metrics: results/stableeval_120_enriched/metrics/prediction_quality_by_platform.csv
  • 507 v6 leaderboard: results/stableeval_507_natural/final_tables/stableeval_507_natural_v6_leaderboard.csv
  • 507 stress summary: results/stableeval_507_natural/stress_audits/stress_recall_summary_from_v6.csv
  • 507 cost/reliability summary: results/stableeval_507_natural/cost_logs/cost_reliability_summary_from_v6.csv

Recommended Use

  • Reproduce paper-facing saved-output evaluations.
  • Compare new agents against the released cases and labels.
  • Study cost/reliability tradeoffs and structured-output validity.
  • Audit stress recall and false alarms under a common benchmark protocol.

This dataset is not validated for live trading, financial advice, regulatory compliance, or automated treasury execution.

Limitations

  • Stress cases are rare, especially in the natural 507-case distribution.
  • The 120-case block is intentionally stress-enriched and is not a deployment prevalence estimate.
  • All released agent results use a common OpenRouter-hosted Qwen backend; this isolates platform/workflow differences but is not a broad foundation-model comparison.
  • Prompt packets and case metadata contain processed/derived market data and require redistribution review.
  • Current agents do not solve stablecoin stress prediction; weak stress recall is a central trustworthiness finding.

Citation

See CITATION.cff. No DOI is assigned in this local package; add one only after a public archival DOI exists.

License And Redistribution

See LICENSE and docs/redistribution_notice.md. Raw vendor snapshots are not included. Processed prompt packets, case metadata, and labels are included for academic reproducibility but may remain subject to upstream provider terms.