datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
proxyquotes_library
The Proxy Quotes (pxyq) library
includes a fuction for calling the cell value with respect to a column and row of the csv dataset table. It calls for proxy stoploss distance, lotsize, and margins with leverages covering a betsize of 1 cash, commissions, swaps, spread, and more. It only have one simple function call, and that is pxyq.column('ASSET').
step 1: make sure you have pxyq.py in your directory. No need pip installations.
step 2: make an import pxyq is written on top of… See the full description on the dataset page: https://huggingface.co/datasets/algorembrant/proxyquotes_library.ecup-2026-matching-eval-proxy-v2
E-CUP 2026 retrospective evaluation proxy v2
Private team dataset for reproducing the checksum-bound retrospective contest transfer diagnostic.
Primary panel
file: soft_sample_pairs.parquet
immutable revision: 183601763fd7f4d1695325b315ef3c7cc98e67c1
immutable download: https://huggingface.co/datasets/mariklolik/ecup-2026-matching-eval-proxy-v2/resolve/183601763fd7f4d1695325b315ef3c7cc98e67c1/soft_sample_pairs.parquet
rows: 250000
SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/mariklolik/ecup-2026-matching-eval-proxy-v2.OR-synthetic-proxy
OR Scheduling Synthetic Proxy Dataset
This synthetic dataset accompanies the paper:
Decision-Focused Learning for Operating Room Scheduling Under Uncertain
Surgery Durations with Two-Stage Stochastic OptimizationAli Elaswad, Rodrigo Carrasco, Nourhan Sakr — ECML-PKDD 2026
Description
A synthetic proxy dataset generated from anonymized aggregate statistics
of a real neurosurgical dataset from the Instituto de Neurocirugia Dr.
Raul Asenjo, Santiago, Chile.… See the full description on the dataset page: https://huggingface.co/datasets/alyelaswad/OR-synthetic-proxy.proxy-mt-translations
Proxy-MT Translations
English→X machine translations generated with vLLM
across 50 open-weight LLMs on three evaluation benchmarks. This dataset holds the
raw model outputs (one CSV per model × dataset × target language); metric scores
(BLEU / chrF / COMET / MetricX) live in proxy-mt-eval-scores.
Layout
flores-200/<model>/eng-<lang>.csv # 119 target languages
ntrex/<model>/eng-<lang>.csv # 87 target languages
wmt24/<model>/eng-<lang>.csv # 51… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-translations.ai-proxy-objective-drift-detection-v0.1
Purpose
Detect when an AI system begins optimizing a proxy metricinstead of the true objective.
This is the most common early alignment failure.
What this dataset tests
proxy metric drift
reward hacking
objective–behavior decoupling
early alignment collapse
Task
Given a scenario:
Identify the true objective
Identify the proxy metric
Detect drift between them
Explain risk
Required outputs
proxy drift detection
alignment risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-proxy-objective-drift-detection-v0.1.proxy-mt-eval-scores
Proxy-MT Eval Scores
Corpus-level MT metrics for 50 open-weight LLMs on the translations in
proxy-mt-translations.
Computed by evaluate_mt.py (BLEU, chrF++, ROUGE-L, METEOR, XCOMET-XL, SSA-COMET).
MetricX is backfilled separately and may still be empty in this snapshot.
Layout
<model>/flores-200.csv
<model>/ntrex.csv
<model>/wmt24.csv
Each CSV has one row per eng-<lang> pair:
column
description
translation-pair
e.g. eng-yor
bleu
sacrebleu corpus BLEU… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-eval-scores.meta-proxy-to-outcome-control-medicine-v0.3
Proxy-to-Outcome Control in Medicine
Meta Dataset v0.3
Purpose
This dataset tests whether a model:
Treats proxies as proxies
Avoids upgrading signals into outcomes
Maintains causal boundaries under incomplete evidence
Resists reassurance based on measurable movement alone
You are testing inference discipline.
Why this matters
Medicine is proxy-dense.
Biomarkers, scores, and early trends move all the time.
Unsafe systems convert that movement… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/meta-proxy-to-outcome-control-medicine-v0.3.sqli-rce-proxy1
