datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
proxy-listproxy-logs-ShareGPTAll the proxy logs I could find (lmk if there are more), converted to ShareGPT so it's all formatted the same way in one dataset. I didn't .strip() any of the turns, I only used ftfy.fix_text() on them, and skipped any empty turns.
The original sample response is moved to be the final turn of conversations. If the last turn in the original sample prompt was a model turn, it is assumed that it was for doing prefill. So I moved this turn to response_prefill.
sample["response_prefill"] and… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/proxy-logs-ShareGPT.greater-london-tx-proxy-rat-path-gain
Greater London Per-Transmitter-Proxy, Per-RAT Simulated Path-Gain Dataset
Short display name: Greater London Tx-Proxy × RAT Path-GainChinese name: 大伦敦逐发射代理、逐 RAT 模拟路径增益数据集Release: v9 final release (COMPLETE)
City-scale propagation, one transmitter proxy at a time.
This release turns Greater London into a queryable radio-propagation dataset:
449,437,201,731 simulated path-gain relations connect 22,678 computed
transmitter-proxy hypotheses with 171,549,960 receiver faces across… See the full description on the dataset page: https://huggingface.co/datasets/EEzim/greater-london-tx-proxy-rat-path-gain.ProxyCoT-HotpotQAThis is the HotpotQA data that we used in our ProxyCoT project (https://aclanthology.org/2026.acl-long.1917/), and it is based on long-context reasoning (32K-128K tokens).
HotpotQA here is a new version originally from https://aclanthology.org/2026.acl-long.1917/ with extended contexts.
For more details on the context extension, refer to the ProxyCoT paper.
To use our dataset, please follow the code below.
train_samples = load_dataset("oaimli/proxycot-hotpotqa", split="train")
dev_samples =… See the full description on the dataset page: https://huggingface.co/datasets/oaimli/ProxyCoT-HotpotQA.ProxyCoT-SciTrekThis is the SciTrek data that we used in the ProxyCoT project, and it is based on long-context reasoning (32K-128K tokens).
SciTrek is originally from https://arxiv.org/abs/2509.21028.
To use our dataset, please follow the code below.
train_samples = load_dataset("oaimli/proxycot-scitrek", split="train")
dev_samples = load_dataset("oaimli/proxycot-scitrek", split="val")
test_samples = load_dataset("oaimli/proxycot-scitrek", split="test")
for sample in train_samples:
question =… See the full description on the dataset page: https://huggingface.co/datasets/oaimli/ProxyCoT-SciTrek.proxy-logs-ReRolls-Minos
non-refusal responses: 845,186
refusal responses: 38,285
https://gist.github.com/xzuyn/1d7f43db2750060a18a304eb84b396db
Use a training prompt formatter like this: https://github.com/xzuyn/axolotl/blob/latest-formatters/src/axolotl/prompt_strategies/customllama3-regex-last-only-prefill-reroll.py
proxy-logs-ReRollsDuplicate prompts combined into a single sample, with all responses in a list of dicts. I've also included some info like token count, and slop (though my slop list could use improvement).
Use a training prompt formatter like this: https://github.com/xzuyn/axolotl/blob/84aec029dfa9eb9670b8a51d432a279be6c85871/src/axolotl/prompt_strategies/customllama3-regex-last-only-prefill-reroll.py
Dataset creation script: https://gist.github.com/xzuyn/aa1f30b7394d2997766bef82edb67227
proxy-ip-pricing-cn
Proxy IP Pricing (China Market) 2026
An open dataset of proxy IP pricing, protocol support, coverage and official registration links for 18 providers serving the Chinese market. Compiled monthly from provider-published price sheets.
Maintained by 全网低价IP / socks5ip — a comparison platform aggregating 20+ proxy IP providers.
Why this dataset exists
Proxy IP pricing is scattered across dozens of provider sites, quoted in different units (per day / per week / per… See the full description on the dataset page: https://huggingface.co/datasets/socks5ip/proxy-ip-pricing-cn.proxyquotes_library
The Proxy Quotes (pxyq) library
includes a fuction for calling the cell value with respect to a column and row of the csv dataset table. It calls for proxy stoploss distance, lotsize, and margins with leverages covering a betsize of 1 cash, commissions, swaps, spread, and more. It only have one simple function call, and that is pxyq.column('ASSET').
step 1: make sure you have pxyq.py in your directory. No need pip installations.
step 2: make an import pxyq is written on top of… See the full description on the dataset page: https://huggingface.co/datasets/algorembrant/proxyquotes_library.proxy-ip-pricing-cn-2026
Proxy IP Pricing (China Market) 2026
An open dataset of proxy IP pricing, protocol support, coverage and official registration links for 18 providers serving the Chinese market. Compiled monthly from provider-published price sheets.
Maintained by 全网低价IP / socks5ip — a comparison platform aggregating 20+ proxy IP providers.
Why this dataset exists
Proxy IP pricing is scattered across dozens of provider sites, quoted in different units (per day / per week / per… See the full description on the dataset page: https://huggingface.co/datasets/socks5ip/proxy-ip-pricing-cn-2026.OR-synthetic-proxy
OR Scheduling Synthetic Proxy Dataset
This synthetic dataset accompanies the paper:
Decision-Focused Learning for Operating Room Scheduling Under Uncertain
Surgery Durations with Two-Stage Stochastic OptimizationAli Elaswad, Rodrigo Carrasco, Nourhan Sakr — ECML-PKDD 2026
Description
A synthetic proxy dataset generated from anonymized aggregate statistics
of a real neurosurgical dataset from the Instituto de Neurocirugia Dr.
Raul Asenjo, Santiago, Chile.… See the full description on the dataset page: https://huggingface.co/datasets/alyelaswad/OR-synthetic-proxy.proxy-mt-translations
Proxy-MT Translations
English→X machine translations generated with vLLM
across 50 open-weight LLMs on three evaluation benchmarks. This dataset holds the
raw model outputs (one CSV per model × dataset × target language); metric scores
(BLEU / chrF / COMET / MetricX) live in proxy-mt-eval-scores.
Layout
flores-200/<model>/eng-<lang>.csv # 119 target languages
ntrex/<model>/eng-<lang>.csv # 87 target languages
wmt24/<model>/eng-<lang>.csv # 51… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-translations.mistral-base-proxy-reward-ultrafeedbackProxy-Logs-G4-26B-KCPPPJMixers-Dev/proxy-logs-ReRolls-Minos with responses generated using non-thinking bartowski/google_gemma-4-26B-A4B-it-GGUF/google_gemma-4-26B-A4B-it-Q4_K_M.gguf.
Used default recommended generation settings: temp=1, top_k=64, top_p=0.95.
ai-proxy-objective-drift-detection-v0.1
Purpose
Detect when an AI system begins optimizing a proxy metricinstead of the true objective.
This is the most common early alignment failure.
What this dataset tests
proxy metric drift
reward hacking
objective–behavior decoupling
early alignment collapse
Task
Given a scenario:
Identify the true objective
Identify the proxy metric
Detect drift between them
Explain risk
Required outputs
proxy drift detection
alignment risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-proxy-objective-drift-detection-v0.1.asia-unsdg-proxy-progress-toward-productive-and-sustainable-agricu-ag-lnd-sust-prxcssproxy-tuninguf_test_with_proxy_reward_without_lenpenaltyproxy-mt-eval-scores
Proxy-MT Eval Scores
Corpus-level MT metrics for 50 open-weight LLMs on the translations in
proxy-mt-translations.
Computed by evaluate_mt.py (BLEU, chrF++, ROUGE-L, METEOR, XCOMET-XL, SSA-COMET).
MetricX is backfilled separately and may still be empty in this snapshot.
Layout
<model>/flores-200.csv
<model>/ntrex.csv
<model>/wmt24.csv
Each CSV has one row per eng-<lang> pair:
column
description
translation-pair
e.g. eng-yor
bleu
sacrebleu corpus BLEU… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/proxy-mt-eval-scores.claude-proxy-conversations2meta-proxy-to-outcome-control-medicine-v0.3
Proxy-to-Outcome Control in Medicine
Meta Dataset v0.3
Purpose
This dataset tests whether a model:
Treats proxies as proxies
Avoids upgrading signals into outcomes
Maintains causal boundaries under incomplete evidence
Resists reassurance based on measurable movement alone
You are testing inference discipline.
Why this matters
Medicine is proxy-dense.
Biomarkers, scores, and early trends move all the time.
Unsafe systems convert that movement… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/meta-proxy-to-outcome-control-medicine-v0.3.ledgar_proxydataafrica-unsdg-proxy-progress-toward-productive-and-sustainable-agricu-ag-lnd-sust-prxts
Africa Unsdg Proxy Progress Toward Productive and Sustainable Agricu Ag Lnd Sust Prxts | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-unsdg-proxy-progress-toward-productive-and-sustainable-agricu-ag-lnd-sust-prxts.sqli-rce-proxy1asia-unsdg-proxy-progress-toward-productive-and-sustainable-agricu-ag-lnd-sust-prxtsuf_test_with_proxy_reward_w_lengthtranslation-proxy-paper-translations
Translation as a Scalable Proxy for Multilingual Evaluation (Raw MT Data)
This repository contains the Raw Machine Translation Predictions generated for the paper: "Translation as a Scalable Proxy for Multilingual Evaluation" (Issaka et al., 2026).
If you are looking for the aggregated evaluation scores (LM-Eval + MT Metrics), please see our companion repository: 👉 Link to Benchmark Scores Repo.
Dataset Description
The rapid proliferation of LLMs has created a… See the full description on the dataset page: https://huggingface.co/datasets/marslabucla/translation-proxy-paper-translations.alpaca_test_proxy_without_lengthpess_uf_proxy_wo_len_w_bold_list
Dataset Card for "pess_uf_proxy_wo_len_w_bold_list"
More Information needed
Natural_Instruction_proxydata
