datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kNN-Targets-wikipedia-mistral
Dataset Overview
This dataset provides k-nearest neighbor (kNN) target distributions for language modeling. Each token in the Wikipedia corpus is associated with a soft probability distribution over its top-k nearest neighbors in the representation space of a frozen language model. These targets can be used to train MLP Memory.
Corresponding Preprocessed Corpus: Rubin-Wei/enwiki-dec2021-preprocessed-mistral
Compatible Model: Mistral-7B-v0.3
Paper: MLP Memory: A Retriever-Pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/kNN-Targets-wikipedia-mistral.mistralai-tekken-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
qrpo-paper-mistral-sft-ultrafeedback-armorm-temp1-ref50-offline-armorm
qrpo-paper-mistral-sft-ultrafeedback-armorm-temp1-ref50-offline-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2best-armorm
qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2best-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offline-armorm
qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offline-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
qrpo-paper-mistral-nosft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
qrpo-paper-mistral-nosft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2random-armorm
qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2random-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
qrpo-paper-mistral-sft-ultrafeedback-armorm-temp1-ref50-offpolicy2best-armorm
qrpo-paper-mistral-sft-ultrafeedback-armorm-temp1-ref50-offpolicy2best-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
qrpo-paper-mistral-nosft-magpieair-armorm-temp1-ref50-offline-armorm
qrpo-paper-mistral-nosft-magpieair-armorm-temp1-ref50-offline-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
qrpo-paper-mistral-sft-ultrafeedback-armorm-temp1-ref50-offpolicy2random-armorm
qrpo-paper-mistral-sft-ultrafeedback-armorm-temp1-ref50-offpolicy2random-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offline-armorm
qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offline-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
openhermes-dev__mistralai_Mixtral-8x7B-Instruct-v0.1__1707245027mistral-675b-eval-logs-and-scoresdetails_aws-prototyping__MegaBeam-Mistral-7B-512k
Dataset Card for Evaluation run of aws-prototyping/MegaBeam-Mistral-7B-512k
Dataset automatically created during the evaluation run of model aws-prototyping/MegaBeam-Mistral-7B-512k.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_aws-prototyping__MegaBeam-Mistral-7B-512k.qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2random-armorm
qrpo-paper-mistral-nosft-ultrafeedback-armorm-temp1-ref50-offpolicy2random-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
10k_prompts_ranked_mistral_large_responses
Description
This dataset contains responses generated for the prompts of the DIBT/10k_prompts_ranked, using distilabel
with mistral-large. The script used for the generation can be seen at the repository: generate_reference_spin.py.
Ultrafeedback-mistral-ddo-selection-iteration2-4-responseslm-eval-results-chihoonlee10-T3Q-Mistral-Orca-Math-DPO-private
Dataset Card for Evaluation run of chihoonlee10/T3Q-Mistral-Orca-Math-DPO
Dataset automatically created during the evaluation run of model chihoonlee10/T3Q-Mistral-Orca-Math-DPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chihoonlee10-T3Q-Mistral-Orca-Math-DPO-private.sl-multi-embeddings-results-40-mistral-selfMMInstruct-GPT4V_mistral-7b_l0_cutlm-eval-results-chlee10-T3Q-Merge-Mistral7B-private
Dataset Card for Evaluation run of chlee10/T3Q-Merge-Mistral7B
Dataset automatically created during the evaluation run of model chlee10/T3Q-Merge-Mistral7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chlee10-T3Q-Merge-Mistral7B-private.statecraft-sft-v3
Statecraft SFT data v3
Statecraft (host-injected behavior snapshot) SFT rows, one file per category. Chinese and English only.
Row = {id, category, messages, tools, prefix_len}; a user message may carry a snapshot object, which
chat_template.jinja expands to <snapshot>{json}</snapshot> in a system block right before that
turn (see handmade_example.jsonl). mix_v4/ is the training mix (with statecraft-general-selfroll),
quality/ the QA reports.
category
rows
what it covers… See the full description on the dataset page: https://huggingface.co/datasets/mistral0105/statecraft-sft-v3.MMInstruct-GPT4V_mistral-7b_cooccur_cutMMInstruct-GPT4V_mistral-7b_cosi_cutqrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
mistral_nemo_base-mmlu-valalqa-multi-embeddings-results-40-mistral-selflm-eval-results-nbeerbower-bophades-mistral-truthy-DPO-7B-private
Dataset Card for Evaluation run of nbeerbower/bophades-mistral-truthy-DPO-7B
Dataset automatically created during the evaluation run of model nbeerbower/bophades-mistral-truthy-DPO-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-bophades-mistral-truthy-DPO-7B-private.mistral8x22b-features-redditdetails_cognitivecomputations__dolphin-2.9.3-mistral-nemo-12b
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.3-mistral-nemo-12b
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.3-mistral-nemo-12b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_cognitivecomputations__dolphin-2.9.3-mistral-nemo-12b.
