datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-supervised-datasetSelf-Supervised_RLThis repository contains the dataset and resources related to the paper Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following.
The paper introduces a self-supervised reinforcement learning (RL) framework that improves instruction following capabilities of reasoning models by leveraging their internal signals, without requiring external supervision. This approach aims to address the trade-off between reasoning and instruction following, offering a… See the full description on the dataset page: https://huggingface.co/datasets/dd12345789/Self-Supervised_RL.gsm8k-supervised-uncertainty-cache
GSM8K supervised uncertainty cache
Frozen feature cache used by the LM-Polygraph supervised-uncertainty seminar.
It is provided so the notebook can run without regenerating model outputs.
Contents
gsm8k_native_all_layer_compact_v1_200_100_100_c4_20_source701.joblib contains one compressed joblib payload
(835485642 bytes; SHA-256 9ed7f5beb05b6055a5a8ba04da16c086bb827b452e681f98ceee4189f8413266) with:
200 GSM8K training records and 100 disjoint GSM8K test records;… See the full description on the dataset page: https://huggingface.co/datasets/ArtemVazhentsev21/gsm8k-supervised-uncertainty-cache.
