datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026.transcoder-adapters.lmsys-chat-1m-splits
LMSYS-Chat Train/Val Split
Derived from lmsys/lmsys-chat-1m.
Methodology
This dataset was created by excluding all LMSYS rows that were used in a
prior training run, then splitting the remaining rows into train and val sets.
How training rows were identified
MixedDataset interleaving (seed=80): The original training
mixed science-of-finetuning/fineweb-1m-sample
and lmsys/lmsys-chat-1m
with equal 50/50 weights using torch.multinomial + per-dataset… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.transcoder-adapters.lmsys-chat-1m-splits.2026.transcoder-adapters.templated_chats.lmsys_lmsys-chat-1mtranscoder-adapters-openthoughts3-stratified-55k
