CoolFace
Datasetpublic

hubnemo/tulu3-sft-mini

This is a subset derived from tulu3 sft mixture limited to 20 samples for each source. The use case for this smaller dataset is to have a short, consistent evaluation dataset over different domains for multi-token prediction. Here's the code for how to derive this dataset: import datasets DATASET = "allenai/tulu-3-sft-mixture" OFFSETS = [ ("ai2-adapt-dev/oasst1_converted", 0, 7131), ("ai2-adapt-dev/flan_v2_converted", 7131, 97113), ("ai2-adapt-dev/tulu_hard_coded_repeated_10"… See the full description on the dataset page: https://huggingface.co/datasets/hubnemo/tulu3-sft-mini.

sourceHugging Faceupdated 24d agoView on Hugging Face
0likes110downloads
6 commits on main
2dca7e324d ago

Update README.md

hubnemo
4eb836924d ago

Update README.md

hubnemo
527b36724d ago

Update README.md

hubnemo
3e1c1e324d ago

Create README.md

hubnemo
2ed6c0125d ago

Upload folder using huggingface_hub

hubnemo
ff554f925d ago

initial commit

hubnemo