hubnemo/tulu3-sft-mini
This is a subset derived from tulu3 sft mixture limited to 20 samples for each source. The use case for this smaller dataset is to have a short, consistent evaluation dataset over different domains for multi-token prediction. Here's the code for how to derive this dataset: import datasets DATASET = "allenai/tulu-3-sft-mixture" OFFSETS = [ ("ai2-adapt-dev/oasst1_converted", 0, 7131), ("ai2-adapt-dev/flan_v2_converted", 7131, 97113), ("ai2-adapt-dev/tulu_hard_coded_repeated_10"… See the full description on the dataset page: https://huggingface.co/datasets/hubnemo/tulu3-sft-mini.
0110
Update README.md
Update README.md
Update README.md
Create README.md
Upload folder using huggingface_hub
initial commit
