hubnemo/tulu3-sft-mini
This is a subset derived from tulu3 sft mixture limited to 20 samples for each source. The use case for this smaller dataset is to have a short, consistent evaluation dataset over different domains for multi-token prediction. Here's the code for how to derive this dataset: import datasets DATASET = "allenai/tulu-3-sft-mixture" OFFSETS = [ ("ai2-adapt-dev/oasst1_converted", 0, 7131), ("ai2-adapt-dev/flan_v2_converted", 7131, 97113), ("ai2-adapt-dev/tulu_hard_coded_repeated_10"… See the full description on the dataset page: https://huggingface.co/datasets/hubnemo/tulu3-sft-mini.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face