ift
Datasets
All datasets matching “ift”general-reasoning-ift-pairs
Reasoning-IFT Pairs (General Domain)
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data.
We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.general-reasoning-ift-pairs
Reasoning-IFT Pairs (General Domain)
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data.
We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query, we… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/general-reasoning-ift-pairs.math-reasoning-ift-pairs
Reasoning-IFT Pairs (Math Domain)
Paper | Project Page
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain).
It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data.
We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.ift-eval-us-real
Eval IFT SD 1.5 — foto asli, US, gender
Gambar hasil generate dan label Gemini untuk arm foto asli (fassabilf/sd15-ift-us-real,
dataset train fassabilf/ift-train-us-real). Bukan run sintetis — itu ada di
fassabilf/ift-test-us-sd15.
split
checkpoint
protokol
n per okupasi
test_ep*
ep5, 10, 15, 20, 25, 30
pilot_sd15 apa adanya: varian framing, --seed-mode paired --seed-base 0, --max-attempts 3, batch 8
100
val_ep*
ep2..ep30 tiap 2 epoch
sama, tapi --max-attempts 1… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/ift-eval-us-real.bimodal-iftAn instruction dataset for speech->text and text->speech.
This speech data is tokenized using the SpeechTokenize approach: https://arxiv.org/abs/2308.16692.
You can do standard finetuning on this dataset using any LLM!
IFT-Data-For-Tabular-Tasks



