datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
general-reasoning-ift-pairs
Reasoning-IFT Pairs (General Domain)
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data.
We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.general-reasoning-ift-pairs
Reasoning-IFT Pairs (General Domain)
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data.
We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query, we… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/general-reasoning-ift-pairs.math-reasoning-ift-pairs
Reasoning-IFT Pairs (Math Domain)
Paper | Project Page
This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain).
It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data.
We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.bimodal-iftAn instruction dataset for speech->text and text->speech.
This speech data is tokenized using the SpeechTokenize approach: https://arxiv.org/abs/2308.16692.
You can do standard finetuning on this dataset using any LLM!
my_data_handsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "fr3",
"total_episodes": 8,
"total_frames": 2730,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iftekher/my_data_hands.my_data_newThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "fr3",
"total_episodes": 5,
"total_frames": 1607,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iftekher/my_data_new.my_data_handThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "fr3",
"total_episodes": 5,
"total_frames": 1227,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iftekher/my_data_hand.copa_iftgsm8k_ift_translated_nllb0dnd-dataset-improved-ift-qaIFT800k_ift
Dataset Card for "800k_ift"
More Information needed
handwriting_forms
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ift/handwriting_forms.long_sr_ift_duplicate_Qwen2.5-7B-Instructcopa_ift_translated_nllb0copa_ift_v03online-retail-II837k_iftpiqa_iftcopa_ift_v02_filteredgsm8k_ift_v02_translatedift_hhrlhf_flan
Dataset Card for "ift_hhrlhf_flan"
My favorite subsets of FLAN with single-turn data filtered from HH RLHF
flan_cats_i_like = {
"arc_challenge_10templates",
"arc_easy_10templates",
"cola_10templates",
"copa_10templates",
"coqa_10templates",
"cosmos_qa_10templates",
"fix_punct_10templates",
"math_dataset_10templates",
"natural_questions_10templates",
"openbookqa_10templates",
"squad_v2_10templates",
"trivia_qa_10templates",
}
sold-dataset-for-mistral7b-iftcopa_ift_v02copa_ift_v03_translated_nllb0self_rewarding_ift_example_raw_data1sold-dataset-for-llama3-iftift-scaffold-datasetarc_easy_ift_v3copa_ift_v02_filtered_translated
