llama-processed
oasst1-processed_Llama-2-7b-chat-hf_KTO_lora_oasst1-processed_randomSub_all_1_0.3_28334_mergedllama-2-7b-td-academy-processedllama-2-7b-chat-dataset-processedDeepSeek-R1-Distill-Llama-70B-medical-o1-reasoning-en-mix-processed-1cecaf7c-a1cf-4978-8f3b-42afLlama-3.1-8B-Instruct-SEvolve2_re_30k_tag5_processed_llamav2_systemllama-3-qlora-vicuna-processed-indicator-0.6llama-3-qlora-wizard-processed-indicator-0.6Llama-3.1-8B-Ins_data-distill_r1_qwen_1p5B_gpt_4o_verify_processed_train_e6_LR-1e-5_6Klen
Llama-HybridDiffusion-processed-data-run1
Llama-HybridDiffusion processed training mixture — run 1
Built with Llama.
This repository preserves the exact Hugging Face Dataset.save_to_disk Arrow snapshot
used by run 1 of a Qwen3.5-2B HybridDiffusion reproduction. The directory names,
dataset_info.json, state.json, and Arrow shard boundaries are retained so the data
can be downloaded and supplied to the existing training configuration without a lossy
format conversion.
Exact snapshot inventory
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arushhh/Llama-HybridDiffusion-processed-data-run1.processed_llama_dataset_2048refine-book-wiki_processed_llama_dataset_2048openweb_processed_llama_dataset_2048Llama-4-Scout-17B-16E-Instruct-FP8-Instruct-FP8-ProcessedOpenAssistantllama_3.1_binary_train_processed
