datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qvhighlight_internvideo2_llama_text_featuredalle3-llama3.2-11b
Dataset Card for dalle3-llama3.2-11b
Dataset Summary
This is 3,577,716 new synthetic captions for the 1,192,572 images found in ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions. The dataset was filtered for duplicates and then re-encoded with JPEGXL lossless or lossy depending on the source. The long captions were produced using meta-llama/Llama-3.2-11B-Vision-Instruct. Medium and short captions were produced from these captions using… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/dalle3-llama3.2-11b.GPQA_verifications_GenRM-Base_Llama-3.3-70B-Instructcharade_sta_internvideo2_llama_text_featureMATH128_verifications_GenRM-FT_Llama-3.1-8B-InstructMATH128_Solutions_Llama-3.1-8B-Instructcompressed_verifications_lcb128_llama-3.3-70b_genrm_baseMATH128_Solutions_Llama-3.3-70B-Instructlcb_llama_70B_256_32LCB128_Llama3.1-8B-Inst_LLM-as-judge_256_32verifications_llama3.1_8b_finetuned_verifiermath128_qwen2p57B_VER-Llama-3.1-8B-Inst_data-qwen_25_7b_gpt_4o_verify_train_e3_LR-5e-7_7Klenlcb128_llama3.3_70b_verificationsverifications_aime24_llama_3.1_8b_finetuned_verifier_4Ktokenslcb_llama70B_128_32MATH128_verifications_Llama-3.3-70B-Instruct_GenRM-Basemath128_qwen2.5-7b_ver-llama3.1-8b_train_gpt_4o_verifications_e3_lr5e-7-31389-merged_4Ktokensaime24_qwen2.5-7b_ver-llama3.1-8b_train_gpt_4o_verifications_e3_lr5e-7-31389-merged_4Ktokensaime24_qwen2.5-7b_ver_Llama-3.1-8B-Instruct_data-qwen_25_7b_gpt_4o_verify_train_e3_LR-5e-7_7Klenlcb128_llama3-8B-instruct_256samples_ver32_temp0-7aime24_verifications_qwen_2.5_7B_solutions_llama_3.1_8b_finetuned_verifier_4Klcb128_llama8b_verifications_4096tokensmmlu-chat-Llama-3.2-Instructbenchmark-chat-Llama-3.2-Instructimagenet-llamagen-cache
