CoolFace
20 results

data-juicer

datajuicer /HumanVBenchA Huggingface leaderboard is coming soon. More details (e.g., evaluation script) can be found in https://github.com/modelscope/data-juicer/tree/HumanVBench Huggingface排行榜即将发布~ 更多细节(如测评代码和方法)请参见 https://github.com/modelscope/data-juicer/tree/HumanVBench video1K<n<10K3 likes511 downloads10mo agoHugging Facedatajuicer /VeriSciQA VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering Paper: VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering Dataset Description VeriSciQA is a large-scale, high-quality dataset for Scientific Visual Question Answering (SVQA), containing 20,272 QA pairs spanning 20 scientific domains, 12 figure types, and 5 question types. The dataset is constructed using a Cross-Modal Verification framework that generates QA pairs from… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/VeriSciQA.imagevisual-question-answering10K<n<100K0 likes267 downloads8mo agoHugging Facedatajuicer /Img-Diff Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models We release Img-Diff, A high-quality synthesis dataset focusing on describing object differences for MLLMs. See more details in our paper and code. Abstract: High-performance Multimodal Large Language Models (MLLMs) rely heavily on data quality. This study introduces a novel dataset named Img-Diff, designed to enhance fine-grained image recognition in MLLMs by leveraging insights from contrastive learning and… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/Img-Diff.6 likes137 downloads10mo agoHugging Facedatajuicer /the-pile-pubmed-abstracts-refined-by-data-juicer The Pile -- PubMed Abstracts (refined by Data-Juicer) A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.texttext-generationn<1K3 likes113 downloads3y agoHugging Facedatajuicer /DetailMaster DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? We introduce DetailMaster, a benchmark designed to evaluate text-to-image generation in long-prompt scenarios, accompanied by a robust fine-grained evaluation protocol. See more details in our paper. Abstract: While recent text-to-image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, their performance significantly degrades when confronted with long, detail-intensive prompts… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/DetailMaster.texttext-to-image1K<n<10K3 likes63 downloads1y agoHugging Facedatajuicer /Trinity-ToolAce-SFT-splittextn<1K0 likes58 downloads1y agoHugging Face