data-juicer
HumanVBenchA Huggingface leaderboard is coming soon. More details (e.g., evaluation script) can be found in https://github.com/modelscope/data-juicer/tree/HumanVBench
Huggingface排行榜即将发布~ 更多细节(如测评代码和方法)请参见 https://github.com/modelscope/data-juicer/tree/HumanVBench
VeriSciQA
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
Paper: VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
Dataset Description
VeriSciQA is a large-scale, high-quality dataset for Scientific Visual Question Answering (SVQA), containing 20,272 QA pairs spanning 20 scientific domains, 12 figure types, and 5 question types. The dataset is constructed using a Cross-Modal Verification framework that generates QA pairs from… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/VeriSciQA.Img-Diff
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
We release Img-Diff, A high-quality synthesis dataset focusing on describing object differences for MLLMs. See more details in our paper and code.
Abstract: High-performance Multimodal Large Language Models (MLLMs) rely heavily on data quality. This study introduces a novel dataset named Img-Diff, designed to enhance fine-grained image recognition in MLLMs by leveraging insights from contrastive learning and… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/Img-Diff.the-pile-pubmed-abstracts-refined-by-data-juicer
The Pile -- PubMed Abstracts (refined by Data-Juicer)
A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G).
Dataset Information
Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.DetailMaster
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
We introduce DetailMaster, a benchmark designed to evaluate text-to-image generation in long-prompt scenarios, accompanied by a robust fine-grained evaluation protocol. See more details in our paper.
Abstract: While recent text-to-image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, their performance significantly degrades when confronted with long, detail-intensive prompts… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/DetailMaster.Trinity-ToolAce-SFT-split
