CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datajuicer /HumanVBenchA Huggingface leaderboard is coming soon. More details (e.g., evaluation script) can be found in https://github.com/modelscope/data-juicer/tree/HumanVBench Huggingface排行榜即将发布~ 更多细节(如测评代码和方法)请参见 https://github.com/modelscope/data-juicer/tree/HumanVBench video1K<n<10K3 likes511 downloads10mo agoHugging Face02datajuicer /VeriSciQA VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering Paper: VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering Dataset Description VeriSciQA is a large-scale, high-quality dataset for Scientific Visual Question Answering (SVQA), containing 20,272 QA pairs spanning 20 scientific domains, 12 figure types, and 5 question types. The dataset is constructed using a Cross-Modal Verification framework that generates QA pairs from… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/VeriSciQA.imagevisual-question-answering10K<n<100K0 likes267 downloads8mo agoHugging Face03datajuicer /Img-Diff Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models We release Img-Diff, A high-quality synthesis dataset focusing on describing object differences for MLLMs. See more details in our paper and code. Abstract: High-performance Multimodal Large Language Models (MLLMs) rely heavily on data quality. This study introduces a novel dataset named Img-Diff, designed to enhance fine-grained image recognition in MLLMs by leveraging insights from contrastive learning and… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/Img-Diff.6 likes137 downloads10mo agoHugging Face04datajuicer /the-pile-pubmed-abstracts-refined-by-data-juicer The Pile -- PubMed Abstracts (refined by Data-Juicer) A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.texttext-generationn<1K3 likes113 downloads3y agoHugging Face05datajuicer /DetailMaster DetailMaster: Can Your Text-to-Image Model Handle Long Prompts? We introduce DetailMaster, a benchmark designed to evaluate text-to-image generation in long-prompt scenarios, accompanied by a robust fine-grained evaluation protocol. See more details in our paper. Abstract: While recent text-to-image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, their performance significantly degrades when confronted with long, detail-intensive prompts… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/DetailMaster.texttext-to-image1K<n<10K3 likes63 downloads1y agoHugging Face06datajuicer /Trinity-ToolAce-SFT-splittextn<1K0 likes58 downloads1y agoHugging Face07datajuicer /the-pile-pubmed-central-refined-by-data-juicer The Pile -- PubMed Central (refined by Data-Juicer) A refined version of PubMed Central dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 83G). Dataset Information Number of samples: 2,694,860 (Keep ~86.96% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-central-refined-by-data-juicer.texttext-generationn<1K2 likes50 downloads3y agoHugging Face08datajuicer /data-juicer-t2v-optimal-data-pool Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development Project description The emergence of large-scale multi-modal generative models has drastically advanced artificial intelligence, introducing unprecedented levels of performance and functionality. However, optimizing these models remains challenging due to historically isolated paths of model-centric and data-centric developments, leading to suboptimal outcomes and inefficient resource… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/data-juicer-t2v-optimal-data-pool.texttext-to-videon<1K0 likes50 downloads2y agoHugging Face09datajuicer /RealMedConv Dataset Description The RealMedConv dataset consists of anonymized, real-world dialogues between licensed pharmacists and users seeking over-the-counter (OTC) medication advice. Each conversation is goal-oriented: the pharmacist gathers sufficient symptom information to provide an online and appropriate recommendation. Dialogues are typically concise, spanning 3–5 turns, reflecting the efficient and expert-driven nature of professional medical consultations. This dataset originates… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/RealMedConv.text1K<n<10K1 likes44 downloads11mo agoHugging Face10datajuicer /MindGYM Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. textquestion-answering1K<n<10K0 likes42 downloads1y agoHugging Face11datajuicer /alpaca-cot-zh-refined-by-data-juicer Alpaca-CoT -- ZH (refined by Data-Juicer) A refined Chinese version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to fine-tune a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 18.7GB). Dataset Information Number of samples: 9,873,214 (Keep ~46.58% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-zh-refined-by-data-juicer.texttext-generationn<1K5 likes37 downloads3y agoHugging Face12datajuicer /llava-pretrain-refined-by-data-juicer LLaVA pretrain -- LCS-558k (refined by Data-Juicer) A refined version of LLaVA pretrain dataset (LCS-558k) by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Multimodal Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 115MB). Dataset Information Number of samples: 500,380 (Keep ~89.65% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/llava-pretrain-refined-by-data-juicer.imageimage-to-textn<1K2 likes30 downloads3y agoHugging Face13datajuicer /the-pile-uspto-refined-by-data-juicer The Pile -- USPTO (refined by Data-Juicer) A refined version of USPTO dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 18G). Dataset Information Number of samples: 4,516,283 (Keep ~46.77% from the original dataset) Refining Recipe #… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-uspto-refined-by-data-juicer.texttext-generationn<1K0 likes28 downloads3y agoHugging Face14datajuicer /geometry_sft Overview This dataset is a Supervised Fine-Tuning (SFT) dataset generated from a subset of the Geometry3K dataset using Qwen2.5-VL. It serves as an example dataset for demonstrating VLM (Vision-Language Model) SFT training in the Trinity-RFT library. imagen<1K0 likes28 downloads11mo agoHugging Face15datajuicer /redpajama-arxiv-refined-by-data-juicer RedPajama -- ArXiv (refined by Data-Juicer) A refined version of ArXiv dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 85GB). Dataset Information Number of samples: 1,655,259 (Keep ~95.99% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-arxiv-refined-by-data-juicer.texttext-generationn<1K2 likes26 downloads3y agoHugging Face16datajuicer /redpajama-cc-2022-05-refined-by-data-juicer RedPajama -- CommonCrawl-2022-05 (refined by Data-Juicer) A refined version of CommonCrawl-2022-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 265GB). Dataset Information Number of samples: 42,648,496 (Keep ~45.34% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2022-05-refined-by-data-juicer.tabulartext-generationn<1K0 likes25 downloads3y agoHugging Face17datajuicer /alpaca-cot-en-refined-by-data-juicer Alpaca-CoT -- EN (refined by Data-Juicer) A refined English version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to fine-tune a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 226GB). Dataset Information Number of samples: 72,855,345 (Keep ~54.48% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-en-refined-by-data-juicer.texttext-generationn<1K0 likes25 downloads3y agoHugging Face18datajuicer /redpajama-pile-stackexchange-refined-by-data-juicer RedPajama & The Pile -- StackExchange (refined by Data-Juicer) A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB). Dataset Information Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.texttext-generationn<1K0 likes20 downloads3y agoHugging Face19datajuicer /redpajama-cc-2019-30-refined-by-data-juicer RedPajama -- CommonCrawl-2019-30 (refined by Data-Juicer) A refined version of CommonCrawl-2019-30 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 240GB). Dataset Information Number of samples: 36,557,283 (Keep ~45.08% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2019-30-refined-by-data-juicer.tabulartext-generationn<1K0 likes19 downloads3y agoHugging Face20datajuicer /redpajama-wiki-refined-by-data-juicer RedPajama -- Wikipedia (refined by Data-Juicer) A refined version of Wikipedia dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 68GB). Dataset Information Number of samples: 26,990,659 (Keep ~90.47% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-wiki-refined-by-data-juicer.texttext-generationn<1K2 likes18 downloads3y agoHugging Face21datajuicer /the-pile-europarl-refined-by-data-juicer The Pile -- EuroParl (refined by Data-Juicer) A refined version of EuroParl dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 2.2GB). Dataset Information Number of samples: 61,601 (Keep ~88.23% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-europarl-refined-by-data-juicer.texttext-generationn<1K0 likes15 downloads3y agoHugging Face22datajuicer /redpajama-stack-code-refined-by-data-juicer RedPajama & TheStack -- Github Code (refined by Data-Juicer) A refined version of Github Code dataset in RedPajama & TheStack by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 232GB). Dataset Information Number of samples: 49,279,344 (Keep ~52.09% from the original… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-stack-code-refined-by-data-juicer.tabulartext-generationn<1K2 likes13 downloads3y agoHugging Face23datajuicer /redpajama-book-refined-by-data-juicer RedPajama -- Book (refined by Data-Juicer) A refined version of Book dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 91GB). Dataset Information Number of samples: 195,983 (Keep ~95.51% from the original dataset) Refining Recipe #… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-book-refined-by-data-juicer.text-generation100K<n<1M5 likes13 downloads3y agoHugging Face24datajuicer /the-pile-hackernews-refined-by-data-juicer The Pile -- HackerNews (refined by Data-Juicer) A refined version of HackerNews dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.8G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-hackernews-refined-by-data-juicer.texttext-generationn<1K0 likes13 downloads3y agoHugging Face25datajuicer /the-pile-philpaper-refined-by-data-juicer The Pile -- PhilPaper (refined by Data-Juicer) A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB). Dataset Information Number of samples: 29,117 (Keep ~88.82% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-philpaper-refined-by-data-juicer.texttext-generationn<1K0 likes12 downloads3y agoHugging Face26datajuicer /redpajama-cc-2020-05-refined-by-data-juicer RedPajama -- CommonCrawl-2020-05 (refined by Data-Juicer) A refined version of CommonCrawl-2020-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 297GB). Dataset Information Number of samples: 42,612,596 (Keep ~46.90% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2020-05-refined-by-data-juicer.tabulartext-generationn<1K0 likes12 downloads3y agoHugging Face27datajuicer /Trinity-ToolAce-RL-splittext1K<n<10K0 likes11 downloads1y agoHugging Face28datajuicer /redpajama-cc-2023-06-refined-by-data-juicer RedPajama -- CommonCrawl-2023-06 (refined by Data-Juicer) A refined version of CommonCrawl-2023-06 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 310GB). Dataset Information Number of samples: 50,643,699 (Keep ~45.46% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2023-06-refined-by-data-juicer.tabulartext-generationn<1K0 likes8 downloads3y agoHugging Face29datajuicer /the-pile-freelaw-refined-by-data-juicer The Pile -- FreeLaw (refined by Data-Juicer) A refined version of FreeLaw dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 45GB). Dataset Information Number of samples: 2,942,612 (Keep ~82.61% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-freelaw-refined-by-data-juicer.texttext-generationn<1K0 likes8 downloads3y agoHugging Face30datajuicer /the-pile-nih-refined-by-data-juicer The Pile -- NIHExPorter (refined by Data-Juicer) A refined version of NIHExPorter dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 2.0G). Dataset Information Number of samples: 858,492 (Keep ~91.36% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-nih-refined-by-data-juicer.texttext-generationn<1K0 likes8 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.