datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VeriSciQA
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
Paper: VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
Dataset Description
VeriSciQA is a large-scale, high-quality dataset for Scientific Visual Question Answering (SVQA), containing 20,272 QA pairs spanning 20 scientific domains, 12 figure types, and 5 question types. The dataset is constructed using a Cross-Modal Verification framework that generates QA pairs from… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/VeriSciQA.the-pile-pubmed-abstracts-refined-by-data-juicer
The Pile -- PubMed Abstracts (refined by Data-Juicer)
A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G).
Dataset Information
Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.DetailMaster
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
We introduce DetailMaster, a benchmark designed to evaluate text-to-image generation in long-prompt scenarios, accompanied by a robust fine-grained evaluation protocol. See more details in our paper.
Abstract: While recent text-to-image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, their performance significantly degrades when confronted with long, detail-intensive prompts… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/DetailMaster.Trinity-ToolAce-SFT-splitthe-pile-pubmed-central-refined-by-data-juicer
The Pile -- PubMed Central (refined by Data-Juicer)
A refined version of PubMed Central dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 83G).
Dataset Information
Number of samples: 2,694,860 (Keep ~86.96% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-central-refined-by-data-juicer.data-juicer-t2v-optimal-data-pool
Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development
Project description
The emergence of large-scale multi-modal generative models has drastically advanced artificial intelligence, introducing unprecedented levels of performance and functionality.
However, optimizing these models remains challenging due to historically isolated paths of model-centric and data-centric developments, leading to suboptimal outcomes and inefficient resource… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/data-juicer-t2v-optimal-data-pool.RealMedConv
Dataset Description
The RealMedConv dataset consists of anonymized, real-world dialogues between licensed pharmacists and users seeking over-the-counter (OTC) medication advice. Each conversation is goal-oriented: the pharmacist gathers sufficient symptom information to provide an online and appropriate recommendation. Dialogues are typically concise, spanning 3–5 turns, reflecting the efficient and expert-driven nature of professional medical consultations.
This dataset originates… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/RealMedConv.MindGYM
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
alpaca-cot-zh-refined-by-data-juicer
Alpaca-CoT -- ZH (refined by Data-Juicer)
A refined Chinese version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to fine-tune a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 18.7GB).
Dataset Information
Number of samples: 9,873,214 (Keep ~46.58% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-zh-refined-by-data-juicer.llava-pretrain-refined-by-data-juicer
LLaVA pretrain -- LCS-558k (refined by Data-Juicer)
A refined version of LLaVA pretrain dataset (LCS-558k) by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Multimodal Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 115MB).
Dataset Information
Number of samples: 500,380 (Keep ~89.65% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/llava-pretrain-refined-by-data-juicer.the-pile-uspto-refined-by-data-juicer
The Pile -- USPTO (refined by Data-Juicer)
A refined version of USPTO dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 18G).
Dataset Information
Number of samples: 4,516,283 (Keep ~46.77% from the original dataset)
Refining Recipe
#… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-uspto-refined-by-data-juicer.geometry_sft
Overview
This dataset is a Supervised Fine-Tuning (SFT) dataset generated from a subset of the Geometry3K dataset using Qwen2.5-VL. It serves as an example dataset for demonstrating VLM (Vision-Language Model) SFT training in the Trinity-RFT library.
redpajama-arxiv-refined-by-data-juicer
RedPajama -- ArXiv (refined by Data-Juicer)
A refined version of ArXiv dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 85GB).
Dataset Information
Number of samples: 1,655,259 (Keep ~95.99% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-arxiv-refined-by-data-juicer.redpajama-cc-2022-05-refined-by-data-juicer
RedPajama -- CommonCrawl-2022-05 (refined by Data-Juicer)
A refined version of CommonCrawl-2022-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 265GB).
Dataset Information
Number of samples: 42,648,496 (Keep ~45.34% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2022-05-refined-by-data-juicer.alpaca-cot-en-refined-by-data-juicer
Alpaca-CoT -- EN (refined by Data-Juicer)
A refined English version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to fine-tune a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 226GB).
Dataset Information
Number of samples: 72,855,345 (Keep ~54.48% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-en-refined-by-data-juicer.redpajama-pile-stackexchange-refined-by-data-juicer
RedPajama & The Pile -- StackExchange (refined by Data-Juicer)
A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB).
Dataset Information
Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.redpajama-cc-2019-30-refined-by-data-juicer
RedPajama -- CommonCrawl-2019-30 (refined by Data-Juicer)
A refined version of CommonCrawl-2019-30 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 240GB).
Dataset Information
Number of samples: 36,557,283 (Keep ~45.08% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2019-30-refined-by-data-juicer.redpajama-wiki-refined-by-data-juicer
RedPajama -- Wikipedia (refined by Data-Juicer)
A refined version of Wikipedia dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 68GB).
Dataset Information
Number of samples: 26,990,659 (Keep ~90.47% from the original dataset)
Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-wiki-refined-by-data-juicer.the-pile-europarl-refined-by-data-juicer
The Pile -- EuroParl (refined by Data-Juicer)
A refined version of EuroParl dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 2.2GB).
Dataset Information
Number of samples: 61,601 (Keep ~88.23% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-europarl-refined-by-data-juicer.redpajama-stack-code-refined-by-data-juicer
RedPajama & TheStack -- Github Code (refined by Data-Juicer)
A refined version of Github Code dataset in RedPajama & TheStack by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 232GB).
Dataset Information
Number of samples: 49,279,344 (Keep ~52.09% from the original… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-stack-code-refined-by-data-juicer.the-pile-hackernews-refined-by-data-juicer
The Pile -- HackerNews (refined by Data-Juicer)
A refined version of HackerNews dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.8G).
Dataset Information
Number of samples: 371,331 (Keep ~99.55% from the original dataset)
Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-hackernews-refined-by-data-juicer.the-pile-philpaper-refined-by-data-juicer
The Pile -- PhilPaper (refined by Data-Juicer)
A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB).
Dataset Information
Number of samples: 29,117 (Keep ~88.82% from the original dataset)
Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-philpaper-refined-by-data-juicer.redpajama-cc-2020-05-refined-by-data-juicer
RedPajama -- CommonCrawl-2020-05 (refined by Data-Juicer)
A refined version of CommonCrawl-2020-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 297GB).
Dataset Information
Number of samples: 42,612,596 (Keep ~46.90% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2020-05-refined-by-data-juicer.Trinity-ToolAce-RL-splitredpajama-cc-2023-06-refined-by-data-juicer
RedPajama -- CommonCrawl-2023-06 (refined by Data-Juicer)
A refined version of CommonCrawl-2023-06 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 310GB).
Dataset Information
Number of samples: 50,643,699 (Keep ~45.46% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2023-06-refined-by-data-juicer.the-pile-freelaw-refined-by-data-juicer
The Pile -- FreeLaw (refined by Data-Juicer)
A refined version of FreeLaw dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 45GB).
Dataset Information
Number of samples: 2,942,612 (Keep ~82.61% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-freelaw-refined-by-data-juicer.the-pile-nih-refined-by-data-juicer
The Pile -- NIHExPorter (refined by Data-Juicer)
A refined version of NIHExPorter dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 2.0G).
Dataset Information
Number of samples: 858,492 (Keep ~91.36% from the original dataset)
Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-nih-refined-by-data-juicer.redpajama-cc-2021-04-refined-by-data-juicer
RedPajama -- CommonCrawl-2021-04 (refined by Data-Juicer)
A refined version of CommonCrawl-2021-04 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 284GB).
Dataset Information
Number of samples: 44,724,752 (Keep ~45.23% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2021-04-refined-by-data-juicer.redpajama-c4-refined-by-data-juicer
RedPajama -- C4 (refined by Data-Juicer)
A refined version of C4 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to pretrain a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 832GB).
Dataset Information
Number of samples: 344,491,171 (Keep ~94.42% from the original dataset)
Refining Recipe
#… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-c4-refined-by-data-juicer.
