datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midjourney-prompts-highquality
Thank you to the Akash Network for sponsoring this project and providing A100s/H100s for compute!
About
A filtered version of the vivym/midjourney-prompts dataset
Filtering criteria
top 10% in length (assuming that longer prompts = more effort and higher quality)
used on an image to be upscaled (assuming that users are more likely to upscale an image that is aesthetically pleasing)
used on midjourney version 5.0+
deduplicated
Run yourself
filter.py script… See the full description on the dataset page: https://huggingface.co/datasets/gaodrew/midjourney-prompts-highquality.nli-high-quality
NLI High-Quality Balanced Dataset
A combined, filtered, and class-balanced natural language inference (NLI)
dataset built from MNLI, SNLI, FEVER-NLI, and ANLI, intended for fine-tuning
NLI models for use in zero-shot text classification via the entailment trick
(hypothesis = "This example is about {label}.").
The goal of this dataset was quality and generalization over raw volume:
rather than concatenating the four source datasets as-is, several filtering
stages were applied to… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/nli-high-quality.InstructQA-Highquality-16k
Dataset Card for Dataset Name
This dataset is a culmination of diverse sources, carefully curated with the intention of constructing a versatile and comprehensive dataset. We have amalgamated high-quality text from various datasets to form this unified dataset, designed to serve as a valuable and multifaceted resource for diverse purposes."
Datasets used to create-
aditijha/instruct_v1_10k
mosaicml/instruct-v3
jondurbin/airoboros-2.2.1
Uses
Can be used to… See the full description on the dataset page: https://huggingface.co/datasets/CrabfishAI/InstructQA-Highquality-16k.high-quality-prompt-dataset
