drr
Datasets
All datasets matching “drr”DRR_dataDRRMDRR-RATEThe DRR-RATE dataset is built upon the recently released CT-RATE[1] dataset, which comprises 25,692 non-contrast chest CT volumes from 21,304 unique patients. Each study is accompanied by a corresponding radiology text report and binary labels for 18 pathology classes. The dataset has been expanded to 50,188 volumes through the modification of the reconstruction matrix extracted from the raw DICOM study. As the dataset was already anonymized, compliance with the Health Insurance Portability… See the full description on the dataset page: https://huggingface.co/datasets/farrell236/DRR-RATE.stocks-DRREDDY-1D-candlescombined-instruct-sft
Combined Instruct SFT Mixture
A curated, globally shuffled mixture of 2,980,737 high-quality instruction-following dialogues harmonized into the standardized OpenAI / ChatML format.
Dataset Mixture Breakdown
Source Dataset
Split & Filtering Restrictions
Harmonized Count
allenai/Dolci-Instruct-SFT
Train split, Safety category samples removed
2,041,725
nvidia/Nemotron-SFT-Instruction-Following-Chat-v2
reasoning_off split, English dialogues only
888,460… See the full description on the dataset page: https://huggingface.co/datasets/DrRiceIO7/combined-instruct-sft.SlopReviewHowdy! This is a curated dataset for training models to distinguish between Slop and Quality Writing. You could also use it to train an LLM to write, but most of the AI responses are cosnidered Slop, so it's not recommended. Might also want to check in with specific models' licenses to see if distillation is allowed.
This dataset was made by feeding 200 prompts from ChaoticNeutrals/Reddit-SFW-Writing_Prompts_ShareGPT into various LLMs.
In v1, I compared the responses with the human generated… See the full description on the dataset page: https://huggingface.co/datasets/DrRiceIO7/SlopReview.
