datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
processed_dataset_orca-math-word-problems-200kDataset Description:
This dataset contains data that has undergone two preprocessing steps:
Removal of Instructions with Less Than 100 Tokens in Response: Instructions with less than 100 tokens in the response have been removed from the dataset. This preprocessing step helps to ensure that the dataset contains substantial and informative responses.
Data Deduplication by Grouping Using Cosine Similarity (Threshold > 0.95): Data deduplication has been performed by grouping similar instances… See the full description on the dataset page: https://huggingface.co/datasets/AryanAnuj/processed_dataset_orca-math-word-problems-200k.orca-math-word-problems-200korca-math-word-problems-tr-short-1math-word-problems
