AryanAnuj/processed_dataset_orca-math-word-problems-200k
Dataset Description: This dataset contains data that has undergone two preprocessing steps: Removal of Instructions with Less Than 100 Tokens in Response: Instructions with less than 100 tokens in the response have been removed from the dataset. This preprocessing step helps to ensure that the dataset contains substantial and informative responses. Data Deduplication by Grouping Using Cosine Similarity (Threshold > 0.95): Data deduplication has been performed by grouping similar instances… See the full description on the dataset page: https://huggingface.co/datasets/AryanAnuj/processed_dataset_orca-math-word-problems-200k.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face