CoolFace
Datasetpublic

MasonMac/WildChat-4M-English-Semantic-Deduplicated

ALERT: This dataset has been superseded by https://huggingface.co/datasets/MasonMac/WildChat-4.8M-EN-Semantic-Deduplicated This is a dataset of all English prompts from the WildChat-4M dataset. It was created by checking that at least 80% non-punctuation characters were in the English alphabet (to remove some more non-English entries). Then, it was deduplicated (ignoring punctuation/whitespace differences) by collecting to a HashSet. Another deduplication pass was done with MinHash. And… See the full description on the dataset page: https://huggingface.co/datasets/MasonMac/WildChat-4M-English-Semantic-Deduplicated.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
0likes35downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face