natong19/lmsys-chat-1m-filtered
Dataset Card for natong19/lmsys-chat-1m-filtered Filtered version of lmsys/lmsys-chat-1m, a collection of one million real-world conversations with various LLMs. Data cleaning process inspired by OpenLeecher/lmsys_chat_1m_clean. Overview of filtering process: 1. Filtering REDACTED Entries Entries that were labeled as REDACTED due to containing Personally Identifiable Information (PII) were removed. 1000000 samples -> 733740 samples 2. Format… See the full description on the dataset page: https://huggingface.co/datasets/natong19/lmsys-chat-1m-filtered.
Dataset Card for natong19/lmsys-chat-1m-filtered
Filtered version of lmsys/lmsys-chat-1m, a collection of one million real-world conversations with various LLMs. Data cleaning process inspired by OpenLeecher/lmsys_chat_1m_clean.
Overview of filtering process:
1. Filtering REDACTED Entries
Entries that were labeled as REDACTED due to containing Personally Identifiable Information (PII) were removed.
1000000 samples -> 733740 samples
2. Format validation
Entries not matching the messages format or containing empty turns were removed.
733740 samples -> 724698 samples
3. Removing Pure Duplicate User Turns
Pure duplicate user turns were removed. This involved removing whitespace and punctuation, then ensuring that if two instructions matched after that, only one was retained.
724698 samples -> 448833 samples
4. MinHash Deduplication
The dataset was deduplicated with MinHash LSH to remove entries with minor variances in the text.
448833 samples -> 317954 samples
5. Semantic Deduplication
The dataset was deduplicated based on text embeddings to remove entries with larger variances in the text but similar semantic meaning.
317954 samples -> 263457 samples
6. Filtering Repetitive Entries
To further clean the dataset, repetitive entries were filtered based on prefix string frequency and ngram frequency.
263457 samples -> 259434 samples
