open-athena/snowball-replay-index
Snowball replay index This dataset is a compact membership and ordering index for an approximate replay of Snowball's 10,372,343,704,053-token data store. It contains no source text or token arrays. The 6,301 Parquet files contain three columns: source_id: logical source key; join it to the source_id field in sources.json document_id: the retained XXH3-128 content hash as 16 bytes bucket_id: domain_cluster * 5 + quality_bucket Document join contract document_id… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/snowball-replay-index.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face