CoolFace
Datasetpublic

Comet/wikipedia-2017-bm25

Wikipedia 2017 BM25 Search Index This dataset provides a production-ready BM25 search index over 5.2 million Wikipedia article abstracts from the 2017 snapshot. Built using the bm25s library with English stemming and optimized Parquet compression, it enables fast, offline information retrieval for research and production AI systems. The corpus is identical to the one used in influential AI research papers including DSPy and GEPA, ensuring reproducible benchmarking and fair… See the full description on the dataset page: https://huggingface.co/datasets/Comet/wikipedia-2017-bm25.

sourceHugging Facecc-by-sa-3.0updated 10mo agoView on Hugging Face
1likes356downloads
7 commits on main
826739610mo ago

Update README.md

vincentkoc
e2da37510mo ago

Update README.md

vincentkoc
7d6a93310mo ago

Update README.md

vincentkoc
c98d79f10mo ago

Update README.md

vincentkoc
8bc89b710mo ago

Upload README.md with huggingface_hub

vincentkoc
8c38a3110mo ago

Upload folder using huggingface_hub

vincentkoc
07db6d610mo ago

initial commit

vincentkoc