CoolFace
Datasetpublic

Comet/wikipedia-2017-bm25

Wikipedia 2017 BM25 Search Index This dataset provides a production-ready BM25 search index over 5.2 million Wikipedia article abstracts from the 2017 snapshot. Built using the bm25s library with English stemming and optimized Parquet compression, it enables fast, offline information retrieval for research and production AI systems. The corpus is identical to the one used in influential AI research papers including DSPy and GEPA, ensuring reproducible benchmarking and fair… See the full description on the dataset page: https://huggingface.co/datasets/Comet/wikipedia-2017-bm25.

sourceHugging Facecc-by-sa-3.0updated 10mo agoView on Hugging Face
1likes356downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Comet/wikipedia-2017-bm25 · CoolFace