prebuilt
Datasets
All datasets matching “prebuilt”prebuilt-indexes-msmarco-v1
Prebuilt Indexes for MS MARCO v1
Available indexes:
Lucene Standard Inverted
msmarco-v1-doc
[readme]
Lucene index of the MS MARCO V1 document corpus.
msmarco-v1-doc-slim
[readme]
Lucene index of the MS MARCO V1 document corpus ('slim' version).
msmarco-v1-doc-full
[readme]
Lucene index of the MS MARCO V1 document corpus ('full' version).
msmarco-v1-doc.d2q-t5
[readme]
Lucene index of the MS MARCO V1 document corpus with doc2query-T5 expansions.… See the full description on the dataset page: https://huggingface.co/datasets/castorini/prebuilt-indexes-msmarco-v1.prebuilt-indexes-beir
Prebuilt Indexes for BEIR
Available indexes:
Lucene Flat
beir-v1.0.0-trec-covid.bge-base-en-v1.5.flat
[readme]
Lucene flat index of BEIR collection 'trec-covid' encoded by BGE-base-en-v1.5.
beir-v1.0.0-bioasq.bge-base-en-v1.5.flat
[readme]
Lucene flat index of BEIR collection 'bioasq' encoded by BGE-base-en-v1.5.
beir-v1.0.0-nfcorpus.bge-base-en-v1.5.flat
[readme]
Lucene flat index of BEIR collection 'nfcorpus' encoded by BGE-base-en-v1.5.
beir-v1.0.0-nq.bge-base-en-v1.5.flat… See the full description on the dataset page: https://huggingface.co/datasets/castorini/prebuilt-indexes-beir.prebuilt-indexes-cacmprebuilt-indexes-bright
Prebuilt Indexes for BRIGHT
Available indexes:
Lucene Standard Inverted
bright-biology
[readme]
Lucene inverted index of BRIGHT: biology.
bright-earth-science
[readme]
Lucene inverted index of BRIGHT: earth-science.
bright-economics
[readme]
Lucene inverted index of BRIGHT: economics.
bright-psychology
[readme]
Lucene inverted index of BRIGHT: psychology.
bright-robotics
[readme]
Lucene inverted index of BRIGHT: robotics.
bright-stackoverflow
[readme]
Lucene inverted index of… See the full description on the dataset page: https://huggingface.co/datasets/castorini/prebuilt-indexes-bright.prebuilt-indexes-m-beirswe-rebench-prebuilt-images
SWE-rebench leaderboard Docker bundles
This repository contains migration-compatible Docker save archives for a deterministic one-task-per-repository sample of the public nebius/SWE-rebench-leaderboard test split.
Cohort
Source revision: 34d5a58864acf91613740a09ec5d205228dcfa39
Source rows: 860
Unique repositories and sampled tasks: 413
Unique Docker images: 413
Bundle archives: 374
Compressed bundle bytes: 751,033,962,288
Repositories are normalized with… See the full description on the dataset page: https://huggingface.co/datasets/dongyuanjushi/swe-rebench-prebuilt-images.
