rul
Datasets
All datasets matching “rul”rule-ling-conceptsrulerMassiveDS-140BWe release the raw passages, embeddings, and index of MassiveDS.
Website: https://retrievalscaling.github.io
We release two versions of MassiveDS:
MassiveDS-1.4T, which contains 1.4T tokens in the datastore.
MassiveDS-140B, which is a subsampled version containing 140B tokens in the datastore.
File structure:
raw_data: plain data in JSONL files.
passages: chunked raw passages with passage IDs. Each passage is chunked to have no more than 256 words.
embeddings: embeddings of the passages… See the full description on the dataset page: https://huggingface.co/datasets/rulins/MassiveDS-140B.ru-llm-judge-dataset
RU-LLM-Judge-Dataset
Датасет собирается инкрементально в течение нескольких сессий бесплатного Google Colab.
Текущий объём: 18,821 суждений (по состоянию на последний запуск).
Прогресс к цели (5,000 суждений)
[████████████████████] 100% (18,821 / 5,000)
История сессий сбора
Сессия
Дата
Добавлено
Итого
1
2026-08-05 08:42
617
617
2
2026-08-06 14:40
583
1,200
3
2026-08-07 19:20
486
1,686
4
2026-08-08 22:34
868
2,554
5
2026-08-13… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-llm-judge-dataset.MassiveDS-1.4T-raw-dataWe release the raw passages, embeddings, and index of MassiveDS.
Website: https://retrievalscaling.github.io
Versions
We release two versions of MassiveDS:
MassiveDS-1.4T, which contains the embeddings and passages of the 1.4T-token datastore.
MassiveDS-1.4T-raw-text, contains the raw text of the 1.4T-token datastore.
MassiveDS-140B, which contains the index, embeddings, passages, and raw text of a subsampled version containing 140B tokens in the datastore.
Note:
Code support to… See the full description on the dataset page: https://huggingface.co/datasets/rulins/MassiveDS-1.4T-raw-data.MasssiveDS-1.4T-raw-data
