standard
Datasets
All datasets matching “standard”standard-chess-games
[!CAUTION]
This dataset is still a work in progress and some breaking changes might occur.
Lichess Rated Standard Chess Games Dataset
Dataset Description
6,771,826,271 standard rated games, played on lichess.org, updated monthly from the database dumps.
This version of the data is meant for data analysis. If you need PGN files you can find those here. That said, once you have a subset of interest, it is trivial to convert it back to PGN as shown in the Dataset Usage… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/standard-chess-games.MegaPairs-Standard
MegaPairs-Standard (Standardized Version)
Dataset Summary
This is a standardized, high-efficiency version of the JUNJIE99/MegaPairs dataset.
Why use this version?
The original dataset is distributed as a massive Tar archive containing millions of images, accompanied by a separate JSONL annotation file.
The Problem: Using the original format requires extracting terabytes of small files (which can exhaust disk inodes) or writing complex logic to read from archives. It… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/MegaPairs-Standard.home-standard-resultsStandard-Pipeline
Standard Pipeline
Environment adaptation, training evidence, latency profiles, raw demonstrations and evaluation traces, organized by environment and experiment stage.
AirRaid: zero-latency and profile-latency experiments.
Only raw demonstrations are distributed. Generate converted training datasets on the training server. Model weights remain in the dedicated model repository and are referenced from the experiment records. Pin a commit revision for reproducible downloads.… See the full description on the dataset page: https://huggingface.co/datasets/latency-sensitive-bench/Standard-Pipeline.pile-standard-pythia-preshuffledstandardebooks
Standard Ebooks Text Dataset
This dataset contains the full text of public domain books sourced from Standard Ebooks. It is intended for use in Natural Language Processing tasks, particularly Large Language Model pretraining, fine-tuning, and research.
Standard Ebooks provides high-quality, carefully formatted, and proofread versions of classic literature, making this a valuable collection of clean text data.
Dataset Structure
The dataset consists of a single split:… See the full description on the dataset page: https://huggingface.co/datasets/Nelathan/standardebooks.
