datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretrain-dataset-T4096-10M
Pretrain Dataset (Tokenized)
This dataset contains tokenized and packed sequences ready for LLM pretraining.
Dataset Details
Property
Value
Sequences
3,237,049
Sequence Length
4096
Tokenizer
./vn_spm_v3_fast2/
Total Tokens
13,258,950,332
Shards
7
Created
2025-12-10
Dataset Structure
Each sample contains:
input_ids: List of token IDs (length: 4096)
attention_mask: Attention mask (1 for real tokens, 0 for padding)
Usage… See the full description on the dataset page: https://huggingface.co/datasets/tvu-vlinhd11/pretrain-dataset-T4096-10M.t4-proof-runs
T4 Proof — Open Model Reality Check
Not estimated. Actually run. This dataset tracks reproducible runs of trending open models on Kaggle's free Dual Tesla T4 environment.
What makes a run verified?
A row may be a candidate, running, failed, or verified. Verified is reserved for a run that has all of the following:
a recorded Kaggle hardware and software environment;
a fixed, representative input and generation configuration;
cold-start, peak-VRAM, and task timing… See the full description on the dataset page: https://huggingface.co/datasets/livadies/t4-proof-runs.mida-autotrain2
RuTurboAlpaca
Dataset of ChatGPT-generated instructions in Russian.
Code: rulm/self_instruct
Code is based on Stanford Alpaca and self-instruct.
29822 examples
Preliminary evaluation by an expert based on 400 samples:
83% of samples contain correct instructions
63% of samples have correct instructions and outputs
Crowdsouring-based evaluation on 3500 samples:
90% of samples contain correct instructions
68% of samples have correct instructions and outputs
Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/t4t455/mida-autotrain2.mida-autotrain
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/t4t455/mida-autotrain.
