CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tvu-vlinhd11 /pretrain-dataset-T4096-10M Pretrain Dataset (Tokenized) This dataset contains tokenized and packed sequences ready for LLM pretraining. Dataset Details Property Value Sequences 3,237,049 Sequence Length 4096 Tokenizer ./vn_spm_v3_fast2/ Total Tokens 13,258,950,332 Shards 7 Created 2025-12-10 Dataset Structure Each sample contains: input_ids: List of token IDs (length: 4096) attention_mask: Attention mask (1 for real tokens, 0 for padding) Usage… See the full description on the dataset page: https://huggingface.co/datasets/tvu-vlinhd11/pretrain-dataset-T4096-10M.text-generation1M<n<10M0 likes110 downloads10mo agoHugging Face02livadies /t4-proof-runs T4 Proof — Open Model Reality Check Not estimated. Actually run. This dataset tracks reproducible runs of trending open models on Kaggle's free Dual Tesla T4 environment. What makes a run verified? A row may be a candidate, running, failed, or verified. Verified is reserved for a run that has all of the following: a recorded Kaggle hardware and software environment; a fixed, representative input and generation configuration; cold-start, peak-VRAM, and task timing… See the full description on the dataset page: https://huggingface.co/datasets/livadies/t4-proof-runs.image-text-to-textn<1K0 likes16 downloads2mo agoHugging Face03t4t455 /mida-autotrain2 RuTurboAlpaca Dataset of ChatGPT-generated instructions in Russian. Code: rulm/self_instruct Code is based on Stanford Alpaca and self-instruct. 29822 examples Preliminary evaluation by an expert based on 400 samples: 83% of samples contain correct instructions 63% of samples have correct instructions and outputs Crowdsouring-based evaluation on 3500 samples: 90% of samples contain correct instructions 68% of samples have correct instructions and outputs Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/t4t455/mida-autotrain2.tabulartext-generation10K<n<100K0 likes7 downloads3mo agoHugging Face04t4t455 /mida-autotrain SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set) Dataset Summary SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/t4t455/mida-autotrain.textquestion-answering10K<n<100K0 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.