CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QuangDuy /FineWeb2-mds-tokenized-40960 likes6.6k downloads11mo agoHugging Face02jeremycochoy /gift-pretrain-full-4096 gift-pretrain-full-4096 Full counterpart to jeremycochoy/gift-pretrain-small-4096: every series of every arrow file in Salesforce/GiftEvalPretrain, cropped into non-overlapping 4096-point windows and globally shuffled. The small companion sub-samples 10 series per sub-dataset; this one keeps everything. 6,376 source arrow files across 152 sub-datasets fully consumed 42,571,692 windows of length 4096 (float32) 4,274 parquet shards, ~619 GB total (zstd) Layout . ├──… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/gift-pretrain-full-4096.timeseriestime-series-forecasting10M<n<100M0 likes6.2k downloads5mo agoHugging Face03QuangDuy /FineWeb2-mds-tokenized-v2-40960 likes4.2k downloads9mo agoHugging Face04fromziro /Cement-409610M<n<100M0 likes2.2k downloads29d agoHugging Face05CyanMonkey /Cold-Large-32-409610M<n<100M0 likes1.1k downloads1mo agoHugging Face06ragrawal36 /multihop_qa_sft_doc4096_seq1024_v2text1M<n<10M0 likes1k downloads20d agoHugging Face07maknee /bioasq_bier_4096_1m BioASQ-BEIR Vector Database Dataset (4096d, 1M) Generated embeddings dataset for vector database training and evaluation. Dataset Summary This dataset contains 1,000,000 text samples with vector embeddings (4096 dimensions) generated from the BioASQ-BEIR dataset using Qwen/Qwen3-Embedding-8B. Dataset Structure Base dataset: 1,000,000 samples with embeddings Embedding dimension: 4096 Repository Structure parquet/ base.parquet - Main… See the full description on the dataset page: https://huggingface.co/datasets/maknee/bioasq_bier_4096_1m.textfeature-extraction1M<n<10M0 likes897 downloads6mo agoHugging Face08Changyeli03 /obelics_100k-tokenized-4image_llava_vicuna-7B_40960 likes864 downloads2y agoHugging Face09ssmits /tokenized-falcon2-dutch-40961M<n<10M0 likes793 downloads2y agoHugging Face10Changyeli03 /obelics_100k-tokenized-4image_llava_vicuna-13B_40960 likes781 downloads2y agoHugging Face11maknee /bioasq_bier_4096_100k BioASQ-BEIR Vector Database Dataset (4096d, 100k) Generated embeddings dataset for vector database training and evaluation. Dataset Summary This dataset contains 100,000 text samples with vector embeddings (4096 dimensions) generated from the BioASQ-BEIR dataset using Qwen/Qwen3-Embedding-8B. Dataset Structure Base dataset: 100,000 samples with embeddings Embedding dimension: 4096 Repository Structure parquet/ base.parquet - Main… See the full description on the dataset page: https://huggingface.co/datasets/maknee/bioasq_bier_4096_100k.textfeature-extraction100K<n<1M0 likes640 downloads7mo agoHugging Face12Changyeli03 /lllavasae_obelics100k-tokenized-4096_all0 likes616 downloads2y agoHugging Face13hoanganhpham /Miriad_Pubmed_metadata_4096_chunkstext10M<n<100M0 likes555 downloads1y agoHugging Face14JackHsieh /dclm-replay.seq-4096.tokens-32B2^35 tokens of replay data from DCLM-baseline, concatenated into 2^23 sequences of 4096 tokens each with <|endoftext|> separators. text1M<n<10M0 likes498 downloads3mo agoHugging Face15allenai /tulu-v3.1-mix-preview-4096-OLMoE OLMoE SFT Mix The SFT mix used is an expanded version of the Tulu v2 SFT mix with new additions for code, CodeFeedback-Filtered-Instruction, reasoning, MetaMathQA, and instruction following, No Robots and a subset of Daring Anteater. Please see the referenced datasets for the multiple licenses used in subsequent data. We do not introduce any new data with this dataset. Config for creation via open-instruct: dataset_mixer: allenai/tulu-v2-sft-mixture-olmo-4096: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v3.1-mix-preview-4096-OLMoE.textquestion-answering100K<n<1M8 likes483 downloads2y agoHugging Face16Changyeli03 /llavasae_obelics3k-tokenized-4096_4image1K<n<10K0 likes471 downloads2y agoHugging Face17nongiga /nt_multispecies_4096tabular100K<n<1M0 likes422 downloads3y agoHugging Face18plantcad /plantcad2-c4096tabular1M<n<10M0 likes420 downloads9mo agoHugging Face19philschmid /llama2-german-corpus-tokenized-llama-chunk-4096 Dataset Card for "llama2-german-corpus-tokenized-llama-chunk-4096" More Information needed 10M<n<100M0 likes357 downloads3y agoHugging Face20kaiwenw /distill-r1-qwen-1.5b-aime-24-4096-with-labels-prmtabular100K<n<1M0 likes347 downloads1y agoHugging Face21DopeorNope /train_group_theory_cpt_chunked_4096text10M<n<100M0 likes331 downloads1y agoHugging Face22QuangDuy /merged-dataset-40960 likes325 downloads11mo agoHugging Face23kaiwenw /distill-r1-qwen-1.5b-hmmt-feb-25-4096-with-bt-model-with-sigmoidtabular100K<n<1M0 likes316 downloads1y agoHugging Face24kaiwenw /distill-r1-qwen-1.5b-aime-24-4096-with-bt-model-wout-sigmoidtabular100K<n<1M0 likes308 downloads1y agoHugging Face25ragrawal36 /science-qa-hard-neg-think-doc4096-seq1024-v2text1M<n<10M0 likes303 downloads20d agoHugging Face26plantcad /opengenome2-metagenomes-plantcad2-c4096 OpenGenome2 Metagenomes PlantCAD2 Subset (4096bp) This dataset is a curated subset of arcinstitute/opengenome2 designed for comparative spectral analysis with plant genomic data. Dataset Description Sequences were randomly sampled from OpenGenome2, filtered and truncated to match the sample sizes per split of the plantcad/Angiosperm_65_genomes_8192bp dataset. Processing Steps Streaming: Records were streamed from the metagenomes subfolder… See the full description on the dataset page: https://huggingface.co/datasets/plantcad/opengenome2-metagenomes-plantcad2-c4096.text1M<n<10M0 likes292 downloads9mo agoHugging Face27fhai50032 /BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65 fhai50032/BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65 Packed pretraining corpus, 2,050,000 rows x 4096 tokens = 8.397B tokens, tokenized with fhai50032/QTK-81K. Format column type notes label list<int32> exactly 4096 tokens, no padding raw_label string label decoded back to text (redundant, for inspection) There is no attention_mask column: the corpus is packed, so every position is a real token and the mask would be all ones on every row.… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/BiBo_corpus-4096-packed-qtk-2.05M-hi35-en65.text1M<n<10M0 likes270 downloads2mo agoHugging Face28aiteamOD /data-artifacts-4096gated10M<n<100M0 likes255 downloads23d agoHugging Face29fhai50032 /English_corpus-4096-packed-qtk-2.0M fhai50032/English_corpus-4096-packed-qtk-2.0M Packed pretraining corpus, 2,004,054 rows x 4096 tokens = 8.209B tokens, tokenized with fhai50032/QTK-81K. Format column type notes label list<int32> exactly 4096 tokens, no padding raw_label string label decoded back to text (redundant, for inspection) There is no attention_mask column: the corpus is packed, so every position is a real token and the mask would be all ones on every row.… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/English_corpus-4096-packed-qtk-2.0M.text1M<n<10M0 likes242 downloads2mo agoHugging Face30kaiwenw /distill-r1-qwen-1.5b-hmmt-feb-24-4096-with-bt-model-wout-sigmoidtabular100K<n<1M0 likes230 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.