CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hlillemark /c4_t5_corrupted_seqlen256 Dataset Card for "c4_t5_corrupted_seqlen256" More Information needed 100M<n<1B0 likes2.5k downloads3y agoHugging Face02embedded-language-flows /openwebtext-t51M<n<10M2 likes2.4k downloads4mo agoHugging Face03s-nlp /Mintaka_Graph_Features_T5-xl-ssm Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm" More Information needed tabular100K<n<1M0 likes1.9k downloads2y agoHugging Face04Birchlabs /c4-t5-ragged C4, T5 tokenized, in ragged array format Processed distribution of Google's C4 dataset: a colossal, cleaned version of Common Crawl's web crawl corpus. Uses the text data from allenai/c4. Includes en subset only. T5 tokenizer was applied to the text.Distributed as a ragged array. Converted via json_to_ragged.py. Download size of all shards: Split Data+Lengths Size Divided across n Shards Typical shard size: data.npy Typical shard size: len.npy Train 293G 1024 344M 1.4M… See the full description on the dataset page: https://huggingface.co/datasets/Birchlabs/c4-t5-ragged.text-generationn<1K1 likes975 downloads3y agoHugging Face05embedded-language-flows /xsum_validation_t50 likes923 downloads4mo agoHugging Face06Hemabhushan /capstone_sakuga_iblip_t5_embeddingstabular10K<n<100K0 likes811 downloads2y agoHugging Face07cloneofsimo /laion-pop-vae-t50 likes746 downloads2y agoHugging Face08hlillemark /c4_t5_pretrain Dataset Card for "c4_t5_pretrain" More Information needed 100M<n<1B0 likes719 downloads3y agoHugging Face09hlillemark /c4_t5 Dataset Card for "c4_t5" More Information needed 100M<n<1B0 likes595 downloads3y agoHugging Face10crumb /flan-t5-large-embed-refinedwebAll of the data together is around 81.3GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-base. Structure: { "encoding": List, shaped (512, 1024) aka (tokens, d_model), "text": String, the original text that was encoded, "attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens } textfeature-extraction1K<n<10K0 likes584 downloads3y agoHugging Face11Tristan /t5-small-october-wikipedia-2022-tokenized-512 Dataset Card for "t5-small-october-wikipedia-2022-tokenized-512" More Information needed 1M<n<10M0 likes495 downloads4y agoHugging Face12princeton-nlp /gtr-t5-xxl-wikipedia-psgs_w100-index0 likes486 downloads3y agoHugging Face13embedded-language-flows /wmt14_de-en_validation_t50 likes453 downloads4mo agoHugging Face14yorkerlin /c4-t5-subset20 likes437 downloads1y agoHugging Face15CLAPv2 /a_t5_page14_batch1audio10K<n<100K0 likes412 downloads1y agoHugging Face16hlillemark /c4_t5_packed Dataset Card for "c4_t5_packed" More Information needed 100M<n<1B0 likes404 downloads3y agoHugging Face17embedded-language-flows /xsum_train_t5tabular100K<n<1M0 likes386 downloads4mo agoHugging Face18CLAPv2 /a_t5_page10_batch2audio10K<n<100K0 likes325 downloads1y agoHugging Face19crumb /flan-t5-base-embed-refinedwebAll of the data together is around 61GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-base. Structure: { "encoding": List, shaped (512, 768) aka (tokens, d_model), "text": String, the original text that was encoded, "attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens } textfeature-extraction1K<n<10K1 likes319 downloads3y agoHugging Face20hle2000 /KGQA_T5-xl-ssm Dataset Card for "KGQA_T5-xl-ssm" More Information needed tabular100K<n<1M1 likes319 downloads2y agoHugging Face21CLAPv2 /a_t5_page16_batch1audio10K<n<100K0 likes317 downloads1y agoHugging Face22CLAPv2 /a_t5_page15_batch1audio10K<n<100K0 likes308 downloads1y agoHugging Face23daruokta /t5gemma2-indonesia-instruct-v1 T5Gemma-2 Indonesian Instruct — Mono-Repo Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia. Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder, setiap config = folder dan berisi split train + validation (80:20) di level percakapan. Struktur (by fungsi) t5gemma2-indonesia-instruct-v1/ ├── README.md ├── manifest.json ├── chat_idx_map.json ├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.imagetext-generation100K<n<1M0 likes277 downloads12d agoHugging Face24CLAPv2 /a_t5_page12_batch1audio10K<n<100K0 likes269 downloads1y agoHugging Face25CLAPv2 /a_t5_page17_batch2audio10K<n<100K0 likes263 downloads1y agoHugging Face26hlillemark /c4_t5_10m Dataset Card for "c4_t5_10m" More Information needed 10M<n<100M0 likes255 downloads3y agoHugging Face27CLAPv2 /a_t5_page9_batch1audio10K<n<100K0 likes249 downloads1y agoHugging Face28CLAPv2 /epidemic_sound_effects_t5_debiasedaudio10K<n<100K0 likes248 downloads2y agoHugging Face29yorkerlin /c4-t5-subset0 likes243 downloads1y agoHugging Face30crumb /flan-t5-small-embed-refinedwebAll of the data together is around 41GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-small. Structure: { "encoding": List, shaped (512, 512) aka (tokens, d_model), "text": String, the original text that was encoded, "attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens } just a tip, you cannot load this with the RAM in the free ver of google colab, not… See the full description on the dataset page: https://huggingface.co/datasets/crumb/flan-t5-small-embed-refinedweb.textfeature-extraction100K<n<1M0 likes239 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.