t5-small
Datasets
All datasets matching “t5-small”t5-small-october-wikipedia-2022-tokenized-512
Dataset Card for "t5-small-october-wikipedia-2022-tokenized-512"
More Information needed
flan-t5-small-embed-refinedwebAll of the data together is around 41GB. It's the last hidden states of 131,072 samples from refinedweb padded/truncated to 512 tokens on the left, fed through google/flan-t5-small.
Structure:
{
"encoding": List, shaped (512, 512) aka (tokens, d_model),
"text": String, the original text that was encoded,
"attention_mask": List, binary mask to pass to your model with encoding to not attend to pad tokens
}
just a tip, you cannot load this with the RAM in the free ver of google colab, not… See the full description on the dataset page: https://huggingface.co/datasets/crumb/flan-t5-small-embed-refinedweb.t5-v1_1-small-k-mktr-improved-flux-prompts-latents
Dataset Card for Prompt Latents from T5-small
Latents from T5-small used for distillation.
Dataset Details
Dataset Description
Curated by: Dave Lage
License: Apache 2
Dataset Sources [optional]
Repository: rockerBOO/t5-distill
Uses
Latents from T5-small used for distillation.
Direct Use
[More Information Needed]
Out-of-Scope Use
[More Information Needed]
Dataset Structure
latents:… See the full description on the dataset page: https://huggingface.co/datasets/rockerBOO/t5-v1_1-small-k-mktr-improved-flux-prompts-latents.gpqa_diamond_deepseek_moe_16b_token_real_and_predicted_patterns_t5-smallaime2024_deepseek_moe_16b_token_real_and_predicted_patterns_t5-smallprocessed_t5_small_context_len_512
Dataset Card for "processed_t5_small_context_len_512"
More Information needed
