umt5
Datasets
All datasets matching “umt5”43.1d_17_128_128_umt5_xxlUsage:
cat 43.1d_17_128_128_umt5_xxl.tar.zst.part-* | zstd -d | tar -xf -
vidprom_16k_umt5_text_embed43.2d_17_256_256_umt5_xxlUsage:
cat 43.2d_17_256_256_umt5_xxl.tar.zst.part-* | zstd -d | tar -xf -
vidprom_filtered_extended_umt5_text_embeddroid_wan_umt5_cache
DROID umT5-XXL text-embedding cache (Wan2.1 / Wan2.2 공용)
source: /data/shared_dataset/DreamZero-DROID-Data — 123,259 raw task strings → 120,260 unique normalized strings (+ '' empty prompt)
encoder: umT5-XXL (Wan-AI/Wan2.2-TI2V-5B text_encoder, bf16; Wan2.1-14B/1.3B와 동일 가중치·tokenizer), tokenizer Wan-AI/Wan2.1-T2V-1.3B
max_sequence_length=512, trimmed=True (유효 토큰만 저장 + mask; 로더가 512까지 zero-pad, mask False), formalize_language=True (text.lower() → re.sub(r"[^\w\s]", "", text))… See the full description on the dataset page: https://huggingface.co/datasets/huiwon/droid_wan_umt5_cache.google__umt5-base-details
Dataset Card for Evaluation run of google/umt5-base
Dataset automatically created during the evaluation run of model google/umt5-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__umt5-base-details.
