datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t5gemma2-indonesia-instruct-v1
T5Gemma-2 Indonesian Instruct — Mono-Repo
Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia.
Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder,
setiap config = folder dan berisi split train + validation (80:20) di level percakapan.
Struktur (by fungsi)
t5gemma2-indonesia-instruct-v1/
├── README.md
├── manifest.json
├── chat_idx_map.json
├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.text-vqa-clip-t5-processedpeanuts-flan-t5-xl
Peanut Comic Strip Dataset (Snoopy & Co.)
This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13.
There are 77,456 panels extracted from 17,816 comic strips.
The dataset size is approximately 4.4G.
Each row in the dataset contains the following fields:
image: PIL.Image containing the extracted panel.
panel_name: unique identifier for the row.
characters: tuple[str, ...] of characters included in the comic strip the panel is part of.
themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-flan-t5-xl.t5-gemma-2-multimodal-embeddingicub_sim_dataset_t5_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 10716,
"total_tasks": 90,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t5_smooth_lerobot.t5_evalt5gemma2-indonesia-vision-formatted
T5Gemma-2 Indonesian Vision & Multimodal Dataset
A high-quality Indonesian language visual instruction tuning and preference alignment dataset, specifically formatted for multimodal sequence-to-sequence (Seq2Seq) models like T5-Gemma-2 Vision (e.g., google/t5gemma-2-270m-270m).
This dataset provides embedded image features (datasets.Image) along with multi-turn conversations in Bahasa Indonesia, facilitating instruction-following and preference alignment training on… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-vision-formatted.t5spiders-1000
Dataset Card for "t5spiders-1000"
More Information needed
verl-lt-merged_T_dline_v3_1-gem3f_med-0522_T55s46ftjs_dlines17tj.deprecatedt5spiders
Dataset Card for "t5spiders"
More Information needed
twitter-zhuboluowuaiha2-2025.05.19-1924461482781082047-mBtnAlzqjTdD_t5D-part1t5twitter-sxiaoliya-2026.03.08-2030534214949732521-HMF_t51NuO9Aeffa-part1t5_allt5_smallzon-ml-t5T579875widelegjeansnetflix-movie-dataset-t5twitter-Rose99952-2025.07.19-1946597420227563856-L0SnUN_t5zClf-t8-part1verl-lt-merged_T_dline_v3_1-gem3f_med-0522_T55s46ftjs_dlines17tj-replace_cntxt-sft_a256tj.dverl-lt-merged_T_dline_v3_1-gem3f_med-0522_T55s54ftjs_dlines17tj-replace_cntxt-base.deprecatedtwitter-ARtuo88-2024.08.11-1822634203189739731-t5YzR1dxBGMVtoI8-part1T-54_tank_Datasettwitter-jiejie_waifu-2026.03.14-2032827004140294421-hAekozv-_q5_T5kw-part1twitter-LiveLifeCAN1-2025.10.15-1978374172230238684-dkSYASY_T5Vf4Z3G-part1
