datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai_handwriting_trio
Thai Handwritten Dataset
This Thai handwritten dataset is curated from three sources:
Ancient scripts [1]
General sentences [2]
Syllables [3]
Dataset Composition
Each canvas contains:
1 ancient script sample
2–5 general sentence samples
4–8 syllable samples
All data were used exactly once, except for the syllable data. Of the 268,056 available syllable samples, only 20,717 were used, balanced across 320 words.
Each canvas has a resolution of 0.4 - 1M px.… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/thai_handwriting_trio.bookmia_lexical_unique_trio_ratio_1.50_adaptive_match_mink_random_7_p0.25_a0.25TrioBench
TrioBench
TrioBench evaluates LLMs as hybrid query planners across three database engines — SQLite (structured facts + aggregation), Milvus (semantic text/image retrieval), and Neo4j (graph constraints + multi-hop reasoning) — on the Yelp Open Dataset.
Given a natural-language question, a planner must orchestrate the retrieval trio and produce two artifacts: (1) an executable multi-step JSON plan, and (2) a fully executable end-to-end Python program. 341 questions were sent to 5… See the full description on the dataset page: https://huggingface.co/datasets/iwei0/TrioBench.olympiads_paraphrased_lexical_unique_trio_ratio_2.0_adaptive_match_minkplus_random_7_p0.25Trio-Image-Audio-Text
Trio
A unified multimodal dataset combining image, audio, and text from diverse public sources.
Usage
This dataset uses Configurations (Subsets) to manage its diverse data sources. You can load specific parts or the entire "filtered" dataset without downloading the NSFW portions.
pip install datasets
1. Load the "filtered" Subset
This configuration loads all 29 safe subsets, excluding the NSFW content.
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Trio-Image-Audio-Text.details_paloalma__Le_Triomphant-ECE-TW3
Dataset Card for Evaluation run of paloalma/Le_Triomphant-ECE-TW3
Dataset automatically created during the evaluation run of model paloalma/Le_Triomphant-ECE-TW3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_paloalma__Le_Triomphant-ECE-TW3.olympiads_lexical_unique_trio_ratio_2.0_adaptive_match_minkplus_augment_random_7_p0.25dolma3-arxiv_paraphrased_unique_trio_ratio_1.50_adaptive_match_random_7_p0.25_a0.25mujoco-simple-structures-visuals1_gemini-r1_lexical_unique_trio_penalty_1.25_seed4220250524_mujoco-robot-descriptions-datasetmujoco-robot-descriptions-dataset-visualtimor-mujoco-robot-descriptions-dataset-with-assetsolympiads_lexical_unique_trio_penalty_2.0_augment_random_7_p0.25s1_deepseek-r1_lexical_unique_trio_penalty_1.25_seed42Shiva_Triologymujoco-robot-descriptions-py-with-assetsdolma3-arxiv_unique_trio_ratio_1.50_adaptive_match_loss_random_7_p0.25_a0.25mujoco_doc_dataset_qwen3_32bmorphology-exampleswikimia24_hard_unique_trio_ratio_1.50_adaptive_match_mink_pair_3_p0.25_a0.25mujoco_pusher_sft_geminitimor-mujoco-robot-descriptions-dataset-with-assets-visualmujoco-simple-structuresmujoco-component-descriptionsdolma3-arxiv_unique_trio_ratio_1.50_random_7_p0.25_a0.25mujoco-robot-descriptions-datasetmujoco-robot-descriptions-dataset-with-assetsmujoco-robot-descriptions-py-with-assets-visualmujoco-robot-descriptions-dataset-with-assets-visual
