datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hitoribocchinoisekaikouryaku
Bangumi Image Base of Hitoribocchi No Isekai Kouryaku
This is the image base of bangumi Hitoribocchi no Isekai Kouryaku, we detected 66 characters, 4987 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/hitoribocchinoisekaikouryaku.FinMTM
Fin Benchmark v1
统一整理的金融多模态 benchmark,共 11,133 条。原始成品文件保持不变,本目录为独立发布副本。
数据构成
类别
文件
数量
单选题
objective/single_choice.jsonl
1,982
多选题
objective/multiple_choice.jsonl
1,982
L1 理解/空间感知
open_ended/L1.jsonl
2,082
L2 多步数值计算
open_ended/L2.jsonl
1,893
L3 自我纠错
open_ended/L3.jsonl
1,210
L4 多页 Memory
open_ended/L4.jsonl
984
Financial Agent
agent/agent.jsonl
1,000
总计
11,133
关键说明
单选和多选逐行对应同一批 1,982 个题目与图片。
L3 中 605/1,210 条含显式错误… See the full description on the dataset page: https://huggingface.co/datasets/HiThink-Research/FinMTM.hit-uav-thermal-human-detectionMATE
Dataset Card for HiTZ/MATE
This dataset provides a benchmark consisting of 5,500 question-answering examples to assess the cross-modal entity linking capabilities of vision-language models (VLMs). The ability to link entities in different modalities is measured in a question-answering setting, where each scene is represented in both the visual modality (image) and the textual one (a list of objects and their attributes in JSON format).
The dataset is provided in two configurations:… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/MATE.MME-FinancePaper |Homepage |Github
🛠️ Usage
Regarding the data, first of all, you should download the MMfin.tsv and MMfin_CN.tsv files, as well as the relevant financial images. The folder structure is shown as follows:
├─ datasets
├─ images
├─ MMfin
...
├─ MMfin_CN
...
│ MMfin.tsv
│ MMfin_CN.tsv
The following is the process of inference and evaluation (Qwen2-VL-2B-Instruct as an example):
export LMUData="The path of the datasets"
python… See the full description on the dataset page: https://huggingface.co/datasets/hithink-ai/MME-Finance.other_regions_datasetpiper_dual_beat_drum_eef_xyz3d_100tool_pathverify_v0piper_dual_weigh_apple_eef_xyz3d_100piper_dual_fold_towel_eef_xyz3d_100CIGEval_sft_data
Dataset Card for CIGEval_sft_data
CIGEval_sft_data is the dataset used for fine-tuning LMMs in the paper CIGEval. It contains data on both tool selection and image evaluation, which can be combined into 2.3k complete evaluation trajectories. The dataset was constructed through the following steps:
Using GPT-4o + CIGEval to evaluate the full ImageHub dataset, generating 4,903 evaluation trajectories.
Randomly selecting 60% of these and filtering out the ones where the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/HIT-TMG/CIGEval_sft_data.tool_rlset_v1.1IllustrisTNG_SKIRT_SDSS
IllustrisTNG SKIRT SDSS
Preprocessed synthetic galaxy images derived from the
IllustrisTNG cosmological simulations.
Raw multi-band FITS images were produced with the
SKIRT Monte Carlo radiative transfer code in SDSS
photometric bands and subsequently processed into 128 × 128 RGB images
ready for machine-learning applications.
The dataset is designed as training data for the
Spherinator /
HiPSter framework, but it is
general-purpose and suitable for any galaxy-morphology task.… See the full description on the dataset page: https://huggingface.co/datasets/HITS-AIN/IllustrisTNG_SKIRT_SDSS.piper_dual_block_drawer_eef_xyz3d_100OpenThinkIMG-Chart-SFT-2942PixMoCapQA_eu
Pixmo-cap-qa-eu (Basque Translation)
📚 Overview
Pixmo-cap-qa-eu is a Basque-language subset of the original Pixmo Captioned QA multimodal question-answering dataset. A random sample of 5,000 English yes/no pairs was translated into Basque with HiTZ/Latxa-Llama-3.1-70B-Instruct
Important: This is not the official dataset. It is an independent community translation aimed at supporting Basque-speaking researchers and practitioners.
✍️ Authors & Acknowledgements… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/PixMoCapQA_eu.piper_dual_stack_cups_eef_xyz3dpiper_dual_stack_cups_ros2_shiftedpiper_dual_clean_table_eef_xyz3dA-OKVQA_eu
A-OK-VQA-eu (Basque Translation • 5 K Sample)
📚 Overview
A-OK-VQA-eu is a Basque-language subset of the original A-OKVQA knowledge-based visual question-answering benchmark. A random sample of 5 000 English QA pairs was translated into Basque with HiTZ/Latxa-Llama-3.1-70B-Instruct; roughly 20 % of those translations were manually post-edited to ensure fluency and adequacy.
Important: This is not the official dataset. It is an independent community translation intended… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/A-OKVQA_eu.hiten_tagpiper_dual_stack_cups_ros2Hitchhiker
Hitchhiker's Guide to the Galaxy
GPT-4-Turbo generations to elicit responses modelled on the Hitchhiker's Guide to the Galaxy.
Add some spice to your LLMs. Enjoy!
piper_dual_stack_cups_eef_xyz3d_100Agri-CM3
Agri-CM3: A Chinese Massive Multi-modal, Multi-level Benchmark for Agricultural Understanding and Reasoning
This is the repository containing evaluation datas, instructions and demonstrations with ACL 2025 paper Agri-CM3: A Chinese Massive Multi-modal, Multi-level Benchmark for Agricultural Understanding and Reasoning (Wang et al., 2025)
Agri-CM3 Benchmark
Design Principal
Existing benchmarks often evaluate complex reasoning tasks as a whole… See the full description on the dataset page: https://huggingface.co/datasets/HIT-Kwoo/Agri-CM3.libero_mem_bowl_two_tasksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 198,
"total_frames": 42857,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:198"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HITdongdong/libero_mem_bowl_two_tasks.HITOpenThinkIMG-Chart-Test-994piper_dual_stack_cups_eef_rot6d_shiftedOpenThinkIMG-Chart-RL-14501
