datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HERBench
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
A challenging benchmark for evaluating multi-evidence integration capabilities of vision-language models
🎉 HERBench has been accepted to CVPR 2026!
🆕 New: Lite-v2 config. We released a refreshed lite_v2 version of the
Lite split (1,971 questions / 68 videos) in which 9 of the 12 tasks were
regenerated and went through additional manual refinement for higher
quality, while TSO, SVA… See the full description on the dataset page: https://huggingface.co/datasets/DanBenAmi/HERBench.Herbarium-2022-FGVC9_masked
Herbarium 2022 FGVC9 Masked
Segmentation masks for the Herbarium 2022 FGVC9 dataset, stored as RLE-encoded masks in a single Parquet file.
Note: This file covers 15,992 images (63 of 400 shards processed so far).
File
File
Description
masks.parquet
15,992 rows — one per image — with RLE mask, score, species label, and file_name
Schema
Column
Type
Description
dataset
str
Always Herbarium-2022-FGVC9
text_prompt
str
Text… See the full description on the dataset page: https://huggingface.co/datasets/kaityc06/Herbarium-2022-FGVC9_masked.arracher_une_mauvaise_herbe_400_front_eyeso101_dataset1_arracher_les_mauvaises_herbesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 42046,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kagyvro48/so101_dataset1_arracher_les_mauvaises_herbes.so101_dataset1_arracher_la_mauvaise_herbeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 61,
"total_frames": 27013,
"total_tasks": 1,
"total_videos": 183,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:61"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kagyvro48/so101_dataset1_arracher_la_mauvaise_herbe.Chinese-Herbal-Medicine-Sentiment
中药情感分析数据集 - 数据说明书
Chinese Herbal Medicine Sentiment Analysis Dataset - Datacard
数据集概述 / Dataset Overview
基本信息 / Basic Information
数据集名称 / Dataset Name: Chinese Herbal Medicine Sentiment Analysis Dataset
版本 / Version: 1.0.0
创建日期 / Created: 2025-08-26
作者 / Author: Xingqiang Chen
许可证 / License: MIT
语言 / Language: 中文 (Chinese)
领域 / Domain: 中药 / 传统中医药 (Traditional Chinese Medicine)
数据规模 / Data Scale
总样本数 / Total Samples: 234,879
唯一产品数 /… See the full description on the dataset page: https://huggingface.co/datasets/OpenModels/Chinese-Herbal-Medicine-Sentiment.africa-synth-herbal-traditional-medicine-safety-all
Herbal & Traditional Medicine Safety (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-herbal-traditional-medicine-safety-all.arracher_une_mauvaise_herbe_250_2cameras360qvr-subsample
Subsample
Representative subsample of the test split for reviewer inspection.
One counterfactual and one non-counterfactual sample were selected per task type from randomly presented candidates drawn from the test split. Samples were accepted or skipped based on whether they passed the same validation criteria used during full dataset annotation, not on answer quality or difficulty.
Contents
qa.parquet - QA samples (17 rows: 9 non-CF + 8 CF across 9 task types)… See the full description on the dataset page: https://huggingface.co/datasets/goldsmith-herbal-clay/360qvr-subsample.360qvr
