datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HERBench
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
A challenging benchmark for evaluating multi-evidence integration capabilities of vision-language models
🎉 HERBench has been accepted to CVPR 2026!
🆕 New: Lite-v2 config. We released a refreshed lite_v2 version of the
Lite split (1,971 questions / 68 videos) in which 9 of the 12 tasks were
regenerated and went through additional manual refinement for higher
quality, while TSO, SVA… See the full description on the dataset page: https://huggingface.co/datasets/DanBenAmi/HERBench.herb-ai-vaultherbHERB
Dataset Card for HERB
Dataset Description
HERB is a benchmark for evaluating LLM agents’ ability to perform Deep Search and Long Context Reasoning. It is generated using a synthetic data pipeline that simulates business workflows across product planning, development, and support stages, generating interconnected content with realistic noise and multi-hop questions with guaranteed ground-truth answers.
Directory Structure
data/
├── metadata/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/HERB.Herbarium_FieldCoVUBench
Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models
Abstract
Large Vision-Language Models (LVLMs), trained on web-scale data, risk memorizing and regenerating copyrighted visual content like characters and logos, creating significant challenges. Machine unlearning offers a path to mitigate these risks by removing specific content post-training, but evaluating its effectiveness, especially in the complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/herbwood27/CoVUBench.Herbarium-2022-FGVC9_masked
Herbarium 2022 FGVC9 Masked
Segmentation masks for the Herbarium 2022 FGVC9 dataset, stored as RLE-encoded masks in a single Parquet file.
Note: This file covers 15,992 images (63 of 400 shards processed so far).
File
File
Description
masks.parquet
15,992 rows — one per image — with RLE mask, score, species label, and file_name
Schema
Column
Type
Description
dataset
str
Always Herbarium-2022-FGVC9
text_prompt
str
Text… See the full description on the dataset page: https://huggingface.co/datasets/kaityc06/Herbarium-2022-FGVC9_masked.arracher_une_mauvaise_herbe_400_front_eyeso101_dataset1_arracher_les_mauvaises_herbesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 42046,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kagyvro48/so101_dataset1_arracher_les_mauvaises_herbes.arracher_une_mauvaise_herbe_250herbso101_dataset1_arracher_la_mauvaise_herbeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 61,
"total_frames": 27013,
"total_tasks": 1,
"total_videos": 183,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:61"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kagyvro48/so101_dataset1_arracher_la_mauvaise_herbe.herb-corpusHerbs
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [washed Ashore Relics]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [internet archives]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Washedashore/Herbs.Chinese-Herbal-Medicine-Sentiment
中药情感分析数据集 - 数据说明书
Chinese Herbal Medicine Sentiment Analysis Dataset - Datacard
数据集概述 / Dataset Overview
基本信息 / Basic Information
数据集名称 / Dataset Name: Chinese Herbal Medicine Sentiment Analysis Dataset
版本 / Version: 1.0.0
创建日期 / Created: 2025-08-26
作者 / Author: Xingqiang Chen
许可证 / License: MIT
语言 / Language: 中文 (Chinese)
领域 / Domain: 中药 / 传统中医药 (Traditional Chinese Medicine)
数据规模 / Data Scale
总样本数 / Total Samples: 234,879
唯一产品数 /… See the full description on the dataset page: https://huggingface.co/datasets/OpenModels/Chinese-Herbal-Medicine-Sentiment.viper-racing-mods
Viper Racing community mods
Twenty-odd years of community-made cars and tracks for Viper Racing (MGI /
Sierra, 1998), preserved as their authors distributed them.
Each file here is an original distribution archive — the same zip someone
uploaded to a fan site, with its readme, its screenshots and whatever else the
author put in it. Nothing has been repacked, renamed, or tidied. That is the
point: the archive is the record, and once it is edited it stops being one.
Browse them… See the full description on the dataset page: https://huggingface.co/datasets/herbfargus/viper-racing-mods.herbier3herbier5herbal-knowledge-embedding-wikipedia
Herbs Domain Knowledge Safety Wikipedia Embeddings
Pre-computed vector embeddings from Wikipedia articles covering Assorted Herbs, spices, and other botanical items used for alternative medicine — ready to drop into your RAG pipeline without any embedding overhead.
Dataset Details
Property
Details
Embedding Model
nomic-ai/nomic-embed-text-v1.5 (135M)
Vector Dimensions
768
Source
Wikipedia
Topics
herbs, plants, spices, and other botanical items for… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/herbal-knowledge-embedding-wikipedia.TCM-Herbal-Dataset-Structured-Sample
Structured TCM Herbal Dataset (50-Herb Sample)
Dataset Description
This dataset is a high-precision, structured collection of Traditional Chinese Medicine (TCM) herbs. It bridges the gap between classical herbal knowledge and modern data engineering.
Total Sample Records: 50
Master Database Size: 2,170+ records
Format: JSON
Fields: English/Latin/Chinese Names, Nature, Taste, Meridians, Toxicity, Chemical Compounds, and Pharmacological Mechanisms.
Use Cases… See the full description on the dataset page: https://huggingface.co/datasets/AdamGoldman/TCM-Herbal-Dataset-Structured-Sample.id-voicemail-dataset-v2
Dataset Card for "id-voicemail-dataset-v2"
More Information needed
africa-synth-herbal-traditional-medicine-safety-all
Herbal & Traditional Medicine Safety (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-herbal-traditional-medicine-safety-all.ritual-agent-workspaceHerbarium-Atlas-Field-Notes
Herbarium Atlas — Field Notes
This card contains field-note transcriptions from regional herbarium surveys.
Review status: pending
Reuse declarations
collection_key
notice
reviewed_on
alpine-ferns
CC BY 4.0
2025-02-14
coastal-mosses
All rights reserved
2025-01-09
desert-seeds
CC0 1.0
2025-03-02
river-reeds
CC BY-NC 4.0
2025-02-20
alpine-ferns
CC BY-NC 4.0
2025-02-10
incomplete-entry
2025-02-01
tdtu_vqa_dataset_herb
TDTU VQA Dataset — Vietnamese Medicinal Herbs 🌿
Dataset Description
TDTU VQA Dataset Herb is a Vietnamese Visual Question Answering (VQA) dataset focused on medicinal plants and herbs. It was developed for scientific research at Ton Duc Thang University (TDTU), with the goal of advancing AI models capable of recognizing and answering questions about Vietnamese medicinal herbs.
Homepage: Hugging Face Dataset
Repository: azan100an/tdtu_vqa_dataset_herb
Point of Contact:… See the full description on the dataset page: https://huggingface.co/datasets/azan100an/tdtu_vqa_dataset_herb.Herbarium-Atlas-Annotations
Herbarium Atlas — Annotations
This card contains normalized annotations prepared for the Atlas release.
Reuse declarations
collection_key
notice
reviewed_on
alpine-ferns
CC BY 4.0
2025-02-14
coastal-mosses
CC BY-NC 4.0
2025-01-12
lake-lichens
CC0 1.0
2025-03-11
river-reeds
CC BY 4.0
2025-02-20
river-reeds
CC BY 4.0
2025-02-20
missing-date
CC BY 4.0
Local-Herbs-Datasetherbal-chatbot-training-data
Dataset Training MedGemma LoRA
Folder ini berisi dataset yang dibuat dari dokumen pada data/source, termasuk PDF lokal dan sumber eksternal terverifikasi yang disimpan di data/source/herbal/external.
File Output
medgemma_lora_train.jsonl: dataset utama untuk supervised fine-tuning MedGemma LoRA.
medgemma_lora_validation.jsonl: dataset validasi.
medical_diagnosis_safety_train.jsonl: subset medical interview untuk diagnosis awal dan safety response.… See the full description on the dataset page: https://huggingface.co/datasets/Tillie2026/herbal-chatbot-training-data.Remem
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
Abstract
While Large Vision-Language Models (LVLMs) offer powerful capabilities, they pose privacy risks by unintentionally memorizing sensitive personal information. Current unlearning benchmarks attempt to mitigate this using fictitious identities but overlook a critical stage 1 failure: models fail to effectively memorize target information initially, rendering… See the full description on the dataset page: https://huggingface.co/datasets/herbwood27/Remem.herbier_mesuem6
