datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-bench-dummy-test-datasettokenizers-test-data
tokenizers-test-data
Test and benchmark fixtures for huggingface/tokenizers,
pulled on demand by the repo Makefiles (make test / make bench / make fixtures
via hf download).
Layout
fixtures/ — multilingual + modality corpora for cross-language encode
benchmarks. Organized, documented, and reproducible: see
fixtures/FIXTURES.md for provenance and
fixtures/fixtures_manifest.json for
exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.dataset-test-1hindi_audio_dataset_testLora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.protein_data_testsplit 1, 2 -> for sequences
split 3, 4 -> for residues
test-datachain-llm-evalW_LSTMix_test_datasetsmart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test.whisperkit-test-dataLong-video-test-dataprotein_data_test_2openadmet-expansionrx-challenge-test-data-blinded
OpenADMET-ExpansionRx Challenge blinded test dataset
This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-test-data-blinded.external_data_test_examplechunking-test-datatestdataffv4_dataset_testthis is a testing dataset for future model testing. you should not use this (yet)
there are multiple datasets,
notebook_defaults
notebook_defaults_ratio0.8_likes10
you can load each like this:
import datasets
# see FFV4.BUILDER_CONFIGS for all possible names
ds = datasets.load_dataset('./dataset_code.py', name='notebook_defaults_ratio0.8_likes10')
then use them like this
ds_real = ds['everything'] # there is no such thing as a train/test split here
one_item = ds_real[0] # grab first story… See the full description on the dataset page: https://huggingface.co/datasets/main-horse/ffv4_dataset_test.test-datasetIUXray-Data-Train-Testtable_rec_test_dataset
表格识别测试集
数据集简介
该数据集包含百度生成工具 20 张有线 20 张无线,wtw 数据集 15, pubnet val 集 20 张,自我零散标注 18 张,共计 93 张表格图片,涵盖多种场景、不同光照条件、不同的图像分辨率。
该数据集可以结合 表格指标评测库-TableRecognitionMetric 使用,快速评测各种表格还原算法。
关于该数据集,欢迎小伙伴贡献更多数据呦!有任何想法,可以前往 issue讨论。
如果遇到标注有误的,还请指出。
数据集支持的任务
可用于自定义数据集下的模型验证和性能评估等。
数据集的格式和结构
数据格式
数据集只有测试集,仅用于客观评估算法表现。
data
└── test
├── images
│ ├── 000cce9ca593055d4618466e823e6d7c.jpg
│ ├── 0aNtiNtRRLqEZ9y6PuShtAAAACMAAQED.jpg
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/SWHL/table_rec_test_dataset.Bread-chatbot-dataset-test
Dataset Card for "Bread-chatbot-dataset-test"
More Information needed
test-big-dataset
Dataset Card for Danish WIT
Dataset Summary
Google presented the Wikipedia Image Text (WIT) dataset in July
2021, a dataset which contains
scraped images from Wikipedia along with their descriptions. WikiMedia released
WIT-Base in September
2021,
being a modified version of WIT where they have removed the images with empty
"reference descriptions", as well as removing images where a person's face covers more
than 10% of the image surface, along with inappropriate images… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/test-big-dataset.audio_test_dataset
Dataset Card for "audio_test_dataset"
This dataset consists of the first 5 samples of mozilla-foundation/common_voice_13_0 and is only used for unit testing.
rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.bioreason-pro-test-data
🧬 BioReason-ProAdvancing Protein Function Prediction withMultimodal Biological Reasoning
BioReason-Pro Test Data
Evaluation dataset for BioReason-Pro. Contains proteins with GO term annotations, GO-GPT predictions, InterPro domains, STRING protein-protein interactions, and protein metadata. Temporal holdout follows the CAFA framework.
Citation
If you find this work useful, please cite our papers:
@article {Fallahpour2026.03.19.712954… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/bioreason-pro-test-data.smart-product-pricing-2025_test_dataprocessed_sroie_donut_dataset_train_test_split
Dataset Card for "processed_sroie_donut_dataset_train_test_split"
More Information needed
test_tsv_datasetssynth_data-test
S-SYNTH
S-SYNTH is an open-source, flexible skin simulation framework to rapidly generate synthetic skin models and images using digital rendering of an anatomically inspired multi-layer, multi-component skin and growing lesion model. It allows for generation of highly-detailed 3D skin models and digitally rendered synthetic images of diverse human skin tones, with full control of underlying parameters and the image formation process.
Framework Details
S-SYNTH… See the full description on the dataset page: https://huggingface.co/datasets/didsr/ssynth_data-test.huatuo26M-testdatasets
Dataset Card for huatuo26M-testdatasets
Dataset Summary
We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper.
We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo26M-testdatasets.
