datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-bench-dummy-test-datasetLora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.Long-video-test-datarlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.huatuo26M-testdatasets
Dataset Card for huatuo26M-testdatasets
Dataset Summary
We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper.
We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo26M-testdatasets.WideSeek-R1-test-data
Testing Dataset
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
test_dataset
This dataset is for testing purposes... blah blah
About the dataset
Mixture of prompt and answer completions taken from Subnet18 ...
code_net_test_final_datasetTest-Dataset
AXIS held-out 20 — a cross-embodiment few-shot adaptation benchmark
20 tasks never seen in pretraining, 20 demonstrations each, LoRA adaptation, rollout evaluation.
This release is the SPECIFICATION and the INDEX, not the demonstration data. It is published
first and on purpose: everything here is what you need to render the benchmark on your own
embodiment, and none of it depends on our video encoding being finished.
Status: the task list is CANDIDATES. The learnability gate… See the full description on the dataset page: https://huggingface.co/datasets/axisrobotics/Test-Dataset.zalo-ai-2025-public-test-data-v2
Zalo AI Challenge 2025 - RoadBuddy Public Test Data V2set
This dataset contains public test data v2 for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025.
Dataset Description
The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam.
Dataset Structure
Files
frames/:… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-public-test-data-v2.ur5fail_test_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.rag_instruct_test_dataset_0.1
Dataset Card for RAG-Instruct-Test-Dataset
Dataset Summary
This is a test dataset for basic "retrieval augmented generation" (RAG) use cases in the enterprise, especially for finance and legal. This test dataset includes 100 samples with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases,
contracts, invoices, technical articles, general news and short texts. The primary use case is to evaluate the effectiveness of an… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_test_dataset_0.1.open_lm_test_databdv2fail_test_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.speechmap-judge-rl-test-data
SpeechMap Judge RL Test Data
This is an experimental training data package for SpeechMap-style judge model
training. It is not the main SpeechMap dataset release and should not be cited
or treated as a canonical benchmark distribution.
The dataset is intended for training and evaluating a judge that labels whether
a candidate model response complies with a user request. Labels are:
COMPLETE: the user's request is handled directly and fulfilled.
EVASIVE: the response avoids… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/speechmap-judge-rl-test-data.Japanese-RP-Bench-testdata-SFW
Japanese-RP-Bench-testdata-SFW
本データセットは、LLMの日本語ロールプレイ能力を計測するベンチマークJapanese-RP-Bench用の評価データセットです。
ベンチマークの詳細については記事を参照してください。
データの概要
本データは以下のようなキーを持ちます。
genre: ロールプレイのジャンル
tag: ロールプレイの年齢区分
world_setting: ロールプレイの世界観設定
scene_setting: ロールプレイのシーン設定
user_setting: ロールプレイのユーザー側キャラクター設定
assistant_setting: ロールプレイのアシスタント側キャラクター設定
dialogue_tone: ロールプレイの対話のトーン
first_user_input: ロールプレイの最初のユーザー発話
response_format: ロールプレイの応答形式
id: データのid
ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-RP-Bench-testdata-SFW.test_datasetWideSeek-R1-test-data
Testing Dataset
🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearchdataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
Acknowledgement
Thanks to WideSearch for providing a… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-test-data.testdata001function-calling-test
function-calling-test
Test split for lm-eval-harness.
1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr
Specifications
Data content
한국어 K12 시험 문제
Amount
약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.toolathlon-test-dataset-9862TestDatasetScratch repo for testing dataset-viewer schema inference. All data is synthetic and
contains no real personal information.
case_bank/ holds a few raw per-case JSON files whose nested mock block is union-typed
across tools (some lists are empty, some hold structs), which breaks single-schema
inference. viewer/cases.jsonl is a flattened table with a stable schema, and the
configs block above points the viewer at it so the raw files are not globbed.
1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.StrikeGPT-R1-Zero-Test-CoT_DataNetwork Security Reasoning Dataset(test)
Obtained through distillation of DeepSeek-R1
think to CoT code:https://github.com/Bouquets-ai/Data-Processing/blob/main/think%20to%20CoT.py
testdatafalcon-small-test-datasetrag_instruct_test_dataset2_financial_0.1
Dataset Card for RAG-Instruct-Financial-Test-Dataset
Dataset Summary
This is a test dataset for "retrieval augmented generation" (RAG) use cases, especially for financial data extraction and analysis, including a series of questions relating to tabular financial data and common-sense math operations (small increments, decrements, sorting and ordering as well as recognizing when information is not included in a particular source). This test dataset includes 100 samples… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_test_dataset2_financial_0.1.LLM_Test_datasetreference from: huggingface: timdettmers/openassistant-guanaco
https://huggingface.co/datasets/timdettmers/openassistant-guanaco
lora-testdataset
