datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MeetingBank_Audio
Overview
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/MeetingBank_Audio.meetingbank
Overview
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/meetingbank.mmlu-sars-cov-2CameraOperator-BlockCam
CameraOperator-BlockCam (Synthetic)
CameraOperator-BlockCam (Synthetic) contains 37,499 repaired-and-audited synthetic annotation-label trajectory records pairing English camera-motion descriptions, 150-frame camera trajectories, and time-varying target-object 3D oriented bounding boxes (OBBs). They are grouped into 5,180 reconstructed source events and 13,121 augmentation families; the 37,499 records should not be interpreted as 37,499 independent scenes or events. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/huuuuuuuuu/CameraOperator-BlockCam.FIOVA
FIOVA: Five-In-One Video Annotations
FIOVA is a benchmark for detailed video captioning with five independent human descriptions per video and a Unified Consensus Groundtruth (UCG). Differences among the five descriptions support analysis of human disagreement, while their shared content provides a reference for event-based evaluation.
Dataset
3,002 videos across 38 themes, with an average duration of 33.6 seconds.
15,010 English human descriptions, five per… See the full description on the dataset page: https://huggingface.co/datasets/huuuuusy/FIOVA.SportsMetrics
SportsMetrics
Benchmark data to evaluate numerical reasoning and information fusion of LLMs.
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL'24), Bangkok, Thailand. Arxiv Paper
Usage
from datasets import load_dataset
def get_task(domain… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsMetrics.CYCLO_INTELLIGENSE_AIWORKER
Task_1_lift_carton_box_MCAP
Created with Cyclo Intelligence by ROBOTIS.
long_squad_v2
Dataset Card for long_squad_v2
long_squad_v2 is a long-context question answering dataset based on the SQuAD v2 format. It was constructed by concatenating multiple SQuAD v2 contexts to significantly increase the average document length, enabling training and evaluation of models on long-range understanding and sparse answer retrieval tasks.
Dataset Details
Uses
To load the dataset using the 🤗 Datasets library:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/huutuan/long_squad_v2.pick_place3_20260706_142028This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HuuKhang2212/pick_place3_20260706_142028.56image-data56image-data-0831vifinqa
Hướng Dẫn Phát Triển Hệ Thống ViFinQA (Text-to-Pandas)
Dự án tự động hóa việc tra cứu và tính toán chỉ số tài chính từ Báo cáo tài chính (BCTC) Việt Nam dạng OCR (.txt) sang mã Pandas thực thi được.
1. Cấu Trúc Thư Mục Dự Án
Road-to-AI/
├── data/
│ ├── raw_vifinqa/ # Kho BCTC dạng văn bản (.txt) và questions.jsonl từ Ban Tổ Chức
│ ├── processed_csv/ # Kết quả bóc tách tự động (.csv) từ file .txt
│ ├── mock_csv/ # Dữ liệu giả cũ… See the full description on the dataset page: https://huggingface.co/datasets/huunghiac/vifinqa.56image-data-072056image-data-081756image-data-0907ViFinQA
ViFinQA Dataset
Dataset Description
ViFinQA is a corpus-level dataset for Vietnamese financial question answering and numerical reasoning over annual financial statements. This public release contains 1,012 Vietnamese questions and 1,973 OCR-extracted reports from 100 Vietnamese listed companies, covering 2015–2025.
The dataset can support document retrieval, retrieval-augmented generation (RAG), financial information extraction, table understanding, and… See the full description on the dataset page: https://huggingface.co/datasets/HuuDong03uet/ViFinQA.pick_place8_20260708_142526This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HuuKhang2212/pick_place8_20260708_142526.56image-data-0810saas-chatbot-v4
SaaS Chatbot V4 Dataset
Multi-industry, multilingual conversational dataset for fine-tuning LLMs as SaaS AI chatbot agents with tool calling.
Stats
Metric
Value
Train
4,043
Test
450
Total messages
64,645
Avg msgs/conv
14.4
Think blocks
29,345 (21% empty)
Tool calls
15,215
Tool responses
15,387
Industries (8)
E-commerce (1,301), Travel (641), Services (504), Food (490), Beauty (478), Healthcare (404), Education (357), Real Estate… See the full description on the dataset page: https://huggingface.co/datasets/huutho13254/saas-chatbot-v4.pick_place4_20260706_171158This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HuuKhang2212/pick_place4_20260706_171158.Wu-kong
Wu-kong Dataset
This dataset provides a comprehensive knowledge base for the game "Black Myth: Wukong". It is derived from detailed game guides (including IGN's guide) and is structured to support Question Answering (QA) and Retrieval-Augmented Generation (RAG) tasks.
The dataset includes walkthroughs, boss strategies, item descriptions, and game mechanics explanations, making it an ideal resource for building game companion agents or testing RAG systems on domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/huuuuuz/Wu-kong.pick_place_20260706_122159This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HuuKhang2212/pick_place_20260706_122159.pick_place6_20260707_170648This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HuuKhang2212/pick_place6_20260707_170648.pick_place7_20260708_115602This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HuuKhang2212/pick_place7_20260708_115602.data_leader_r35_37pen_holder_20260626_145717This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HuuKhang2212/pen_holder_20260626_145717.pick_place5_20260707_132345This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/HuuKhang2212/pick_place5_20260707_132345.CYCLO_INTELLIGENSE_SH5
Task_1_lift_carton_box_MCAP
Created with Cyclo Intelligence by ROBOTIS.
56image-data-080356image-data-0914
