datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meetingbank
Overview
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/meetingbank.ViFinQA
ViFinQA Dataset
Dataset Description
ViFinQA is a corpus-level dataset for Vietnamese financial question answering and numerical reasoning over annual financial statements. This public release contains 1,012 Vietnamese questions and 1,973 OCR-extracted reports from 100 Vietnamese listed companies, covering 2015–2025.
The dataset can support document retrieval, retrieval-augmented generation (RAG), financial information extraction, table understanding, and… See the full description on the dataset page: https://huggingface.co/datasets/HuuDong03uet/ViFinQA.saas-chatbot-v4
SaaS Chatbot V4 Dataset
Multi-industry, multilingual conversational dataset for fine-tuning LLMs as SaaS AI chatbot agents with tool calling.
Stats
Metric
Value
Train
4,043
Test
450
Total messages
64,645
Avg msgs/conv
14.4
Think blocks
29,345 (21% empty)
Tool calls
15,215
Tool responses
15,387
Industries (8)
E-commerce (1,301), Travel (641), Services (504), Food (490), Beauty (478), Healthcare (404), Education (357), Real Estate… See the full description on the dataset page: https://huggingface.co/datasets/huutho13254/saas-chatbot-v4.Wu-kong
Wu-kong Dataset
This dataset provides a comprehensive knowledge base for the game "Black Myth: Wukong". It is derived from detailed game guides (including IGN's guide) and is structured to support Question Answering (QA) and Retrieval-Augmented Generation (RAG) tasks.
The dataset includes walkthroughs, boss strategies, item descriptions, and game mechanics explanations, making it an ideal resource for building game companion agents or testing RAG systems on domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/huuuuuz/Wu-kong.test-codeact-pretrainingDeFineTest set and Data Resources for analogical reasoning with earnings call transcripts in research: DeFine: Decision-Making with Analogical Reasoning over Factor Profiles Yebowen Hu, Xiaoyang Wang, Wenlin Yao, Yiming Lu, Daoan Zhang, Hassan Foroosh, Dong Yu, Fei Liu Accepted to findings of ACL 2025, Vienna, Austria, USA 📄 Arxiv Paper
🏠 Home Page
🐙 Github
Abstract
LLMs are ideal for decision-making thanks to their ability to reason over long contexts. However, challenges… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/DeFine.firstatomic_synthhumanoid_inverse_kinematicshuu-ontocord__wide_3b_orpo_stage1.1-ss1-orpo3-details
Dataset Card for Evaluation run of huu-ontocord/wide_3b_orpo_stage1.1-ss1-orpo3
Dataset automatically created during the evaluation run of model huu-ontocord/wide_3b_orpo_stage1.1-ss1-orpo3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huu-ontocord__wide_3b_orpo_stage1.1-ss1-orpo3-details.humanoid_gait_phaseshuman-daily-instructionsITSupport_TCIS_fine_tunningtest_2_record
