CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huuuyeah /meetingbank Overview MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/meetingbank.textsummarization1K<n<10K36 likes796 downloads1y agoHugging Face02HuuDong03uet /ViFinQA ViFinQA Dataset Dataset Description ViFinQA is a corpus-level dataset for Vietnamese financial question answering and numerical reasoning over annual financial statements. This public release contains 1,012 Vietnamese questions and 1,973 OCR-extracted reports from 100 Vietnamese listed companies, covering 2015–2025. The dataset can support document retrieval, retrieval-augmented generation (RAG), financial information extraction, table understanding, and… See the full description on the dataset page: https://huggingface.co/datasets/HuuDong03uet/ViFinQA.textquestion-answering1K<n<10K0 likes47 downloads2mo agoHugging Face03huutho13254 /saas-chatbot-v4 SaaS Chatbot V4 Dataset Multi-industry, multilingual conversational dataset for fine-tuning LLMs as SaaS AI chatbot agents with tool calling. Stats Metric Value Train 4,043 Test 450 Total messages 64,645 Avg msgs/conv 14.4 Think blocks 29,345 (21% empty) Tool calls 15,215 Tool responses 15,387 Industries (8) E-commerce (1,301), Travel (641), Services (504), Food (490), Beauty (478), Healthcare (404), Education (357), Real Estate… See the full description on the dataset page: https://huggingface.co/datasets/huutho13254/saas-chatbot-v4.texttext-generation1K<n<10K0 likes40 downloads6mo agoHugging Face04huuuuuz /Wu-kong Wu-kong Dataset This dataset provides a comprehensive knowledge base for the game "Black Myth: Wukong". It is derived from detailed game guides (including IGN's guide) and is structured to support Question Answering (QA) and Retrieval-Augmented Generation (RAG) tasks. The dataset includes walkthroughs, boss strategies, item descriptions, and game mechanics explanations, making it an ideal resource for building game companion agents or testing RAG systems on domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/huuuuuz/Wu-kong.imagequestion-answeringn<1K1 likes37 downloads9mo agoHugging Face05huu-ontocord /test-codeact-pretrainingtext1K<n<10K0 likes32 downloads5mo agoHugging Face06huuuyeah /DeFineTest set and Data Resources for analogical reasoning with earnings call transcripts in research: DeFine: Decision-Making with Analogical Reasoning over Factor Profiles Yebowen Hu, Xiaoyang Wang, Wenlin Yao, Yiming Lu, Daoan Zhang, Hassan Foroosh, Dong Yu, Fei Liu Accepted to findings of ACL 2025, Vienna, Austria, USA 📄 Arxiv Paper    🏠 Home Page    🐙 Github Abstract LLMs are ideal for decision-making thanks to their ability to reason over long contexts. However, challenges… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/DeFine.textfeature-extractionn<1K1 likes24 downloads1y agoHugging Face07huunam /firsttexttext-classification100K<n<1M1 likes23 downloads3y agoHugging Face08huu-ontocord /atomic_synthtext1K<n<10K0 likes11 downloads7mo agoHugging Face09HuuNguyen1513 /humanoid_inverse_kinematicstextn<1K0 likes7 downloads9mo agoHugging Face10open-llm-leaderboard /huu-ontocord__wide_3b_orpo_stage1.1-ss1-orpo3-detailsgated Dataset Card for Evaluation run of huu-ontocord/wide_3b_orpo_stage1.1-ss1-orpo3 Dataset automatically created during the evaluation run of model huu-ontocord/wide_3b_orpo_stage1.1-ss1-orpo3 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/huu-ontocord__wide_3b_orpo_stage1.1-ss1-orpo3-details.tabular10K<n<100K0 likes4 downloads2y agoHugging Face11HuuNguyen1513 /humanoid_gait_phasestextn<1K0 likes4 downloads9mo agoHugging Face12HuuNguyen1513 /human-daily-instructionstextn<1K0 likes3 downloads9mo agoHugging Face13huunghia0695 /ITSupport_TCIS_fine_tunningtextn<1K0 likes1 downloads1y agoHugging Face14huunghia0695 /test_2_recordtextn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.