CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huuuyeah /meetingbank Overview MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/meetingbank.textsummarization1K<n<10K36 likes796 downloads1y agoHugging Face02huuuyeah /SportsMetrics SportsMetrics Benchmark data to evaluate numerical reasoning and information fusion of LLMs. SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL'24), Bangkok, Thailand. Arxiv Paper Usage from datasets import load_dataset def get_task(domain… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsMetrics.textquestion-answering1K<n<10K4 likes88 downloads2y agoHugging Face03huutho13254 /saas-chatbot-v4 SaaS Chatbot V4 Dataset Multi-industry, multilingual conversational dataset for fine-tuning LLMs as SaaS AI chatbot agents with tool calling. Stats Metric Value Train 4,043 Test 450 Total messages 64,645 Avg msgs/conv 14.4 Think blocks 29,345 (21% empty) Tool calls 15,215 Tool responses 15,387 Industries (8) E-commerce (1,301), Travel (641), Services (504), Food (490), Beauty (478), Healthcare (404), Education (357), Real Estate… See the full description on the dataset page: https://huggingface.co/datasets/huutho13254/saas-chatbot-v4.texttext-generation1K<n<10K0 likes40 downloads6mo agoHugging Face04huuuyeah /SportsGenDataset and scripts for sports analyzing tasks proposed in research: When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Wenlin Yao, Hassan Foroosh, Dong Yu, Fei Liu Accepted to main conference of EMNLP 2024, Miami, Florida, USA Arxiv Paper Abstract Reasoning is most powerful when an LLM accurately aggregates relevant information. We examine the critical role of information aggregation in… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsGen.textquestion-answering10K<n<100K6 likes33 downloads2y agoHugging Face05huunam /firsttexttext-classification100K<n<1M1 likes23 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.