huu
Datasets
All datasets matching “huu”MeetingBank_Audio
Overview
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/MeetingBank_Audio.meetingbank
Overview
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/meetingbank.mmlu-sars-cov-2CameraOperator-BlockCam
CameraOperator-BlockCam (Synthetic)
CameraOperator-BlockCam (Synthetic) contains 37,499 repaired-and-audited synthetic annotation-label trajectory records pairing English camera-motion descriptions, 150-frame camera trajectories, and time-varying target-object 3D oriented bounding boxes (OBBs). They are grouped into 5,180 reconstructed source events and 13,121 augmentation families; the 37,499 records should not be interpreted as 37,499 independent scenes or events. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/huuuuuuuuu/CameraOperator-BlockCam.FIOVA
FIOVA: Five-In-One Video Annotations
FIOVA is a benchmark for detailed video captioning with five independent human descriptions per video and a Unified Consensus Groundtruth (UCG). Differences among the five descriptions support analysis of human disagreement, while their shared content provides a reference for event-based evaluation.
Dataset
3,002 videos across 38 themes, with an average duration of 33.6 seconds.
15,010 English human descriptions, five per… See the full description on the dataset page: https://huggingface.co/datasets/huuuuusy/FIOVA.SportsMetrics
SportsMetrics
Benchmark data to evaluate numerical reasoning and information fusion of LLMs.
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL'24), Bangkok, Thailand. Arxiv Paper
Usage
from datasets import load_dataset
def get_task(domain… See the full description on the dataset page: https://huggingface.co/datasets/huuuyeah/SportsMetrics.
