CoolFace
29 results

AV

Avelina /smollm-corpus SmolLM-Corpus: Now shuffled and sharded! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23 containing 1000 shards… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus.text-generation100M<n<1B5 likes20k downloads2y agoHugging FaceProgramComputer /avspeech-visual-audio AVSpeech Video + Audio This repository is a media-bearing reconstruction of the public AVSpeech annotations. Each row represents an already-trimmed segment and keeps the original source-video timing and target-face-center metadata. Dataset structure clip_id: identifier derived as {youtube_id}_{start_sec:.3f}_{end_sec:.3f}. avspeech_metadata: JSON containing youtube_id, start_sec, end_sec, x_center, and y_center from the AVSpeech annotation. video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.audio1M<n<10M5 likes18k downloads1mo agoHugging FaceAvelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Facetsinghua-ee /AVUTBenchmark Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text Shortcut. Code Repository: https://github.com/lark-png/AVUT Paper: https://arxiv.org/pdf/2503.19951 Introduction The Audio-centric Video Understanding Benchmark (AVUT) aims to evaluate the video comprehension capabilities of multimodal Large Language Models (LLMs), with a particular focus on auditory information. Audio… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/AVUTBenchmark.videovideo-text-to-text1K<n<10K2 likes9.8k downloads1y agoHugging Faceavery00 /MomentSeekerMomentSeeker: A Comprehensive Benchmark and A Strong Baseline For Moment Retrieval Within Long Videos This repo contains the annotation data for the paper "MomentSeeker: A Comprehensive Benchmark and A Strong Baseline For Moment Retrieval Within Long Videos". 🔔 News: 🥳 2025/03/07: We have released the MomentSeeker Benchmark and Paper! 🔥 Introduction We present MomentSeeker, a comprehensive benchmark to… See the full description on the dataset page: https://huggingface.co/datasets/avery00/MomentSeeker.question-answering8 likes9k downloads8mo agoHugging Faceavalab /Allo-AVAaudion>1T3 likes8.3k downloads2y agoHugging Face