CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jazzypajamas /mytown-local-gov-meetings MyTown — open dataset of US & Canadian local-government meetings The documents themselves, not just the metadata. Most civic datasets publish meeting titles, dates and links. This one publishes 2,109,683 full text extractions of the primary documents — the actual agendas and minutes, pulled out of the PDFs — alongside 11,949,495 per-member roll-call votes and 61,661,080 campaign-finance transactions, all joinable on the same keys. That combination is the point: you can go from… See the full description on the dataset page: https://huggingface.co/datasets/jazzypajamas/mytown-local-gov-meetings.summarization1M<n<10M1 likes1.4k downloads8d agoHugging Face02Jazzcharles /audioverse_for_annotation_soundlyaudion<1K0 likes128 downloads6mo agoHugging Face03LeData /jazz-music-archivestexttext-classification10K<n<100K2 likes124 downloads2y agoHugging Face04TheMindExpansionNetwork /jimsky-radio-jazz-lab Jimsky Radio Jazz Lab One original character theme and two historical jazz melody-conditioned cover tests, generated on an RTX 4070 using the supplied ComfyUI workflows. Three audio tests and all three planned artwork pieces succeeded (artwork generated via Comfy Cloud with the supplied workflow nodes; the desktop variant remains blocked by desktop API-node authorization). Listen Raw Raw on the Radio — 120.00s · MP3 · FLAC Livery Stable Blues — glitch-hop cover… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/jimsky-radio-jazz-lab.audio-to-audio0 likes107 downloads8d agoHugging Face05Jazzcharles /Egoinstructor_downstream_metadata 📙 Overview This repo contains the metafiles to evaluate the performance of EgoInstructor, including Epic-Kitchen multi-instance retrieval YouCook videoclip-text retrieval Charadesego egovideo-exovideo retrieval EgoLearn egovideo-exovideo retrieval Ego4d Summarization multiple choice question YouCook video-text retrieval The following metafiles are used for retrieval-augmented egocentric video captioning, including E4DOL_crossview_train_instructions.json… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/Egoinstructor_downstream_metadata.video-classification10B<n<100B0 likes86 downloads2y agoHugging Face06Jazzresin /JoyceCacheimagen<1K1 likes77 downloads2mo agoHugging Face07jazzysnake01 /oasst1-en-hun-gemini Open assistant 1 dataset hungarian translation (english subset) This dataset contains hungarian translations for the oasst1 dataset's english subset. The translations were done via gemini pro and the model was instructed to keep stlye, meaning and english entites as they are. I think this produced a higher quality translation than google translate, but even this version is far from perfect. The exact code used for creating the dataset can be found here. license:… See the full description on the dataset page: https://huggingface.co/datasets/jazzysnake01/oasst1-en-hun-gemini.tabular10K<n<100K2 likes71 downloads3y agoHugging Face08jazznilky /ftextn<1K0 likes62 downloads22d agoHugging Face09Jazzcharles /EgoHOD2 likes45 downloads2y agoHugging Face10AlekseyKorshuk /gpt4all-jazzy-chatml Dataset Card for "gpt4all-jazzy-chatml" More Information needed text100K<n<1M4 likes42 downloads3y agoHugging Face11Jazzcharles /ego4d_videomae_L14_feature_fps8 📙 Overview Ego4d video features extracted by VideoMAE_L14 at 8 fps. It contains 9645 files, each file (e.g. fffbaeef-577f-45f0-baa9-f10cabf62dfb.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓 Citation @article{xu2024retrieval, title={Retrieval-augmented egocentric video captioning}, author={Xu, Jilan and Huang, Yifei and Hou, Junlin and Chen, Guo… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_videomae_L14_feature_fps8.textvideo-classification1K<n<10K0 likes39 downloads2y agoHugging Face12jazzysnake01 /quizgen-chat-mdtext10K<n<100K2 likes33 downloads2y agoHugging Face13Jazzres /SchaferEveResearchdocumentn<1K1 likes32 downloads2mo agoHugging Face14Jazzcharles /youcook2_internvideo_MM_L14_features_fps8 📙 Overview YouCook2 video features extracted by InternVideo_MM_L14 at 8 fps. It is used for evaluating the video-text retrieval ability of EgoInstructor. Each file (e.g. 10dZTHlkb8w.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓 Citation @article{xu2024retrieval, title={Retrieval-augmented egocentric video captioning}, author={Xu, Jilan and Huang… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/youcook2_internvideo_MM_L14_features_fps8.textvideo-classificationn<1K0 likes29 downloads2y agoHugging Face15jspr /symbolic-jazz-standardsgated Symbolic Jazz Standards A symbolic-domain music dataset of jazz standards, transcribed stem by stem from the audio domain into the symbolic domain. The dataset contains the equivalent of 10,000 minutes of audio from ~200 public-domain well-known songs. Methodology To create this dataset, recordings of public-domain jazz standards were downloaded and separated into their component stems using the venerable Demucs source separation library in 4-stem mode. The resulting… See the full description on the dataset page: https://huggingface.co/datasets/jspr/symbolic-jazz-standards.textaudio-to-audion<1K6 likes26 downloads3y agoHugging Face16Jazzcharles /HowTo100M_llama3_refined_caption 📙 Overview The metadata for HowTo100M. The original ASR is refined by LLAMA-3 language model. Each sample represents a short video clip, which consists of vid: the initial video id. uid: a given unique id to index the clip. start_second: the timestamp of the narration. end_second: the end timestamp of the narration (which is simply set to start + 1). text: the original ASR transcript. noun: a list containing the index of nouns in the noun vocabulary. verb: a list containing the… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/HowTo100M_llama3_refined_caption.video-classification1B<n<10B2 likes25 downloads2y agoHugging Face17Jazzcharles /ego4d_train_pair_howto100m 📙 Overview The metadata for Ego4d training set, with paired howto100m video clips. The ego-exo pair is constructed by choosing the ones with shared nouns/verbs. Each sample represents a short video clip, which consists of vid: the initial video id. start_second: the start timestamp of the narration. end_second: the end timestamp of the narration. text: the original narration. noun: a list containing the index of nouns in the Ego4d noun vocabulary. verb: a list containing the… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_train_pair_howto100m.video-classification1B<n<10B0 likes23 downloads2y agoHugging Face18Jazzcharles /egolearn_videomae_internvideo_features 📙 Overview Egolearn video features. egocentric videos are extracted by VideoMAE_L14 at 8 fps. exocentric videos are extracted by InternVideo_MM_L14 at 8 fps. They are used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor. Each file (e.g. 2cad4224-56c4-11ee-88ee-80615f12b59e.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/egolearn_videomae_internvideo_features.textvideo-classificationn<1K0 likes23 downloads2y agoHugging Face19FrantzesE /jazz-solos Dataset description This dataset contains the jazz solos from the Weimar Jazz Database (https://jazzomat.hfm-weimar.de/dbformat/dboverview.html) that have been processed in various ways to be musically rhythmically accurate to the transcription sheet music they provide. The solos have been converted into the SCAMP format with the addition of the current chord to give context. Format Instruction Arbitrary text requesting a jazz solo be created from a… See the full description on the dataset page: https://huggingface.co/datasets/FrantzesE/jazz-solos.texttext-generationn<1K1 likes23 downloads2y agoHugging Face20jazza234234 /david-datasetaudio1K<n<10K5 likes23 downloads1y agoHugging Face21linoyts /wan_jazz_handsThis dataset contains videos generated using Wan 2.1 T2V 14B. textn<1K0 likes23 downloads1y agoHugging Face22Jazz1508 /tokenized-punjabi210M<n<100M0 likes23 downloads1y agoHugging Face23sebascorreia /jazz-set Dataset Card for "edm_wavset" More Information needed image1K<n<10K0 likes20 downloads3y agoHugging Face24ChaoticEconomist /Jazz-Blues-Music-Dataset_SFT-or-LoRA Jazz & Blues Music Dataset (SFT / LoRA Ready) A structured dataset covering 82 iconic Jazz and Blues songs, 21 artist profiles, and 41 historical events, expanded into 1,219 instruction-tuning rows across 7 task types. Designed for fine-tuning LLMs on music knowledge, cultural history, artist biography, and domain-specific Q&A tasks. Overview Property Value Domain Jazz & Blues Music Total rows 1,219 Train split 1,036 (85%) Validation split 91 (~7.5%)… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Jazz-Blues-Music-Dataset_SFT-or-LoRA.texttext-generation1K<n<10K0 likes20 downloads5mo agoHugging Face25jazzysnake01 /oasst-1-hun-openaitext1K<n<10K1 likes19 downloads2y agoHugging Face26PlutoG99001 /MusicGen-Jazzaudion<1K0 likes19 downloads2y agoHugging Face27PRAIG /JAZZMUSgated Optical Music Recognition of Jazz Lead Sheets 📄 Conf. materials | 🖥️ Slides | 🎶 Poster Dataset How to use Check the following code: import ast from datasets import load_dataset from PIL import ImageDraw DATASET_NAME = "PRAIG/JAZZMUS" ds = load_dataset(DATASET_NAME) image = ds["train"][0]["image"] # list of systems, with bounding boxes and encoding systems = ast.literal_eval(ds["train"][0]["annotation"])["systems"] # full page encodings encoding… See the full description on the dataset page: https://huggingface.co/datasets/PRAIG/JAZZMUS.imageimage-to-textn<1K3 likes19 downloads7mo agoHugging Face28LeData /media-metadata-jazz-artists TigreGotico/media-metadata-jazz-artists Rich entity dataset scraped by metadatarr scraper jazz_artists. Rows: 7,941 Fields artist_slug name genres genre country bio url n_albums Source Generated by scrapers/jazz_artists.py. See the metadatarr repo for the full pipeline and scraper source code. text1K<n<10K0 likes18 downloads3mo agoHugging Face29eigenben /jazz-harmony-embeddings Jazz Harmony Embeddings — 6,900 tune vectors One 128-dimensional vector per jazz standard, from a small transformer trained from scratch so that tunes with related harmony — transpositions, alternate charts, contrafacts — land close together. Produced by the 3-seed ensemble released at eigenben/jazz-harmony-embeddings; code and full experiment records at github.com/eigenben/jazz-harmony-embeddings. Files embeddings.npz — embeddings: (6900, 128) float32… See the full description on the dataset page: https://huggingface.co/datasets/eigenben/jazz-harmony-embeddings.tabular1K<n<10K0 likes18 downloads2mo agoHugging Face30Jazzcharles /charadesego_videomae_L14_feature_fps8 📙 Overview CharadesEgo video features extracted by VideoMAE_L14 at 8 fps. It is used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor. It contains 7860 files, each file (e.g. 005BUEGO.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768. 🏋️ How-To-Use Please refer to code EgoInstructor for details. 🎓 Citation @article{xu2024retrieval, title={Retrieval-augmented egocentric video captioning}… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/charadesego_videomae_L14_feature_fps8.textvideo-classification1K<n<10K0 likes16 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.