CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Narsil /image_dummy\audion<1K0 likes146k downloads5y agoHugging Face02hf-internal-testing /librispeech_asr_dummyaudion<1K11 likes101k downloads2y agoHugging Face03klieret /swe-bench-dummy-test-datasettextn<1K0 likes75k downloads1y agoHugging Face04BramVanroy /fineweb-2-duckdbs DuckDB datasets for (dump, id) querying on FineWeb 2 This repo contains some DuckDB databases to check whether a given WARC UID exists in a FineWeb-2 dump. Usage example is given below, but note especially that if you are using URNs (likely, if you are working with CommonCrawl data), then you first have to extract the UID (the id column is of type UUID in the databases). Download All files: huggingface-cli download BramVanroy/fineweb-2-duckdbs --local-dir… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/fineweb-2-duckdbs.0 likes64k downloads1y agoHugging Face05echodict /KakologArchives_duplicate ニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA などの家電への対応が軒並み終了する中、当時の生の声が詰まった約11年分の過去ログも同時に失われることとなってしまいました。 そこで 5ch の DTV 板の住民が中心となり、旧ニコニコ実況が終了するまでに11年分の全チャンネルの過去ログをアーカイブする計画が立ち上がりました。紆余曲折あり Nekopanda 氏が約11年分のラジオや BS も含めた全チャンネルの過去ログを完璧に取得してくださったおかげで、11年分の過去ログが電子の海に消えていく事態は回避できました。しかし、旧 API が廃止されてしまったため過去ログを API… See the full description on the dataset page: https://huggingface.co/datasets/echodict/KakologArchives_duplicate.text-classification0 likes36k downloads6mo agoHugging Face06patrickvonplaten /librispeech_asr_dummyLibriSpeech is a corpus of approximately 1000 hours of read English speech with sampling rate of 16 kHz, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned. Note that in order to limit the required storage for preparing this dataset, the audio is stored in the .flac format and is not converted to a float32 array. To convert, the audio file to a float32 array, please make use of the `.map()` function as follows: ```python import soundfile as sf def map_to_array(batch): speech_array, _ = sf.read(batch["file"]) batch["speech"] = speech_array return batch dataset = dataset.map(map_to_array, remove_columns=["file"]) ```1 likes21k downloads5y agoHugging Face07dungpham30833 /dungpham308330 likes19k downloads19d agoHugging Face08duainsan /KEGUNAAN0 likes16k downloads3y agoHugging Face09qualialabsAI /DuplexConv DuplexConv DuplexConv is a large-scale Chinese multi-channel conversational speech dataset with LLM-assisted annotations, developed by ASLP@NPU and QualiaLabs as part of the SmoothConv–DuplexConv corpus family. Companion dataset: SmoothConv on HuggingFace (100 hours, expert human annotation). DuplexConv and SmoothConv share the same conversational domains and a unified data design. SmoothConv focuses on high-quality human annotations for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/qualialabsAI/DuplexConv.audio100K<n<1M13 likes16k downloads3mo agoHugging Face10benoit-dufumier /openBHBOpenBHB: a Multi-Site Brain MRI Dataset for Age Prediction and Debiasing The Open Big Healthy Brains (OpenBHB) dataset is a large (N>5000) multi-site 3D brain MRI dataset gathering 10 public datasets (IXI, ABIDE 1, ABIDE 2, CoRR, GSP, Localizer, MPI-Leipzig, NAR, NPC, RBP) of T1 images acquired across 93 different centers, spread worldwide (North America, Europe and China). Only healthy controls have been included in OpenBHB with age ranging from 6 to 88 years old, balanced between males and… See the full description on the dataset page: https://huggingface.co/datasets/benoit-dufumier/openBHB.text1K<n<10K7 likes15k downloads1y agoHugging Face11dugthemanjana888 /storageimage0 likes13k downloads2y agoHugging Face12hf-internal-testing /dummy-audio-samplesaudion<1K0 likes12k downloads13h agoHugging Face13dungle51248 /dungle512487 likes12k downloads18d agoHugging Face14hf-internal-testing /dummy_image_text_data Dataset Card for "dummy_image_text_data" More Information needed imagen<1K1 likes12k downloads4y agoHugging Face15BramVanroy /wikipedia_culturax_dutch Filtered CulturaX + Wikipedia for Dutch This is a combined and filtered version of CulturaX and Wikipedia, only including Dutch. It is intended for the training of LLMs. Different configs are available based on the number of tokens (see a section below with an overview). This can be useful if you want to know exactly how many tokens you have. Great for using as a streaming dataset, too. Tokens are counted as white-space tokens, so depending on your tokenizer, you'll likely end up… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/wikipedia_culturax_dutch.texttext-generation1B<n<10B6 likes12k downloads2y agoHugging Face16dubsta /xmasimagen<1K1 likes11k downloads2mo agoHugging Face17dungngo59568 /dungngo595681 likes9.3k downloads19d agoHugging Face18duy19955 /duy199559 likes9.1k downloads2mo agoHugging Face19BangumiBase /dungeonnideaiwomotomerunowamachigatteirudaroukavhoujounomegamihen Bangumi Image Base of Dungeon Ni Deai Wo Motomeru No Wa Machigatteiru Darou Ka V: Houjou No Megami-hen This is the image base of bangumi Dungeon ni Deai wo Motomeru no wa Machigatteiru Darou ka V: Houjou no Megami-hen, we detected 132 characters, 6223 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/dungeonnideaiwomotomerunowamachigatteirudaroukavhoujounomegamihen.image1K<n<10K0 likes8.7k downloads1y agoHugging Face20Narsil /asr_dummySelf-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV). The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various tasks with minimal adaptation. However, the speech processing community lacks a similar setup to systematically explore the paradigm. To bridge this gap, we introduce Speech processing Universal PERformance Benchmark (SUPERB). SUPERB is a leaderboard to benchmark the performance of a shared model across a wide range of speech processing tasks with minimal architecture changes and labeled data. Among multiple usages of the shared model, we especially focus on extracting the representation learned from SSL due to its preferable re-usability. We present a simple framework to solve SUPERB tasks by learning task-specialized lightweight prediction heads on top of the frozen shared model. Our results demonstrate that the framework is promising as SSL representations show competitive generalizability and accessibility across SUPERB tasks. We release SUPERB as a challenge with a leaderboard and a benchmark toolkit to fuel the research in representation learning and general speech processing. Note that in order to limit the required storage for preparing this dataset, the audio is stored in the .flac format and is not converted to a float32 array. To convert, the audio file to a float32 array, please make use of the `.map()` function as follows: ```python import soundfile as sf def map_to_array(batch): speech_array, _ = sf.read(batch["file"]) batch["speech"] = speech_array return batch dataset = dataset.map(map_to_array, remove_columns=["file"]) ```0 likes7.8k downloads2y agoHugging Face21ducthanh1995 /ducthanh19959 likes7.4k downloads18d agoHugging Face22dungho50022 /dungho500220 likes7.1k downloads42m agoHugging Face23INS-IntelligentNetworkSolutions /Waste-Dumpsites-DroneImagery Dataset for Waste/Dumpsite Detection using drone imagery Contains 2115 drone images of illegal waste dumpsites 1280 x 1280 px resolution Nadir perspective (camera pointing straight down at a 90-degree angle to the ground) Annotations and Images train | valid | test actual images COCO - annotations_coco.json files in each split directory .parquet files in data directory with embeded images The dataset was collected as part of the [ Raven Scan ] project, more… See the full description on the dataset page: https://huggingface.co/datasets/INS-IntelligentNetworkSolutions/Waste-Dumpsites-DroneImagery.imageobject-detection10K<n<100K8 likes7k downloads2y agoHugging Face24otoearth /otoSpeech-full-duplex-turn-104hgated Dataset Card for otoSpeech-full-duplex-turn-104h Contact Website: https://oto.earthEmail: agent@oto.earth Dataset Summary otoSpeech-full-duplex-turn-104h is an English, full-duplex conversational speech dataset for research on turn-taking and related spoken-dialogue phenomena. It contains 420 two-speaker conversations totaling approximately 104.94 hours. Each conversation includes time-aligned, channel-separated audio, a stereo combined recording… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-turn-104h.audioaudio-to-audio1K<n<10K11 likes6.6k downloads24d agoHugging Face25dungcoco /9router-data11 likes6.2k downloads7m agoHugging Face26dungcoco /9router-data27 likes6.2k downloads2m agoHugging Face27dungcoco /9router-data38 likes6.2k downloads9m agoHugging Face28RedTachyon /tmlr-md-dumpimage1K<n<10K0 likes5.9k downloads2y agoHugging Face29duskstride /dts0 likes5.7k downloads1y agoHugging Face30DuplexGen /duplexgen-spoken DuplexGen Spoken Rendered spoken audio for the DuplexGen turn-taking dialogues — the exact set of clips used to fine-tune the full-duplex model (PP-DG) in DuplexGen: Adaptive Synthesis of Human–AI Turn-Taking Dialogues. Each clip is a full render of one generated dialogue variation: the mixed two-speaker dialogue audio, the isolated per-turn utterances, and the inserted backchannel clips, plus per-clip metadata. Audio is synthesized with Chatterbox TTS; the dialogue text it… See the full description on the dataset page: https://huggingface.co/datasets/DuplexGen/duplexgen-spoken.text-to-speech10K<n<100K4 likes5.6k downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.