CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Fred808 /datatabular10K<n<100K0 likes11k downloads10mo agoHugging Face02freds0 /TAGARELA TAGARELA: A Portuguese Speech Dataset From Podcasts TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page: https://huggingface.co/datasets/freds0/TAGARELA.audioautomatic-speech-recognition1M<n<10M10 likes9k downloads2mo agoHugging Face03freddyaboulton /bucket0 likes8.6k downloads6mo agoHugging Face04freddyaboulton /gradio-theme-subdomains2 likes6.3k downloads2y agoHugging Face05GabrieleMagrini /FRED ⭐ FRED: Florence RGB-Event Drone dataset ⭐ Official repository for the FRED dataset, a large-scale multimodal dataset specifically designed for drone detection, tracking, and trajectory forecasting, with spatiotemprally synchronized RGB and event data. It includes train and test splits with zipped subfolders for each sequence. The dataset can also be downloaded from here. The dataset splits in .txt format, along with the alternative challenging split, can be found here. Demos… See the full description on the dataset page: https://huggingface.co/datasets/GabrieleMagrini/FRED.image1M<n<10M9 likes5.6k downloads1y agoHugging Face06Fredithefish /Nemotron-CC-HQ-20B Nemotron-CC-HQ-20B This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007. For more information about Nemotron-CC check the Paper by Nvidia Disclaimer: Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed. Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.texttext-generation10M<n<100M1 likes2.2k downloads6mo agoHugging Face07FredyRivera-dev /LLaDA-Sample-10BT Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.text-generation1B<n<10B2 likes1.9k downloads3mo agoHugging Face08yatsbm /FREDData extracted from FRED. All data in this datasets comes from FRED series which are tagged as "Public Domain: Citation Requested". 0 likes1.2k downloads2y agoHugging Face09freds0 /OpenBHB-Dataset0 likes1.1k downloads8mo agoHugging Face10FredZhang7 /toxi-text-3MThis is a large multilingual toxicity dataset with 3M rows of text data from 55 natural languages, all of which are written/sent by humans, not machine translation models. The preprocessed training data alone consists of 2,880,667 rows of comments, tweets, and messages. Among these rows, 416,529 are classified as toxic, while the remaining 2,463,773 are considered neutral. Below is a table to illustrate the data composition: Toxic Neutral Total multilingual-train-deduplicated.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/toxi-text-3M.texttext-classification1M<n<10M32 likes671 downloads1y agoHugging Face11fredzzp /uniref-50-foldseek-v1text10M<n<100M1 likes601 downloads10mo agoHugging Face12freds0 /BRSpeech BRSpeech BRSpeech is a single-speaker Brazilian Portuguese speech dataset extracted and curated specifically for Text-to-Speech (TTS) and voice modeling tasks. It corresponds directly to speaker 2961 from the multi-speaker BRSpeech-TTS dataset, which represents the speaker with the highest volume of recorded audio/hours in the entire corpus. Dataset Summary Language: Portuguese (pt-BR) Speaker ID: 2961 (from BRSpeech-TTS) Task: Single-speaker Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/freds0/BRSpeech.audiotext-to-speech10K<n<100K1 likes510 downloads27d agoHugging Face13FredZhang7 /malicious-website-features-2.4MImportant Notice: A subset of the URL dataset is from Kaggle, and the Kaggle datasets contained 10%-15% mislabelled data. See this dicussion I opened for some false positives. I have contacted Kaggle regarding their erroneous "Usability" score calculation for these unreliable datasets. The feature extraction methods shown here are not robust at all in 2023, and there're even silly mistakes in 3 functions: not_indexed_by_google, domain_registration_length, and age_of_domain. The features… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/malicious-website-features-2.4M.text-classification1M<n<10M6 likes508 downloads3y agoHugging Face14freds0 /cml_tts_dataset_spanishaudio100K<n<1M3 likes468 downloads2y agoHugging Face15F-Red /shared-ethical-memory-sem-2063Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference text0 likes468 downloads6mo agoHugging Face16frederickwang99 /SoAyBench SoAyBench by WangYC We've based SoAyBench creation on AMiner. To really understand how well LLMs can use SoAPI, we need to make AMiner's basic SoAPIs available for LLMs to use. We also need a test set made up of academic (question, solution, answer) triplets for checking how they're doing. The tricky part is, academic data keeps changing fast – stuff like info on scholars and their publications. So, keeping a test set with fixed answers is tough. To tackle this, what we've done is… See the full description on the dataset page: https://huggingface.co/datasets/frederickwang99/SoAyBench.0 likes409 downloads2y agoHugging Face17FredyRivera-dev /LLaDA-Sample-ES Dataset: LLaDA-Sample-ES Base: crscardellino/spanish_billion_words Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~ 652,089 Shards: 65… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-ES.text-generation100M<n<1B1 likes407 downloads3mo agoHugging Face18freds0 /cml_tts_dataset_portugueseaudio10K<n<100K3 likes381 downloads2y agoHugging Face19freds0 /cml_tts_dataset_frenchaudio100K<n<1M2 likes381 downloads2y agoHugging Face20freddyaboulton /gradio-reviewstabularn<1K2 likes374 downloads20h agoHugging Face21Fred808 /BG1image0 likes369 downloads1y agoHugging Face22FredyRivera-dev /Flux2-Image Flux2-Image This dataset is used for the Flux2-from-scratch project to train the Flux2 transformer from scratch. Dataset structure . ├── images_1/ │ └── a4ad6f1d46fb4ef3bcda56f05eacc537.png... ├── images_2/ │ └── 000121ce53464a63ad43ff134e979c7a.png... ├── images_3/ │ └── a47dc3302bb64b2ebfc7d525c8f0a88e.png... ├── images_4/ │ └── 0053ea2a8c764123b263d88c217f2995.png... ├── images_5/ │ └── 23fb74d40d8e4635914b4a399ee27b71.png... └── data.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/Flux2-Image.imagetext-to-image10K<n<100K2 likes364 downloads7mo agoHugging Face23fredzzp /esm-teddymer-pseudodimers ESM-Teddymer pseudo-dimers 60,177,402 intra-chain domain pairs ("pseudo-dimers") cut out of ESM Metagenomic Atlas monomers at Chainsaw/TED domain boundaries. Each row is a target domain and a binder domain that were adjacent in one real folded chain, so the pair comes with a real interface without anyone having to dock anything. Built to train target-conditioned binder-design models. All-atom structures for both chains ship alongside as foldcomp. What is in here… See the full description on the dataset page: https://huggingface.co/datasets/fredzzp/esm-teddymer-pseudodimers.tabularother100M<n<1B0 likes362 downloads23d agoHugging Face24fredericowieser /arc-agi-3-wm-traces ARC-AGI-3 World Model Traces This dataset contains ARC-AGI-3 transition traces in the same parquet schema used by HHazard/arc-agi-3. Each row is one environment transition: state, game_id, level_id, action_id, action_args, next_state, level_done, frame_idx, origin, transformation, player state and next_state are 64x64 ARC grids stored as nested integer arrays. action_id is the ARC-AGI-3 action kind; click actions use action_args.x and action_args.y. Splits… See the full description on the dataset page: https://huggingface.co/datasets/fredericowieser/arc-agi-3-wm-traces.tabularreinforcement-learning10M<n<100M0 likes351 downloads3mo agoHugging Face25fredxlpy /LuxAlign Dataset Card for LuxAlign Loading the Dataset The dataset is currently at version v3, which can be loaded as: from datasets import load_dataset ds = load_dataset("fredxlpy/LuxAlign", name="lb-en") # or "lb-fr" If you want to reproduce the results from the paper (v1) or use any previous version, you can specify the version folder: # Load version v1 (as used in the paper) ds_v1 = load_dataset("fredxlpy/LuxAlign", data_dir="data/v1", data_files={"train": "lb_en.json"})… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/LuxAlign.textsentence-similarity100K<n<1M1 likes333 downloads1y agoHugging Face26fredxlpy /ParaLux Dataset Card for ParaLux Benchmark Dataset Summary ParaLux is a Luxembourgish paraphrase detection benchmark that requires models to identify the correct paraphrase from two candidates for a given anchor sentence: one representing a valid paraphrase and the other an adversarial not_paraphrase. The dataset, consisting of 312 examples, is sourced from news articles published by RTL.lu and was introduced in LuxEmbedder: A Cross-Lingual Approach to Enhanced Luxembourgish… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/ParaLux.textsentence-similarityn<1K2 likes316 downloads2y agoHugging Face27Fred808 /helium0 likes299 downloads10mo agoHugging Face28freds0 /cml_tts_dataset_germanaudio100K<n<1M3 likes290 downloads2y agoHugging Face29freds0 /cml_tts_dataset_italianaudio10K<n<100K4 likes289 downloads2y agoHugging Face30fredzzp /mixproteintabular100M<n<1B0 likes288 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.