CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Fred808 /datatabular10K<n<100K0 likes9.1k downloads11mo agoHugging Face02freds0 /TAGARELA TAGARELA: A Portuguese Speech Dataset From Podcasts TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page: https://huggingface.co/datasets/freds0/TAGARELA.audioautomatic-speech-recognition1M<n<10M10 likes7.5k downloads2mo agoHugging Face03fredzzp /uniref-50-foldseek-v1text10M<n<100M1 likes736 downloads10mo agoHugging Face04fredzzp /esm-teddymer-pseudodimers ESM-Teddymer pseudo-dimers 60,177,402 intra-chain domain pairs ("pseudo-dimers") cut out of ESM Metagenomic Atlas monomers at Chainsaw/TED domain boundaries. Each row is a target domain and a binder domain that were adjacent in one real folded chain, so the pair comes with a real interface without anyone having to dock anything. Built to train target-conditioned binder-design models. All-atom structures for both chains ship alongside as foldcomp. What is in here… See the full description on the dataset page: https://huggingface.co/datasets/fredzzp/esm-teddymer-pseudodimers.tabularother100M<n<1B0 likes491 downloads27d agoHugging Face05freds0 /BRSpeech BRSpeech BRSpeech is a single-speaker Brazilian Portuguese speech dataset extracted and curated specifically for Text-to-Speech (TTS) and voice modeling tasks. It corresponds directly to speaker 2961 from the multi-speaker BRSpeech-TTS dataset, which represents the speaker with the highest volume of recorded audio/hours in the entire corpus. Dataset Summary Language: Portuguese (pt-BR) Speaker ID: 2961 (from BRSpeech-TTS) Task: Single-speaker Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/freds0/BRSpeech.audiotext-to-speech10K<n<100K1 likes483 downloads1mo agoHugging Face06freds0 /cml_tts_dataset_spanishaudio100K<n<1M3 likes468 downloads2y agoHugging Face07freds0 /cml_tts_dataset_portugueseaudio10K<n<100K3 likes384 downloads2y agoHugging Face08freds0 /cml_tts_dataset_frenchaudio100K<n<1M2 likes354 downloads2y agoHugging Face09freds0 /cml_tts_dataset_germanaudio100K<n<1M3 likes289 downloads2y agoHugging Face10fredzzp /mixproteintabular100M<n<1B0 likes283 downloads1y agoHugging Face11Fred808 /helium_memory Try gpt-oss · Guides · Model card · OpenAI blog Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters)gpt-oss-20b — for lower latency, and local or… See the full description on the dataset page: https://huggingface.co/datasets/Fred808/helium_memory.tabularn<1K0 likes256 downloads1y agoHugging Face12fredericowieser /arc-agi-3-wm-traces ARC-AGI-3 World Model Traces This dataset contains ARC-AGI-3 transition traces in the same parquet schema used by HHazard/arc-agi-3. Each row is one environment transition: state, game_id, level_id, action_id, action_args, next_state, level_done, frame_idx, origin, transformation, player state and next_state are 64x64 ARC grids stored as nested integer arrays. action_id is the ARC-AGI-3 action kind; click actions use action_args.x and action_args.y. Splits… See the full description on the dataset page: https://huggingface.co/datasets/fredericowieser/arc-agi-3-wm-traces.tabularreinforcement-learning10M<n<100M0 likes253 downloads3mo agoHugging Face13fredzzp /uniref50-sorted-structure-tokentabular10M<n<100M1 likes201 downloads10mo agoHugging Face14fredzzp /fine_code Fine Code A collection of high-quality code dataset. Consists of the following: https://huggingface.co/datasets/OpenCoder-LLM/opc-annealing-corpus https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA text10M<n<100M5 likes200 downloads1y agoHugging Face15freds0 /cml_tts_dataset_dutchaudio100K<n<1M1 likes193 downloads2y agoHugging Face16freds0 /BRSpeech-TTSaudio10K<n<100K0 likes154 downloads1y agoHugging Face17fredzzp /proteingym_substitutionstext1M<n<10M1 likes151 downloads1y agoHugging Face18fredhugg /Pink_pen-Yellow_2_20260824_224048This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/fredhugg/Pink_pen-Yellow_2_20260824_224048.tabularrobotics1K<n<10K0 likes115 downloads1mo agoHugging Face19takara-ai /FRED-CONVERTEDimage100K<n<1M1 likes111 downloads10mo agoHugging Face20fredhugg /Pink_pen-Yellow_20260824_211746This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/fredhugg/Pink_pen-Yellow_20260824_211746.tabularrobotics1K<n<10K0 likes107 downloads1mo agoHugging Face21fredhugg /Pink_pen_20260823_225553This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/fredhugg/Pink_pen_20260823_225553.tabularroboticsn<1K0 likes105 downloads1mo agoHugging Face22freds0 /cml_tts_dataset_italianaudio10K<n<100K4 likes102 downloads2y agoHugging Face23fredhugg /Pink_pen-v2_20260824_085530This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/fredhugg/Pink_pen-v2_20260824_085530.tabularrobotics1K<n<10K0 likes100 downloads1mo agoHugging Face24fredzzp /Uniref50 Uniref50: Uniref Sequences clustered at 50% sequence identity ~40M Protein Sequences. Split into train val and test. Usage from datasets import load_dataset # Step 1: Load the dataset from HuggingFace Hub dataset = load_dataset("zhangzhi/Uniref50") # Step 2: Access a specific split (e.g., "train", "validation", "test") train_split = dataset["train"] print(f"Number of sequences in the train split: {len(train_split)}") text10M<n<100M0 likes87 downloads2y agoHugging Face25fredzzp /spcas9tabularn<1K0 likes75 downloads1y agoHugging Face26fredguth /cgu__notas_fiscais Dataset Card: cgu_notas_fiscais Data from electronic invoices for federal government purchases made available by Comptroller General of the Union (Controladoria-Geral da União), which is a Brazilian federal government agency responsible for oversight and transparency. Dataset Details Dataset Description Curated by: Fred Guth (@fredguth) Funded by: World Bank Language(s) (NLP): pt-br License: CC-BY 4.0 Dataset Sources The source of this datasets… See the full description on the dataset page: https://huggingface.co/datasets/fredguth/cgu__notas_fiscais.tabulartabular-classification10M<n<100M0 likes74 downloads2y agoHugging Face27fredericlin /CaMiT CaMiT: Car Models in Time CaMiT (Car Models in Time) is a large-scale, fine-grained, time-aware dataset of car images collected from Flickr. It is designed to support research on temporal adaptation in visual models, continual learning, and time-aware generative modeling. Dataset Highlights Labeled Subset: 787,000 samples 190 car models 2007–2023 Unlabeled Pretraining Subset: 5.1 million samples 2005–2023 Metadata includes: Image URLs (not the images themselves)… See the full description on the dataset page: https://huggingface.co/datasets/fredericlin/CaMiT.text1M<n<10M3 likes70 downloads1y agoHugging Face28FreddyFazbear0209 /muong_voice_textaudio1K<n<10K0 likes67 downloads10mo agoHugging Face29fredzzp /OMG_prot50tabular100M<n<1B0 likes66 downloads1y agoHugging Face30Fredithefish /edushortstabular1K<n<10K0 likes62 downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.