CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pulmo /ncbi-genbank-complete Dataset Card for NCBI GenBank Complete Dataset Summary GenBank® is the NIH genetic sequence database, an annotated collection of all publicly available DNA sequences. GenBank is part of the International Nucleotide Sequence Database Collaboration (INSDC), which comprises the DNA DataBank of Japan (DDBJ), the European Nucleotide Archive (ENA), and GenBank at NCBI. These three organizations exchange data on a daily basis. This dataset has been processed into a… See the full description on the dataset page: https://huggingface.co/datasets/pulmo/ncbi-genbank-complete.n>1T4 likes276k downloads4mo agoHugging Face02RidheshBhati /Complete_Data_Source_100K_HOURS Multi-Language Audio Collection (100K Hours) This repository is physically reorganized for Absolute 100% Data Visibility. 🏗️ Global Consolidator Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here. audio1M<n<10M4 likes16k downloads5mo agoHugging Face03huggingworld /ncbi-refseq-complete Dataset Card for NCBI RefSeq Complete Dataset Summary The NCBI Reference Sequence (RefSeq) complete dataset provides a comprehensive, integrated, non-redundant, well-annotated set of sequences, including genomic DNA, transcripts, and proteins. It serves as a stable reference for genome annotation, gene identification and characterization, mutation and polymorphism analysis, expression studies, and comparative analyses. This dataset has been processed into a… See the full description on the dataset page: https://huggingface.co/datasets/huggingworld/ncbi-refseq-complete.n>1T2 likes15k downloads5mo agoHugging Face04secemp9 /arxiv-complete arXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions reported with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.tabulartext-generation100M<n<1B369 likes9.9k downloads2d agoHugging Face05GeorgeDaDude /Jailbreak_Complete_DS_labeledtext10K<n<100K1 likes8.1k downloads2y agoHugging Face06omnicad-lab-L3d /Omni-CAD-Subset-Completeimage0 likes8k downloads10mo agoHugging Face07Spirit-26 /BraTS-2024-Complete BraTS 2024 Complete Prepared Dataset Brain Tumor Segmentation (Leave a like 💖 if this helped you) Dataset Description This is an organized and verified version of the BraTS 2024 challenge datasets, including three tumor types. Included Datasets Dataset Type Cases Source BraTS-GLI Glioma 1,809 Synapse (Dec 2024) BraTS-MEN-RT Meningioma + RT 571 Synapse (Feb 2025) BraTS-PED Pediatric 348 Cancer Imaging Archive… See the full description on the dataset page: https://huggingface.co/datasets/Spirit-26/BraTS-2024-Complete.imageimage-segmentation10K<n<100K5 likes2.1k downloads6mo agoHugging Face08dementor-research /dementor-complete-experiment-results Dementor complete experiment results Audited outputs for the configuration-defined Dementor completion campaign. Audited scope Behavioral imitation adapters: 1,104 total (528 SFT, 528 DPO, 48 self-SFT controls). Behavioral-fidelity evaluation: 1,104 adapters on 200 held-out prompts, with embedding and primary LLM-judge scores, plus 48 target-reference response sets. Activation steering: 29 models, seven benchmarks, and two operators (original and fpall), totaling… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-complete-experiment-results.text-generation0 likes1.8k downloads1mo agoHugging Face09OpenVideo /pexel-0808-complete-final-testGithub Page: https://github.com/UmiMarch/OpenVideo license: cc-by-4.0 task_categories: - video-text-to-text size_categories: - 100K<n<1M text100K<n<1M6 likes1.7k downloads2y agoHugging Face10acroitoru /features_mavos_complete0 likes1.4k downloads11mo agoHugging Face11Crownelius /Complete-FABLE.5-traces-2M license: mit pretty_name: Claude Library — Fable 5 · Opus · Sonnet annotations_creators: machine-generated language: en language_creators: found machine-generated multilinguality: monolingual size_categories: 10K<n<100K task_categories: text-generation task_ids: language-modeling tags: agent-traces claude claude-fable-5 claude-opus claude-sonnet chain-of-thought tool-use coding-agents content-verified maintained-mirror deduplicated parquet configs: config_name:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Complete-FABLE.5-traces-2M.tabular100K<n<1M151 likes1.4k downloads2mo agoHugging Face12PNW-GM /tvelve_map_complete_datasetimage10K<n<100K0 likes1.1k downloads4mo agoHugging Face13bubblebird /zcache-results-completetext0 likes990 downloads1d agoHugging Face14AITRADER /dutch-tts-labeled-complete Dutch TTS Dataset - Complete Labeled A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of audio data. Quick Preview The default config shows a 100-row sample for the dataset viewer. To access the full dataset, use the full config. Dataset Description This dataset contains Dutch speech recordings with rich metadata including: Emotion labels (neutral, happy, sad, angry) Speaker IDs (239,388 unique speakers)… See the full description on the dataset page: https://huggingface.co/datasets/AITRADER/dutch-tts-labeled-complete.audiotext-to-speech100K<n<1M0 likes712 downloads8mo agoHugging Face15gokuls /wiki_book_corpus_complete_processed_bert_dataset Dataset Card for "wiki_book_corpus_complete_processed_bert_dataset" More Information needed 1M<n<10M0 likes415 downloads4y agoHugging Face16PerRing /coco_captioning_complete_formatimage100K<n<1M0 likes414 downloads11mo agoHugging Face17Heera-fdp2025 /COMPLETE_COAL_DATASETgeospatialn<1K0 likes387 downloads4mo agoHugging Face18whiskwhite /leetcode-complete Complete LeetCode Problems Dataset This dataset contains a comprehensive collection of LeetCode problems (including premium) with AI-generated solutions in JSONL format. It is regularly updated to include new problems as they are added to LeetCode. Splits The dataset is divided into the following splits: train: Contains approximately 80% of the problems for training validation: Contains approximately 10% of the problems for validation test: Contains approximately… See the full description on the dataset page: https://huggingface.co/datasets/whiskwhite/leetcode-complete.tabulartext-generation1K<n<10K1 likes384 downloads9d agoHugging Face19ubaada /booksum-complete-cleaned Description: This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization . This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.textsummarization1K<n<10K23 likes380 downloads2y agoHugging Face20gsri-18 /ISEAR-dataset-completetext1K<n<10K0 likes369 downloads2y agoHugging Face21gokuls /wiki_book_corpus_complete_raw_dataset Dataset Card for "wiki_book_corpus_complete_raw_dataset" More Information needed text10M<n<100M0 likes325 downloads4y agoHugging Face22WithinUsAI /GOD_Coder_Complete_DataSet GOD_Coder_Complete_DataSet Subtitle A large-scale complete-project coding dataset by gss1147 / WithIn Us AI, built to train language models into stronger professional software-engineering assistants. Dataset Summary GOD_Coder_Complete_DataSet is a large synthetic supervised fine-tuning dataset designed to help turn a general language model into a professional complete-project AI coder. The dataset focuses on teaching models how to: diagnose… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GOD_Coder_Complete_DataSet.text-generation100K<n<1M4 likes315 downloads6mo agoHugging Face23Abeyankar /Visdrone_fisheye-v51-complete Dataset Card for visdrone_fisheye-v51-complete This is a FiftyOne dataset with 20942 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Abeyankar/Visdrone_fisheye-v51-complete") # Launch the App session =… See the full description on the dataset page: https://huggingface.co/datasets/Abeyankar/Visdrone_fisheye-v51-complete.imageobject-detection10K<n<100K3 likes311 downloads2y agoHugging Face24Aff77 /BraTS-2024-Complete BraTS 2024 Complete Prepared Dataset Brain Tumor Segmentation Dataset Description This is an organized and verified version of the BraTS 2024 challenge datasets, including three tumor types. Included Datasets Dataset Type Cases Source BraTS-GLI Glioma 1,809 Synapse (Dec 2024) BraTS-MEN-RT Meningioma + RT 571 Synapse (Feb 2025) BraTS-PED Pediatric 348 Cancer Imaging Archive Total: 2,728 multi-parametric MRI cases… See the full description on the dataset page: https://huggingface.co/datasets/Aff77/BraTS-2024-Complete.imageimage-segmentation10K<n<100K0 likes297 downloads6mo agoHugging Face25Nyanmero /vie-speech-corpus-completeaudio100K<n<1M0 likes289 downloads2y agoHugging Face26DanhVuiVe /Benetech_PlotQa_DVQA_combined_matcha_completeimage100K<n<1M0 likes275 downloads2y agoHugging Face27penfever /meta-llama_Llama-3.1-8B-Instruct-jdgfct-Completenesstext100K<n<1M0 likes274 downloads5mo agoHugging Face28Abrak /wikipedia-paragraph-embeddings-en-gist-complete Dataset Summary Paragraph embeddings for every article in English Wikipedia (not the Simple English version). Based on wikimedia/wikipedia, 20231101.en. Embeddings were generated with avsolatorio/GIST-small-Embedding-v0 and are quantized to int8. You can load the data with the following: from datasets import load_dataset ds = load_dataset(path="Abrak/wikipedia-paragraph-embeddings-en-gist-complete", data-dir="20231101.en") Dataset Structure The structure of the… See the full description on the dataset page: https://huggingface.co/datasets/Abrak/wikipedia-paragraph-embeddings-en-gist-complete.text10M<n<100M1 likes260 downloads2y agoHugging Face29morka17 /sat-mathematics-complete-questions0 likes250 downloads1y agoHugging Face30SilencioNetwork /complete-voiceai-speech-dataset Silencio Voice AI Sample Dataset Speaker-attributed spontaneous speech. 363 labelled contributors across 111 self-reported origin varieties, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment. Hours 16.05 Clips 1,305 Speakers 363 Origin varieties 111 Languages 21 Configs 44 Speaker metadata origin region / variety, mother tongue, gender, device, OS… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset.audioautomatic-speech-recognition1K<n<10K1 likes250 downloads4h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.