CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01monology /pile-uncopyrighted Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The Pile with all copyrighted content removed.Please consider using this dataset to train your future LLMs, to respect authors and abide by copyright law.Creating an uncopyrighted version of a larger dataset (ie RedPajama) is planned, with no ETA. MethodologyCleaning was performed by removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2… See the full description on the dataset page: https://huggingface.co/datasets/monology/pile-uncopyrighted.text100M<n<1B175 likes80k downloads3y agoHugging Face02IgnisCogitationis /quantum-like-attention-framework-1.3b-untuned-validation Quantum Like Attention Framework (Q.L.A.F) 1.3b untuned This repository contains the model checkpoints, downstream evaluation scores, and pretraining convergence logs for the Quantum Like Attention Framework (Q.L.A.F) 1.3B configuration. Key Specifications & Architecture Model Name: Q.L.A.F 1.3b untuned (Quantum Like Attention Framework - Hybrid Architecture) Parameters: 1.3B parameters total configuration (327M active parameter student subset) Layer Count: 12… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.textn<1K3 likes33k downloads6m agoHugging Face03h8st6ptv /turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format: All Universities in Turkey Dataset Description This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities. Fields 1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.imagen<1K2 likes22k downloads2y agoHugging Face04UnipatAI /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.imagetext-generationn<1K2 likes14k downloads5mo agoHugging Face05Aeala /ShareGPT_Vicuna_unfiltered Dataset Card This is a reupload of this dataset that was further cleaned by gozfarb. text100K<n<1M52 likes9.7k downloads3y agoHugging Face06mteb /cqadupstack-unix CQADupstackUnixRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web, Programming Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackUnixRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-unix.texttext-retrieval10K<n<100K0 likes6.3k downloads1y agoHugging Face07ulamai /UnsolvedMath🌐 Browse UnsolvedMath online ✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark UnsolvedMath Dataset A comprehensive curated collection of 15,458 open and partially solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com. Paper: "Open Mathematical Problems as an AI Reasoning Benchmark" Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.documentquestion-answering10K<n<100K80 likes5.8k downloads10d agoHugging Face08LSX-UniWue /LLaMmlein-DatasetThis dataset is a strict subset of the RedPajama V2 dataset and therefore retains all licenses from RedPajama V2. More details in our preprint! Data Take Down texttext-generation100M<n<1B5 likes4.6k downloads11mo agoHugging Face09UnFaZeD07 /Music-AVQAtabular10K<n<100K0 likes4.2k downloads7mo agoHugging Face10unsloth /alpaca-cleaned Dataset Card for Alpaca-Cleaned Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.texttext-generation10K<n<100K25 likes3.4k downloads9mo agoHugging Face11UCB-team /unclickbait-synthetic-27b-trajectories Unclickbait Synthetic 27B Trajectories Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline. Contents : Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates). : 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring). texttext-generationn<1K0 likes2.5k downloads12d agoHugging Face12Goekdeniz-Guelmez /Function_Calling_Unfilteredtext100K<n<1M4 likes2.2k downloads3y agoHugging Face13unireo /sn38-submission-r10textn<1K0 likes1.9k downloads19d agoHugging Face14unireo /sn38-submission-r12textn<1K0 likes1.8k downloads5d agoHugging Face15UniverseTBD /arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.texttext-generation1M<n<10M7 likes1.7k downloads3y agoHugging Face16unireo /sn38-submission-r11textn<1K0 likes1.3k downloads12d agoHugging Face17unireo /sn38-submissiontextn<1K0 likes1.2k downloads26d agoHugging Face18globis-university /aozorabunko-clean Overview This dataset provides a convenient and user-friendly format of data from Aozora Bunko (青空文庫), a website that compiles public-domain books in Japan, ideal for Machine Learning applications. [For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f Methodology The code to reproduce this dataset is made available on GitHub: globis-org/aozorabunko-exctractor. 1. Data collection We firstly downloaded the CSV file that… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-clean.texttext-generation10K<n<100K48 likes1.2k downloads3y agoHugging Face19marin-dna /gpn-star-p-uniform-v1-cds marin-dna/gpn-star-p-uniform-v1-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-cds.tabular10M<n<100M0 likes1k downloads1mo agoHugging Face20marin-dna /phylop-uniform-v1-cds marin-dna/phylop-uniform-v1-cds Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the cds region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-cds.tabular10M<n<100M0 likes1k downloads29d agoHugging Face21marin-dna /gpn-star-p-uniform-v1-enhancer-arm-a marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.tabular100M<n<1B0 likes966 downloads1mo agoHugging Face22DataProvenanceInitiative /Commercial_or_unspecified_licenses_and_terms Dataset Card for Data Provenance Initiative - Commercial-Or-Unspecified-Licenses-and-Terms Legal Disclaimer / Notice Collected License Information is NOT Legal Advice. It is important to note we collect self-reported licenses, from the papers and repositories that released these datasets, and categorize them according to our best efforts, as a volunteer research and transparency initiative. The information provided by any of our works and any outputs of the Data… See the full description on the dataset page: https://huggingface.co/datasets/DataProvenanceInitiative/Commercial_or_unspecified_licenses_and_terms.text10M<n<100M0 likes952 downloads2y agoHugging Face23semran1 /synth-cc-unfilteredtext100M<n<1B1 likes940 downloads1y agoHugging Face24unireo /sn38-submission-r13textn<1K0 likes921 downloads5d agoHugging Face25ChrisDing1105 /unified-agent-trajectories Unified Benchmark Agent Trajectories Dataset release: v2.1.1 (2026-09-18)Record format: unified-agent-sft-v1 A growing collection of benchmark agent execution trajectories converted into one transparent, multimodal, tool-aware representation. These are complete recorded benchmark runs—not ordinary chat transcripts—including benchmark tasks, model reasoning and answers, tool calls, tool observations, runtime status, and benchmark scores when available. The directory layout is… See the full description on the dataset page: https://huggingface.co/datasets/ChrisDing1105/unified-agent-trajectories.imagetext-generation1K<n<10K3 likes859 downloads7d agoHugging Face26xzuyn /open-instruct-uncensored-alpacaOriginal dataset page from ehartford. 810,102 entries. Sourced from open-instruct-uncensored.jsonl. Converted the jsonl to a json which can be loaded into something like LLaMa-LoRA-Tuner. I've also included smaller datasets that includes less entries depending on how much memory you have to work with. Each one is randomized before being converted, so each dataset is unique in order. Count of each Dataset: code_alpaca: 19991 unnatural_instructions: 68231 baize: 166096 self_instruct: 81512… See the full description on the dataset page: https://huggingface.co/datasets/xzuyn/open-instruct-uncensored-alpaca.text1M<n<10M7 likes813 downloads3y agoHugging Face27Skywork /unipic_nano_2images Skywork/unipic_nano_2images: A Multi-Image Composition Dataset ⚡ Quick Start The image archive is split into multiple parts for easier downloading. To reconstruct and extract: # Step 1: Concatenate split files into a single zip cat nano-banana-2image_part_* > nano-banana-2images.zip # Step 2: Extract the images unzip nano-banana-2images.zip 📖 Overview UniPic-Nano-2Images is a high-quality multi-image composition dataset containing 41,812 samples designed for… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/unipic_nano_2images.textimage-to-image10K<n<100K6 likes798 downloads8mo agoHugging Face28marin-dna /gpn-star-p-uniform-v1-background marin-dna/gpn-star-p-uniform-v1-background Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.tabular10M<n<100M0 likes786 downloads1mo agoHugging Face29rtsikynyantsa /unet-jerymotro UNet Jery Motro — Burned Area Dataset Dataset de patches satellitaires multi-canaux pour la détection et la segmentation sémantique des zones brûlées à Madagascar dans le cadre du projet Jery Motro. Structure des Données (.npz) Chaque fichier .npz contient les tenseurs et métadonnées suivants : image : Tenseur de forme (256, 256, 6) de type float32 (canaux spectraux / features). mask : Masque binaire de segmentation de forme (256, 256) de type uint8 (0 = sain, 1… See the full description on the dataset page: https://huggingface.co/datasets/rtsikynyantsa/unet-jerymotro.tabularimage-segmentationn<1K0 likes782 downloads1mo agoHugging Face30marin-dna /phylop-uniform-v1-enhancer-arm-a marin-dna/phylop-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.tabular10M<n<100M0 likes761 downloads29d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.