CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01meshllm /catalog Mesh-LLM Catalog This dataset is the Hugging Face-backed catalog for Mesh-LLM. The runtime catalog entries live under entries/**/*.json. The Dataset Viewer uses catalog_rows.jsonl, a flat generated table with one row per model variant. The catalog deliberately excludes raw blob URLs. Entries should resolve to Hugging Face repositories and canonical Mesh refs. tabularn<1K0 likes50k downloads4d agoHugging Face02cschell /xr-motion-dataset-catalogue XR Motion Dataset Catalogue Overview The XR Motion Dataset Catalogue, accompanying our paper "Navigating the Kinematic Maze: A Comprehensive Guide to XR Motion Dataset Standards," standardizes and simplifies access to Extended Reality (XR) motion datasets. The catalogue represents our initiative to streamline the usage of kinematic data in XR research by aligning various datasets to a consistent format and structure. Dataset Specifications All datasets in this… See the full description on the dataset page: https://huggingface.co/datasets/cschell/xr-motion-dataset-catalogue.8 likes30k downloads2y agoHugging Face03v-bible /catholic-resources Vietnamese Catholic resources by v-bible Data Structure calendar: Generated Liturgical calendars using v-bible/js-sdk. misc/proper-names.json: Name translation from ktcgkpv.org, generated by v-bible/bible-scraper. liturgical: Liturgical data from The Lectionary for Mass (1998/2002 USA Edition), compiled by Felix Just, S.J., Ph.D., and generated by v-bible/bible-scraper. books/bible: Generated Bible markdown data. books/catechism-books: Official catechism… See the full description on the dataset page: https://huggingface.co/datasets/v-bible/catholic-resources.image10K<n<100K1 likes27k downloads18d agoHugging Face04catherinearnett /montok MonTok: A Suite of Monolingual Tokenizers This is a set of monolingual tokenizers for 98 languages. For each language, there are Unigram, BPE, and SuperBPE tokenizers, ranging in vocabulary size from around 6k to over 200k. Training Details Training Data All tokenizers are trained on samples of the data used to the train the Goldfish language models. The tokenizers were either trained on scaled or unscaled data. This refers to whether the models are trained on… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/montok.4 likes25k downloads1y agoHugging Face05huggingface /cats-imageimagen<1K5 likes12k downloads1y agoHugging Face06Central-Cat /vbvr-latent-cache-832x832x33f-t2v-only VBVR Latent Cache (832×832 × 33f, Wan2.2-TI2V-5B VAE + UMT5-XXL) Pre-encoded latent cache for the Video-Reason/VBVR-Dataset geometric / logical reasoning video corpus, prepared for Equilibrium Matching (EqM) post-training of Wan-AI/Wan2.2-TI2V-5B-Diffusers on AWS Trainium2. This is a working cache, not a primary dataset. It exists to skip the ~5 s/sample VAE+T5 encode cost during training. The original videos + prompts live in the upstream VBVR-Dataset repo. Source →… See the full description on the dataset page: https://huggingface.co/datasets/Central-Cat/vbvr-latent-cache-832x832x33f-t2v-only.text-to-video0 likes7.3k downloads5mo agoHugging Face07datania /ine-catalog INE Este repositorio contiene todas las tablas¹ del Instituto Nacional de Estadística exportadas a ficheros Parquet. Puedes encontrar cualquiera de las tablas o sus metadatos en la carpeta tablas. Cada tabla está identificado un una ID. Puedes encontrar la ID de la tabla tanto en el INE (es el número que aparece en la URL) or en el archivo tablas.jsonl de este repositorio que puedes explorar en el Data Viewer. Por ejemplo, la tabla de Índices nacionales de clases se corresponde al… See the full description on the dataset page: https://huggingface.co/datasets/datania/ine-catalog.tabular1K<n<10K4 likes5.4k downloads5mo agoHugging Face08SuhxsReddy /cati-singapore-dataset CATI Singapore Expressway Traffic Dataset Real-time vehicle detection data collected from Singapore's 90 LTA traffic cameras using CATI (Context-Aware Traffic Intelligence) — a novel FiLM-conditioned YOLOv11 detector that adapts to environmental conditions in real time. Dataset Description This dataset contains per-camera vehicle detection results collected continuously from Singapore's Land Transport Authority (LTA) expressway camera network. Each record captures… See the full description on the dataset page: https://huggingface.co/datasets/SuhxsReddy/cati-singapore-dataset.imageobject-detectionn<1K3 likes5k downloads3d agoHugging Face09CATMuS /medieval Dataset Card for CATMuS Medieval Join our Discord to ask questions about the dataset: Dataset Details Handwritten Text Recognition (HTR) has emerged as a crucial tool for converting manuscripts images into machine-readable formats, enabling researchers and scholars to analyse vast collections efficiently. Despite significant technological progress, establishing consistent ground truth across projects for HTR tasks, particularly for complex and heterogeneous… See the full description on the dataset page: https://huggingface.co/datasets/CATMuS/medieval.imageimage-to-text100K<n<1M28 likes5k downloads2y agoHugging Face10CatManga /Cat-Mangaaudion<1K1 likes5k downloads1mo agoHugging Face11altaidevorg /fineweb-2-turkish-categorized What is this THis is the categorized version of the Turkish subset of the fineweb-2 dataset. It is an ongoing effort, and the details will be added soon with the rest of the dataset. tabular10M<n<100M15 likes4.7k downloads2y agoHugging Face12projecte-aina /CATalog Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and Leaderboards Fill-Mask Text Generation other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.textfill-mask10M<n<100M8 likes4.1k downloads1y agoHugging Face13hf-internal-testing /cats_vs_dogs_sampleimagen<1K1 likes3k downloads1y agoHugging Face14microsoft /cats_vs_dogs Dataset Card for Cats Vs. Dogs Dataset Summary A large set of images of cats and dogs. There are 1738 corrupted images that are dropped. This dataset is part of a now-closed Kaggle competition and represents a subset of the so-called Asirra dataset. From the competition page: The Asirra data set Web services are often protected with a challenge that's supposed to be easy for people to solve, but difficult for computers. Such a challenge is often called a CAPTCHA… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/cats_vs_dogs.imageimage-classification10K<n<100K73 likes2.9k downloads2y agoHugging Face15CathleenTico /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/CathleenTico/stack-v3-train.tabulartext-generation100M<n<1B0 likes2.4k downloads2mo agoHugging Face16softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face17mnemoraorg /usgs-global-earthquake-catalog USGS Global Earthquake Catalog Provides historical data on global seismic events, sourced directly from the U.S. Geological Survey (USGS) Earthquake Hazards Program via its FDSN Event Web Service. Each record represents a single seismic event (primarily earthquakes) and contains detailed information, including: Event Time & Location: Precise timestamp, geographic coordinates (latitude, longitude), and depth of the event. Magnitude: The magnitude of the event (mag) and the method… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/usgs-global-earthquake-catalog.tabulartext-classification1M<n<10M1 likes2.3k downloads11mo agoHugging Face18shahmirshabir /Products-Catalogimage10K<n<100K0 likes1.9k downloads2mo agoHugging Face19Gramscii-IT /european-open-data-catalogue European Open Data Catalogue This repository publishes independently versioned metadata and licensed source snapshots: A discovery catalogue with 15565 dataset entries from ISTAT, Eurostat, OECD, ILO, DoveVannoINostriSoldi (DVNS) and Cruscotto Italia. 3 independently pinned availability indexes with 911,795 joint combinations across 35 datasets, built from complete source responses within the explicitly declared scope. Licensed Cruscotto source snapshots, stored separately from… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.text100K<n<1M0 likes1.8k downloads2d agoHugging Face20catch-a-vlm /catch-a-vlm-embeddings0 likes1.8k downloads4mo agoHugging Face21brighter-dataset /BRIGHTER-emotion-categories BRIGHTER Emotion Categories Dataset This dataset contains the emotion categories data from the BRIGHTER paper: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages. Dataset Description The BRIGHTER Emotion Categories dataset is a comprehensive multi-language, multi-label emotion classification dataset with separate configurations for each language. It represents one of the largest human-annotated emotion datasets across multiple… See the full description on the dataset page: https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-categories.tabular100K<n<1M18 likes1.8k downloads10mo agoHugging Face22Voxel51 /cats-vs-dogs-sample Dataset Card for Dataset Name Subset of https://huggingface.co/datasets/microsoft/cats_vs_dogs, converted into FiftyOne dataset format. This is a FiftyOne dataset with 5000 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/cats-vs-dogs-sample.imageimage-classification1K<n<10K1 likes1.7k downloads2y agoHugging Face23CATMuS /medieval-segmentation Dataset Card for CATMuS Medieval (Segmentation Version) Join our Discord to ask questions about the dataset: Dataset Details CATMuS Medieval Segmentation (Consistent Approaches to Transcribing Manuscripts) is a specialized dataset designed for layout analysis of medieval manuscripts using the SegmOnto vocabulary for region and line classification. This dataset addresses the challenges associated with establishing consistent ground truth in layout analysis tasks… See the full description on the dataset page: https://huggingface.co/datasets/CATMuS/medieval-segmentation.imageimage-segmentation1K<n<10K7 likes1.7k downloads2y agoHugging Face24heegyu /news-category-datasetDataset from https://www.kaggle.com/datasets/rmisra/news-category-dataset text100K<n<1M4 likes1.6k downloads4y agoHugging Face25CATIE-AQ /ssl-checkpoints ssl-checkpoints The code to load the checkpoints to follow... The repository is organised as follows: Each folder corresponds to the data set used for our experiment. Each subfolder represents the corresponding SSL technique used. These subfolders contain the checkpoints for each transformation/pretext task considered. The five checkpoint files correspond to the transformation Baseline, SimClr, Orthogonality, LoRot and DCL, respectively, described in the blog. 0 likes1.5k downloads3y agoHugging Face26aryashah00 /CatVision CatVision: Human–Cat Vision Frame Pairs Official dataset for the paper: Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTsArya Shah et al. · arXiv:2511.02404 This dataset contains 346,400 paired video frames rendered under human vision and simulated cat vision optics. It was used to benchmark cross-species representational alignment across CNNs, supervised ViTs, windowed transformers, and self-supervised ViTs (DINO) using CKA and… See the full description on the dataset page: https://huggingface.co/datasets/aryashah00/CatVision.imageimage-to-image100K<n<1M0 likes1.4k downloads7mo agoHugging Face27GEM /wiki_cat_sumSummarise the most important facts of a given entity in the Film, Company, and Animal domains from a cluster of related documents.textsummarization100K<n<1M4 likes1.4k downloads4y agoHugging Face28Voxel51 /cats-vs-dogs-imbalanced Dataset Card for cats-vs-dogs-imbalanced This is a FiftyOne dataset with 2551 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = fouh.load_from_hub("Voxel51/cats-vs-dogs-imbalanced") # Launch the App session = fo.launch_app(dataset) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/cats-vs-dogs-imbalanced.imageimage-classification1K<n<10K2 likes1.4k downloads2y agoHugging Face29catherinearnett /monolingual-tokenizer-dataTodo: add language to metadata cite source and explain sampling text100M<n<1B1 likes1.4k downloads1y agoHugging Face30catherinearnett /apertus_multiblimp0 likes1.3k downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.