CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01philippesaade /wikidata Wikidata Entities Connected to Wikipedia This dataset is a multilingual, JSON-formatted version of the Wikidata dump from May 7, 2026. It contains 73,769,737 entities after filtering out scholarly articles from the original 120,182,414 entity dump. Curated by: Jonathan Fraine & Philippe Saadé, Wikimedia Deutschland Funded by: Wikimedia Deutschland Language(s) (NLP): All Wikidata Languages License: CC0-1.0 Dataset Structure Each row in this dataset represents a… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/wikidata.text10M<n<100M21 likes13k downloads2mo agoHugging Face02phields /a-share-l2-trades China A-share Level 2 Trades Canonical Level 2 trade records for China A-shares, stored as one fact table. Coverage Date range: 2026-04-01 to 2026-09-24 Trading days: 119 Rows: 18730990496 Parquet files: 842 Compressed local size: 149.49 GiB Layout data/l2_trades/ trade_date=YYYY-MM-DD/ code_prefix=00/ part-00000.parquet code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68. Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.tabular10B<n<100B1 likes7k downloads2d agoHugging Face03philippesaade /Wikidata_Vectors_0.2 Wikidata Entity Embeddings 0.2 Dataset Summary Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.textfeature-extraction10M<n<100M3 likes5.5k downloads1mo agoHugging Face04kierth /retail-products-philippinesimage1K<n<10K1 likes5.2k downloads5mo agoHugging Face05phiyodr /InpaintCOCO InpaintCOCO - Fine-grained multimodal concept understanding (for color, size, and COCO objects) Dataset Summary A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object. Many multimodal tasks, such as Vision-Language Retrieval and Visual Question Answering, present results in terms of overall performance. Unfortunately, this approach overlooks more nuanced concepts, leaving us unaware… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/InpaintCOCO.imageimage-to-text1K<n<10K5 likes5k downloads2y agoHugging Face06PhisherJR /Truecallertext100M<n<1B1 likes4.8k downloads1mo agoHugging Face07philschmid /dolly-15k-oai-style Dataset Card for "dolly-15k-oai-style" More Information needed text10K<n<100K7 likes2.9k downloads3y agoHugging Face08PhilipMay /stsb_multi_mt Dataset Card for STSb Multi MT Dataset Summary STS Benchmark comprises a selection of the English datasets used in the STS tasks organized in the context of SemEval between 2012 and 2017. The selection of datasets include text from image captions, news headlines and user forums. (source) These are different multilingual translations and the English original of the STSbenchmark dataset. Translation has been done with deepl.com. It can be used to train sentence embeddings… See the full description on the dataset page: https://huggingface.co/datasets/PhilipMay/stsb_multi_mt.texttext-classification10K<n<100K68 likes2.3k downloads2y agoHugging Face09philschmid /guanaco-sharegpt-style Dataset Card for "guanaco-sharegpt-style" More Information needed text1K<n<10K49 likes2.3k downloads3y agoHugging Face10PhisherJR /ULP-logstext10M<n<100M0 likes2k downloads10d agoHugging Face11phiyodr /coco2017 coco2017 Image-text pairs from MS COCO2017. Data origin Data originates from cocodataset.org While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy. phiyodr/coco2017: One row corresponds one image with several sentences. phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.imageimage-to-text100K<n<1M29 likes2k downloads3y agoHugging Face12phields /a-share-l2-market-depth China A-share Level 2 Market Depth Canonical order-event and ten-level snapshot data for China A-shares. Canonical trade records remain in the separate phields/a-share-l2-trades dataset. Coverage Date range: 2026-07-24 to 2026-07-24 Trading days: 1 Table Rows Parquet files Compressed size l2_orders 249,705,486 10 2.14 GiB l2_snapshots 20,279,887 4 0.91 GiB Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.tabular10B<n<100B0 likes1.9k downloads2mo agoHugging Face13open-phi /textbooks Textbooks Are All You Need Leveraging Large Language Models (LLMs), there's an opportunity to create a comprehensive open-source repository reminiscent of the historic Library of Alexandria. This initiative represents a preliminary attempt at producing high-quality books covering an extensive range of subjects. The source of these samples varies: Some generated using the RAG model, referencing Wikipedia or other search data. Some are completely synthetically generated. Some created… See the full description on the dataset page: https://huggingface.co/datasets/open-phi/textbooks.text1K<n<10K97 likes1.8k downloads3y agoHugging Face14bettergovph /raw-philippine-data Raw Philippine Data This repository contains raw data about Philippine politicians, public officials, and legislative documents collected from various sources. The data is intended for research, analysis, and civic technology purposes. Dataset Overview This dataset currently contains: Persons 45,424 person records of Philippine politicians and public officials with: ID: Unique identifier (ULID format) First Name: Person's first name Last Name: Person's last… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/raw-philippine-data.text100K<n<1M0 likes1.6k downloads11mo agoHugging Face15PhisherJR /phonebooktext1B<n<10B0 likes1.6k downloads10d agoHugging Face16TAUR-Lab /Taur_CoT_Analysis_Project___microsoft__Phi-3-small-8k-instructtext10K<n<100K0 likes915 downloads2y agoHugging Face17PhisherJR /850M-India-datatext100M<n<1B0 likes905 downloads2mo agoHugging Face18saidutta69 /PhishTrap PhishTrap Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours. Priorities: Quality > Ease of Access > Quantity Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published. Dataset Overview PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.tabular10K<n<100K0 likes823 downloads10h agoHugging Face19AiresPucrs /stanford-encyclopedia-philosophy Stanford Encyclopedia Philosophy (Teeny-Tiny Castle) This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research. How to Use from datasets import load_dataset dataset = load_dataset("AiresPucrs/stanford-encyclopedia-philosophy", split = 'train') texttext-classification100K<n<1M53 likes751 downloads2y agoHugging Face20philgzl /ears EARS: Expressive Anechoic Recordings of Speech This is a mirror of the Expressive Anechoic Recordings of Speech (EARS) dataset. The original files were converted from WAV to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: Train: 92 hours, 15939 utterances, speakers p001 to p099 Validation: 2 hours, 322 utterances, speakers p100 and p101 Test: 6 hours, 966 utterances, speakers p102 to p107 License: CC BY-NC 4.0 Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/ears.audio10K<n<100K0 likes663 downloads1y agoHugging Face21philgzl /wham WHAM!48kHz noise dataset This is a mirror of the WHAM!48kHz noise dataset. The original files were segmented and converted from WAV to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 2 Format: Opus Splits: Train: 59 hours, 21216 segments, files 000 to 188 Validation: 12 hours, 4444 segments, files 189 to 225 Test: 7 hours, 2613 segments, files 226 to 249 License: CC BY-NC 4.0 Source: http://wham.whisper.ai/ Paper: WHAM!: Extending Speech… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/wham.audio10K<n<100K0 likes604 downloads1y agoHugging Face22philgzl /fsd50k FSD50K: An open dataset of human-labeled sound events This is a mirror of the FSD50K sound event dataset. The original files were converted from WAV to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: Dev: 80 hours, 40966 clips. Eval: 28 hours, 10231 clips. License: FSD50K is released under CC-BY. However, each clip has its own licence. Clip licenses include CC0, CC-BY, CC-BY-NC and CC Sampling+. Clip licenses are specified… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/fsd50k.audio10K<n<100K0 likes603 downloads1y agoHugging Face23puyang2025 /seven-phishing-email-datasets Dataset Card for Seven Phishing/Spam Email Datasets Dataset Summary This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks. Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label). Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.tabulartext-classification100K<n<1M1 likes570 downloads8mo agoHugging Face24if001 /hle_math_category_phi4textn<1K0 likes544 downloads1y agoHugging Face25philgzl /dns5 DNS5 Challenge data This is a mirror of the DNS5 Challenge data. The original files were converted from WAV to Opus to reduce the size and accelerate streaming. ⚠️ Only the LibriVox, AudioSet, Freesound, OpenSLR26, and OpenSLR28 data is included. The VCTK, VocalSet, CREMA-D, VoxCeleb2, and DEMAND data is excluded. ⚠️ Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: speech_english: 245 hours, 186743 files speech_french: 95 hours, 60454 files speech_german: 137 hours, 119175… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/dns5.audio100K<n<1M1 likes515 downloads5mo agoHugging Face26ENSEONG /full-math-private-n256-Phi-4-mini-instruct-bontabular100K<n<1M0 likes513 downloads1mo agoHugging Face27kmack /Phishing_urls Dataset Card for "Phishing_urls" More Information needed text100K<n<1M5 likes460 downloads2y agoHugging Face28phicoltan /robommeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "panda", "total_episodes": 1600, "total_frames": 768897, "total_tasks": 116, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:1600" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/phicoltan/robomme.tabularrobotics100K<n<1M0 likes455 downloads1mo agoHugging Face29philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes447 downloads2y agoHugging Face30pirocheto /phishing-url Dataset Description The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs. Features are from three different classes: 56 extracted from the structure and syntax of URLs 24 extracted from the content of their correspondent pages 7 are extracetd by querying external services. The… See the full description on the dataset page: https://huggingface.co/datasets/pirocheto/phishing-url.tabulartext-classification10K<n<100K14 likes395 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.