CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ESA-philab /OceanDepths OceanDepths GeoTIFF Raster and Aligned ARGO Dataset This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.image1M<n<10M0 likes47k downloads1mo agoHugging Face02philschmid /mt-benchtextn<1K4 likes18k downloads3y agoHugging Face03PhillyMac /Corpus_Gap_Logtext1K<n<10K0 likes15k downloads7d agoHugging Face04philippesaade /wikidata Wikidata Entities Connected to Wikipedia This dataset is a multilingual, JSON-formatted version of the Wikidata dump from May 7, 2026. It contains 73,769,737 entities after filtering out scholarly articles from the original 120,182,414 entity dump. Curated by: Jonathan Fraine & Philippe Saadé, Wikimedia Deutschland Funded by: Wikimedia Deutschland Language(s) (NLP): All Wikidata Languages License: CC0-1.0 Dataset Structure Each row in this dataset represents a… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/wikidata.text10M<n<100M21 likes13k downloads2mo agoHugging Face05phields /a-share-l2-trades China A-share Level 2 Trades Canonical Level 2 trade records for China A-shares, stored as one fact table. Coverage Date range: 2026-04-01 to 2026-09-24 Trading days: 119 Rows: 18730990496 Parquet files: 842 Compressed local size: 149.49 GiB Layout data/l2_trades/ trade_date=YYYY-MM-DD/ code_prefix=00/ part-00000.parquet code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68. Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.tabular10B<n<100B1 likes6.7k downloads14h agoHugging Face06philippesaade /Wikidata_Vectors_0.2 Wikidata Entity Embeddings 0.2 Dataset Summary Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.textfeature-extraction10M<n<100M3 likes5.6k downloads1mo agoHugging Face07kierth /retail-products-philippinesimage1K<n<10K1 likes5.2k downloads5mo agoHugging Face08phiyodr /InpaintCOCO InpaintCOCO - Fine-grained multimodal concept understanding (for color, size, and COCO objects) Dataset Summary A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object. Many multimodal tasks, such as Vision-Language Retrieval and Visual Question Answering, present results in terms of overall performance. Unfortunately, this approach overlooks more nuanced concepts, leaving us unaware… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/InpaintCOCO.imageimage-to-text1K<n<10K5 likes5.1k downloads2y agoHugging Face09PhisherJR /Truecallertext100M<n<1B1 likes5k downloads1mo agoHugging Face10phishdestroy /destroylist PhishDestroy Blocklist Dataset Real-time feed of phishing, crypto drainer, and scam domains detected by PhishDestroy. Updated hourly from GitHub. Statistics Metric Count Total Domains 204,669 DNS Active 126,539 Content Active 88,358 Dead Domains 78,090 Community Blocklist 1,082,988 Added Today 4 Added This Week 4 Last updated: 2026-09-13 06:30 UTC Files File Description list.json Full domain list (JSON array)… See the full description on the dataset page: https://huggingface.co/datasets/phishdestroy/destroylist.texttext-classification1M<n<10M5 likes3.6k downloads12d agoHugging Face11philschmid /dolly-15k-oai-style Dataset Card for "dolly-15k-oai-style" More Information needed text10K<n<100K7 likes2.9k downloads3y agoHugging Face12Philip-MIT /sole_training_data This is the training dataset for SOLE-R1-8B SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning. This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.image1M<n<10M0 likes2.9k downloads4mo agoHugging Face13philschmid /trl-test-instructiontextn<1K0 likes2.8k downloads3y agoHugging Face14LingoIITGN /PHINCAbstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities, it is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to communicate efficiently with the target audience. But, the noisy user-generated code-mixed text adds to the challenge of processing and understanding natural language to a much larger extent. Machine translation from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PHINC.texttranslation10K<n<100K1 likes2.6k downloads2y agoHugging Face15philschmid /guanaco-sharegpt-style Dataset Card for "guanaco-sharegpt-style" More Information needed text1K<n<10K49 likes2.3k downloads3y agoHugging Face16PhilipMay /stsb_multi_mt Dataset Card for STSb Multi MT Dataset Summary STS Benchmark comprises a selection of the English datasets used in the STS tasks organized in the context of SemEval between 2012 and 2017. The selection of datasets include text from image captions, news headlines and user forums. (source) These are different multilingual translations and the English original of the STSbenchmark dataset. Translation has been done with deepl.com. It can be used to train sentence embeddings… See the full description on the dataset page: https://huggingface.co/datasets/PhilipMay/stsb_multi_mt.texttext-classification10K<n<100K68 likes2.2k downloads2y agoHugging Face17phiyodr /coco2017 coco2017 Image-text pairs from MS COCO2017. Data origin Data originates from cocodataset.org While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy. phiyodr/coco2017: One row corresponds one image with several sentences. phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.imageimage-to-text100K<n<1M29 likes2k downloads3y agoHugging Face18PhisherJR /ULP-logstext10M<n<100M0 likes1.9k downloads9d agoHugging Face19open-phi /textbooks Textbooks Are All You Need Leveraging Large Language Models (LLMs), there's an opportunity to create a comprehensive open-source repository reminiscent of the historic Library of Alexandria. This initiative represents a preliminary attempt at producing high-quality books covering an extensive range of subjects. The source of these samples varies: Some generated using the RAG model, referencing Wikipedia or other search data. Some are completely synthetically generated. Some created… See the full description on the dataset page: https://huggingface.co/datasets/open-phi/textbooks.text1K<n<10K97 likes1.8k downloads3y agoHugging Face20phields /a-share-l2-market-depth China A-share Level 2 Market Depth Canonical order-event and ten-level snapshot data for China A-shares. Canonical trade records remain in the separate phields/a-share-l2-trades dataset. Coverage Date range: 2026-07-24 to 2026-07-24 Trading days: 1 Table Rows Parquet files Compressed size l2_orders 249,705,486 10 2.14 GiB l2_snapshots 20,279,887 4 0.91 GiB Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.tabular10B<n<100B0 likes1.8k downloads2mo agoHugging Face21bettergovph /raw-philippine-data Raw Philippine Data This repository contains raw data about Philippine politicians, public officials, and legislative documents collected from various sources. The data is intended for research, analysis, and civic technology purposes. Dataset Overview This dataset currently contains: Persons 45,424 person records of Philippine politicians and public officials with: ID: Unique identifier (ULID format) First Name: Person's first name Last Name: Person's last… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/raw-philippine-data.text100K<n<1M0 likes1.6k downloads11mo agoHugging Face22PhisherJR /phonebooktext1B<n<10B0 likes1.5k downloads9d agoHugging Face23zefang-liu /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K38 likes1.3k downloads3y agoHugging Face24philschmid /markdown-documentation-transformers Hugging Face Transformers documentation as markdown dataset This dataset was created using Clipper.js. Clipper is a Node.js command line tool that allows you to easily clip content from web pages and convert it to Markdown. It uses Mozilla's Readability library and Turndown under the hood to parse web page content and convert it to Markdown. This dataset can be used to create RAG applications, which want to use the transformers documentation. Example document:… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/markdown-documentation-transformers.textn<1K11 likes917 downloads3y agoHugging Face25PhisherJR /850M-India-datatext100M<n<1B0 likes914 downloads2mo agoHugging Face26TAUR-Lab /Taur_CoT_Analysis_Project___microsoft__Phi-3-small-8k-instructtext10K<n<100K0 likes907 downloads2y agoHugging Face27saidutta69 /PhishTrap PhishTrap Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours. Priorities: Quality > Ease of Access > Quantity Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published. Dataset Overview PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.tabular10K<n<100K0 likes835 downloads9h agoHugging Face28ealvaradob /phishing-datasetDataset designed for phishing classification tasks in various data types.texttext-classification10K<n<100K63 likes829 downloads3y agoHugging Face29philschmid /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. tabularn<1K0 likes789 downloads2y agoHugging Face30AiresPucrs /stanford-encyclopedia-philosophy Stanford Encyclopedia Philosophy (Teeny-Tiny Castle) This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research. How to Use from datasets import load_dataset dataset = load_dataset("AiresPucrs/stanford-encyclopedia-philosophy", split = 'train') texttext-classification100K<n<1M53 likes752 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.