datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OceanDepths
OceanDepths GeoTIFF Raster and Aligned ARGO Dataset
This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and
salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order
to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable
reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.a-share-l2-trades
China A-share Level 2 Trades
Canonical Level 2 trade records for China A-shares, stored as one fact table.
Coverage
Date range: 2026-04-01 to 2026-09-24
Trading days: 119
Rows: 18730990496
Parquet files: 842
Compressed local size: 149.49 GiB
Layout
data/l2_trades/
trade_date=YYYY-MM-DD/
code_prefix=00/
part-00000.parquet
code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68.
Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.coco2017
coco2017
Image-text pairs from MS COCO2017.
Data origin
Data originates from cocodataset.org
While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy.
phiyodr/coco2017: One row corresponds one image with several sentences.
phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.a-share-l2-market-depth
China A-share Level 2 Market Depth
Canonical order-event and ten-level snapshot data for China A-shares. Canonical
trade records remain in the separate phields/a-share-l2-trades dataset.
Coverage
Date range: 2026-07-24 to 2026-07-24
Trading days: 1
Table
Rows
Parquet files
Compressed size
l2_orders
249,705,486
10
2.14 GiB
l2_snapshots
20,279,887
4
0.91 GiB
Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.PhishTrap
PhishTrap
Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours.
Priorities: Quality > Ease of Access > Quantity
Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published.
Dataset Overview
PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training!
This is the Benchmark of AIME from year 1983~2023, and 2024(part 2).
Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime.
seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.full-math-private-n256-Phi-4-mini-instruct-bonrobommeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1600,
"total_frames": 768897,
"total_tasks": 116,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:1600"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/phicoltan/robomme.phishing-url
Dataset Description
The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs.
Features are from three different classes:
56 extracted from the structure and syntax of URLs
24 extracted from the content of their correspondent pages
7 are extracetd by querying external services.
The… See the full description on the dataset page: https://huggingface.co/datasets/pirocheto/phishing-url.The-Philosophy-Data-Project
About dataset
The Philosophy Data Project is a corpus and a set of anaylsis based philosophy texts, totaling over 50 texts and 30 authors, made by Kourosh Alizadeh.
school: Broad categorization of which school of thought each book belongs to. Sometimes, this classification can be vague or depend on interpretation. Thankfully, texts in this corpus are all distinctive examples of respective school of thought, so at leat here they are reasonable.
sentence_spacy and sentence_str:… See the full description on the dataset page: https://huggingface.co/datasets/yjkim27/The-Philosophy-Data-Project.full-aime_2026-n256-Phi-4-mini-instruct-bond1_science_load_in_phi_temp40d1_science_load_in_phi_temp20PHI-SPIKE-C172x-Community-Dataset-v1.0
PHI-SPIKE C172X Community Dataset v1.0
Dataset Summary
PHI-SPIKE C172X Community Dataset v1.0 is a simulation-based aerospace Prognostics and Health Management (PHM) dataset and training-artifact release developed from the PHI-SPIKE C172X research campaign.
The release provides:
JSBSim C172X reference telemetry;
benchmark metadata;
training histories;
trained PyTorch model checkpoints;
per-run evaluation metrics; and
five-seed campaign summaries.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-SPIKE-C172x-Community-Dataset-v1.0.d1_science_load_in_phi_temp2picko-v7This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 10488,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Philmat/picko-v7.d1_code_load_in_phi_temp_2d1_science_load_in_phi_temp10d1_math_load_in_phi_temp2Magpie-Phi3-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/Magpie-Phi3-Pro-1M-v0.1.preprocessed-full-MATH-500-n256-Phi-4-mini-instruct-bondetails_EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200.preprocessed-full-gsm8k-private-n256-Phi-4-mini-instruct-bontitanicThe legendary Titanic dataset from this Kaggle competition
fed-phishing-urls
Dataset Card for Federated Phishing URLs
Dataset Summary
This dataset is a federated, non-IID phishing URL classification benchmark derived from two public Hugging Face datasets:
ealvaradob/phishing-dataset, using the urls.json file.
kmack/Phishing_urls, using the merged train+test+valid splits.
The resulting dataset contains URL strings, binary phishing labels, and a client_id field assigning each example to one of 100 simulated clients.
Client assignment is… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/fed-phishing-urls.ipfs_philippines_laws_ir
Philippines legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_philippines_laws (revision 2f7ee80e0ff2a2d687d339bb6a5e45733db2e6ec) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Philippines prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_philippines_laws_ir.full-MATH-500-n256-Phi-4-mini-instruct-bond1_math_load_in_phi_temp4e1_math_all_phi_temp40
