datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
a-share-l2-trades
China A-share Level 2 Trades
Canonical Level 2 trade records for China A-shares, stored as one fact table.
Coverage
Date range: 2026-04-01 to 2026-09-24
Trading days: 119
Rows: 18730990496
Parquet files: 842
Compressed local size: 149.49 GiB
Layout
data/l2_trades/
trade_date=YYYY-MM-DD/
code_prefix=00/
part-00000.parquet
code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68.
Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.coco2017
coco2017
Image-text pairs from MS COCO2017.
Data origin
Data originates from cocodataset.org
While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy.
phiyodr/coco2017: One row corresponds one image with several sentences.
phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.a-share-l2-market-depth
China A-share Level 2 Market Depth
Canonical order-event and ten-level snapshot data for China A-shares. Canonical
trade records remain in the separate phields/a-share-l2-trades dataset.
Coverage
Date range: 2026-07-24 to 2026-07-24
Trading days: 1
Table
Rows
Parquet files
Compressed size
l2_orders
249,705,486
10
2.14 GiB
l2_snapshots
20,279,887
4
0.91 GiB
Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.PhishTrap
PhishTrap
Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours.
Priorities: Quality > Ease of Access > Quantity
Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published.
Dataset Overview
PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.full-math-private-n256-Phi-4-mini-instruct-bonrobommeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1600,
"total_frames": 768897,
"total_tasks": 116,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:1600"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/phicoltan/robomme.phishing-url
Dataset Description
The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs.
Features are from three different classes:
56 extracted from the structure and syntax of URLs
24 extracted from the content of their correspondent pages
7 are extracetd by querying external services.
The… See the full description on the dataset page: https://huggingface.co/datasets/pirocheto/phishing-url.full-aime_2026-n256-Phi-4-mini-instruct-bond1_science_load_in_phi_temp40d1_science_load_in_phi_temp20d1_science_load_in_phi_temp2picko-v7This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 10488,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Philmat/picko-v7.d1_code_load_in_phi_temp_2d1_science_load_in_phi_temp10d1_math_load_in_phi_temp2Magpie-Phi3-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/Magpie-Phi3-Pro-1M-v0.1.preprocessed-full-MATH-500-n256-Phi-4-mini-instruct-bondetails_EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200.preprocessed-full-gsm8k-private-n256-Phi-4-mini-instruct-bonfed-phishing-urls
Dataset Card for Federated Phishing URLs
Dataset Summary
This dataset is a federated, non-IID phishing URL classification benchmark derived from two public Hugging Face datasets:
ealvaradob/phishing-dataset, using the urls.json file.
kmack/Phishing_urls, using the merged train+test+valid splits.
The resulting dataset contains URL strings, binary phishing labels, and a client_id field assigning each example to one of 100 simulated clients.
Client assignment is… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/fed-phishing-urls.ipfs_philippines_laws_ir
Philippines legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_philippines_laws (revision 2f7ee80e0ff2a2d687d339bb6a5e45733db2e6ec) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Philippines prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_philippines_laws_ir.full-MATH-500-n256-Phi-4-mini-instruct-bond1_math_load_in_phi_temp4e1_math_all_phi_temp40modified-swiss-dwellings-enriched
Modified Swiss Dwellings (MSD), enriched
Floor plans of medium-to-large multi-apartment building complexes (ECCV 2024 benchmark
MSD), each linking three modalities of one floor plan:
image, geometry, and access graph.
1. Why this is here & what was done
Hosted on Hugging Face for reach and one-line loading by the ML community. The Swiss
Dwellings (SD) license (CC BY 4.0) permits redistribution with attribution — so this
also enriches the public MSD release, which… See the full description on the dataset page: https://huggingface.co/datasets/philippds/modified-swiss-dwellings-enriched.e1_math_all_phi_temp10e1_code_fasttext_phi_temp2e1_science_longest_phipicko-v6
picko-v6
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
