datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikidata
Wikidata Entities Connected to Wikipedia
This dataset is a multilingual, JSON-formatted version of the Wikidata dump from May 7, 2026. It contains 73,769,737 entities after filtering out scholarly articles from the original 120,182,414 entity dump.
Curated by: Jonathan Fraine & Philippe Saadé, Wikimedia Deutschland
Funded by: Wikimedia Deutschland
Language(s) (NLP): All Wikidata Languages
License: CC0-1.0
Dataset Structure
Each row in this dataset represents a… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/wikidata.a-share-l2-trades
China A-share Level 2 Trades
Canonical Level 2 trade records for China A-shares, stored as one fact table.
Coverage
Date range: 2026-04-01 to 2026-09-24
Trading days: 119
Rows: 18730990496
Parquet files: 842
Compressed local size: 149.49 GiB
Layout
data/l2_trades/
trade_date=YYYY-MM-DD/
code_prefix=00/
part-00000.parquet
code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68.
Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.Wikidata_Vectors_0.2
Wikidata Entity Embeddings 0.2
Dataset Summary
Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata.
The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.retail-products-philippinesInpaintCOCO
InpaintCOCO - Fine-grained multimodal concept understanding (for color, size, and COCO objects)
Dataset Summary
A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object.
Many multimodal tasks, such as Vision-Language Retrieval and Visual Question Answering, present results in terms of overall performance.
Unfortunately, this approach overlooks more nuanced concepts, leaving us unaware… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/InpaintCOCO.Truecallerdolly-15k-oai-style
Dataset Card for "dolly-15k-oai-style"
More Information needed
stsb_multi_mt
Dataset Card for STSb Multi MT
Dataset Summary
STS Benchmark comprises a selection of the English datasets used in the STS tasks organized
in the context of SemEval between 2012 and 2017. The selection of datasets include text from
image captions, news headlines and user forums. (source)
These are different multilingual translations and the English original of the STSbenchmark dataset. Translation has been done with deepl.com. It can be used to train sentence embeddings… See the full description on the dataset page: https://huggingface.co/datasets/PhilipMay/stsb_multi_mt.guanaco-sharegpt-style
Dataset Card for "guanaco-sharegpt-style"
More Information needed
ULP-logscoco2017
coco2017
Image-text pairs from MS COCO2017.
Data origin
Data originates from cocodataset.org
While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy.
phiyodr/coco2017: One row corresponds one image with several sentences.
phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.a-share-l2-market-depth
China A-share Level 2 Market Depth
Canonical order-event and ten-level snapshot data for China A-shares. Canonical
trade records remain in the separate phields/a-share-l2-trades dataset.
Coverage
Date range: 2026-07-24 to 2026-07-24
Trading days: 1
Table
Rows
Parquet files
Compressed size
l2_orders
249,705,486
10
2.14 GiB
l2_snapshots
20,279,887
4
0.91 GiB
Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.textbooks
Textbooks Are All You Need
Leveraging Large Language Models (LLMs), there's an opportunity to create a comprehensive open-source repository reminiscent of the historic Library of Alexandria.
This initiative represents a preliminary attempt at producing high-quality books covering an extensive range of subjects. The source of these samples varies:
Some generated using the RAG model, referencing Wikipedia or other search data.
Some are completely synthetically generated.
Some created… See the full description on the dataset page: https://huggingface.co/datasets/open-phi/textbooks.raw-philippine-data
Raw Philippine Data
This repository contains raw data about Philippine politicians, public officials, and legislative documents collected from various sources. The data is intended for research, analysis, and civic technology purposes.
Dataset Overview
This dataset currently contains:
Persons
45,424 person records of Philippine politicians and public officials with:
ID: Unique identifier (ULID format)
First Name: Person's first name
Last Name: Person's last… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/raw-philippine-data.phonebookTaur_CoT_Analysis_Project___microsoft__Phi-3-small-8k-instruct850M-India-dataPhishTrap
PhishTrap
Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours.
Priorities: Quality > Ease of Access > Quantity
Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published.
Dataset Overview
PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.stanford-encyclopedia-philosophy
Stanford Encyclopedia Philosophy (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/stanford-encyclopedia-philosophy", split = 'train')
ears
EARS: Expressive Anechoic Recordings of Speech
This is a mirror of the Expressive Anechoic Recordings of Speech (EARS) dataset.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
Train: 92 hours, 15939 utterances, speakers p001 to p099
Validation: 2 hours, 322 utterances, speakers p100 and p101
Test: 6 hours, 966 utterances, speakers p102 to p107
License: CC BY-NC 4.0
Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/ears.wham
WHAM!48kHz noise dataset
This is a mirror of the WHAM!48kHz noise dataset.
The original files were segmented and converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 2
Format: Opus
Splits:
Train: 59 hours, 21216 segments, files 000 to 188
Validation: 12 hours, 4444 segments, files 189 to 225
Test: 7 hours, 2613 segments, files 226 to 249
License: CC BY-NC 4.0
Source: http://wham.whisper.ai/
Paper: WHAM!: Extending Speech… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/wham.fsd50k
FSD50K: An open dataset of human-labeled sound events
This is a mirror of the FSD50K sound event dataset.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
Dev: 80 hours, 40966 clips.
Eval: 28 hours, 10231 clips.
License: FSD50K is released under CC-BY. However, each clip has its own licence. Clip licenses include CC0, CC-BY, CC-BY-NC and CC Sampling+. Clip licenses are specified… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/fsd50k.seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.hle_math_category_phi4dns5
DNS5 Challenge data
This is a mirror of the DNS5 Challenge data.
The original files were converted from WAV to Opus to reduce the size and accelerate streaming.
⚠️ Only the LibriVox, AudioSet, Freesound, OpenSLR26, and OpenSLR28 data is included. The VCTK, VocalSet, CREMA-D, VoxCeleb2, and DEMAND data is excluded. ⚠️
Sampling rate: 48 kHz
Channels: 1
Format: Opus
Splits:
speech_english: 245 hours, 186743 files
speech_french: 95 hours, 60454 files
speech_german: 137 hours, 119175… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/dns5.full-math-private-n256-Phi-4-mini-instruct-bonPhishing_urls
Dataset Card for "Phishing_urls"
More Information needed
robommeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1600,
"total_frames": 768897,
"total_tasks": 116,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:1600"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/phicoltan/robomme.gretel-synthetic-text-to-sql
Fork of gretelai/synthetic_text_to_sql
The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance.
phishing-url
Dataset Description
The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs.
Features are from three different classes:
56 extracted from the structure and syntax of URLs
24 extracted from the content of their correspondent pages
7 are extracetd by querying external services.
The… See the full description on the dataset page: https://huggingface.co/datasets/pirocheto/phishing-url.
