datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
candidate-matching-synthetic
💼 Candidate-Job Matching Synthetic Dataset
A large-scale, high-quality synthetic dataset for resume-job matching and candidate retrieval tasks, generated using state-of-the-art LLMs (Qwen 2.5).
📊 Dataset Overview
This dataset contains 10,000 synthetic resumes, 2,500 job postings, and 2,500 ground truth matching records designed for training and evaluating candidate-job matching systems.
Figure 1: Balanced distribution across seniority levels - Supply (Resumes)… See the full description on the dataset page: https://huggingface.co/datasets/michaelozon/candidate-matching-synthetic.p2-etf-score-matching-resultsarxiv-abstract-matching
Dataset Card for "arxiv-abstract-matching"
More Information needed
image-matching-test-datasetEMID-Emotion-Matching
EMID-Emotion-Matching
orrzohar/EMID-Emotion-Matching is a derived dataset built on top of
the Emotionally paired Music and Image Dataset (EMID) from ECNU (ecnu-aigc/EMID).
It is designed for music ↔ image emotion matching with Qwen-Omni–style models.
Each example contains:
audio: mono waveform stored as datasets.Audio (HF Hub preview can play it)
sampling_rate: sampling rate used when decoding (typically 16 kHz)
image: a single image (datasets.Image)
same: bool, whether the audio… See the full description on the dataset page: https://huggingface.co/datasets/orrzohar/EMID-Emotion-Matching.meta-flow-matching-organoid-data-preprocessedStorage for Meta Flow Matching preprocessed drug organoid dataset.
Download preprocessed data here.
This data is sourced from Ramos Zapatero et al. Trellis tree-based analysis reveals stromal regulation of patient-derived organoid drug
responses. Cell. 2023.
The raw data can be downloaded here: Data (biological data download link)
For more information checkout: lazaratan/meta-flow-matching
matchinglaplacian-matching-theorem-frontier-v16
Laplacian Matching Theorem Frontier v16: A Sharp One-KE Bound
Produced by the Ouroboros AI Research System, under human direction.
Continuation of v15
This is the immutable v16 continuation of
cjc0013/laplacian-matching-theorem-frontier-v15,
pinned at revision fc3d7290908806cf06f04a27edf34477402f50c2. The v15 repository is preserved
unchanged. V15 isolated a difficult graph-local Hall-core envelope and left its
larger-core extension open. V16 does not silently… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/laplacian-matching-theorem-frontier-v16.GLAMI-Entity-Matching-Dataset
GLAMI Duplication Detection
Product-duplicate detection over GLAMI e-commerce listings: ~1.3M product
images plus multilingual titles, descriptions and attributes, with labelled
groups of items that do or do not refer to the same physical product.
Released under the Apache License 2.0 — see LICENSE.
TODO: describe how the labels were produced.
Structure
Config
Files
Contents
images
images/shard-*.parquet
itemId → image bytes, one row per product image… See the full description on the dataset page: https://huggingface.co/datasets/zidcenek/GLAMI-Entity-Matching-Dataset.image-matching-datasetshape_matching_uThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 47,
"total_frames": 22374,
"total_tasks": 1,
"total_videos": 94,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:47"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jccj/shape_matching_u.so100_train_move_six_blocks_tray_to_matching_dishes_flipThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 11246,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lt-s/so100_train_move_six_blocks_tray_to_matching_dishes_flip.so100_train_move_six_blocks_tray_to_matching_dishesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 11829,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lt-s/so100_train_move_six_blocks_tray_to_matching_dishes.abo-matching
ABO subset — product matching (teaching)
Subset of Amazon Berkeley Objects
(CC BY 4.0), repackaged for a practical on vision-language product matching:
« are these two listings the same product? »
⚠️ Positive pairs are simulated. ABO has no cross-seller duplicates, so a
positive pair is two catalog views of the same listing. Real marketplace
duplicates are two different sellers photographing the same object — harder, and
what the industrial pipeline (embedding for recall… See the full description on the dataset page: https://huggingface.co/datasets/bpiwowar/abo-matching.shape_matching2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 57,
"total_frames": 26946,
"total_tasks": 1,
"total_videos": 114,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:57"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jccj/shape_matching2.matching_data
Matching datasets
This "dataset" contains most commons matching datasets, with their oriented versions, all in one place.
The oriented version are use in zero-shot scenario such as in Diffumatch
rsj2025_train_resolve_any_anomalies_and_move_all_blocks_to_matching_dishesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 40,
"total_frames": 25212,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lt-s/rsj2025_train_resolve_any_anomalies_and_move_all_blocks_to_matching_dishes.beyond-pattern-matchingtask1285_kpa_keypoint_matching
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1285_kpa_keypoint_matching
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1285_kpa_keypoint_matching.laplacian-matching-theorem-frontier-v18
Laplacian Matching Theorem Frontier v18
This immutable successor to v17 is independently reproduced end to end. It packages the complete computational dependency
closure behind the one-Konig-Egervary theorem archive: exact input rows, symbolic checks, Z3,
cvc5, Lean 4.32.0, graph-atlas replay, adversarial replay, Hall-polynomial replay, paper source,
and SHA-backed expected outputs.
Run the pinned workflow with the command in REPRODUCE.md. Reproduction verifies the archive's… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/laplacian-matching-theorem-frontier-v18.flow-matching-landing-episodesshape_matching
shape_matching
This dataset was generated using the phospho cli
More information on robots.phospho.ai.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
pattern-matching-suppressionREmatch-Investment-Matching-Dataset
REmatch Investment Matching Dataset
Many people want to invest in real estate but do not know how to analyze markets, compare property data, or identify which type of property fits the way they want to invest.
Different investors have different budgets, risk preferences, liquidity needs, financing preferences, and investment goals. Therefore, the same property may be suitable for one investor and unsuitable for another.
REmatch addresses this problem by using a short… See the full description on the dataset page: https://huggingface.co/datasets/omershahar/REmatch-Investment-Matching-Dataset.diagnosis-matching-project-data
Diagnosis Matching Dataset
Dataset for matching pairs of free-text Russian medical diagnoses. Part of the kilimanj4r0/diagnosis-matching-project.
Files
File
Size
Description
final-test.csv
55 MB
Raw diagnosis pairs (272K rows)
final-test-preprocessed.csv
77 MB
Cleaned diagnosis pairs
rag-diagnosis2icd.json
2.2 MB
Diagnosis → ICD-10 mapping (~16K entries)
rag-bergamot-index.faiss
51 MB
FAISS embedding index for RAG
rag-bergamot-index.pkl
3.5 MB… See the full description on the dataset page: https://huggingface.co/datasets/sm1rk/diagnosis-matching-project-data.theorem-matching
TheoremGraph Matching
Formal–informal theorem matches from the TheoremGraph paper. Each row pairs a
Lean declaration with the most similar natural-language statement from arXiv,
found by cosine similarity over slogan embeddings, and labeled by an LLM judge
as exact, inexact, or wrong (the first two count as a match).
The file contains every candidate pair at cosine similarity 0.80 and above:
100,831 pairs. Our primary judge, GPT-5.4, labels 47,952 of them as matches; a
second… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-matching.mediasum-summary-matching
Dataset Card for "mediasum-summary-matching"
More Information needed
pubmed-abstract-matching
Dataset Card for "pubmed-abstract-matching"
More Information needed
resume-matching-dataset-v2
📄 Dataset Card - Resume Matching Dataset v2
Overview
This dataset is designed for training and evaluating large language models (LLMs) on resume-job matching tasks, specifically in AI and software engineering domains.
All data samples were generated using GPT-4o-mini.
The dataset exclusively contains synthetic data — no real resumes, self-introductions, or job postings are used.
This dataset targets three roles:
AI/LLM Developer
Frontend Developer
Backend Developer… See the full description on the dataset page: https://huggingface.co/datasets/Divyanandh/resume-matching-dataset-v2.deepswe-verifier-only-matching-pairs-v1
