datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
candidate-matching-synthetic
💼 Candidate-Job Matching Synthetic Dataset
A large-scale, high-quality synthetic dataset for resume-job matching and candidate retrieval tasks, generated using state-of-the-art LLMs (Qwen 2.5).
📊 Dataset Overview
This dataset contains 10,000 synthetic resumes, 2,500 job postings, and 2,500 ground truth matching records designed for training and evaluating candidate-job matching systems.
Figure 1: Balanced distribution across seniority levels - Supply (Resumes)… See the full description on the dataset page: https://huggingface.co/datasets/michaelozon/candidate-matching-synthetic.arxiv-abstract-matching
Dataset Card for "arxiv-abstract-matching"
More Information needed
EMID-Emotion-Matching
EMID-Emotion-Matching
orrzohar/EMID-Emotion-Matching is a derived dataset built on top of
the Emotionally paired Music and Image Dataset (EMID) from ECNU (ecnu-aigc/EMID).
It is designed for music ↔ image emotion matching with Qwen-Omni–style models.
Each example contains:
audio: mono waveform stored as datasets.Audio (HF Hub preview can play it)
sampling_rate: sampling rate used when decoding (typically 16 kHz)
image: a single image (datasets.Image)
same: bool, whether the audio… See the full description on the dataset page: https://huggingface.co/datasets/orrzohar/EMID-Emotion-Matching.matchinglaplacian-matching-theorem-frontier-v16
Laplacian Matching Theorem Frontier v16: A Sharp One-KE Bound
Produced by the Ouroboros AI Research System, under human direction.
Continuation of v15
This is the immutable v16 continuation of
cjc0013/laplacian-matching-theorem-frontier-v15,
pinned at revision fc3d7290908806cf06f04a27edf34477402f50c2. The v15 repository is preserved
unchanged. V15 isolated a difficult graph-local Hall-core envelope and left its
larger-core extension open. V16 does not silently… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/laplacian-matching-theorem-frontier-v16.abo-matching
ABO subset — product matching (teaching)
Subset of Amazon Berkeley Objects
(CC BY 4.0), repackaged for a practical on vision-language product matching:
« are these two listings the same product? »
⚠️ Positive pairs are simulated. ABO has no cross-seller duplicates, so a
positive pair is two catalog views of the same listing. Real marketplace
duplicates are two different sellers photographing the same object — harder, and
what the industrial pipeline (embedding for recall… See the full description on the dataset page: https://huggingface.co/datasets/bpiwowar/abo-matching.matching_data
Matching datasets
This "dataset" contains most commons matching datasets, with their oriented versions, all in one place.
The oriented version are use in zero-shot scenario such as in Diffumatch
Freelancer-Project-Matching
license: apache-2.0
task1285_kpa_keypoint_matching
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1285_kpa_keypoint_matching
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1285_kpa_keypoint_matching.pattern-matching-suppressionREmatch-Investment-Matching-Dataset
REmatch Investment Matching Dataset
Many people want to invest in real estate but do not know how to analyze markets, compare property data, or identify which type of property fits the way they want to invest.
Different investors have different budgets, risk preferences, liquidity needs, financing preferences, and investment goals. Therefore, the same property may be suitable for one investor and unsuitable for another.
REmatch addresses this problem by using a short… See the full description on the dataset page: https://huggingface.co/datasets/omershahar/REmatch-Investment-Matching-Dataset.theorem-matching
TheoremGraph Matching
Formal–informal theorem matches from the TheoremGraph paper. Each row pairs a
Lean declaration with the most similar natural-language statement from arXiv,
found by cosine similarity over slogan embeddings, and labeled by an LLM judge
as exact, inexact, or wrong (the first two count as a match).
The file contains every candidate pair at cosine similarity 0.80 and above:
100,831 pairs. Our primary judge, GPT-5.4, labels 47,952 of them as matches; a
second… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-matching.mediasum-summary-matching
Dataset Card for "mediasum-summary-matching"
More Information needed
SO101-ma_sort_RGBblock_to_matchingplate_30fpsThis dataset was created using LeRobot.
Dataset Description
SO-101 robot dataset collected with the MA autonomous forward/reset pipeline. The task is to sort red, green, and blue RGB blocks onto plates of the matching color. The dataset contains 100 successful episodes at 30 FPS with top and left-wrist RGB camera streams, robot state/action features, end-effector pose, gripper state, and skill/subtask annotations.
Homepage: https://huggingface.co/CoRL2026-CSI
Paper: [More… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/SO101-ma_sort_RGBblock_to_matchingplate_30fps.pubmed-abstract-matching
Dataset Card for "pubmed-abstract-matching"
More Information needed
flow-matching-landing-episodesresume-matching-dataset-v2
📄 Dataset Card - Resume Matching Dataset v2
Overview
This dataset is designed for training and evaluating large language models (LLMs) on resume-job matching tasks, specifically in AI and software engineering domains.
All data samples were generated using GPT-4o-mini.
The dataset exclusively contains synthetic data — no real resumes, self-introductions, or job postings are used.
This dataset targets three roles:
AI/LLM Developer
Frontend Developer
Backend Developer… See the full description on the dataset page: https://huggingface.co/datasets/Divyanandh/resume-matching-dataset-v2.irish-speech-dataset_shunya_801010_matchingdeepswe-verifier-only-matching-pairs-v1candidate-matching-synthetic
💼 Candidate-Job Matching Synthetic Dataset
A large-scale, high-quality synthetic dataset for resume-job matching and candidate retrieval tasks, generated using state-of-the-art LLMs (Qwen 2.5).
📊 Dataset Overview
This dataset contains 10,000 synthetic resumes, 2,500 job postings, and 2,500 ground truth matching records designed for training and evaluating candidate-job matching systems.
Figure 1: Balanced distribution across seniority levels - Supply… See the full description on the dataset page: https://huggingface.co/datasets/Gehan77/candidate-matching-synthetic.resume-matching-dataset-v2ecommerce-retail-product-matching-workflow-dataset
Ecommerce Retail Product Matching Workflow Dataset
This dataset is a public-facing sanitized workflow preview for managed ecommerce and retail product matching. It shows how candidate retrieval, UPC/model/brand/title/image evidence, customer-visible URL validation, confidence bands, review buckets, and rejection reasons can be structured for pricing intelligence, merchandising, data engineering, and AI-assisted product matching workflows.
Use this dataset to evaluate product… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/ecommerce-retail-product-matching-workflow-dataset.resume-matching-dataset-v2
📄 Dataset Card - Resume Matching Dataset v2
Overview
This dataset is designed for training and evaluating large language models (LLMs) on resume-job matching tasks, specifically in AI and software engineering domains.
All data samples were generated using GPT-4o-mini.
The dataset exclusively contains synthetic data — no real resumes, self-introductions, or job postings are used.
This dataset targets three roles:
AI/LLM Developer
Frontend Developer
Backend Developer… See the full description on the dataset page: https://huggingface.co/datasets/jminc/resume-matching-dataset-v2.ror-matching-train-validation-test
ROR Affiliation Matching (DataCite + Crossref + AffRoDB + Synthetic OpenAlex)
Raw author-affiliation strings paired with the ROR (Research Organization
Registry) identifiers they should resolve to, prepared for training and
evaluating affiliation matching and entity-linking systems.
The dataset ships six subsets (loadable as Hugging Face configs), each
split into train/validation/test:
Subset
Records
Source
Empty-label rows
crossref (default)
3,000
Crossref-derived… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/ror-matching-train-validation-test.SO101-cap_sort_RGBblock_to_matchingplate_10fps
SO101 CAP Sort RGB Blocks to Matching Plates
This dataset contains 100 LeRobot v3.0 demonstration episodes for an SO101 follower robot. The task is: Sort red, green, and blue blocks onto plates of the matching color. The dataset was collected at 10 Hz and includes paired top-view and wrist-view RGB videos, robot state/action trajectories, and CAP skill annotations.
Dataset Details
Field
Value
Repository… See the full description on the dataset page: https://huggingface.co/datasets/CoRL2026-CSI/SO101-cap_sort_RGBblock_to_matchingplate_10fps.ecommerce-visual-matching-dataset
E-commerce Visual Matching Dataset
Candidate match workflow dataset for product identity resolution, visual similarity, and review decision fields.
A public-safe workflow preview of how Octoparse structures AI-assisted product matching pipelines for e-commerce and pricing teams. Every row represents a candidate pair evaluation — the same structure delivered to production clients.
Built by Octoparse Managed Data Service — managed web data pipelines for pricing intelligence and… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/ecommerce-visual-matching-dataset.celeb-face-matching-datacandidate-matching-synthetic
💼 Candidate-Job Matching Synthetic Dataset
A large-scale, high-quality synthetic dataset for resume-job matching and candidate retrieval tasks, generated using state-of-the-art LLMs (Qwen 2.5).
📊 Dataset Overview
This dataset contains 10,000 synthetic resumes, 2,500 job postings, and 2,500 ground truth matching records designed for training and evaluating candidate-job matching systems.
Figure 1: Balanced distribution across seniority levels - Supply… See the full description on the dataset page: https://huggingface.co/datasets/shreyad2806/candidate-matching-synthetic.UR7e_CaP_sort_RGBblock_to_matchingplate_10fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur7e",
"total_episodes": 121,
"total_frames": 68120,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:121"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CoRL2026-CSI/UR7e_CaP_sort_RGBblock_to_matchingplate_10fps.product-matching
Dataset Description
This repository offers an ideal ground for evaluating product matching algorithms and clustering/classification models.
The dataset contains e-commerce data; that is, product IDs, their titles, and their corresponding category. However, they can easily be applied to
any problem which involves text/short-text mining.
The data originates from PriceRunner, a popular product comparison platform. It includes 35,311 products from 10 categories,
provided by 306… See the full description on the dataset page: https://huggingface.co/datasets/lakritidis/product-matching.
