datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us-patentsThe us-patents dataset is a collection of ~ 8M US patent grants and applications from 1976-2025, cleaned, filtered, and formatted for pre-training of language models.
Document Format
corpus_id: Unique integer key with no semantic value.
filing_date: The filing date of the grant or application. In case of duplicates, earliest filing date from the duplicate cluster.
patent_type: The type of patent.
text: The text content of the concatenated title, abstract, and specification.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/us-patents.Patent-searchdesign-patents-not-in-impact
US Design Patents Not Included in IMPACT (2008-2026)
Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are
absent from the AI4Patents/IMPACT dataset.
IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that
IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period,
plus 4,824 patents from years IMPACT does cover but did not include. There is no patent
overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.patents-publications-datasetgoogle-patents-data-previewpatentspatent-spec-xmlpatents_claims_1.5m_traim_testus-patents-citation-company-dataset
US Patents, Citations & Assignee Graph (USPTO / PatentsView)
9.1M US patents (1976–present) with the full citation graph, assignees, inventors and CPC classifications — 255M+ rows in the full-graph edition. Built from official USPTO/PatentsView bulk data.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/us-patents-citation-company-dataset
Formats & how to… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-patents-citation-company-dataset.01-3-cognition-river-patents
千年鹿认知体系 · 技术专利与合规档案库|QNLOO Cognition System · Patents & Compliance Archive
1. 本仓库定位 / Repository Profile
性质:全球官方唯一《认知之河》技术专利与合规授权档案库,为定稿、只读、存证级官方仓库,用于全套发明专利文书的公开备案与在先技术证据固化。
Nature: The world's official archive for River of Cognition technical patents and compliance authorization documents. A finalized, read-only repository dedicated to the public filing of complete invention patent documents and prior-art evidence preservation.
资产形态 / Asset Format… See the full description on the dataset page: https://huggingface.co/datasets/QNLOO/01-3-cognition-river-patents.my-patents-data
Global Patent Publications with Chinese Patent-Family Quality Measures
This dataset repository contains two Parquet files.
Files
data/cn_family_quality_master_gft.parquet
One row per Chinese focal invention patent family for 2003–2019.
It contains the cumulative patent-quality pipeline, including:
semantic knowledge recombination (semantic KI);
strict historical semantic novelty;
IPC-based KI robustness measure;
three-year and five-year forward… See the full description on the dataset page: https://huggingface.co/datasets/niban73/my-patents-data.apple-patents-embeddingsApple patent embeddings for https://docs.sutro.sh/examples/large-scale-embeddings
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
dataset_info:
features:
- name: text
dtype: large_string
- name: job-76844041-b2bf-4248-9603-b7f750231b34
large_list: float64
splits:
- name: train
num_bytes: 37048620333
num_examples: 4039988
download_size: 7413609880
dataset_size: 37048620333
license: mit
French-Patents-2020-2026-Raw
🇫🇷 Brevets français 2020–2026 — RAW 🇫🇷
Dataset de brevets français publiés entre 2020 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet).
Format : Parquet, prêt pour chargement streaming / distribué.
Source
Données issues de documents publics de brevets français (A1).Extraction réalisé de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI).
Génération
468 000 fichiers XML… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/French-Patents-2020-2026-Raw.DLT-Patents
DLT-Patents
Paper | Code
Dataset Description
Dataset Summary
DLT-Patents is a comprehensive corpus of patent documents related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, innovation studies, and patent analysis in the DLT domain.
The dataset contains 49,023 patent documents with 1,296 million tokens (1.296 billion tokens), spanning patents from 1990 to 2025. All documents… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Patents.us-patentspatentsgoogle-patents-chem-translation-pairspatents-classified-2106-gpt5-minikl3m-sft-patentstrundle_patents-pocdistilroberta-base_tokenized_trundle-patentspatent-strategist-bench-v0.1
Patent-Strategist Bench v0.1
A 200-question, seven-shape benchmark for patent-prosecution reasoning, anchored
to three public sources (USPTO MPEP, HPI-Naumann PatentMatch, BIGPATENT) with
oracle context attached to every row. Built to evaluate whether a small open
LLM can perform the day-to-day reasoning tasks of a patent practitioner.
Companion artifact to two methodology articles:
Patent-Strategist v1 baseline on Spark — establishes the first tri-mode (closed-book / retrieval /… See the full description on the dataset page: https://huggingface.co/datasets/Orionfold/patent-strategist-bench-v0.1.distilroberta-base_tokenized_english_patentsmodel_df_patentSBERTapatents_claims_1.5m_traim_test_embeddingsenglish-patents_titlespatents_edtechWe have collected a representative sample of patent data from various BRICS countries.
We used a limited interpretation of the BRICS member countries and conducted an analysis of patent activity in the following countries: Brazil, India, China, Russia, and South Africa.
As data sources, we used the websites patents.google.com and patentscope.wipo.int.
Since the data in these sources is presented unevenly, we simultaneously used data from different sources to achieve maximum completeness in… See the full description on the dataset page: https://huggingface.co/datasets/visualcomments/patents_edtech.patents-green-50k
Green Patent Claims — 50k Balanced Dataset
A balanced binary-classification dataset of 50,000 US patent first claims labelled as green technology (1) or not green (0), plus 100 human-verified gold labels from a targeted HITL review.
Created for the Applied Deep Learning (AAU, Spring 2025) exam assignment on active learning, multi-agent silver labelling, and human-in-the-loop verification.
Dataset Details
Property
Value
Total examples
50,000
Label balance… See the full description on the dataset page: https://huggingface.co/datasets/CTB2001/patents-green-50k.distilroberta-base_tokenized_english-patents_2africa-egypt-capmas-patents-and-trademarks-efed9124
Patents and Trademarks | Africa (CAPMAS Egypt Open Data)
5,615 rows - 1 Africa country/area - 2010-2023 - 26 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 5,615 rows from CAPMAS Egypt Open Data, covering Patents and Trademarks. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Economic datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-egypt-capmas-patents-and-trademarks-efed9124.
