CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /us-patentsThe us-patents dataset is a collection of ~ 8M US patent grants and applications from 1976-2025, cleaned, filtered, and formatted for pre-training of language models. Document Format corpus_id: Unique integer key with no semantic value. filing_date: The filing date of the grant or application. In case of duplicates, earliest filing date from the duplicate cluster. patent_type: The type of patent. text: The text content of the concatenated title, abstract, and specification.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/us-patents.texttext-generation1M<n<10M14 likes2.2k downloads9mo agoHugging Face02Mouuns /Patent-search0 likes1.7k downloads11mo agoHugging Face03SoichiOnozuka /design-patents-not-in-impact US Design Patents Not Included in IMPACT (2008-2026) Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are absent from the AI4Patents/IMPACT dataset. IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period, plus 4,824 patents from years IMPACT does cover but did not include. There is no patent overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.text1M<n<10M0 likes1.1k downloads2mo agoHugging Face04labofsahil /patents-publications-datasettabular100M<n<1B0 likes1k downloads8mo agoHugging Face05nbettencourt /google-patents-data-previewtabular100K<n<1M0 likes777 downloads1y agoHugging Face06dvdmrs09 /patentstabular10M<n<100M3 likes305 downloads2y agoHugging Face07Yehoon /patent-spec-xml0 likes295 downloads6mo agoHugging Face08AI-Growth-Lab /patents_claims_1.5m_traim_testtabular1M<n<10M10 likes252 downloads4y agoHugging Face09zalizedata /us-patents-citation-company-dataset US Patents, Citations & Assignee Graph (USPTO / PatentsView) 9.1M US patents (1976–present) with the full citation graph, assignees, inventors and CPC classifications — 255M+ rows in the full-graph edition. Built from official USPTO/PatentsView bulk data. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/us-patents-citation-company-dataset Formats & how to… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-patents-citation-company-dataset.tabulartext-classification100M<n<1B0 likes225 downloads2mo agoHugging Face10QNLOO /01-3-cognition-river-patents 千年鹿认知体系 · 技术专利与合规档案库|QNLOO Cognition System · Patents & Compliance Archive 1. 本仓库定位 / Repository Profile 性质:全球官方唯一《认知之河》技术专利与合规授权档案库,为定稿、只读、存证级官方仓库,用于全套发明专利文书的公开备案与在先技术证据固化。 Nature: The world's official archive for River of Cognition technical patents and compliance authorization documents. A finalized, read-only repository dedicated to the public filing of complete invention patent documents and prior-art evidence preservation. 资产形态 / Asset Format… See the full description on the dataset page: https://huggingface.co/datasets/QNLOO/01-3-cognition-river-patents.0 likes152 downloads3mo agoHugging Face11niban73 /my-patents-data Global Patent Publications with Chinese Patent-Family Quality Measures This dataset repository contains two Parquet files. Files data/cn_family_quality_master_gft.parquet One row per Chinese focal invention patent family for 2003–2019. It contains the cumulative patent-quality pipeline, including: semantic knowledge recombination (semantic KI); strict historical semantic novelty; IPC-based KI robustness measure; three-year and five-year forward… See the full description on the dataset page: https://huggingface.co/datasets/niban73/my-patents-data.tabular-classification0 likes152 downloads2mo agoHugging Face12sutro /apple-patents-embeddingsApple patent embeddings for https://docs.sutro.sh/examples/large-scale-embeddings configs: - config_name: default data_files: - split: train path: data/train-* dataset_info: features: - name: text dtype: large_string - name: job-76844041-b2bf-4248-9603-b7f750231b34 large_list: float64 splits: - name: train num_bytes: 37048620333 num_examples: 4039988 download_size: 7413609880 dataset_size: 37048620333 license: mit text1M<n<10M0 likes146 downloads1y agoHugging Face13INPI-France /French-Patents-2020-2026-Raw 🇫🇷 Brevets français 2020–2026 — RAW 🇫🇷 Dataset de brevets français publiés entre 2020 et 2026, extrait depuis les XML d’origine, avec un document = une ligne (texte complet). Format : Parquet, prêt pour chargement streaming / distribué. Source Données issues de documents publics de brevets français (A1).Extraction réalisé de manière indépendante grâce à un accès aux API / FTP PI (sur demande à l’INPI). Génération 468 000 fichiers XML… See the full description on the dataset page: https://huggingface.co/datasets/INPI-France/French-Patents-2020-2026-Raw.text100K<n<1M3 likes135 downloads6mo agoHugging Face14ExponentialScience /DLT-Patents DLT-Patents Paper | Code Dataset Description Dataset Summary DLT-Patents is a comprehensive corpus of patent documents related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, innovation studies, and patent analysis in the DLT domain. The dataset contains 49,023 patent documents with 1,296 million tokens (1.296 billion tokens), spanning patents from 1990 to 2025. All documents… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Patents.textfeature-extraction10K<n<100K0 likes100 downloads6mo agoHugging Face15osanseviero /us-patentstext10K<n<100K0 likes87 downloads4y agoHugging Face16Opscidia /patentstext10K<n<100K0 likes85 downloads2y agoHugging Face17BASF-AI /google-patents-chem-translation-pairs0 likes78 downloads26d agoHugging Face18istat-ai /patents-classified-2106-gpt5-minitabular1K<n<10K1 likes65 downloads1y agoHugging Face19alea-institute /kl3m-sft-patentstext1M<n<10M4 likes59 downloads2y agoHugging Face20kennbyee25 /trundle_patents-pocn<1K0 likes53 downloads4y agoHugging Face21kennbyee25 /distilroberta-base_tokenized_trundle-patentsn<1K0 likes47 downloads4y agoHugging Face22Orionfold /patent-strategist-bench-v0.1 Patent-Strategist Bench v0.1 A 200-question, seven-shape benchmark for patent-prosecution reasoning, anchored to three public sources (USPTO MPEP, HPI-Naumann PatentMatch, BIGPATENT) with oracle context attached to every row. Built to evaluate whether a small open LLM can perform the day-to-day reasoning tasks of a patent practitioner. Companion artifact to two methodology articles: Patent-Strategist v1 baseline on Spark — establishes the first tri-mode (closed-book / retrieval /… See the full description on the dataset page: https://huggingface.co/datasets/Orionfold/patent-strategist-bench-v0.1.textquestion-answeringn<1K0 likes47 downloads4mo agoHugging Face23kennbyee25 /distilroberta-base_tokenized_english_patents1K<n<10K0 likes46 downloads4y agoHugging Face24csmoilis /model_df_patentSBERTatabular10K<n<100K0 likes42 downloads7mo agoHugging Face25AI-Growth-Lab /patents_claims_1.5m_traim_test_embeddings0 likes41 downloads4y agoHugging Face26kennbyee25 /english-patents_titlestext10K<n<100K3 likes40 downloads4y agoHugging Face27visualcomments /patents_edtechWe have collected a representative sample of patent data from various BRICS countries. We used a limited interpretation of the BRICS member countries and conducted an analysis of patent activity in the following countries: Brazil, India, China, Russia, and South Africa. As data sources, we used the websites patents.google.com and patentscope.wipo.int. Since the data in these sources is presented unevenly, we simultaneously used data from different sources to achieve maximum completeness in… See the full description on the dataset page: https://huggingface.co/datasets/visualcomments/patents_edtech.0 likes39 downloads2y agoHugging Face28CTB2001 /patents-green-50k Green Patent Claims — 50k Balanced Dataset A balanced binary-classification dataset of 50,000 US patent first claims labelled as green technology (1) or not green (0), plus 100 human-verified gold labels from a targeted HITL review. Created for the Applied Deep Learning (AAU, Spring 2025) exam assignment on active learning, multi-agent silver labelling, and human-in-the-loop verification. Dataset Details Property Value Total examples 50,000 Label balance… See the full description on the dataset page: https://huggingface.co/datasets/CTB2001/patents-green-50k.tabulartext-classification10K<n<100K0 likes39 downloads7mo agoHugging Face29kennbyee25 /distilroberta-base_tokenized_english-patents_210K<n<100K0 likes38 downloads4y agoHugging Face30electricsheepafrica /africa-egypt-capmas-patents-and-trademarks-efed9124 Patents and Trademarks | Africa (CAPMAS Egypt Open Data) 5,615 rows - 1 Africa country/area - 2010-2023 - 26 indicators - Engineered by Electric Sheep Africa TL;DR This dataset contains 5,615 rows from CAPMAS Egypt Open Data, covering Patents and Trademarks. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures Economic datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-egypt-capmas-patents-and-trademarks-efed9124.tabulartabular-regression1K<n<10K0 likes38 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.