datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
big_patent
Dataset Card for Big Patent
Dataset Summary
BIGPATENT, consisting of 1.3 million records of U.S. patent documents along with human written abstractive summaries.
Each US patent application is filed under a Cooperative Patent Classification (CPC) code.
There are nine such classification categories:
a: Human Necessities
b: Performing Operations; Transporting
c: Chemistry; Metallurgy
d: Textiles; Paper
e: Fixed Constructions
f: Mechanical Engineering; Lightning; Heating;… See the full description on the dataset page: https://huggingface.co/datasets/NortheasternUniversity/big_patent.Birds_of_North_America_Fullnorthwind_invocies
Northwind Invoices and Related Documents
This dataset contains a collection of invoices and related documents from the Northwind database, a sample database used by Microsoft for demonstrating database functionalities.
The invoices include information about the customer, the salesperson, the order date, order ID, product IDs, product names, quantities, unit prices, and total prices. The related documents include shipping documents and stock documents.
This dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/northwind_invocies.northwind_PurchaseOrders
Purchase Orders Dataset
This dataset consists of purchase orders from various companies. It was created by CHERGUELAINE Ayoub & BOUBEKRI Faycal with the help of ChatGPT for the purpose of document classification and analytics.
Description
The dataset contains a collection of purchase orders from different companies. Each purchase order consists of the following fields:
order_id: The unique identifier for the purchase order.
order_date: The date on which the purchase… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/northwind_PurchaseOrders.mmu_ssl_legacysurvey_north
mmu_ssl_legacysurvey_north HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_ssl_legacysurvey_north.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_ssl_legacysurvey_north.datagen-stack-v1-joint-5cam
datagen-stack-v1-joint-5cam
Auto-generated SFT dataset for the stack_retrieve family — cuRobo-planned, physics- &
LTL-safety-checked demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 28 stack_retrieve base tasks.
Per task: 40 success + LTL-safe trajectories → 1120 episodes.
Contents
Episodes
1120 (28 base tasks × 40)
Frames
2,652,083
Unique language tasks
8… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-stack-v1-joint-5cam.datagen-cabinet-v1-joint-5cam
datagen-cabinet-v1-joint-5cam
Auto-generated SFT dataset for the cabinet drawer pick-and-place family — cuRobo-planned, physics- &
LTL-safety-checked demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 35 cabinet_pickup base tasks.
Per task: 40 success + LTL-safe trajectories → 1400 episodes.
Contents
Episodes
1400 (35 base tasks × 40)
Frames
4,172,962
Unique… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-cabinet-v1-joint-5cam.northwind-Stock_rapport
Northwind Stock Report Dataset
This dataset was created by CHERGUELAINE Ayoub & BOUBEKRI Faycal for the purpose of document classification and analytics. The dataset contains monthly stock reports and monthly stock reports by category, extracted from the Northwind dataset.
The Northwind dataset is a sample database that comes with Microsoft Access, and is commonly used as a demo database for learning SQL. The dataset contains data on a fictional company called "Northwind Traders"… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/northwind-Stock_rapport.northwind_Shipping_orders
Northwind Shipping Orders and Related Documents
This dataset contains a collection of Shipping Orders and related documents from the Northwind database, a sample database used by Microsoft for demonstrating database functionalities.
The Shipping Orders include information about the ship name, Address , Region, postal code ,country, customer ,employee shipped date product names, quantities, unit prices, and total prices. The related documents include shipping documents and stock… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/northwind_Shipping_orders.datagen-lid-v1-joint-5cam
datagen-lid-v1-joint-5cam
Auto-generated SFT dataset for the lid_transport family (place a lid on a container, then
transport the lidded container into the goal region) - cuRobo-planned, physics- & LTL-safety-checked
demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench - all 30 lid_transport base tasks.
Per task: 40 success + LTL-safe trajectories -> 1200 episodes.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-lid-v1-joint-5cam.datagen-dusty-v1-joint-5cam
datagen-dusty-v1-joint-5cam
Auto-generated SFT dataset for the dusty_transfer family (wipe a dusty container clean with a
sponge, then transfer a target object into it) — cuRobo-planned, physics- & LTL-safety-checked demos,
converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 26 dusty_transfer base tasks.
Per task: 40 success + LTL-safe trajectories -> 1040 episodes.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-dusty-v1-joint-5cam.datagen-jar-v1-joint-5cam
datagen-jar-v1-joint-5cam
Auto-generated SFT dataset for the jar_transport family (close an articulated hinged jar's lid,
then side-grasp the closed jar and carry it to a goal region) — cuRobo-planned, physics- &
LTL-safety-checked demos, converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — all 26 jar_transport base tasks.
Per task: 40 success + LTL-safe trajectories → 1040 episodes.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-jar-v1-joint-5cam.datagen-clutter-v1-joint-5cam
datagen-clutter-v1-joint-5cam
Auto-generated SFT dataset for the clutter (pick-out-of-clutter → place-in-goal) family — a
cuRobo-planned, physics- & LTL-safety-checked demonstration set, already converted to LeRobot v2.1.
Creator: yypeng666 (IDEAS-Lab-Northwestern)
Source bench: IDEAS-Lab-Northwestern/ManiGuard-Bench — collected on all 55 clutter_pickup base tasks.
Per task: 40 success + LTL-safe trajectories → 2,200 episodes total.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/datagen-clutter-v1-joint-5cam.ManiGuard-Bench
ManiGuard-Bench
A safety-aware robotic manipulation benchmark. 200 everyday tabletop tasks
across 6 families, each rendered as one in-distribution task plus 4 out-of-distribution
perturbations, in high-fidelity physics simulation. Unlike a plain success benchmark,
every task carries formal safety constraints — a policy is measured not only on
whether it finishes the task, but on whether it stays safe while doing it (never
knocking objects over, dropping them, spilling, or… See the full description on the dataset page: https://huggingface.co/datasets/IDEAS-Lab-Northwestern/ManiGuard-Bench.SWE-bench-Live
A brand-new, continuously updated SWE-bench-like dataset powered by an automated curation pipeline.
For the official data release page, please see microsoft/SWE-bench-Live.
Dataset Summary
SWE-bench-Live is a live benchmark for issue resolving, designed to evaluate an AI system’s ability to complete real-world software engineering tasks. Thanks to our automated dataset curation pipeline, we plan to update SWE-bench-Live on a monthly basis to provide the… See the full description on the dataset page: https://huggingface.co/datasets/Northy1717/SWE-bench-Live.mmu-norm-legacy-north
Legacy Survey DR9 North image cutouts — L1 (release v1)
This L1 repository contains 2,191,927 objects matched across the release, in 726 shards (about 700 GB). Each object has 152×152-pixel cutouts at 0.262″ per pixel in three bands.
Band names: DR9 North imaging uses BASS g/r and MzLS z. This repository uses bass-g, bass-r, and mzls-z; the original des-g/des-r/des-z tokens are kept in native_band_tokens.
Schema
image struct: band (3), flux (3×152×152… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-legacy-north.northern-kurdish-raw-audio
Northern Kurdish Raw Audio Collection
Overview
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The collection was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-supervised Learning (SSL)
Spoken Language Understanding (SLU)
The dataset contains more than 2,000 hours… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-raw-audio.northwind_PurchaseOrders
Purchase Orders Dataset
This dataset consists of purchase orders from various companies. It was created by CHERGUELAINE Ayoub & BOUBEKRI Faycal with the help of ChatGPT for the purpose of document classification and analytics.
Description
The dataset contains a collection of purchase orders from different companies. Each purchase order consists of the following fields:
order_id: The unique identifier for the purchase order.
order_date: The date on which the purchase… See the full description on the dataset page: https://huggingface.co/datasets/boussad/northwind_PurchaseOrders.northwind_purchase_requisitionsukmt_senior_2024north-carolina-layoffs-warn-act-notices-daily
North Carolina WARN Act layoff notices — every filing we hold since 2014, one CSV, rebuilt daily
1,086 North Carolina WARN notices — every one this dataset holds, back to 2014 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-16
· state source last checked 2026-09-24T14:06Z · official source: North Carolina Department of Commerce — WARN notices.
North Carolina employers must file a WARN Act notice with the state before a… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/north-carolina-layoffs-warn-act-notices-daily.ENADE_Brazilian_national_university_examination_MCQ_483northern-kurdish-pseudolabel
Northern Kurdish Raw Audio Collection
Dataset Summary
This repository contains a large collection of raw Northern Kurdish (Kurmanji Kurdish) speech recordings gathered from publicly available Kurdish media sources.
The corpus was assembled to support research and development in:
Automatic Speech Recognition (ASR)
Speech Translation (ST)
Text-to-Speech (TTS)
Self-Supervised Learning (SSL)
Spoken Language Understanding (SLU)
Low-Resource Speech Processing
The… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/northern-kurdish-pseudolabel.ipfs_northmacedonia_laws
North Macedonia Official Gazette (Sluzben vesnik)
Research snapshot of official national legislation from Sluzben vesnik (slvesnik.com.mk) — live FreeIssue API + Issues PDFs; Wayback of official URLs.
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-26
Coverage
partial catalog (resume 2026-09-25 batches I/J ok=250+250/+500; tip∪local @996ecbbe)
Source
Sluzben vesnik… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_northmacedonia_laws.magicdata-dialect-northeastern-chinese-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/magicdata-dialect-northeastern-chinese-tts-lite.magicdata-dialect-northeastern-chinese-tts-lite
MagicData-Dialect-Northeastern Chinese-TTS-Lite
MAGIC DATA OPEN-SOURCE LICENSE
Dataset Overview
Item
Information
Dataset Type
N/A
Language
Chinese Dialect
Speech Style
Scripted
Content
N/A
Audio Parameters
48 kHz, 16 bits
File Format
WAV (PCM)
Recording Equipment
microphone
Recording Environment
quiet indoor environment
License
Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/magicdata-dialect-northeastern-chinese-tts-lite.sciscinet-v2
📢🚨📣 Sciscinet-v2
Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support research in the science of science domain. It combines scientific publications with their network of relationships to funding sources, patents, citations, and institutional affiliations, creating a rich ecosystem for analyzing scientific productivity, impact, and innovation. Know more.
About Sciscinet-v2
The newer version Sciscinet-v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v2.UFAL_Parallel_Corpus_of_North_Levantine_1.0
[!NOTE]
Dataset origin: https://zenodo.org/records/4012218
UFAL Parallel Corpus of North Levantine 1.0
March 10, 2023
Authors
Shadi Saleh <saleh@ufal.mff.cuni.cz>
Hashem Sellat <sellat@ufal.mff.cuni.cz>
Mateusz Krubiński <krubinski@ufal.mff.cuni.cz>
Adam Posppíšil <adam.pospisil@ff.cuni.cz>
Petr Zemánek <petr.zemanek@ff.cuni.cz>
Pavel Pecina <pecina@ufal.mff.cuni.cz>
Overview
This is the first release of the UFAL Parallel Corpus of North Levantine… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/UFAL_Parallel_Corpus_of_North_Levantine_1.0.northstar-datads424-geologic-map-north-america
DS-424 Geologic Map of North America — GeoParquet Edition
A standardized, AI/API-ready GeoParquet conversion of the USGS Database of the Geologic Map of North America (DS-424, 2009). This dataset modernizes the original GIS distribution into a format built for programmatic access — no geological reinterpretation, only representation, validation, and metadata enhancement on top of the authoritative USGS/GSA source.
Quick Links
📦 Dataset on… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/ds424-geologic-map-north-america.
