datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omega-overture-addressesmixed-address-parsing
mixed-address-parsing "📫"
Overview
The mixed-address-parsing dataset is designed to simulate the challenges encountered when processing real-world address inputs. It contains paired examples of noisy address strings (simulating user input) and their corresponding, clean, structured JSON responses. The dataset was generated by extracting components from open geocoding data and deliberately injecting multiple types of noise to mimic common human errors and input… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/mixed-address-parsing.us-addresses-synthetic-v3-llmaddress_standardizationCrypto-Address-Annotation-10K
Codatta Crypto Address Annotations (Sample)
Overview
This dataset is a 10,000-row sample of the comprehensive Codatta Crypto Address Annotations database. The full database serves as a massive repository of over 500 million labeled address pairs across multiple blockchains.
The data provides critical metadata aimed at solving the problem of fragmented and siloed blockchain information. It includes entity names, functional categories (e.g., Exchanges, DeFi, Scam)… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Crypto-Address-Annotation-10K.ofac-sdn-crypto-addresses
AI DECISIONS — OFAC SDN crypto addresses (primary-source TagPack)
949 cryptocurrency addresses designated on the US Treasury OFAC Specially Designated Nationals (SDN) list,
unified by the open-source openlabels attribution pipeline and
emitted in GraphSense TagPack format. The same pack was
submitted upstream as graphsense/graphsense-tagpacks#53.
Every row carries the URI of the primary public source it was taken from. No third-party or licence-restricted
attribution sources are… See the full description on the dataset page: https://huggingface.co/datasets/ai-decisions/ofac-sdn-crypto-addresses.tts_russian_addresses_rhvoice_4voicesworldwide-addresses
Dataset Card for worldwide-addresses
Dataset description
Dataset Summary
This dataset is a collection of annotated international addresses containing over 750,000,000 addresses from 240 countries in over 100 languages. It has been created from the data gathered and provided by libpostal, an international street address parsing package. The original purpose of this dataset was to develop a state-of-the-art neural network-based international address parser named… See the full description on the dataset page: https://huggingface.co/datasets/deepparse/worldwide-addresses.deepparse_address_mutations_comb_3
DeepParse Address Mutations (up to 3 mutations)
This dataset, deepparse_address_mutations_comb_3, provides a robust collection of address mutations designed to improve address matching tasks derived from the Deepparse Address Dataset. The dataset was generated using combinations of mutators applied to 100,000 annotated addresses, creating variations that simulate real-world inconsistencies, typos, and formatting differences.
Dataset Details
Mutation… See the full description on the dataset page: https://huggingface.co/datasets/jarredparrett/deepparse_address_mutations_comb_3.address-geography-postal
Veygrit public geography and postal candidates
Version 1.0.0. A partial pilot, not a worldwide address directory,
postal-authority verification, ownership check, geocoder, or trained model.
Derived exclusively from the cited GeoNames CC BY 4.0 public downloads.
Attribution: GeoNames.
Coverage
Country
Administrative areas
Localities
Neighborhoods
Streets
Postal codes
AD
7
7
0
0
7
AE
435
392
1983
0
0
HK
18
1330
6
0
0
MO
8
9
0
0
0
SG
0
136
218
2
0… See the full description on the dataset page: https://huggingface.co/datasets/veygrit/address-geography-postal.state-of-the-union-addressesThis dataset includes all recorded (spoken and written) addresses to the United States Congress from the President of the United States of America.
The addresses span all Presidents, but differ in their modality. From 1801 to 1913 the addresses were provided in writing to Congress with the remainder provided in person in the form of a speech.
Each address was scraped from The American Presidency Project, a website supported the UC Santa Barbara. Their mission statement is to, "be recognized… See the full description on the dataset page: https://huggingface.co/datasets/jsulz/state-of-the-union-addresses.gyeongsan_address_firestation_ko_14000hrindian-addresses-raw
Indian Addresses (Raw)
Raw, unstructured Indian address strings — the unlabeled corpus behind
gagan1985/qwen3-0.6b-indian-address-parser.
For the labeled/parsed version, see
gagan1985/indian-addresses-gold.
Sources
Two raw source types, distinguished by source_type:
mca — Indian Ministry of Corporate Affairs company-registration addresses
(public disclosure data; CIN numbers are part of India's public company registry)
bank — bank/business-correspondent branch… See the full description on the dataset page: https://huggingface.co/datasets/gagan1985/indian-addresses-raw.Bitcoin_Wallet_AddressesLatest Bitcoin All Address Wallet Type , (Sort by Upper Balance To Lower Balance).
Latest All Bitcoin Wallet Addresses (.tsv file / compressed .gz) : Latest_Bitcoin_Addresses_Mmdrza.tsv.gz
2024-09 : Latest Rich Bitcoin Wallet Addresses:
Latest All Rich Bitcoin Wallet Address : Latest_Rich_Bitcoin_Address_Mmdrza.txt.gz
Per Line a Only Address (Sorted with Balance) | Address Type : All | Main File : .txt | Compressed : .gz
Latest Rich Bitcoin P2PKH Address (start with 1) :… See the full description on the dataset page: https://huggingface.co/datasets/pymmdrza/Bitcoin_Wallet_Addresses.DoD-Instruction-8410-DNS-IP-Address-Use-And-Approval
🌐 DoD Internet Domain Name and IP Address Resource Question-Answer Dataset
Source: DoD Instruction 8410.01
Source Effective Date: December 4, 2015
Change Incorporated: Change 1, effective June 4, 2021
Source Organization: Office of the DoD Chief Information Officer
Source Ownership: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Internet Domain Name and IP Address Resource Question-Answer Dataset is a structured… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8410-DNS-IP-Address-Use-And-Approval.crypto-addresses-dbus-addresses-synthetic-v3-llm-pairedtagged_addresses
Dataset Card for "tagged_addresses"
More Information needed
address-benchmark-v1
Indian Address Benchmark Dataset v1
Mixed benchmark dataset for Indian-address tagging built from synthetic data plus public upstream datasets.
Repository
Dataset repo: mukuls9971/address-benchmark-v1
Train split: 26728
Validation split: 6158
Test split: 1410
Files
train.jsonl
validation.jsonl
test.jsonl
report.json
Notes
Generated and published by the pii-model-oss workflow.
Upstream datasets used to assemble benchmark variants retain their own… See the full description on the dataset page: https://huggingface.co/datasets/mukuls9971/address-benchmark-v1.decompilebench-with-addressesmyanmar_village_based_72k_addresses
Myanmar Bilingual Village Address Directory (72k+ Entries)
Dataset Description
This dataset provides a comprehensive, bilingual (English and Myanmar) list of over 72,000 rural village addresses across Myanmar. Each entry is a clean, human-readable address string containing the village, village tract, township, district, and state/region.
The dataset is ideal for a wide range of NLP and data science tasks, including Named Entity Recognition (NER), address parsing, machine… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_village_based_72k_addresses.synth-addresses-ner-mk3us-address-standardization
US Address Standardization (PostGIS stdaddr schema)
Chat-format instruction data that maps a raw US address string to a strict JSON
object matching the PostGIS address_standardizer stdaddr type (camelCase keys,
USPS-abbreviated values). Used to fine-tune the qwen35-address-std model family.
Each row has three flat fields (system / user / assistant) for easy reading
and grepping; rebuild the chat messages list from them at train time:
{
"system": "<standardization… See the full description on the dataset page: https://huggingface.co/datasets/davidr99/us-address-standardization.Korea_addressasia-owid-legal-frameworks-addressing-gender-equality-in-employment-and-economic-benefits
Legal Frameworks Addressing Gender Equality In Employment And Economic Benefits | Asia (Our World in Data)
🌏 91 observations · 31 Asia countries · 2018–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 91 observations of Legal Frameworks Addressing Gender Equality In Employment And Economic Benefits data across 31 Asia countries, spanning 2018–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-legal-frameworks-addressing-gender-equality-in-employment-and-economic-benefits.CENSUS-NER-Name-Email-Address-PhoneDataset Summary
The CENSUS-NER-Name-Email-Address-Phone dataset is a processed and structured version of the FMCSA (Federal Motor Carrier Safety Administration) CENSUS1 2016Sep dataset. It is designed to assist in training language models for tasks such as Named Entity Recognition (NER), address parsing, and information extraction from unstructured text. The dataset contains records that include information such as name, email, phone number, and address, extracted from the original dataset and… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/CENSUS-NER-Name-Email-Address-Phone.legal-service-method-address-proof-coherence-risk-v0.1What this dataset does
You receive
rule context
required method
address for service
actual steps
proof
timing
You decide
coherent
or
incoherent
Daily use
service QC
re-serve trigger
timing breach flag
extension trigger
short_address_instructionsner-address-standard-datasetThis dataset contains fully auto‑annotated Vietnamese administrative addresses synthesized from the national unit hierarchy (old 3-level wards/districts/provinces and the new 2-level schema).
Each sample is a tokenized address string accompanied by BIO tags for four entity types: STREET, WARD, DISTRICT, and PROVINCE. The generator blends authentic administrative names (including all official aliases and legacy→modern mappings) with realistic street templates, connectors, abbreviations… See the full description on the dataset page: https://huggingface.co/datasets/dathuynh1108/ner-address-standard-dataset.address_dataset
Dataset Card for "address_dataset"
More Information needed
