datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mixed-address-parsing
mixed-address-parsing "📫"
Overview
The mixed-address-parsing dataset is designed to simulate the challenges encountered when processing real-world address inputs. It contains paired examples of noisy address strings (simulating user input) and their corresponding, clean, structured JSON responses. The dataset was generated by extracting components from open geocoding data and deliberately injecting multiple types of noise to mimic common human errors and input… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/mixed-address-parsing.Crypto-Address-Annotation-10K
Codatta Crypto Address Annotations (Sample)
Overview
This dataset is a 10,000-row sample of the comprehensive Codatta Crypto Address Annotations database. The full database serves as a massive repository of over 500 million labeled address pairs across multiple blockchains.
The data provides critical metadata aimed at solving the problem of fragmented and siloed blockchain information. It includes entity names, functional categories (e.g., Exchanges, DeFi, Scam)… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/Crypto-Address-Annotation-10K.myanmar_village_based_72k_addresses
Myanmar Bilingual Village Address Directory (72k+ Entries)
Dataset Description
This dataset provides a comprehensive, bilingual (English and Myanmar) list of over 72,000 rural village addresses across Myanmar. Each entry is a clean, human-readable address string containing the village, village tract, township, district, and state/region.
The dataset is ideal for a wide range of NLP and data science tasks, including Named Entity Recognition (NER), address parsing, machine… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_village_based_72k_addresses.CENSUS-NER-Name-Email-Address-PhoneDataset Summary
The CENSUS-NER-Name-Email-Address-Phone dataset is a processed and structured version of the FMCSA (Federal Motor Carrier Safety Administration) CENSUS1 2016Sep dataset. It is designed to assist in training language models for tasks such as Named Entity Recognition (NER), address parsing, and information extraction from unstructured text. The dataset contains records that include information such as name, email, phone number, and address, extracted from the original dataset and… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/CENSUS-NER-Name-Email-Address-Phone.legal-service-method-address-proof-coherence-risk-v0.1What this dataset does
You receive
rule context
required method
address for service
actual steps
proof
timing
You decide
coherent
or
incoherent
Daily use
service QC
re-serve trigger
timing breach flag
extension trigger
CENSUS-NER-Name-Email-Address-PhoneDataset Summary
The CENSUS-NER-Name-Email-Address-Phone dataset is a processed and structured version of the FMCSA (Federal Motor Carrier Safety Administration) CENSUS1 2016Sep dataset. It is designed to assist in training language models for tasks such as Named Entity Recognition (NER), address parsing, and information extraction from unstructured text. The dataset contains records that include information such as name, email, phone number, and address, extracted from the original dataset and… See the full description on the dataset page: https://huggingface.co/datasets/Usmannn001/CENSUS-NER-Name-Email-Address-Phone.address-neraddressaddress-fixed-4vi-address-correctionmelissa-orange-county-addresseslegal-service-proof-address-method-deadline-coherence-risk-v0.1What this dataset does
You receive
document summary
rules summary
deadline summary
method and address summary
service date
deemed service calculation summary
proof summary
defect flags
You decide
coherent
or
incoherent
Daily use
service QC before filing certificate
avoid strike out
avoid adjournment
avoid wasted costs
address_std_11Address-Classifyeraddress_normalizeraddress_parsingtrain_addressurdu_Roman_Urdu_Addresses_Dataset
