nwchang/sea-product-attribute-extraction-sample
SEA Multilingual Product Attribute Extraction Sample This public sample contains 1,000 synthetic, AI-generated marketplace-style records for product attribute extraction and catalog normalization. Languages English Chinese Malay Indonesian Formats CSV JSONL Intended Use Use this sample for inspection, evaluation, catalog normalization prototypes, search-filtering experiments, and multilingual data-quality testing.… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-product-attribute-extraction-sample.
SEA Multilingual Product Attribute Extraction Sample
This public sample contains 1,000 synthetic, AI-generated marketplace-style records for product attribute extraction and catalog normalization.
Languages
- English
- Chinese
- Malay
- Indonesian
Formats
- CSV
- JSONL
Intended Use
Use this sample for inspection, evaluation, catalog normalization prototypes, search-filtering experiments, and multilingual data-quality testing.
Important Limitations
This is synthetic data. It is not verified marketplace data or verified product specification data. Validate downstream outputs before using them in a real catalog.
Full Dataset
The paid release contains 10,000 records:
View the full Product Attribute Extraction Dataset on Gumroad
Launch offer: use code LAUNCH20 for 20% off. Limited to the first 20 uses across SEA Data Lab products.
License
See SAMPLE_LICENSE.txt.
