CoolFace
Datasetpublic

nwchang/sea-product-attribute-extraction-sample

SEA Multilingual Product Attribute Extraction Sample This public sample contains 1,000 synthetic, AI-generated marketplace-style records for product attribute extraction and catalog normalization. Languages English Chinese Malay Indonesian Formats CSV JSONL Intended Use Use this sample for inspection, evaluation, catalog normalization prototypes, search-filtering experiments, and multilingual data-quality testing.… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-product-attribute-extraction-sample.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes18downloads
Dataset Card

SEA Multilingual Product Attribute Extraction Sample

This public sample contains 1,000 synthetic, AI-generated marketplace-style records for product attribute extraction and catalog normalization.

Languages

  • —English
  • —Chinese
  • —Malay
  • —Indonesian

Formats

  • —CSV
  • —JSONL

Intended Use

Use this sample for inspection, evaluation, catalog normalization prototypes, search-filtering experiments, and multilingual data-quality testing.

Important Limitations

This is synthetic data. It is not verified marketplace data or verified product specification data. Validate downstream outputs before using them in a real catalog.

Full Dataset

The paid release contains 10,000 records:

View the full Product Attribute Extraction Dataset on Gumroad

Launch offer: use code LAUNCH20 for 20% off. Limited to the first 20 uses across SEA Data Lab products.

License

See SAMPLE_LICENSE.txt.