CoolFace
Datasetpublic

timashan/amazon-scrape-4-llm

Amazon Scrape 4 llm Purpose? Feed an LLM raw html to identify products from an ecommerce platform.These datasets contain the extracted innerTexts of all HTML nodes from different ecommerce product pages.The cleaning process significantly reduces the token size from ex: 450k -> 6k Quickstart from datasets import load_dataset data_train = load_dataset("timashan/amazon-scrape-4-llm", "phones") data_test = load_dataset("timashan/amazon-scrape-4-llm"… See the full description on the dataset page: https://huggingface.co/datasets/timashan/amazon-scrape-4-llm.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes26downloads

timashan/amazon-scrape-4-llm · main · files are served by the source, never re-hosted here