CoolFace
Datasetpublic

timashan/amazon-scrape-4-llm

Amazon Scrape 4 llm Purpose? Feed an LLM raw html to identify products from an ecommerce platform.These datasets contain the extracted innerTexts of all HTML nodes from different ecommerce product pages.The cleaning process significantly reduces the token size from ex: 450k -> 6k Quickstart from datasets import load_dataset data_train = load_dataset("timashan/amazon-scrape-4-llm", "phones") data_test = load_dataset("timashan/amazon-scrape-4-llm"… See the full description on the dataset page: https://huggingface.co/datasets/timashan/amazon-scrape-4-llm.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes26downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
timashan/amazon-scrape-4-llm · CoolFace