website
Jahid05_-_llama-3.2-3b-6epoch-website-prompt-ggufJahid05_-_Gemma-2-2b-website-prompt-generation-v2-ggufqwen-3-panda-agi-websites-2-i1-GGUFJahid05_-_Gemma-2-2b-original-website-prompt-generator-ggufJahid05_-_llama-3.2-1b-website-prompt-generator-ggufJahid05_-_Gemma-2-2b-website-prompt-generator-10epoch-ggufJahid05_-_llama-3.2-3b-8epoch-lora-r8-website-prompt-generator-ggufJahid05_-_llama-3.2-3b-r-16-alpha-64-website-prompt-gguf
Datasets
All datasets matching “website”Web_site_legacycrawl-my-websiteosm-polygon-website-tag
OSM Polygon Website Dataset
OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files.
At a glance
Polygons
1,726,474
With extracted text
1,192,980
Words of text
407,685,655
Languages
397
Regional sources
386 / 386
Duplicate objects removed
104,927
Candidates rejected
868,905,743
Status
In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.causvid_website1k_Website_Screenshots_and_Metadata
Dataset Card for 1000 Website Screenshots with Metadata
Dataset Summary
Silatus is sharing, for free, a segment of a dataset that we are using to train a generative AI model for text-to-mockup conversions. This dataset was collected in December 2022 and early January 2023, so it contains very recent data from 1,000 of the world's most popular websites. You can get our larger 10,000 website dataset for free at: https://silatus.com/datasets
This dataset includes:
High-res… See the full description on the dataset page: https://huggingface.co/datasets/silatus/1k_Website_Screenshots_and_Metadata.mind2web_multimodal_test_website
Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split
Note: This dataset is the test split of the Cross-Website dataset introduced in the paper.
This is a FiftyOne dataset with 1019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.
