datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
osm-polygon-website-tag
OSM Polygon Website Dataset
OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files.
At a glance
Polygons
1,726,474
With extracted text
1,192,980
Words of text
407,685,655
Languages
397
Regional sources
386 / 386
Duplicate objects removed
104,927
Candidates rejected
868,905,743
Status
In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.HTML-CSS-Website# Dataset
This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant.
# Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com)
website_screenshots_image_dataset
Website Screenshots Image Dataset
This dataset is obtainable here from roboflow..
Dataset Details
Dataset Description
Language(s) (NLP): [English]
License: [MIT]
Dataset Sources
Source: [https://universe.roboflow.com/roboflow-gw7yv/website-screenshots/dataset/1]
Uses
From the roboflow website:
Annotated screenshots are very useful in Robotic Process Automation. But they can be expensive to label. This dataset would cost over… See the full description on the dataset page: https://huggingface.co/datasets/Zexanima/website_screenshots_image_dataset.website-screenshots-blip-largewebsite-html-2kThis dataset was a subset code tasks 33k filtered exclusively for high-quality websites.
The filtering includes:
Minimum of 100 lines
Includes CSS or Javascript content
The filtering was fairly light, as for stricter cleaning I only had 1,100 examples, which I considered to be low. Having 900 more examples is worth it in my opinion.
I have also included a python file which grabs the HTML content from random responses and saves those as files 1.html ==> 5.html to view them.
osm-polygon-website-tag-eunis
OSM Polygon Website Dataset
OpenStreetMap closed ways and polygon relations carrying a non-empty website OR contact:website tag, with full main-page text extracted using Trafilatura. Every statistic below is regenerated from the current upload-acknowledged Parquet artifacts.
Snapshot
Metric
Value
What it means
Snapshot status
In progress
Current published snapshot
Regional PBFs included
386 / 386
Published source shards / expected source PBFs… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag-eunis.jusci-website
Dataset Card for "jusci-website"
This dataset was scrape jusci.net. JuSci is science website news and is part of Blognone Content Network. It use CC-BY 3.0 license.
goethe-website
Dataset Card for "goethe-website"
This dataset collect the article from https://www.goethe.de/ins/th. (CC BY-SA only)
CIRCL_website_subset
Dataset Card for Dataset Name
Dataset Summary
task_categories:
- image-classification
pretty_name: Subset of circl-ail-dataset-01
size_categories:
- 1K<n<10K
This is a subset of circl-ail-dataset-01 dataset with these labels ["marketplace","forum","general"] each label has 1000 images
circl-ail-dataset-01
This dataset is named circl-ail-dataset-01 and is composed of AIL’s scraped onion websites. Around 37500 pictures are in this dataset to date.
Only one… See the full description on the dataset page: https://huggingface.co/datasets/Abhilashvj/CIRCL_website_subset.inquiringmind-website
Dataset Card for "inquiringmind-website"
This dataset collect web page from Inquiring Mind (ครูไทย หัวใจสืบเสาะ). Ihe license is cc-by-sa-3.0.
merged-coder-website-reasoningsearch-results-websites
Dataset Card for General
This instruction tuning dataset has been created with SPInO and has 2720 rows of textual data related to General.
from datasets import load_dataset
dataset = load_dataset("1rsh/search-results-websites")
print(dataset)
website_categorieswebsite-screenshots-git-largeUSCIS-knowledge-base-full-website
A comprehensive dataset of 99,489 content chunks from 4,666 pages on the USCIS website, with pre-computed OpenAI text-embedding-ada-002 embeddings (1536 dimensions).
Built for RAG (Retrieval-Augmented Generation), semantic search, and GraphRAG applications focused on U.S. immigration law and policy.
🔗 GitHub: github.com/0xrphl/USCIS-knowledge-base-full-website🔥 Scraped with: Firecrawl — The open-source web scraping API for AI🍎 Visualized with: Embedding Atlas — Interactive embedding… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/USCIS-knowledge-base-full-website.website-screenshotsml6-website-ragwebsite-title-descriptionThis dataset is designed for training small models. It primarily consists of webpages from The New York Times and GitHub. Key information is extracted from the HTML and converted into text parameters, which are then summarized into 1 to 4 words using Claude 3.5 by Anthropic.
Website_Segmentation
Dataset Card for "Website_Segmentation"
More Information needed
short_synthetic_website
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
fw-darija-websitesimnet1k_web_site_website_internet_site_siteModern_Website_Code_GeneratorDental_website_scrapingvi-medical-website-dedupamenity-website-imagesmedical-website-rawwebsite_qa_qwen_generatedwebsite-code-datasetsynthetic_website
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
