CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NoeFlandre /osm-polygon-website-tag OSM Polygon Website Dataset OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files. At a glance Polygons 1,726,474 With extracted text 1,192,980 Words of text 407,685,655 Languages 397 Regional sources 386 / 386 Duplicate objects removed 104,927 Candidates rejected 868,905,743 Status In progress… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.tabular1M<n<10M2 likes2.4k downloads3d agoHugging Face02MAsad789565 /HTML-CSS-Website# Dataset This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant. # Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com) textn<1K4 likes499 downloads3y agoHugging Face03Zexanima /website_screenshots_image_dataset Website Screenshots Image Dataset This dataset is obtainable here from roboflow.. Dataset Details Dataset Description Language(s) (NLP): [English] License: [MIT] Dataset Sources Source: [https://universe.roboflow.com/roboflow-gw7yv/website-screenshots/dataset/1] Uses From the roboflow website: Annotated screenshots are very useful in Robotic Process Automation. But they can be expensive to label. This dataset would cost over… See the full description on the dataset page: https://huggingface.co/datasets/Zexanima/website_screenshots_image_dataset.imageobject-detection1K<n<10K22 likes131 downloads3y agoHugging Face04naorm /website-screenshots-blip-largeimage1K<n<10K1 likes113 downloads3y agoHugging Face05Sweaterdog /website-html-2kThis dataset was a subset code tasks 33k filtered exclusively for high-quality websites. The filtering includes: Minimum of 100 lines Includes CSS or Javascript content The filtering was fairly light, as for stricter cleaning I only had 1,100 examples, which I considered to be low. Having 900 more examples is worth it in my opinion. I have also included a python file which grabs the HTML content from random responses and saves those as files 1.html ==> 5.html to view them. text1K<n<10K0 likes73 downloads8mo agoHugging Face06NoeFlandre /osm-polygon-website-tag-eunis OSM Polygon Website Dataset OpenStreetMap closed ways and polygon relations carrying a non-empty website OR contact:website tag, with full main-page text extracted using Trafilatura. Every statistic below is regenerated from the current upload-acknowledged Parquet artifacts. Snapshot Metric Value What it means Snapshot status In progress Current published snapshot Regional PBFs included 386 / 386 Published source shards / expected source PBFs… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag-eunis.tabular1M<n<10M0 likes55 downloads6d agoHugging Face07pythainlp /jusci-website Dataset Card for "jusci-website" This dataset was scrape jusci.net. JuSci is science website news and is part of Blognone Content Network. It use CC-BY 3.0 license. texttext-generation1K<n<10K4 likes54 downloads2y agoHugging Face08pythainlp /goethe-website Dataset Card for "goethe-website" This dataset collect the article from https://www.goethe.de/ins/th. (CC BY-SA only) texttext-generationn<1K0 likes48 downloads3y agoHugging Face09Abhilashvj /CIRCL_website_subset Dataset Card for Dataset Name Dataset Summary task_categories: - image-classification pretty_name: Subset of circl-ail-dataset-01 size_categories: - 1K<n<10K This is a subset of circl-ail-dataset-01 dataset with these labels ["marketplace","forum","general"] each label has 1000 images circl-ail-dataset-01 This dataset is named circl-ail-dataset-01 and is composed of AIL’s scraped onion websites. Around 37500 pictures are in this dataset to date. Only one… See the full description on the dataset page: https://huggingface.co/datasets/Abhilashvj/CIRCL_website_subset.image1K<n<10K1 likes43 downloads3y agoHugging Face10pythainlp /inquiringmind-website Dataset Card for "inquiringmind-website" This dataset collect web page from Inquiring Mind (ครูไทย หัวใจสืบเสาะ). Ihe license is cc-by-sa-3.0. texttext-generationn<1K1 likes40 downloads2y agoHugging Face11usernamebetter /merged-coder-website-reasoningtext10K<n<100K0 likes39 downloads2mo agoHugging Face121rsh /search-results-websites Dataset Card for General This instruction tuning dataset has been created with SPInO and has 2720 rows of textual data related to General. from datasets import load_dataset dataset = load_dataset("1rsh/search-results-websites") print(dataset) text1K<n<10K0 likes37 downloads2y agoHugging Face13massimilianowosz /website_categoriestabular10K<n<100K3 likes36 downloads3y agoHugging Face14naorm /website-screenshots-git-largeimage1K<n<10K5 likes30 downloads3y agoHugging Face150xrphl /USCIS-knowledge-base-full-website A comprehensive dataset of 99,489 content chunks from 4,666 pages on the USCIS website, with pre-computed OpenAI text-embedding-ada-002 embeddings (1536 dimensions). Built for RAG (Retrieval-Augmented Generation), semantic search, and GraphRAG applications focused on U.S. immigration law and policy. 🔗 GitHub: github.com/0xrphl/USCIS-knowledge-base-full-website🔥 Scraped with: Firecrawl — The open-source web scraping API for AI🍎 Visualized with: Embedding Atlas — Interactive embedding… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/USCIS-knowledge-base-full-website.tabulartext-retrieval100K<n<1M0 likes29 downloads3mo agoHugging Face16naorm /website-screenshotsimage1K<n<10K2 likes28 downloads3y agoHugging Face17nielsr /ml6-website-ragtext1K<n<10K0 likes27 downloads3y agoHugging Face18wgcv /website-title-descriptionThis dataset is designed for training small models. It primarily consists of webpages from The New York Times and GitHub. Key information is extracted from the HTML and converted into text parameters, which are then summarized into 1 to 4 words using Claude 3.5 by Anthropic. textsummarization1K<n<10K2 likes25 downloads2y agoHugging Face19miss-swan /Website_Segmentation Dataset Card for "Website_Segmentation" More Information needed imagen<1K1 likes24 downloads3y agoHugging Face20justinsunqiu /short_synthetic_website Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. textn<1K0 likes22 downloads2y agoHugging Face21sawalni-ai /fw-darija-websitestabular1K<n<10K2 likes21 downloads2y agoHugging Face22mlnomad /imnet1k_web_site_website_internet_site_siteimage1K<n<10K0 likes20 downloads1y agoHugging Face23VortexHunter23 /Modern_Website_Code_Generatortextn<1K0 likes19 downloads9mo agoHugging Face24Utshav /Dental_website_scrapingtextn<1K0 likes17 downloads2y agoHugging Face25codin-research /vi-medical-website-dedupgatedtext10K<n<100K0 likes17 downloads1y agoHugging Face26raphael0202 /amenity-website-imagesimage1K<n<10K0 likes17 downloads3mo agoHugging Face27codin-research /medical-website-rawgatedtext10K<n<100K0 likes16 downloads1y agoHugging Face28johngraph /website_qa_qwen_generatedtext100K<n<1M0 likes14 downloads9mo agoHugging Face29InstateLabs /website-code-datasettextn<1K0 likes14 downloads5mo agoHugging Face30justinsunqiu /synthetic_website Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. textn<1K0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.