CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Intelligent-Internet /II-Medical-Reasoning-SFT II-Medical-Reasoning-SFT II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice. The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.text1M<n<10M57 likes3.1k downloads1y agoHugging Face02Intelligent-Internet /pd12m PD12M This is a curated PD12M dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Public Domain 12M image collection, refined by filtering for active image links. EXIF data was extracted, and images underwent preprocessing and feature extraction using SigLIP 2. All vector embeddings are normalized 16-bit half-precision vectors optimized for L2 indexing with vectorchord.… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/pd12m.imagefeature-extraction10M<n<100M8 likes2.4k downloads1y agoHugging Face03sailor2 /sea-internettext10M<n<100M1 likes1.5k downloads2y agoHugging Face04jupiternull /dead-internet-observatorytext1M<n<10M0 likes1.4k downloads2d agoHugging Face05Intelligent-Internet /wikipedia_en wikipedia_en This is a curated Wikipedia English dataset for use with the II-Commons project. Dataset Details Dataset Description This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.tabularfeature-extraction10M<n<100M2 likes1.3k downloads1y agoHugging Face06Intelligent-Internet /ii-agent_gaia-benchmark_validationtextn<1K8 likes926 downloads1y agoHugging Face07Intelligent-Internet /GAIA-Subset-Benchmark GAIA Benchmark Subset Model Card This dataset is a subset of the GAIA benchmark, containing 44 web-search-based questions from the validation set. It evaluates multiple AI models on their ability to retrieve and process real-time information using web search and browser tools. Performance metrics include success indicators and detailed reports for each model. A comparative chart summarizing the results will be provided separately. Benchmark Results textn<1K3 likes893 downloads1y agoHugging Face08Intelligent-Internet /II-Thought-RL-v0 II-Thought RL v0: A Large-Scale Curated Dataset for Reinforcement Learning See our blog here for additional details. We introduce II-Thought RL v0, the first large-scale, multi-task dataset designed for Reinforcement Learning. This dataset consists of high-quality question-answer pairs that have undergone a rigorous multi-step filtering process, leveraging Gemini 2.0 Flash and Qwen 32B as quality evaluators. In this initial release, we have curated and refined publicly available… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Thought-RL-v0.text100K<n<1M54 likes381 downloads1y agoHugging Face09meettilavat /InternetArchive_1899_Large Internet Archive Historical Texts (0001-1899) TL;DR 711,680 cleaned public-domain style documents harvested from the Internet Archive via a high-throughput text-to-parquet pipeline. Coverage targets items that contain textual content dated between 0001 and 1899, ranked by download counts; ~715k IDs were attempted, ~4.1k were filtered during preprocessing. Stored in 620 Zstandard-compressed Parquet shards (shard_00000.parquet ... shard_00619.parquet) occupying ~240 GB on… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Large.texttext-generation100K<n<1M0 likes243 downloads11mo agoHugging Face10meettilavat /InternetArchive_1899_Chunked Internet Archive Historical Texts - Chunked (0001-1899) TL;DR 163 million text chunks extracted from historical public-domain documents sourced from the Internet Archive Content dated 0001-1899, sorted by download popularity to prioritize high-quality, frequently accessed materials 2,445 Zstandard-compressed Parquet shards totaling ~217 GB on disk, ~594 billion characters uncompressed Optimized chunk size of ~3,600 characters (target: 4,000) for efficient language model… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Chunked.texttext-generation100M<n<1B0 likes168 downloads11mo agoHugging Face11letrinhan /vn-provinces-household-internet-rate Vietnam household internet connection rate Vietnam household internet connection rate. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (63 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions (6 rows) data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-household-internet-rate.tabularn<1K0 likes93 downloads2d agoHugging Face12Intelligent-Internet /ChatDoctor-RL Intelligent-Internet/ChatDoctor-Improved-Answer Dataset This dataset represents a carefully curated subset derived from the original ChatDoctor-HealthCareMagic-100k[lavita/ChatDoctor-HealthCareMagic-100k] dataset, where we have undertaken significant improvements to enhance the quality and depth of the responses. The answers have been thoroughly refined to provide greater detail, clarity, and precision, while incorporating a heightened focus on safety awareness to ensure responsible… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/ChatDoctor-RL.text10K<n<100K15 likes91 downloads1y agoHugging Face13NikoThePig /internet-prompts-benchmark Internet Prompts Benchmark Viral internet prompts, memes, and tests that AI historically failed at. Popular ones like counting letters in the word strawberry and nicher ones that test other important capabilities.Mostly made by GPT-5.6 Sol. It is designed for many types of models to participate, small and large, not only transformers. It has prompts from the early days of AI to the very latest. It will be actively updated to preserve various prompts for as long as I can afford… See the full description on the dataset page: https://huggingface.co/datasets/NikoThePig/internet-prompts-benchmark.texttext-generationn<1K0 likes91 downloads13d agoHugging Face14letrinhan /vn-provinces-internet-users-share Vietnam internet users share Vietnam internet users share. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (63 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions (6 rows) data/regions.csv data/regions.dta data/regions.xlsx… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-internet-users-share.tabularn<1K0 likes90 downloads2d agoHugging Face15Intelligent-Internet /OpenAI-HealthBench-II-Medical-8B-GPT-4.1text1K<n<10K1 likes82 downloads1y agoHugging Face16Intelligent-Internet /II-Medical-RL Overview The MedReason-RL dataset is a refined version of the original MedReason dataset, specifically curated for training reinforcement learning (RL) models to enhance reasoning abilities. It has been proven to be the best dataset for improving model reasoning through RL training. Source This dataset is derived from the original MedReason dataset, which focuses on medical reasoning tasks. However, the original dataset contained significant overlap with benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-RL.text10K<n<100K11 likes67 downloads1y agoHugging Face17burpheart /Internet-background-noise Internet Background Noise Dataset (Unlabeled Raw Data) This dataset contains HTTP internet noise data collected by an internet honeypot. It consists of raw, unlabeled network packets, including metadata, payloads, and header information. This data is suitable for training and evaluating machine learning models for network intrusion detection, cybersecurity, and traffic analysis. HoneyPot repository: hachimi on GitHub. Dataset Overview The Internet Background Noise… See the full description on the dataset page: https://huggingface.co/datasets/burpheart/Internet-background-noise.tabulartext-classification1M<n<10M6 likes57 downloads2y agoHugging Face18Intelligent-Internet /II-Thought-RL-v0-Math-50Ktext10K<n<100K3 likes54 downloads1y agoHugging Face19Intelligent-Internet /II-Search-CIR-SFTtext10K<n<100K5 likes54 downloads1y agoHugging Face20electricsheepafrica /africa-worldbank-individuals-using-the-internet-of-population-it-net-user-zs Individuals using the Internet (% of population) | Africa (World Bank — Gender Statistics) | Africa (World Bank) Size category: 1K<n<10K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-individuals-using-the-internet-of-population-it-net-user-zs.tabulartabular-classification1K<n<10K0 likes53 downloads1mo agoHugging Face21Intelligent-Internet /Vietnamese-Entrance-Exam Vietnamese Entrance Exam Dataset The Vietnamese Entrance Exam dataset is a collection of 432 problems derived from Vietnamese University entrance examinations. The dataset aims to provide a novel benchmark for testing reasoning capabilities of language models in several low resource domains specifically designed to minimize potential data contamination from pre-training or post-training exposure. Domain Count Physics 95 Chemistry 94 Math 243 Data… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/Vietnamese-Entrance-Exam.imagen<1K1 likes44 downloads1y agoHugging Face22Intelligent-Internet /OpenAI-HealthBench-II-Medical-8B-1706-GPT-4.1text1K<n<10K2 likes44 downloads1y agoHugging Face23isaquecerqueira /millan_internet_traffic Milan Internet Traffic Dataset This dataset contains information about hourly internet traffic in Milan between 2013-11-01 and 2014-01-01. text1K<n<10K0 likes42 downloads3y agoHugging Face24electricsheepafrica /africa-synth-telecom-massive-internet-of-things-traffic-nigeria Africa Synth Telecom Massive Internet of Things Traffic Nigeria | Africa (Electric Sheep Africa metadata inventory) Size category: 100K<n<1M - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-telecom-massive-internet-of-things-traffic-nigeria.tabulartabular-classification100K<n<1M1 likes36 downloads1mo agoHugging Face25electricsheepafrica /africa-owid-number-of-internet-users Number Of Internet Users | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-number-of-internet-users.tabulartabular-classification1K<n<10K0 likes29 downloads1mo agoHugging Face26Apixhed /Nonsense-Internet-Niches Dataset Card for Dataset Name [this dataset is basically for internet-cultured LLMs or making a LLM more 21st century humane speech] cool dataset that has SynthV, Vocaloids, Blender Niches, GTA, literature, Art, ill update it later to have more things soon!1!! only have like 50s example because only 1 person operating this, the needed components are in file and versions, but you can add others and not mine if you want In ListsForThings.txt, yes you can actually use it on serious… See the full description on the dataset page: https://huggingface.co/datasets/Apixhed/Nonsense-Internet-Niches.textquestion-answeringn<1K0 likes26 downloads1y agoHugging Face27electricsheepasia /asia-owid-landline-internet-subscriptions Landline Internet Subscriptions | Asia (Our World in Data) 🌏 990 observations · 47 Asia countries · 1998–2023 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 990 observations of Landline Internet Subscriptions data across 47 Asia countries, spanning 1998–2023. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Landline Internet Subscriptions Geographic coverage 47… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-landline-internet-subscriptions.tabulartabular-classificationn<1K0 likes25 downloads4mo agoHugging Face28electricsheepafrica /africa-cote-d-ivoire-marche-de-l-internet-et-de-la-telephonie-mobile-de-2010-a-fec81c58 Marche De L Internet Et De La Telephonie Mobile De 2010 a | Africa (Cote d'Ivoire DataFair) 263 rows - 1 Africa country/area - 2010-2016 - 1 indicator - Engineered by Electric Sheep Africa TL;DR This dataset contains 263 rows from Cote d'Ivoire DataFair, covering Marche De L Internet Et De La Telephonie Mobile De 2010 a. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cote-d-ivoire-marche-de-l-internet-et-de-la-telephonie-mobile-de-2010-a-fec81c58.tabulartabular-regressionn<1K0 likes25 downloads1mo agoHugging Face29internetov1488 /internetovai-identitytextn<1K0 likes25 downloads7h agoHugging Face30Intelligent-Internet /II-Search-Benchmark-Details Inspect-Search-Models-Benchmarking-Result Overall result Qwen 4B Jan 4B WebSailor-3B II-Search-4B II-Search-CIR-4B OpenAI/SimpleQA 76.8 80.1 81.8 91.8 91.8 Google/Frames 30.7 24.8 34.0 67.5 72.2 Seal_0 6.31 2.7 1.8 22.5 26.4 Simple QA (SerpDev) Qwen 4B Jan 4B WebSailor-3B II-Search-4B II-Search-CIR-4B Pass rate % 76.8 80.1 81.8 91.8 91.8 # Search 1.0 0.9 2.1 2.2 2.5 # Visit 0.1 1.9 6.4 3.5 5.3 # Tool used 1.1 2.8 8.5 5.7 7.8 Frames… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Search-Benchmark-Details.tabular10K<n<100K2 likes24 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.