internet
Datasets
All datasets matching “internet”webfiddle-internet-raw-cache-datasetA dataset of different files that robots tried to crawl through webfiddle.net
Mostly html files but other files too pdfs, images, binary- i have no idea what is in here at this stage - but gives an interesting idea of what crawlers like to visit and could be the basis of interesting SEO or coding LLM reasearch.
Collected as part of my work on web simulators.
https://webfiddle.net JS/CSS editor for the web, https://websim.netwrck.com Coding Editor for the web.
https://x.com/leeleepenkman
Its… See the full description on the dataset page: https://huggingface.co/datasets/lee101/webfiddle-internet-raw-cache-dataset.II-Medical-Reasoning-SFT
II-Medical-Reasoning-SFT
II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice.
The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.pd12m
PD12M
This is a curated PD12M dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Public Domain 12M image collection, refined by filtering for active image links. EXIF data was extracted, and images underwent preprocessing and feature extraction using SigLIP 2. All vector embeddings are normalized 16-bit half-precision vectors optimized for L2 indexing with vectorchord.… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/pd12m.dead-internet-observatorysea-internetwikipedia_en
wikipedia_en
This is a curated Wikipedia English dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.
