datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sea-internetOpenAI-HealthBench-II-Medical-8B-GPT-4.1internet-prompts-benchmark
Internet Prompts Benchmark
Viral internet prompts, memes, and tests that AI historically failed at. Popular ones like counting letters in the word strawberry and nicher ones that test other important capabilities.Mostly made by GPT-5.6 Sol.
It is designed for many types of models to participate, small and large, not only transformers.
It has prompts from the early days of AI to the very latest.
It will be actively updated to preserve various prompts for as long as I can afford… See the full description on the dataset page: https://huggingface.co/datasets/NikoThePig/internet-prompts-benchmark.internetovai-identitymillan_internet_traffic
Milan Internet Traffic Dataset
This dataset contains information about hourly internet traffic in Milan between 2013-11-01 and 2014-01-01.
en-si-translation-wmt-internet-1300
En Si Translation Wmt Internet 1300
Dataset Summary
English-Sinhala Web Forum and Internet Translation dataset containing ~1,300 high-quality sentences capturing internet slang and long-form narrative text from WMT20.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 1300
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-wmt-internet-1300.internet-assesment-ujian-akhir-jsonl-cleaninternet_archive_azerbaijaniinternet_capability_hallucinationA dataset to evaluate if models are more likely to be mistaken about their capabilities if the task is morally questionable.
Each entry has an action that requires using the internet (e.g. "Do a Google search"), a specific prompt that asks the model to do the action for some "nice" purpose, and a prompt that asks the model to do it for a slightly-evil purpose.
The hypothesis to test here is that some models are likely to hallucinate their internet capabilities and say "I've searched google and… See the full description on the dataset page: https://huggingface.co/datasets/scale-safety-research/internet_capability_hallucination.tgl-blogInternet-Forum-Logs-1-Sharegptinternet_community_style
