datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
II-Medical-Reasoning-SFT
II-Medical-Reasoning-SFT
II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice.
The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.dead-internet-observatoryii-agent_gaia-benchmark_validationGAIA-Subset-Benchmark
GAIA Benchmark Subset Model Card
This dataset is a subset of the GAIA benchmark, containing 44 web-search-based questions from the validation set. It evaluates multiple AI models on their ability to retrieve and process real-time information using web search and browser tools. Performance metrics include success indicators and detailed reports for each model. A comparative chart summarizing the results will be provided separately.
Benchmark Results
II-Thought-RL-v0
II-Thought RL v0: A Large-Scale Curated Dataset for Reinforcement Learning
See our blog here for additional details.
We introduce II-Thought RL v0, the first large-scale, multi-task dataset designed for Reinforcement Learning. This dataset consists of high-quality question-answer pairs that have undergone a rigorous multi-step filtering process, leveraging Gemini 2.0 Flash and Qwen 32B as quality evaluators.
In this initial release, we have curated and refined publicly available… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Thought-RL-v0.InternetArchive_1899_Large
Internet Archive Historical Texts (0001-1899)
TL;DR
711,680 cleaned public-domain style documents harvested from the Internet Archive via a high-throughput text-to-parquet pipeline.
Coverage targets items that contain textual content dated between 0001 and 1899, ranked by download counts; ~715k IDs were attempted, ~4.1k were filtered during preprocessing.
Stored in 620 Zstandard-compressed Parquet shards (shard_00000.parquet ... shard_00619.parquet) occupying ~240 GB on… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Large.television_por_internetInternetArchive_1899_Chunked
Internet Archive Historical Texts - Chunked (0001-1899)
TL;DR
163 million text chunks extracted from historical public-domain documents sourced from the Internet Archive
Content dated 0001-1899, sorted by download popularity to prioritize high-quality, frequently accessed materials
2,445 Zstandard-compressed Parquet shards totaling ~217 GB on disk, ~594 billion characters uncompressed
Optimized chunk size of ~3,600 characters (target: 4,000) for efficient language model… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Chunked.doom-e1-internet-gameplayI've used the Inverse Dynamic Model, I've previously trained on manually recorded gameplay, on pure gameplay YouTube videos.
This dataset is in public domain, use it however you want.
ChatDoctor-RL
Intelligent-Internet/ChatDoctor-Improved-Answer Dataset
This dataset represents a carefully curated subset derived from the original ChatDoctor-HealthCareMagic-100k[lavita/ChatDoctor-HealthCareMagic-100k] dataset, where we have undertaken significant improvements to enhance the quality and depth of the responses. The answers have been thoroughly refined to provide greater detail, clarity, and precision, while incorporating a heightened focus on safety awareness to ensure responsible… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/ChatDoctor-RL.Vietnamese-Entrance-Exam
Vietnamese Entrance Exam Dataset
The Vietnamese Entrance Exam dataset is a collection of 432 problems derived from Vietnamese University entrance examinations. The dataset aims to provide a novel benchmark for testing reasoning capabilities of language models in several low resource domains specifically designed to minimize potential data contamination from pre-training or post-training exposure.
Domain
Count
Physics
95
Chemistry
94
Math
243
Data… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/Vietnamese-Entrance-Exam.II-Medical-RL
Overview
The MedReason-RL dataset is a refined version of the original MedReason dataset, specifically curated for training reinforcement learning (RL) models to enhance reasoning abilities. It has been proven to be the best dataset for improving model reasoning through RL training.
Source
This dataset is derived from the original MedReason dataset, which focuses on medical reasoning tasks. However, the original dataset contained significant overlap with benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-RL.II-Search-CIR-SFTafrica-worldbank-individuals-using-the-internet-of-population-it-net-user-zs
Individuals using the Internet (% of population) | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: 1K<n<10K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-individuals-using-the-internet-of-population-it-net-user-zs.II-Thought-RL-v0-Math-50KOpenAI-HealthBench-II-Medical-8B-1706-GPT-4.1Internet-background-noise
Internet Background Noise Dataset (Unlabeled Raw Data)
This dataset contains HTTP internet noise data collected by an internet honeypot. It consists of raw, unlabeled network packets, including metadata, payloads, and header information. This data is suitable for training and evaluating machine learning models for network intrusion detection, cybersecurity, and traffic analysis.
HoneyPot repository: hachimi on GitHub.
Dataset Overview
The Internet Background Noise… See the full description on the dataset page: https://huggingface.co/datasets/burpheart/Internet-background-noise.africa-synth-telecom-massive-internet-of-things-traffic-nigeria
Africa Synth Telecom Massive Internet of Things Traffic Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-telecom-massive-internet-of-things-traffic-nigeria.africa-owid-number-of-internet-users
Number Of Internet Users | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-number-of-internet-users.test_recordThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grievous_client",
"total_episodes": 10,
"total_frames": 2987,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/InternetSandwich33/test_record.II-Search-Benchmark-Details
Inspect-Search-Models-Benchmarking-Result
Overall result
Qwen 4B
Jan 4B
WebSailor-3B
II-Search-4B
II-Search-CIR-4B
OpenAI/SimpleQA
76.8
80.1
81.8
91.8
91.8
Google/Frames
30.7
24.8
34.0
67.5
72.2
Seal_0
6.31
2.7
1.8
22.5
26.4
Simple QA (SerpDev)
Qwen 4B
Jan 4B
WebSailor-3B
II-Search-4B
II-Search-CIR-4B
Pass rate %
76.8
80.1
81.8
91.8
91.8
# Search
1.0
0.9
2.1
2.2
2.5
# Visit
0.1
1.9
6.4
3.5
5.3
# Tool used
1.1
2.8
8.5
5.7
7.8
Frames… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Search-Benchmark-Details.africa-cote-d-ivoire-marche-de-l-internet-et-de-la-telephonie-mobile-de-2010-a-fec81c58
Marche De L Internet Et De La Telephonie Mobile De 2010 a | Africa (Cote d'Ivoire DataFair)
263 rows - 1 Africa country/area - 2010-2016 - 1 indicator - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 263 rows from Cote d'Ivoire DataFair, covering Marche De L Internet Et De La Telephonie Mobile De 2010 a. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cote-d-ivoire-marche-de-l-internet-et-de-la-telephonie-mobile-de-2010-a-fec81c58.africa-population-and-internet-users-statistics
Africa - Population and Internet users statistics | Africa (original)
Size category: n<1K - Formats: parquet - Sector: humanitarian_development - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-population-and-internet-users-statistics.africa-worldbank-internet-users-per-100-people-it-net-user-p2
Internet users (per 100 people) | Africa (World Bank — Education Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Education… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-internet-users-per-100-people-it-net-user-p2.swebench-pro-claude-sonnet-4.5-ii-agent-trajectoriesimnet1k_web_site_website_internet_site_siteasia-owid-number-of-internet-users
Number Of Internet Users | Asia (Our World in Data)
🌏 1,305 observations · 48 Asia countries · 1990–2021 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 1,305 observations of Number Of Internet Users data across 48 Asia countries, spanning 1990–2021.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Number Of Internet Users
Geographic coverage
48 Asia countries ·… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-number-of-internet-users.africa-worldbank-proportion-of-secondary-schools-with-access-to-internet-for-pedagogical-purpose
Proportion of secondary schools with access to Internet for pedagogical purposes (%) | Africa (World Bank — Education Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-proportion-of-secondary-schools-with-access-to-internet-for-pedagogical-purpose.asia-owid-landline-internet-subscriptions
Landline Internet Subscriptions | Asia (Our World in Data)
🌏 990 observations · 47 Asia countries · 1998–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 990 observations of Landline Internet Subscriptions data across 47 Asia countries, spanning 1998–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Landline Internet Subscriptions
Geographic coverage
47… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-landline-internet-subscriptions.asia-owid-landline-internet-subscriptions-per-100-people-by-speed
Landline Internet Subscriptions Per 100 People By Speed | Asia (Our World in Data)
🌏 984 observations · 47 Asia countries · 2000–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 984 observations of Landline Internet Subscriptions Per 100 People By Speed data across 47 Asia countries, spanning 2000–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Landline Internet… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-landline-internet-subscriptions-per-100-people-by-speed.
