datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YFCC14M_subset_webdatasetopen-web-math-minhash
Dataset Card for "open-web-math-minhash"
An attempt at a "high quality sample" of open-web-math/open-web-math by aggressively applying minhash from text-dedup. The result is 1.82M rows down from the original 6M:
DatasetDict({
train: Dataset({
features: ['url', 'text', 'date', 'metadata'],
num_rows: 1820241
})
})
Usage
Unless you need the metadata, load the text-only config which is only 1.4 GB/5 shards:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/open-web-math-minhash.task1728_web_nlg_data_to_text
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.GBC1M_webdataglm52-datagen-r11-19-knowledge-web-search-mcqa-tracesCantonese-Web-Data
Dataset Summary
Cantonese has been a low-resource language in NLP. This dataset is a major step towards changing that.
To our knowledge, this is the first large-scale, properly curated, and deduplicated web dataset built specifically for Cantonese. It was created by filtering years of Common Crawl data and a Cantonese language detector, followed by a deduplication process using MinHash.
The result is a high-quality collection of ~250K unique documents containing ~150 million words.… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese-Web-Data.web-scraper-datasetdataset-web-attack-newstool-reasoning-sft-TOOLS-toolmind-web-qa-sft-tool-use-data-cleaned-rectified-5.2k
ToolMind-Web-QA — Hermes Reasoning Format
Filtered and restructured version of Nanbeige/ToolMind-Web-QA.
Filters applied: valid role transitions only · known tools only · non-empty user + answer required
Size: 5,274 examples (from 5,624 original trajectories, 350 dropped)
Source
The original dataset contains 5,624 complex multi-hop QA trajectories grounded in Wikipedia
entity-relation graphs. Each trajectory has an average of ~138 turns with multiple tool calls
across… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolmind-web-qa-sft-tool-use-data-cleaned-rectified-5.2k.web_prm_processed_dataweb.facebook.com-ru4Ak2ui-scraped-data-Final-Evalyouthless-homeless-shelter-web-scrape-datasetyouthless-homeless-shelter-web-scrape-dataset-largesmall-open-web-math-dataset-v2# Small Open Web Math Dataset v2
A 10k-sample shuffled subset of OpenWebMath, ensuring randomized selection of high-quality mathematical text.
small-open-web-math-dataset# Small Open Web Math Dataset
A 10k-sample subset of OpenWebMath, focused on high-quality mathematical text.
larkin-web-scrape-dataset-qa-formattedafrica-web-attack-dataset
Web Application Attacks (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-web-attack-dataset.dataset-web-attackweb_text_synthetic_dataset_50kweb_scraper_datasetFine_web_lmsys_datasetsap-web-dataysa-web-scrape-dataset-qa-formatted-small-versionwebdataflan_combined_task1728_web_nlg_data_to_textlarkin-web-scrape-dataset-qa-formatted-small-versionNew-Testament-World-English-Web-Dataset-V1
New-Testament-World-English-Web-Dataset-V1
Made with ❤️ using 🦥 Unsloth Studio
New Testament World English Web Dataset V1 was generated with Unsloth Recipe Studio. It contains 700 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("Markie77/New-Testament-World-English-Web-Dataset-V1", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 700
📋 Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/Markie77/New-Testament-World-English-Web-Dataset-V1.youthless-homeless-shelter-web-scrape-dataset-qa-formattedai_web_scraping_datasetafrica-dark-web-data-trading
Dark Web African Data Trading | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-dark-web-data-trading.
