datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YFCC14M_subset_webdatasetopen-web-math-minhash
Dataset Card for "open-web-math-minhash"
An attempt at a "high quality sample" of open-web-math/open-web-math by aggressively applying minhash from text-dedup. The result is 1.82M rows down from the original 6M:
DatasetDict({
train: Dataset({
features: ['url', 'text', 'date', 'metadata'],
num_rows: 1820241
})
})
Usage
Unless you need the metadata, load the text-only config which is only 1.4 GB/5 shards:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/open-web-math-minhash.task1728_web_nlg_data_to_text
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1728_web_nlg_data_to_text
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1728_web_nlg_data_to_text.GBC1M_webdataglm52-datagen-r11-19-knowledge-web-search-mcqa-tracesweb-agent-graph-dataset
Web Agent Grouped Graph Dataset
This dataset contains web navigation tasks in grouped graph format with full history and candidate actions for training reward models.
Data Format
Each line in graph_dataset.jsonl represents a single step with all candidate actions grouped together:
{
"task_id": "...",
"goal": "Find product X and add to cart",
"domain": "shopping",
"step_index": 3,
"history": [
{"state_id": "S0", "screenshot": "...", "url": "...", "obs":… See the full description on the dataset page: https://huggingface.co/datasets/Anish13/web-agent-graph-dataset.Cantonese-Web-Data
Dataset Summary
Cantonese has been a low-resource language in NLP. This dataset is a major step towards changing that.
To our knowledge, this is the first large-scale, properly curated, and deduplicated web dataset built specifically for Cantonese. It was created by filtering years of Common Crawl data and a Cantonese language detector, followed by a deduplication process using MinHash.
The result is a high-quality collection of ~250K unique documents containing ~150 million words.… See the full description on the dataset page: https://huggingface.co/datasets/jed351/Cantonese-Web-Data.web-scraper-datasetdataset-web-attack-newstool-reasoning-sft-TOOLS-toolmind-web-qa-sft-tool-use-data-cleaned-rectified-5.2k
ToolMind-Web-QA — Hermes Reasoning Format
Filtered and restructured version of Nanbeige/ToolMind-Web-QA.
Filters applied: valid role transitions only · known tools only · non-empty user + answer required
Size: 5,274 examples (from 5,624 original trajectories, 350 dropped)
Source
The original dataset contains 5,624 complex multi-hop QA trajectories grounded in Wikipedia
entity-relation graphs. Each trajectory has an average of ~138 turns with multiple tool calls
across… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolmind-web-qa-sft-tool-use-data-cleaned-rectified-5.2k.web_prm_processed_dataweb.facebook.com-ru4Ak2ui-scraped-data-Final-Evalyouthless-homeless-shelter-web-scrape-datasetyouthless-homeless-shelter-web-scrape-dataset-largesmall-open-web-math-dataset-v2# Small Open Web Math Dataset v2
A 10k-sample shuffled subset of OpenWebMath, ensuring randomized selection of high-quality mathematical text.
small-open-web-math-dataset# Small Open Web Math Dataset
A 10k-sample subset of OpenWebMath, focused on high-quality mathematical text.
larkin-web-scrape-dataset-qa-formattedafrica-web-attack-dataset
Web Application Attacks (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-web-attack-dataset.dataset-web-attackweb_text_synthetic_dataset_50kweb_scraper_datasetFine_web_lmsys_datasetsap-web-dataysa-web-scrape-dataset-qa-formatted-small-versionwebdataflan_combined_task1728_web_nlg_data_to_textlarkin-web-scrape-dataset-qa-formatted-small-versionNew-Testament-World-English-Web-Dataset-V1
New-Testament-World-English-Web-Dataset-V1
Made with ❤️ using 🦥 Unsloth Studio
New Testament World English Web Dataset V1 was generated with Unsloth Recipe Studio. It contains 700 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("Markie77/New-Testament-World-English-Web-Dataset-V1", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 700
📋 Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/Markie77/New-Testament-World-English-Web-Dataset-V1.youthless-homeless-shelter-web-scrape-dataset-qa-formattedai_web_scraping_dataset
