datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phone-ai-contention-bench
Phone AI Contention Bench
Phone AI Contention Bench is an open, real-device evaluation seed for the issues likely to define competition in phone AI:
agent permissions and cross-app control;
indirect prompt injection and screen-perception attacks;
local-versus-cloud privacy boundaries;
structured tool reliability under quantization;
cold start, memory, context, thermals, and energy;
backend and device fragmentation;
offline resilience; and
phone-to-robot command safety.
The… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/phone-ai-contention-bench.wikimedia-enterprise-structured-contents-enwiki
enwiki_namespace_0
Structured Contents snapshot of enwiki_namespace_0 from the
Wikimedia Enterprise API, converted to Parquet.
Source
Upstream: Wikimedia Enterprise Structured Contents API
Snapshot identifier: enwiki_namespace_0
Format at source: .tar.gz containing sharded .ndjson
Shards in this release: 3
Processing
Downloaded the snapshot tarball from the Wikimedia Enterprise API.
Streamed each .ndjson shard through a normalization pass:
JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/chuckreynolds/wikimedia-enterprise-structured-contents-enwiki.runpod_multi_model_think_content_casestudieshtml-description-content
html-description-content
Warning: This dataset is under development and its content is subject to change!
📜 Dataset Summary
This dataset provides a collection of web pages, pairing full raw HTML content with its corresponding ground-truth plaintext content.
A key feature of this dataset is the addition of a LLM-generated (synthetic) query column. This query is a short (1-2 sentence) description of the page's content, designed to be used as a prompt or query for… See the full description on the dataset page: https://huggingface.co/datasets/williambrach/html-description-content.web-contentContent-Articles
Content-Articles Dataset
Overview
The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.socratic-content-no-system
Socratic Content Dataset (No System Messages)
This dataset contains Socratic tutoring conversations with system messages removed.
Data Structure
Each line contains a JSON object with the following structure:
{
"messages": [
{
"role": "user",
"content": "User's question or statement"
},
{
"role": "assistant",
"content": "Assistant's Socratic response (typically a question)"
}
]
}
Dataset Statistics
Total… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/socratic-content-no-system.Web-Content
BYU-Idaho Web Content Dataset (NLP-Enhanced)
State-of-the-art university web content dataset with full NLP enrichment: entity extraction, acronym detection, domain terminology, and semantic features. Enterprise-ready for advanced RAG, semantic search, and AI applications.
Dataset Description
Records: 2,448 ultra-high-quality pages
Source: byui.edu and subdomains
Format: Markdown + NLP metadata (JSON fields)
Quality: 40.2% filtered + 91.5/100 avg score + Full NLP… See the full description on the dataset page: https://huggingface.co/datasets/BYU-Idaho/Web-Content.
