datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Content-Articles
Content-Articles Dataset
Overview
The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.Web-Content
BYU-Idaho Web Content Dataset (NLP-Enhanced)
State-of-the-art university web content dataset with full NLP enrichment: entity extraction, acronym detection, domain terminology, and semantic features. Enterprise-ready for advanced RAG, semantic search, and AI applications.
Dataset Description
Records: 2,448 ultra-high-quality pages
Source: byui.edu and subdomains
Format: Markdown + NLP metadata (JSON fields)
Quality: 40.2% filtered + 91.5/100 avg score + Full NLP… See the full description on the dataset page: https://huggingface.co/datasets/BYU-Idaho/Web-Content.
