CoolFace
Datasetpublic

EssentialAI/eai-taxonomy-stem-w-dclm

πŸ”¬ EAI-Taxonomy STEM w/ DCLM πŸ† Website | πŸ–₯️ Code | πŸ“– Paper A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content. 🎯 Dataset Overview This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
6likes2.7kdownloads
Dataset Card

πŸ”¬ EAI-Taxonomy STEM w/ DCLM

πŸ† Website | πŸ–₯️ Code | πŸ“– Paper

A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content.

🎯 Dataset Overview

This dataset is part of the **Essential-Web** project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require complex domain-specific pipelines, our approach leverages a 12-category taxonomy to efficiently identify and extract high-quality STEM content.

πŸ§ͺ EAI-Taxonomy STEM w/ DCLM (1742B tokens): Documents targeting science, engineering, medical, and computer science content that exhibit reasoning, combined with the DCLM classifier to filter for instruction-dense documents.

πŸ† Performance

Our taxonomy-based approach achieves superior results with significantly less curation effort:

DatasetMMLU-STEMCuration Complexity
DCLM-baseline27.7%General web filtering
FineWeb-Edu26.7%Educational filtering
EAI-Taxonomy STEM29.1%Simple semantic filter
EAI-Taxonomy STEM w/ DCLM34.5%+ DCLM classifier

Results show +24.5% improvement over DCLM and +29.2% improvement over FineWeb-Edu.

πŸ” Key Findings

  • β€”Strong STEM Performance: Outperforms baseline and educational datasets beyond standard error
  • β€”Efficient Curation: Achieves superior results without complex domain-specific pipelines
  • β€”Broad Coverage: Encompasses science, engineering, medical, and computer science domains
  • β€”Quality Focus: Selects high-quality document types and filters for reasoning content

Dataset Schema Documentation

Overview

This dataset contains web-crawled text data with comprehensive metadata, quality signals, and taxonomic classifications. Each record represents a document extracted from web archives with detailed provenance tracking and quality assessment metrics.

Core Fields

FieldTypeDescriptionPath
idInt64Unique identifier based on document hashid
textStringThe main textual content of the documenttext

EAI Taxonomy Classification

Comprehensive hierarchical classification system with primary and secondary labels - the most important feature of this dataset. The taxonomy is designed to provide detailed subject categorization, document type identification, content quality assessment, and extraction quality indicators.

<details> <summary><strong>Free Decimal Correspondence (FDC)</strong></summary>

A Dewey Decimal-inspired classification system with 3-level hierarchical labels. The FDC provides nested categories where each successive level refines its parent category. It's designed to be compatible with the Dewey Decimal System for library cataloging.

Level Structure:

  • β€”Level 1: Top-level categories (0-9) covering broad subject areas like General works, Philosophy, Religion, Social Sciences, etc.
  • β€”Level 2: Sub-divisions (00-99) that refine Level 1 categories
  • β€”Level 3: Specific categories (000-999) that further refine Level 2 categories
ComponentDescriptionPath
Primary CodeMain classification codeeai_taxonomy.free_decimal_correspondence.primary.code
Primary Level 1Top-level category (0=General works, 1=Philosophy, 2=Religion, 3=Social Sciences, 4=Language, 5=Science, 6=Technology, 7=Arts, 8=Literature, 9=History/Geography)eai_taxonomy.free_decimal_correspondence.primary.labels.level_1
Primary Level 2Mid-level categoryeai_taxonomy.free_decimal_correspondence.primary.labels.level_2
Primary Level 3Specific categoryeai_taxonomy.free_decimal_correspondence.primary.labels.level_3
Secondary CodeAlternative classification codeeai_taxonomy.free_decimal_correspondence.secondary.code
Secondary Level 1Alternative top-level categoryeai_taxonomy.free_decimal_correspondence.secondary.labels.level_1
Secondary Level 2Alternative mid-level categoryeai_taxonomy.free_decimal_correspondence.secondary.labels.level_2
Secondary Level 3Alternative specific categoryeai_taxonomy.free_decimal_correspondence.secondary.labels.level_3

We recommend this viewer for easily navigating the FDC categories when curating filters: https://www.librarything.com/mds

</details>

<details> <summary><strong>Bloom's Taxonomy Integration</strong></summary>

Based on Anderson and Krathwohl's 2001 revision of Bloom's Taxonomy of Educational Objectives, providing two complementary categorization dimensions for educational content analysis.

Knowledge Domain

Categorizes the type of knowledge demonstrated in the document:

ComponentDescriptionPath
Primary CodeMain knowledge domain codeeai_taxonomy.bloom_knowledge_domain.primary.code
Primary LabelMain knowledge domain labeleai_taxonomy.bloom_knowledge_domain.primary.label
Secondary CodeAlternative knowledge domain codeeai_taxonomy.bloom_knowledge_domain.secondary.code
Secondary LabelAlternative knowledge domain labeleai_taxonomy.bloom_knowledge_domain.secondary.label

Possible Values: | Code | Label | Description | |------|-------|-------------| | -1 | Abstain | Unable to determine | | 1 | Factual | Basic elements to learn or solve problems | | 2 | Conceptual | Interrelationships between basic elements within larger context | | 3 | Procedural | Methods and techniques in the discipline | | 4 | Metacognitive | Awareness of how learning works in relation to oneself |

Cognitive Processing Level

Assesses the learning and thinking skill levels demonstrated by the document author:

ComponentDescriptionPath
Primary CodeMain cognitive process codeeai_taxonomy.bloom_cognitive_process.primary.code
Primary LabelMain cognitive process labeleai_taxonomy.bloom_cognitive_process.primary.label
Secondary CodeAlternative cognitive process codeeai_taxonomy.bloom_cognitive_process.secondary.code
Secondary LabelAlternative cognitive process labeleai_taxonomy.bloom_cognitive_process.secondary.label

Possible Values: | Code | Label | Description | |------|-------|-------------| | -1 | Abstain | Unable to determine | | 1 | Remember | Retrieve relevant knowledge from memory | | 2 | Understand | Determine meaning of instructional messages | | 3 | Apply | Use a procedure in a given situation | | 4 | Analyze | Break materials into components and determine relationships | | 5 | Evaluate | Make judgments based on criteria and standards | | 6 | Create | Create new or original work |

</details>

<details> <summary><strong>Document Characteristics</strong></summary>

Document Type v1

In-house classification of common web document types and formats:

ComponentDescriptionPath
Primary CodeMain document type codeeai_taxonomy.document_type_v1.primary.code
Primary LabelMain document type labeleai_taxonomy.document_type_v1.primary.label
Secondary CodeAlternative document type codeeai_taxonomy.document_type_v1.secondary.code
Secondary LabelAlternative document type labeleai_taxonomy.document_type_v1.secondary.label

Possible Values: | Code | Label | Examples | |------|-------|----------| | -1 | Abstain | Unable to classify | | 1 | News/Editorial | CNN articles, opinion columns | | 2 | Academic/Research | ArXiv papers, research articles | | 3 | Reference/Encyclopedic/Educational | FAQs, Wikipedia entries | | 4 | Code/Software | GitHub repos, code examples | | 5 | Social/Forum | Conversation threads, Q&A boards | | 6 | Promotional/Advertisement | Product pages, calls to action | | 7 | Search/Directory/Bibliography | Link pages, search results | | 8 | Adult/Pornographic | Adult content | | 9 | Personal/Misc | Blogs, user profiles | | 10 | Machine-Generated | Lorem ipsum, garbled text | | 11 | Legal/Regulatory | Contracts, terms of service | | 12 | Government/Political | Legislation, press releases | | 13 | Literary/Creative | Poems, short stories | | 14 | Reviews/Critiques | Film critiques, product reviews | | 15 | E-Commerce/Marketplace | eBay listings, Amazon pages | | 16 | Images/Videos/Audio | YouTube videos, Imgur pages | | 17 | Other/Unclassified | Documents that resist classification |

Document Type v2

Updated classification based on WebOrganizer taxonomy with refined categories for improved document classification accuracy:

ComponentDescriptionPath
Primary CodeMain document type code (v2)eai_taxonomy.document_type_v2.primary.code
Primary LabelMain document type label (v2)eai_taxonomy.document_type_v2.primary.label
Secondary CodeAlternative document type code (v2)eai_taxonomy.document_type_v2.secondary.code
Secondary LabelAlternative document type label (v2)eai_taxonomy.document_type_v2.secondary.label

Complete Value Mapping: | Code | Label | Examples | |------|-------|----------| | -1 | Abstain | Documents requiring human review | | 1 | About (Org.) | Company about pages, mission statements | | 2 | About (Personal) | Personal bios, LinkedIn profiles | | 3 | Academic Writing | Research papers, abstracts, dissertations | | 4 | Audio Transcript | Interview transcripts, court records, captions | | 5 | Comment Section | Reddit threads, blog comments | | 6 | Content Listing | Site maps, product catalogs, directory listings | | 7 | Creative Writing | Song lyrics, novel excerpts, poetry | | 8 | Documentation | API docs, README files, user manuals | | 9 | FAQ | FAQ pages, Q&A lists | | 10 | Knowledge Article | Wikipedia articles, Britannica entries | | 11 | Legal Notices | Privacy policies, license agreements, terms of service | | 12 | Listicle | Buzzfeed-style articles, "Top 10" lists | | 13 | News (Org.) | Government blog posts, corporate announcements | | 14 | News Article | Newspaper articles, CNN content, breaking news | | 15 | Nonfiction Writing | Editorials, obituaries, memoirs, opinion pieces | | 16 | Personal Blog | Personal journals, diary entries, lifestyle blogs | | 17 | Product Page | Product descriptions, course offerings, sales pages | | 18 | Q&A Forum | Quora posts, Stack Exchange discussions | | 19 | Spam / Ads | SEO keyword stuffing, promotional spam | | 20 | Structured Data | Datasheets, glossaries, JSON files, databases | | 21 | Customer Support | Help articles, troubleshooting guides | | 22 | Truncated | Paywalled sites, image galleries, partial content | | 23 | Tutorial | Cooking recipes, WikiHow pages, step-by-step guides | | 24 | User Review | Yelp reviews, TripAdvisor feedback, product reviews | | 25 | Other/Unclassified | Miscellaneous documents not fitting other categories |

Extraction Artifacts

Assessment of technical extraction quality, identifying issues from HTML-to-text conversion:

ComponentDescriptionPath
Primary CodeMain extraction artifact codeeai_taxonomy.extraction_artifacts.primary.code
Primary LabelMain extraction artifact labeleai_taxonomy.extraction_artifacts.primary.label
Secondary CodeAlternative extraction artifact codeeai_taxonomy.extraction_artifacts.secondary.code
Secondary LabelAlternative extraction artifact labeleai_taxonomy.extraction_artifacts.secondary.label

Possible Values: | Code | Label | Description | |------|-------|-------------| | -1 | Abstain | Unable to determine | | 0 | No Artifacts | Clean text with no leftover HTML or irrelevant elements | | 1 | Leftover HTML | HTML/code artifacts remaining after extraction | | 2 | Text Extraction Errors | Broken math expressions, encoding errors, improperly parsed tables | | 3 | Irrelevant Content | Headers, footers, nav menus extracted by mistake | | 4 | Indeterminate | Insufficient content to judge |

Missing Content

Assessment of content completeness and extraction success:

ComponentDescriptionPath
Primary CodeMain missing content codeeai_taxonomy.missing_content.primary.code
Primary LabelMain missing content labeleai_taxonomy.missing_content.primary.label
Secondary CodeAlternative missing content codeeai_taxonomy.missing_content.secondary.code
Secondary LabelAlternative missing content labeleai_taxonomy.missing_content.secondary.label

Possible Values: | Code | Label | Description | |------|-------|-------------| | -1 | Abstain | Unable to determine | | 0 | No Missing Content | Complete and coherent text | | 1 | Truncated Snippets | Obvious "...", incomplete paragraphs, cut-off text | | 2 | Click Here References | "Download here", "Click here" without linked content | | 3 | Incoherent Flow | Unreadable or illogical flow due to missing context | | 4 | Missing Images or Figures | Placeholders or references to missing visual content | | 5 | Missing Referenced Data | References to absent tables/datasets (e.g., "See Table 3") | | 6 | Indeterminate | Insufficient content to judge |

Text Structure Information

FieldTypeDescriptionPath
Line Start IndicesList[Int32]Starting indices of each lineline_start_n_end_idx.line_start_idx
Line End IndicesList[Int32]Ending indices of each lineline_start_n_end_idx.line_end_idx

</details>

<details> <summary><strong>Content Quality Dimensions</strong></summary>

Quality assessment inspired by NaturalReasoning and FineWeb efforts to categorize web data by information sophistication.

Reasoning Depth

Assesses the complexity and sophistication of logical reasoning in the document:

ComponentDescriptionPath
Primary CodeMain reasoning depth codeeai_taxonomy.reasoning_depth.primary.code
Primary LabelMain reasoning depth labeleai_taxonomy.reasoning_depth.primary.label
Secondary CodeAlternative reasoning depth codeeai_taxonomy.reasoning_depth.secondary.code
Secondary LabelAlternative reasoning depth labeleai_taxonomy.reasoning_depth.secondary.label

Possible Values: | Code | Label | Description | |------|-------|-------------| | -1 | Abstain | Unable to determine | | 1 | No Reasoning | Facts present but no evidence of reasoning | | 2 | Basic Reasoning | Basic analysis with minimal explanation and summarization | | 3 | Intermediate Reasoning | Some logical steps connecting ideas and structured thinking | | 4 | Advanced Reasoning | Multi-step reasoning and thorough analysis with well-developed explanations | | 5 | Exceptional Reasoning | Novel abstractions, theoretical frameworks, long chain-of-thought, original insights, or proofs | | 6 | Indeterminate | Insufficient context to judge |

Technical Correctness

Evaluates the accuracy and precision of technical information:

ComponentDescriptionPath
Primary CodeMain technical correctness codeeai_taxonomy.technical_correctness.primary.code
Primary LabelMain technical correctness labeleai_taxonomy.technical_correctness.primary.label
Secondary CodeAlternative technical correctness codeeai_taxonomy.technical_correctness.secondary.code
Secondary LabelAlternative technical correctness labeleai_taxonomy.technical_correctness.secondary.label

Possible Values: | Code | Label | Description | |------|-------|-------------| | -1 | Abstain | Unable to determine | | 1 | Technically Flawed | Significant errors undermining content validity | | 2 | Partially Correct | Some correctness but contains flaws, omissions, or errors | | 3 | Mostly Correct | Technical correctness with minor flaws or incomplete explanations | | 4 | Highly Correct | High technical correctness with precise definitions and clear explanations | | 5 | Exceptionally Correct | Exceptional technical correctness with formal proofs and flawless content | | 6 | Not Applicable/Indeterminate | No technical content or insufficient context |

Education Level

Assesses the appropriate educational background required to comprehend the content:

ComponentDescriptionPath
Primary CodeMain education level codeeai_taxonomy.education_level.primary.code
Primary LabelMain education level labeleai_taxonomy.education_level.primary.label
Secondary CodeAlternative education level codeeai_taxonomy.education_level.secondary.code
Secondary LabelAlternative education level labeleai_taxonomy.education_level.secondary.label

Possible Values: | Code | Label | Description | |------|-------|-------------| | -1 | Abstain | Unable to determine | | 1 | General Audience | Accessible to anyone with basic literacy; simple terms | | 2 | High School Level | Requires high school education; specialized terminology explained for non-experts | | 3 | Undergraduate Level | Requires college education; uses specialized terminology and assumes background knowledge | | 4 | Graduate/Expert Level | Requires graduate education or domain expertise; assumes deep background knowledge | | 5 | Indeterminate | Insufficient content to judge educational level |

</details>

<details> <summary><strong>Metadata</strong></summary>

Metadata Structure

The metadata field contains a nested structure with web archive information:

FieldTypeDescriptionPath
URL Information
URLStringOriginal URL of the documentmetadata.url
Source DomainStringDomain name of the sourcemetadata.source_domain
Snapshot IDStringIdentifier for the web archive snapshotmetadata.snapshot_id
WARC MetadataWARC (Web ARChive) format metadata
Content LengthStringSize of the contentmetadata.warc_metadata.Content-Length
Content TypeStringMIME type of the contentmetadata.warc_metadata.Content-Type
Block DigestStringChecksum of the WARC blockmetadata.warc_metadata.WARC-Block-Digest
Concurrent ToStringRelated WARC recordsmetadata.warc_metadata.WARC-Concurrent-To
DateStringTimestamp of the crawlmetadata.warc_metadata.WARC-Date
IP AddressStringSource server IP addressmetadata.warc_metadata.WARC-IP-Address
Payload TypeStringIdentified content typemetadata.warc_metadata.WARC-Identified-Payload-Type
Payload DigestStringChecksum of the payloadmetadata.warc_metadata.WARC-Payload-Digest
Record IDStringUnique WARC record identifiermetadata.warc_metadata.WARC-Record-ID
Target URIStringOriginal target URLmetadata.warc_metadata.WARC-Target-URI
TruncatedStringTruncation statusmetadata.warc_metadata.WARC-Truncated
TypeStringWARC record typemetadata.warc_metadata.WARC-Type
Warcinfo IDStringAssociated warcinfo recordmetadata.warc_metadata.WARC-Warcinfo-ID
Additional Info
WARC InfoStringAdditional WARC informationmetadata.warc_info

</details>

<details> <summary><strong>Quality Signals</strong></summary>

The dataset includes two comprehensive quality assessment frameworks:

Red Pajama v2 Quality Metrics

Text quality indicators derived from the Red Pajama v2 filtering pipeline:

Content Structure Metrics

MetricDescriptionPath
Original LengthOriginal document lengthquality_signals.red_pajama_v2.ccnet_original_length
Original LinesNumber of lines in original documentquality_signals.red_pajama_v2.ccnet_original_nlines
Sentence CountTotal sentence countquality_signals.red_pajama_v2.rps_doc_num_sentences
Word CountTotal word countquality_signals.red_pajama_v2.rps_doc_word_count
Mean Word LengthAverage word lengthquality_signals.red_pajama_v2.rps_doc_mean_word_length

Language Quality Metrics

MetricDescriptionPath
Stop Word FractionProportion of stop wordsquality_signals.red_pajama_v2.rps_doc_stop_word_fraction
Unique Words FractionFraction of unique wordsquality_signals.red_pajama_v2.rps_doc_frac_unique_words
All Caps WordsFraction of words in all capitalsquality_signals.red_pajama_v2.rps_doc_frac_all_caps_words
Non-Alphabetic WordsFraction of non-alphabetic wordsquality_signals.red_pajama_v2.rps_doc_frac_no_alph_words
Unigram EntropyEntropy measure of word distributionquality_signals.red_pajama_v2.rps_doc_unigram_entropy

Content Pattern Analysis

MetricDescriptionPath
Curly Bracket DensityCurly bracket density (code indicator)quality_signals.red_pajama_v2.rps_doc_curly_bracket
Symbol-to-Word RatioSymbol-to-word ratioquality_signals.red_pajama_v2.rps_doc_symbol_to_word_ratio
Ellipsis Line EndingsLines ending with ellipsisquality_signals.red_pajama_v2.rps_doc_frac_lines_end_with_ellipsis
Lorem Ipsum DetectionLorem ipsum text detectionquality_signals.red_pajama_v2.rps_doc_lorem_ipsum
Offensive ContentPotentially offensive content detectionquality_signals.red_pajama_v2.rps_doc_ldnoobw_words
UT1 BlacklistUT1 blacklist filtering scorequality_signals.red_pajama_v2.rps_doc_ut1_blacklist

Duplication Detection

MetricDescriptionPath
5-gram DuplicationCharacter-level duplication for 5-gramsquality_signals.red_pajama_v2.rps_doc_frac_chars_dupe_5grams
6-gram DuplicationCharacter-level duplication for 6-gramsquality_signals.red_pajama_v2.rps_doc_frac_chars_dupe_6grams
7-gram DuplicationCharacter-level duplication for 7-gramsquality_signals.red_pajama_v2.rps_doc_frac_chars_dupe_7grams
8-gram DuplicationCharacter-level duplication for 8-gramsquality_signals.red_pajama_v2.rps_doc_frac_chars_dupe_8grams
9-gram DuplicationCharacter-level duplication for 9-gramsquality_signals.red_pajama_v2.rps_doc_frac_chars_dupe_9grams
10-gram DuplicationCharacter-level duplication for 10-gramsquality_signals.red_pajama_v2.rps_doc_frac_chars_dupe_10grams
Top 2-gram CoverageMost frequent 2-gram coveragequality_signals.red_pajama_v2.rps_doc_frac_chars_top_2gram
Top 3-gram CoverageMost frequent 3-gram coveragequality_signals.red_pajama_v2.rps_doc_frac_chars_top_3gram
Top 4-gram CoverageMost frequent 4-gram coveragequality_signals.red_pajama_v2.rps_doc_frac_chars_top_4gram

Domain Importance Scores

MetricDescriptionPath
Books ImportanceSimilarity to book contentquality_signals.red_pajama_v2.rps_doc_books_importance
Books Importance (Length Corrected)Length-corrected books similarityquality_signals.red_pajama_v2.rps_doc_books_importance_length_correction
OpenWebText ImportanceSimilarity to OpenWebTextquality_signals.red_pajama_v2.rps_doc_openwebtext_importance
OpenWebText Importance (Length Corrected)Length-corrected OpenWebText similarityquality_signals.red_pajama_v2.rps_doc_openwebtext_importance_length_correction
Wikipedia ImportanceSimilarity to Wikipediaquality_signals.red_pajama_v2.rps_doc_wikipedia_importance
Wikipedia Importance (Length Corrected)Length-corrected Wikipedia similarityquality_signals.red_pajama_v2.rps_doc_wikipedia_importance_length_correction

FastText Classification Scores

Domain and content type classification probabilities:

MetricDescriptionPath
DCLM ScoreDataComp-LM classifier scorequality_signals.fasttext.dclm
English ConfidenceEnglish language confidencequality_signals.fasttext.english
Educational ContentEducational content approximationquality_signals.fasttext.fineweb_edu_approx
General MathGeneral mathematics contentquality_signals.fasttext.eai_general_math
Web MathOWM Web-based mathematics contentquality_signals.fasttext.eai_open_web_math
Code ContentCode content detectionquality_signals.fasttext.eai_web_code

</details>

How to Load the Dataset

This section provides examples of how to load the EssentialAI/eai-taxonomy-stem-w-dclm dataset using different Python libraries and frameworks.

Using Hugging Face Datasets (Standard Method)

The simplest way to load the dataset is using the Hugging Face datasets library:

python
from datasets import load_dataset

# Load the entire dataset
dataset = load_dataset("EssentialAI/eai-taxonomy-stem-w-dclm")

# View dataset structure
print(dataset)
print(f"Number of examples: {len(dataset['train'])}")

You can also load the dataset in streaming mode to avoid downloading the entire dataset at once:

python
from datasets import load_dataset

# Load in streaming mode
dataset = load_dataset("EssentialAI/eai-taxonomy-stem-w-dclm", streaming=True)
data_stream = dataset["train"]

# Iterate through examples
for example in data_stream.take(5):
    print(example)

Using PySpark

For large-scale distributed processing, you can load the dataset using PySpark with the pyspark_huggingface library:

python
# First install the required library:
# pip install pyspark_huggingface

import pyspark_huggingface
from pyspark.sql import SparkSession

# Initialize Spark session
spark = SparkSession.builder.appName("EAI-Taxonomy-STEM-w-DCLM").getOrCreate()

# Load the dataset using the "huggingface" data source
df = spark.read.format("huggingface").load("EssentialAI/eai-taxonomy-stem-w-dclm")

# Basic dataset exploration
print(f"Dataset shape: {df.count()} rows, {len(df.columns)} columns")
df.show(10)
df.printSchema()

# Load only specific columns for efficiency
df_subset = (
    spark.read.format("huggingface")
    .option("columns", '["column1", "column2"]')  # Replace with actual column names
    .load("EssentialAI/eai-taxonomy-stem-w-dclm")
)

# Run SQL queries on the dataset
df.createOrReplaceTempView("eai_taxonomy_stem_w_dclm_dataset")
result = spark.sql("""
    SELECT COUNT(*) as total_examples
    FROM eai_taxonomy_stem_w_dclm_dataset
""")
result.show()

Using Daft

Daft provides a modern DataFrame library optimized for machine learning workloads. You can load the dataset directly from Hugging Face:

python
import daft

# Load the entire dataset
df = daft.read_parquet("hf://datasets/EssentialAI/eai-taxonomy-stem-w-dclm")

# Basic exploration
print("Dataset schema:")
df.schema()

print("First 5 rows:")
df.show(5)

If you need to access private datasets or use authentication:

python
import daft
from daft.io import IOConfig, HTTPConfig

io_config = IOConfig(http=HTTPConfig(bearer_token="your_token"))
df = daft.read_parquet("hf://datasets/EssentialAI/eai-taxonomy-stem-w-dclm", io_config=io_config)

Installation Requirements

Make sure you have the required libraries installed:

bash
# For Hugging Face datasets
pip install datasets

# For PySpark with Hugging Face integration
pip install pyspark_huggingface

# For Daft
pip install daft

πŸ“œ License

Essential-Web-v1.0 contributions are made available under the ODC attribution license; however, users should also abide by the Common Crawl - Terms of Use. We do not alter the license of any of the underlying data.

πŸ“ Citation

bibtex
@misc{ai2025essentialwebv1024ttokens,
      title={Essential-Web v1.0: 24T tokens of organized web data}, 
      author={Essential AI and : and Andrew Hojel and Michael Pust and Tim Romanski and Yash Vanjani and Ritvik Kapila and Mohit Parmar and Adarsh Chaluvaraju and Alok Tripathy and Anil Thomas and Ashish Tanwer and Darsh J Shah and Ishaan Shah and Karl Stratos and Khoi Nguyen and Kurt Smith and Michael Callahan and Peter Rushton and Philip Monk and Platon Mazarakis and Saad Jamal and Saurabh Srivastava and Somanshu Singla and Ashish Vaswani},
      year={2025},
      eprint={2506.14111},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2506.14111}, 
}