CoolFace
Datasetpublic

m-a-p/FineFineWeb-sample

FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
4likes57kdownloads
Dataset Card

FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus

arXiv: Coming Soon

Project Page: Coming Soon

Blog: Coming Soon

Data Statistics

Domain (#tokens/#samples)Iteration 1 TokensIteration 2 TokensIteration 3 TokensTotal TokensIteration 1 CountIteration 2 CountIteration 3 CountTotal Count
aerospace5.77B261.63M309.33M6.34B910000068850561103410399539
agronomy13.08B947.41M229.04M14.26B15752828271179064940419114022
artistic178.25B5.79B3.75B187.80B314279703161135129957104340350319
astronomy5.20B134.39M54.66M5.38B75965213576471458328100000
atmospheric_science2.80B102.04M259.25M3.16B57095372677895259696503295
automotive36.72B436.34M911.65M38.07B602396791166729153588262942290
beauty19.10B671.88M1.01B20.78B347873761808382220181038797568
biology85.84B371.29M776.99M86.99B81413569995384135034883759301
celebrity9.63B706.41M4.22B14.56B198311881803788794924029584216
chemistry27.80B588.92M131.46M28.52B31188189149908532803833015312
christianity47.72B403.68M732.55M48.86B550131471349874202145858384479
civil_engineering8.85B1.27B402.91M10.52B13591632268394094074217216314
communication_engineering9.21B3.60B327.66M13.14B13001767595952674649519707788
computerscienceand_technology194.46B3.95B4.76B203.16B278420434102635218654255297338210
design96.58B3.80B450.00M100.82B190275603166535882090515209019706
dramaandfilm19.12B10.86B206.27M30.19B331174781844325956425152124988
economics205.01B1.23B2.63B208.87B26396508538740915505880273345056
electronic_science30.19B7.76B482.62M38.43B4274576712572747111560556434119
entertainment152.92B1.67B5.06B159.65B25693514458010819648023272384248
environmental_science56.98B1.48B920.77M59.37B845003933557056196673190024180
fashion18.72B977.27M264.01M19.96B534656283926500134698858739116
finance146.39B327.45M1.13B147.85B18779776412958933058801192152458
food56.10B136.32M978.91M57.22B964858386138753051981100151694
gamble30.12B696.52M158.48M30.98B2490903777054016416825843745
game43.47B2.36B2.68B48.51B656806994670033372070074071432
geography110.18B1.16B192.67M111.53B1616772143835932559447166072593
health191.20B427.93M18.43B210.06B215747152129121523975955241014322
history45.27B1.56B1.69B48.52B557104324167508346303363340973
hobby150.23B42.78B44.05B237.06B2766363628136089371407735429404990
hydraulic_engineering57.36M75.40M3.65M136.41M13507916329913453311831
instrument_science5.35B2.02B165.43M7.54B8307736290427446225611674266
journalismandmedia_communication440.98B21.00B1.55B463.53B645801807506576684909008701368483
landscape_architecture3.07B557.66M64.76M3.70B561314111384091665266918076
law128.58B455.19M2.38B131.42B16647320516609446145032174279181
library57.16B5.01B36.56M62.21B865923051044099115301497186310
literature71.07B7.01B67.53B145.61B711910751324780654760578139199459
materials_science17.79B1.11B303.66M19.20B22136519166337670838424508279
mathematics5.87B50.33M261.65M6.18B1013193317959265305010964575
mechanical_engineering86.13B1.24B129.96M87.49B1117788133201605428714115409132
medical140.03B813.46M4.97B145.81B14959463422664778527901160389012
mining_engineering7.26B206.05M529.02M8.00B55406312361454684586245234
movie13.09B639.20M124.67M13.86B22938808157757651188225028266
musicanddance15.42B10.38B618.46M26.42B2956655420233446199827251798272
news328.47B12.37B11.34B352.18B5085677683320670923482422565256899
nuclear_science559.05M79.89M78.79M717.72M7848471702821335981088727
ocean_science2.36B537.82M229.43M3.13B37000008530524257924978844
optical_engineering2.33B253.06M263.99M2.85B35108365350264003714446233
painting374.41M429.63M96.57M900.61M8757838242173362032036203
pet12.12B154.14M307.28M12.58B1962468845763577897020861293
petroleumandnaturalgasengineering950.08M515.05M121.56M1.59B16694478998602378432807150
philosophy47.99B121.26M335.77M48.44B50396964505275103040551932644
photo6.56B1.74B41.44M8.34B16194329390159817960720275534
physics21.56B372.21M191.17M22.12B2464037384350847375825957639
politics79.52B253.26M930.96M80.70B9740360310263152504127100934045
psychology51.53B688.50M2.56B54.78B588299171881452406666764778036
public_administration100.13B5.54B716.81M106.39B160247751106577681785347172690866
relationship21.87B3.69B129.60M25.69B28153321679477432126835269363
sociology76.34B3.59B8.88B88.82B106447186783689613040695127324777
sports118.64B379.18M1.79B120.80B17324363112867184212540178742889
statistics19.59B1.15B1.75B22.49B299587262746797339060636096129
systems_science24.58B11.30B163.99M36.05B328792491512075147000148470001
textile_science2.59B2.89B94.56M5.57B8018141802200145666816496810
topicality34.87M5.22M040.09M137789135060151295
transportation_engineering12.80B6.61B972.50M20.38B2359562411005933202781236629369
travel78.87B584.78M957.26M80.41B12725019518513422430704131532241
urban_planning12.13B2.93B53.24M15.12B20040937617610420196326419004
weapons_science80.62M3.32B140.89M3.54B21554456951543695416280239
Grand Total4010.76B206.51B208.02B4425.30B57817640554423879643119208606536072879

Data Construction Workflow

[image]

The data construction workflow can be summarized as follows:

  1. 1.Deduplicate: The FineWeb dataset is deduplicated using exact deduplication and MinHash techniques to remove redundant data.
  2. 2.URL Labeling: Root URLs from FineWeb are counted, and the top 1 million URLs are labeled using GPT-4. This step generates DoI (Domain-of-Interest) Coarse-Grained URLs and DoNI (Domain-of-Non-Interest) Coarse-Grained URLs as seed data sources.
  3. 3.Coarse Recall:

a. Based on the labeled root URLs, data is sampled for each domain.

b. The sampled data is labeled using Qwen2-7B-Instruct, producing 500K DoI Positive Data and 500K DoI Negative Data (note that for N>1 iterations, each 500K samples are composed of 250K sampled original seed data and 250K refined data after Fine Recall).

c. A binary FastText model is trained per domain using the labeled data.

d. The FastText model performs coarse recall on FineWeb, generating Coarse DoI Data.

  1. 1.Fine Recall:

a. The Coarse DoI Data is labeled using Qwen2-72B-Instruct to produce 100K DoI Positive Data and 50K DoI Negative Data, with the latter further augmented with 50K negative samples from earlier FastText training.

b. A BERT model is trained using this labeled data.

c. The BERT model performs fine recall on the Coarse DoI Data, producing a refined dataset, which is the DoI subset of FineFineWeb.

  1. 1.Coarse-Fine Recall Iteration: The workflow of coarse and fine recall iterates for 3 rounds with the following adjustments:

a. FastText is re-trained using updated seed data, which combines BERT-recalled samples, BERT-dropped samples, and previously labeled seed data.

b. The BERT model keeps frozen during subsequent iterations.

c. Steps for training FastText, coarse recall, and fine recall are repeated without re-labeling data with Qwen2-Instruct models.

Domain-Domain Similarity Analysis

  1. 1.Perform proportional weighted sampling of the domain subsets based on the sample size of each domain, with a total of 1 billion tokens sampled from the domain subsets.
  2. 2.Use the BGE-M3 model to compute the embeddings of the samples in each domain subset, referred to as domain embeddings.
  3. 3.Use the BGE-M3 model to compute the embeddings of the samples in each benchmark, referred to as benchmark embeddings (bench embeddings).
  4. 4.Calculate the MMD distance and the Wasserstein distance between the domain embeddings and the benchmark embeddings.

[image]

The results above reveal the following observations:

  1. 1.The two code-related benchmarks, MBPP and HumanEval, exhibit relatively large distances from nearly all domains, indicating that the proportion of code data in the training set is relatively small. Notably, their distance to the mathematics domain is comparatively smaller, suggesting a certain degree of overlap between mathematics data and code data.
  2. 2.Benchmarks such as Hellaswag, ARC, MMLU, and BoolQ have distances that are close to almost all domains, except for the gamble domain. This indicates that the samples in these benchmarks involve synergetic effects across multiple domains of knowledge, with a wide distribution.
  3. 3.GSM8K and TriviaQA show significant discrepancies with a small number of domains, suggesting that the distribution differences between domains are more pronounced for samples involving grade-school mathematics and fact-based question answering. Some domains contain a substantial amount of this type of data, while others do not.
  4. 4.The gamble domain exhibits substantial differences from other domains and has large distances from all benchmarks, indicating that pretraining data related to gambling provides limited benefits for these benchmarks.

Domain-Domain Duplication

Let \\(D1, D2, \dots, DN\\) represent \\(N\\) distinct domains, where we select top-20 URLs for each domain \\(Di\\), denoted as \\(\{U{i1}, U{i2}, \dots, U_{i20}\}\\),. The total set of URLs across all domains is represented as \\(\mathcal{U}\\), and the total number of URLs is \\(M = |\mathcal{U}|\\).

For each URL \\(Uk \in \mathcal{U}\\), the term frequency (TF) is defined as the proportion of \\(Uk\\) in the total set of URLs:

\\(\text{TF}(Uk) = \frac{\text{count}(Uk)}{M}\\)

where \\(\text{count}(Uk)\\) is the number of times \\(Uk\\) appears in \\(\mathcal{U}\\). Additionally, the document frequency \\(Kk\\) of \\(Uk\\) is the number of domains in which \\(U_k\\) appears. Based on this, the inverse document frequency (IDF) is calculated as:

\\(\text{IDF}(Uk) = \log(\frac{N}{Kk})\\)

The TF-IDF value for each URL \\(U{ij}\\) in a specific domain \\(Di\\) is then computed as:

\\(\text{TF-IDF}(U{ij}) = \text{TF}(U{ij}) \times \text{IDF}(U_{ij})\\)

[image]

Using the TF-IDF values of all URLs within a domain, the domain-domain duplicate rate can be analyzed by comparing the distribution of TF-IDF values across domains. If a domain has many URLs with high TF-IDF values, it indicates that the domain’s URLs are relatively unique and significant within the entire set of URLs. Conversely, if a domain has many URLs with low TF-IDF values, it suggests that the domain's URLs are more common across other domains. Analyzing these values helps assess how similar or redundant a domain's content is in relation to others based on its URL composition.

As shown in the figure, most domains have low duplication rates, except for topicality, pet, and atmospheric science.

Domain-Benchmark BPC-Acc Correlation

Experimental method: Using 28 models (see the paper), we first calculate BPC for all domains to obtain a model ranking \\(RD\\). Similarly, we compute scores across all benchmarks to obtain a model ranking \\(RM\\). We then calculate the Spearman correlation between \\(RD\\) and \\(RM\\).

[image]

  • For benchmarks like ARC, MMLU, GSM8K, HumanEval, and MBPP, STEM-related domains show higher correlation rankings, particularly mathematics, physics, and systems science.
  • For TriviaQA, which emphasizes factual knowledge over reasoning, domains rich in world knowledge such as literature, history, and library science demonstrate higher correlation rankings.

Bibtex

bibtex
@misc{
title={FineFineWeb: A Comprehensive Study on Fine-grained Domain Web Corpus},
url={[https://huggingface.co/datasets/m-a-p/FineFineWeb](https://huggingface.co/datasets/m-a-p/FineFineWeb)},
author = {M-A-P, Ge Zhang*, Xinrun Du*, Zhimiao Yu*, Zili Wang*, Zekun Wang, Shuyue Guo, Tianyu Zheng, Kang Zhu, Jerry Liu, Shawn Yue, Binbin Liu, Zhongyuan Peng, Yifan Yao, Jack Yang, Ziming Li, Bingni Zhang, Minghao Liu, Tianyu Liu, Yang Gao, Wenhu Chen, Xiaohuan Zhou, Qian Liu, Taifeng Wang+, Wenhao Huang+},
publisher={huggingface},
verision={v0.1.0},
month={December},
year={2024}
}