datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
forge-3b-pretrain-data
FORGE-3B Pretraining Data
Tokenized and packed pretraining data for the FORGE-3B language model.
Stats
Total tokens: 51.4070B
Domains: 10/10
Sequence length: 2048 tokens
Format: .npy shards of shape (N, 2048) with dtype uint32
Tokenizer: CRAYON (xerv-crayon, standard profile)
Domain Breakdown
Domain
Weight
Tokens (B)
Status
fineweb_edu
30%
15.0008
✓
thestack
16%
8.0011
✓
wikipedia
8%
4.2791
✓
openwebmath
8%
3.9654
✓
books
7%… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-pretrain-data.dataforge-sft-trajectories
DataForge SFT Trajectories
This dataset contains chunk-level expert_v1, versioned expert_v2,
inferability-audited expert_v3, and contract-repair expert_v4
supervised-fine-tuning records for the DataForge warmup model. The current
milestone is built from
split-safe dirty/clean CSV diffs (oracle_from_clean_diff) so model training is
anchored to audited labels rather than teacher guesses.
The earlier v0-smoke checkpoint proved the Kaggle-to-Hugging-Face handoff. It
is not a… See the full description on the dataset page: https://huggingface.co/datasets/Praneshrajan15/dataforge-sft-trajectories.forge-3b-sft-data
FORGE-3B SFT Data
Tokenized, chat-templated, loss-masked SFT data for the FORGE-3B language model.
Stats
Total tokens (incl. pad): 1.4007B
Domains: 6/6
Sequence length: 4096 tokens
Format: .npz shards with input_ids (uint32) and loss_mask (uint8), shape (N, 4096)
Chat template: <|SYS|>...<|/SYS|> <|USR|>...<|/USR|> <|ASST|>...<|/ASST|>
Tokenizer: CRAYON (xerv-crayon, standard profile) or fallback HF tokenizer
Domain Breakdown
Domain
Weight… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-sft-data.dataforge2-dataset
Real Estate Faq Routing Benchmark Bilingual AI Buyer Fine Tuning Dataset
DataForge commercial dataset pack.
10000 rows with 9.40/10.
Formats: CSV.
Why this pack
Built for buyer intent and practical AI deployment
Fine-tuning friendly structure for fast iteration
Includes machine-readable metadata for immediate validation
Included files
Dataset bundle (CSV/JSONL and optional Parquet)
Schema and quality artifacts
Readme for deployment and usage guidance… See the full description on the dataset page: https://huggingface.co/datasets/lucasgd123/dataforge2-dataset.forge-3b-dpo-data
FORGE-3B DPO Preference Data
Tokenized (prompt, chosen, rejected) preference triples for DPO post-training
of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2.
This is data preparation output only — no model was trained to produce this.
Stats
Total pairs: 0 (paper target: ~200,000)
Domains: 0/4
Context length: 4096 tokens (paper Appendix A.2, DPO block)
Format: unpacked — one (prompt, chosen, rejected) triple per training example
Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.s1K-1.1-dataforge-testing-20251216-123019
Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-123019
Dataset Summary
Synthetic data generated by DataForge:
Model: Qwen/Qwen3-4B-Instruct-2507 (main)
Source dataset: simplescaling/s1K-1.1 (train split).
Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768
Speculative decoding: disabled
System prompt: None
User prompt: Column question
The run produced 1,000 samples and generated 3,406,836 (~3.4M) tokens.
You can… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-123019.amir-patch-forge-data
PatchForge data
This dataset contains the large data/ directory for the PatchForge project branch:
https://github.com/PGCodeLLM/CodeFoundry/tree/amir-patch-forge
The data is stored as one .tar.zst archive per top-level data/ subdirectory.
Each archive preserves paths like data/<directory>/... when extracted.
Restore
hf download PGCodeLLM/amir-patch-forge-data --repo-type dataset --local-dir patchforge-data
cd patchforge-data
sha256sum -c SHA256SUMS
for f in… See the full description on the dataset page: https://huggingface.co/datasets/PGCodeLLM/amir-patch-forge-data.s1K-1.1-dataforge-testing-20251216-142704
Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-142704
Dataset Summary
Synthetic data generated by DataForge:
Model: Qwen/Qwen3-4B-Instruct-2507 (main)
Source dataset: simplescaling/s1K-1.1 (train split).
Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768
Speculative decoding: disabled
System prompt: None
User prompt: Column question
The run produced 10 samples and generated 30,174 tokens.
You can load the… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-142704.dataforge-cleaned
