CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Phase-Technologies /forge-3b-pretrain-data FORGE-3B Pretraining Data Tokenized and packed pretraining data for the FORGE-3B language model. Stats Total tokens: 51.4070B Domains: 10/10 Sequence length: 2048 tokens Format: .npy shards of shape (N, 2048) with dtype uint32 Tokenizer: CRAYON (xerv-crayon, standard profile) Domain Breakdown Domain Weight Tokens (B) Status fineweb_edu 30% 15.0008 ✓ thestack 16% 8.0011 ✓ wikipedia 8% 4.2791 ✓ openwebmath 8% 3.9654 ✓ books 7%… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-pretrain-data.text-generation10B<n<100B0 likes3.6k downloads3mo agoHugging Face02Praneshrajan15 /dataforge-sft-trajectories DataForge SFT Trajectories This dataset contains chunk-level expert_v1, versioned expert_v2, inferability-audited expert_v3, and contract-repair expert_v4 supervised-fine-tuning records for the DataForge warmup model. The current milestone is built from split-safe dirty/clean CSV diffs (oracle_from_clean_diff) so model training is anchored to audited labels rather than teacher guesses. The earlier v0-smoke checkpoint proved the Kaggle-to-Hugging-Face handoff. It is not a… See the full description on the dataset page: https://huggingface.co/datasets/Praneshrajan15/dataforge-sft-trajectories.tabulartext-generation1K<n<10K1 likes278 downloads4mo agoHugging Face03Phase-Technologies /forge-3b-sft-data FORGE-3B SFT Data Tokenized, chat-templated, loss-masked SFT data for the FORGE-3B language model. Stats Total tokens (incl. pad): 1.4007B Domains: 6/6 Sequence length: 4096 tokens Format: .npz shards with input_ids (uint32) and loss_mask (uint8), shape (N, 4096) Chat template: <|SYS|>...<|/SYS|> <|USR|>...<|/USR|> <|ASST|>...<|/ASST|> Tokenizer: CRAYON (xerv-crayon, standard profile) or fallback HF tokenizer Domain Breakdown Domain Weight… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-sft-data.text-generation1B<n<10B0 likes111 downloads3mo agoHugging Face04lucasgd123 /dataforge2-dataset Real Estate Faq Routing Benchmark Bilingual AI Buyer Fine Tuning Dataset DataForge commercial dataset pack. 10000 rows with 9.40/10. Formats: CSV. Why this pack Built for buyer intent and practical AI deployment Fine-tuning friendly structure for fast iteration Includes machine-readable metadata for immediate validation Included files Dataset bundle (CSV/JSONL and optional Parquet) Schema and quality artifacts Readme for deployment and usage guidance… See the full description on the dataset page: https://huggingface.co/datasets/lucasgd123/dataforge2-dataset.tabulartext-generation100K<n<1M0 likes98 downloads7mo agoHugging Face05Phase-Technologies /forge-3b-dpo-data FORGE-3B DPO Preference Data Tokenized (prompt, chosen, rejected) preference triples for DPO post-training of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2. This is data preparation output only — no model was trained to produce this. Stats Total pairs: 0 (paper target: ~200,000) Domains: 0/4 Context length: 4096 tokens (paper Appendix A.2, DPO block) Format: unpacked — one (prompt, chosen, rejected) triple per training example Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.texttext-generation100K<n<1M0 likes80 downloads3mo agoHugging Face06lewtun /s1K-1.1-dataforge-testing-20251216-123019 Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-123019 Dataset Summary Synthetic data generated by DataForge: Model: Qwen/Qwen3-4B-Instruct-2507 (main) Source dataset: simplescaling/s1K-1.1 (train split). Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: None User prompt: Column question The run produced 1,000 samples and generated 3,406,836 (~3.4M) tokens. You can… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-123019.texttext-generation1K<n<10K0 likes50 downloads9mo agoHugging Face07PGCodeLLM /amir-patch-forge-data PatchForge data This dataset contains the large data/ directory for the PatchForge project branch: https://github.com/PGCodeLLM/CodeFoundry/tree/amir-patch-forge The data is stored as one .tar.zst archive per top-level data/ subdirectory. Each archive preserves paths like data/<directory>/... when extracted. Restore hf download PGCodeLLM/amir-patch-forge-data --repo-type dataset --local-dir patchforge-data cd patchforge-data sha256sum -c SHA256SUMS for f in… See the full description on the dataset page: https://huggingface.co/datasets/PGCodeLLM/amir-patch-forge-data.text-generation0 likes38 downloads3mo agoHugging Face08lewtun /s1K-1.1-dataforge-testing-20251216-142704 Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-142704 Dataset Summary Synthetic data generated by DataForge: Model: Qwen/Qwen3-4B-Instruct-2507 (main) Source dataset: simplescaling/s1K-1.1 (train split). Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: None User prompt: Column question The run produced 10 samples and generated 30,174 tokens. You can load the… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-142704.texttext-generationn<1K0 likes11 downloads9mo agoHugging Face09ReDiX /dataforge-cleanedgatedtexttext-generation10K<n<100K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.