CoolFace
20 results

LLaDA

FredyRivera-dev /LLaDA-Sample-10BT Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.text-generation1B<n<10B2 likes1.1k downloads3mo agoHugging Faced3LLM /trajectory_data_llada_32 d3LLM Trajectory Data Paper | GitHub | Blog | Demo This repository contains the pseudo-trajectory distillation data presented in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation". Introduction d3LLM (pseuDo-Distilled Diffusion LLM) is a novel framework for building ultra-fast diffusion language models with negligible accuracy degradation. This dataset provides the pseudo-trajectory data extracted from teacher models, enabling the… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_llada_32.texttext-generation10K<n<100K2 likes825 downloads4mo agoHugging Faceshubhpatel6001 /llada-planner-2026-02-050 likes597 downloads2d agoHugging FaceFredyRivera-dev /LLaDA-Sample-ES Dataset: LLaDA-Sample-ES Base: crscardellino/spanish_billion_words Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~ 652,089 Shards: 65… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-ES.text-generation100M<n<1B1 likes384 downloads3mo agoHugging FaceNielOk /LLaDA_8B_folio_collected_logits_dataset LLaDA 8B FOLIO Collected Logits Dataset This dataset contains logits collected from the GSAI-ML/LLaDA-8B-Instruct model on the training set of the FOLIO dataset. It is intended for use in latent decomposition of token dynamics using sparse autoencoders, to enable semantic interpretability in masked denoising diffusion inference, specifically for use with the LLaDA model. Contents For each prompt, we record the following fields: prompt_id: unique prompt directory… See the full description on the dataset page: https://huggingface.co/datasets/NielOk/LLaDA_8B_folio_collected_logits_dataset.texttext-generation100K<n<1M0 likes227 downloads1y agoHugging Facedevichand /llada-pcfg-traces0 likes207 downloads2d agoHugging Face