CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01duongttr /vi-dataset-for-pretrain Dataset Card for "vi-dataset-for-pretrain" This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc. The dataset consists of: vietgpt/covid_19_news_vi hieunguyen1053/binhvq-news-corpus oscar (unshuffled_deduplicated_vi) vietgpt/wikipedia_vi Dataset info Splits N.o examples Size Train 23,891,116 77.36 GB Validation 1,257,428 4.06 GB Total 25,148,544 81.43 GB texttext-generation10M<n<100M5 likes572 downloads3y agoHugging Face02duoduoyeah /SimpleStories 📘📕 SimpleStories 📙📗 SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository. If you'd like to commission other languages or story formats, feel free to send mail. When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/SimpleStories.tabulartext-generation1M<n<10M0 likes70 downloads9mo agoHugging Face03DuoNeural /ml-ai-engineer-sft DuoNeural ML/AI Engineer SFT Dataset A synthetic instruction-tuning dataset for training an LLM to be a useful pairing partner on ML/AI engineering work — debugging training runs, reasoning about architecture and infra choices, reviewing experiment design, and explaining core ML concepts with the specificity of someone who's actually run the experiments. Why this dataset exists Most general instruction-tuning data treats ML engineering questions the same as any… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/ml-ai-engineer-sft.texttext-generation1K<n<10K1 likes61 downloads3mo agoHugging Face04duoduoyeah /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/TinyStories.texttext-generation1M<n<10M0 likes47 downloads9mo agoHugging Face05DuoNeural /cot-reasoning-2k DuoNeural CoT Reasoning Dataset (2K) A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces. Benchmark Results Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090): Metric Baseline Post-SFT Δ Absolute Δ Relative GSM8K (flexible-extract) 0.3177 0.4890 +17.1pp +53.9% GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.texttext-generation1K<n<10K1 likes36 downloads5mo agoHugging Face06duoduoyeah /gsm8k-promise GSM8K-Promise A transform of openai/gsm8k (the main subset) into a "promise" format: the original chain-of-thought is kept in answer, and an added promise_answer rewrites the solution so the model emits only operations and operands — every computed value is a var that an external tool resolves. The model itself never computes, compares, or rounds, so the reasoning is machine-translation-safe (numbers don't get mangled in translation) and tool-verifiable. train: 7,473 rows ·… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/gsm8k-promise.texttext-generation1K<n<10K0 likes35 downloads3mo agoHugging Face07onekq-ai /WebApp1K-Duo-React-Generations WebApp1K-Duo-React-Generations A comprehensive evaluation dataset containing React component generations from 32 state-of-the-art AI models on 1,000 paired web application scenarios. Dataset Description This dataset extends the original WebApp1K-Duo-React benchmark by including actual code generations from major AI models. Each row contains a paired web application scenario (combining two functionalities) along with generated React components from 32 different models and… See the full description on the dataset page: https://huggingface.co/datasets/onekq-ai/WebApp1K-Duo-React-Generations.texttext-generation1K<n<10K0 likes28 downloads1y agoHugging Face08duongttr /LPM-24-extendgated LPM-24 dataset This dataset was used in paper Mol2Lang-VLM: Vision- and Text-Guided Generative Pre-trained Language Models for Advancing Molecule Captioning through Multimodal Fusion DOI: https://doi.org/10.18653/v1/2024.langmol-1.12 GitHub: https://github.com/nhattruongpham/mol-lang-bridge/tree/mol2lang This dataset contains: SELFIES strings (converted by selfies package) SMILES strings Molecular images (converted by RDKit) Molecular captions. Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/duongttr/LPM-24-extend.imagetext-generation100K<n<1M1 likes13 downloads2y agoHugging Face09duongttr /chebi-20-newgatedThis dataset was used in paper Mol2Lang-VLM: Vision- and Text-Guided Generative Pre-trained Language Models for Advancing Molecule Captioning through Multimodal Fusion DOI: https://doi.org/10.18653/v1/2024.langmol-1.12 GitHub: https://github.com/nhattruongpham/mol-lang-bridge/tree/mol2lang This dataset contains: SELFIES strings (converted by selfies package) SMILES strings Molecular images (converted by RDKit) Molecular captions. Citation If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/duongttr/chebi-20-new.imagetext-generation10K<n<100K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.