CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gugu8 /Scraped-Datagated Scraped-Data — Building to 1 TB A continuously-growing, cleaned text corpus scraped, parsed, cleaned, losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline. Current repo size: ~26 GB → target: 1 TB (incremental batch uploads). Current files File Size Contents corpus_batch1.clean.zip 4.80 GB ~2.9M Wikipedia docs (stored, from first scrape) corpus_batch2.clean.zip 4.85 GB ~2.5M docs corpus_batch3.clean.zip 4.91 GB ~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.texttext-generation100M<n<1B0 likes1.8k downloads19d agoHugging Face02puschinka /Gugager 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/puschinka/Gugager.texttoken-classification10K<n<100K0 likes50 downloads27d agoHugging Face03Gugu8 /Pattern-Recognition Pattern Completion Dataset A 30 GB synthetic dataset of numeric sequence‑completion prompts and their next values, designed to teach large language models how to recognize and extrapolate patterns. Each row contains a prompt (the sequence with a ? indicating the missing next element) and a completion (the correct next number). Dataset Structure Format: CSV (no header row) Columns: prompt – "Find the next number in the sequence: a,b,c,... ,?" completion – the… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Pattern-Recognition.texttext-generation100M<n<1B0 likes48 downloads2mo agoHugging Face04squarelike /OpenOrca-gugugo-ko OpenOrca 한국어 번역 데이터셋 Gugugo-koen-7B-V1.1을 이용하여 OpenOrca데이터셋을 번역하고 있습니다. 번역 진행상황은 아래를 참고해 주십시오. 진행상황 GPT4 생성물 약 100만 개 중 약 64만 개 번역완료 GPT3.5 생성물 약 350만 개 중 약 159만 개 번역완료 데이터셋 사용 후 출처표기는 제작자에게 큰 힘이 됩니다. Original dataset card: OpenOrca 🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has… See the full description on the dataset page: https://huggingface.co/datasets/squarelike/OpenOrca-gugugo-ko.texttext-classification1M<n<10M35 likes40 downloads3y agoHugging Face05gugett /DeepMath-103K DeepMath-103K 🔥 News May 8, 2025: We found that 48 samples contained hints that revealed the answers. The relevant questions have now been revised to remove the leaked answers. April 14, 2025: We release DeepMath-103K, a large-scale dataset featuring challenging, verifiable, and decontaminated math problems tailored for RL… See the full description on the dataset page: https://huggingface.co/datasets/gugett/DeepMath-103K.texttext-generation100K<n<1M0 likes32 downloads2mo agoHugging Face06gugarosa /synthetic-pretraining-transformers-v1 Synthetic Pre-training Transformers v1.0.0 Dataset Description This is a synthetic pre-training dataset generated from transformer architecture patterns. It contains paraphrased, augmented, and interpolated content derived from validated seed data about neural sequence modeling and attention mechanisms. Dataset Summary Total Samples: 100 Total Tokens: 6,084 Average Tokens per Sample: 60.84 Format: Parquet Version: 1.0.0 License: CC-BY-4.0 Supported… See the full description on the dataset page: https://huggingface.co/datasets/gugarosa/synthetic-pretraining-transformers-v1.tabulartext-generationn<1K0 likes26 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.