CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gugu8 /Scraped-Datagated Scraped-Data — Building to 1 TB A continuously-growing, cleaned text corpus scraped, parsed, cleaned, losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline. Current repo size: ~26 GB → target: 1 TB (incremental batch uploads). Current files File Size Contents corpus_batch1.clean.zip 4.80 GB ~2.9M Wikipedia docs (stored, from first scrape) corpus_batch2.clean.zip 4.85 GB ~2.5M docs corpus_batch3.clean.zip 4.91 GB ~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.texttext-generation100M<n<1B0 likes1.8k downloads19d agoHugging Face02Gugu8 /Te-Reo-Maori Te Reo Māori Multi-Format Training Dataset Dataset Description A 3GB multi-format training dataset for te reo Māori language models, containing approximately 1.5–3 million unique sentences generated using rule-based grammar with authentic Māori vocabulary. ⚠️ Important: This dataset is synthetically generated. It contains programmatically constructed Māori sentences using real vocabulary and grammatical patterns, not natural human-written text. See Limitations… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Te-Reo-Maori.text10M<n<100M0 likes71 downloads18d agoHugging Face03Gugu8 /LOOM LOOM: Language-Only Operational Microworlds LOOM is a synthetic natural-language reasoning dataset designed to teach language models the deep structures behind code and math without exposing source code, formal equations, or symbolic programming syntax. Instead of showing code or math notation, LOOM trains models on ordinary-language microworlds where the hidden logic is algorithmic: state changes, causal chains, conditionals, invariants, iteration, and reverse reasoning. The… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/LOOM.text1M<n<10M0 likes52 downloads1mo agoHugging Face04Gugu8 /Coding-Corpus-Benchgated Coding-Corpus-Bench A benchmark dataset for evaluating language-semantics reasoning across systems programming and low-level programming languages. Overview Coding-Corpus-Bench contains 100 curated programming-language questions designed to test whether a model can reason precisely about language semantics rather than rely on superficial pattern matching or observed behavior. The benchmark covers: Rust Go C C++ Zig V CUDA Questions focus on subtle semantic rules… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Coding-Corpus-Bench.textn<1K0 likes43 downloads27d agoHugging Face05squarelike /OpenOrca-gugugo-ko OpenOrca 한국어 번역 데이터셋 Gugugo-koen-7B-V1.1을 이용하여 OpenOrca데이터셋을 번역하고 있습니다. 번역 진행상황은 아래를 참고해 주십시오. 진행상황 GPT4 생성물 약 100만 개 중 약 64만 개 번역완료 GPT3.5 생성물 약 350만 개 중 약 159만 개 번역완료 데이터셋 사용 후 출처표기는 제작자에게 큰 힘이 됩니다. Original dataset card: OpenOrca 🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has… See the full description on the dataset page: https://huggingface.co/datasets/squarelike/OpenOrca-gugugo-ko.texttext-classification1M<n<10M35 likes40 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.