datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.Te-Reo-Maori
Te Reo Māori Multi-Format Training Dataset
Dataset Description
A 3GB multi-format training dataset for te reo Māori language models, containing approximately 1.5–3 million unique sentences generated using rule-based grammar with authentic Māori vocabulary.
⚠️ Important: This dataset is synthetically generated. It contains programmatically constructed Māori sentences using real vocabulary and grammatical patterns, not natural human-written text. See Limitations… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Te-Reo-Maori.LOOM
LOOM: Language-Only Operational Microworlds
LOOM is a synthetic natural-language reasoning dataset designed to teach language models the deep structures behind code and math without exposing source code, formal equations, or symbolic programming syntax.
Instead of showing code or math notation, LOOM trains models on ordinary-language microworlds where the hidden logic is algorithmic: state changes, causal chains, conditionals, invariants, iteration, and reverse reasoning.
The… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/LOOM.Coding-Corpus-Bench
Coding-Corpus-Bench
A benchmark dataset for evaluating language-semantics reasoning across systems programming and low-level programming languages.
Overview
Coding-Corpus-Bench contains 100 curated programming-language questions designed to test whether a model can reason precisely about language semantics rather than rely on superficial pattern matching or observed behavior.
The benchmark covers:
Rust
Go
C
C++
Zig
V
CUDA
Questions focus on subtle semantic rules… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Coding-Corpus-Bench.OpenOrca-gugugo-ko
OpenOrca 한국어 번역 데이터셋
Gugugo-koen-7B-V1.1을 이용하여 OpenOrca데이터셋을 번역하고 있습니다.
번역 진행상황은 아래를 참고해 주십시오.
진행상황
GPT4 생성물 약 100만 개 중 약 64만 개 번역완료
GPT3.5 생성물 약 350만 개 중 약 159만 개 번역완료
데이터셋 사용 후 출처표기는 제작자에게 큰 힘이 됩니다.
Original dataset card: OpenOrca
🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has… See the full description on the dataset page: https://huggingface.co/datasets/squarelike/OpenOrca-gugugo-ko.
