datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.kr-vocab-synth-20260327-v4
KoalaReads German Vocabulary Trainer - Synthetic Conversations
Overview
This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning.
The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference)
levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary acquisition.… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v4.kr-vocab-synth-20260327-v2
KoalaReads German Vocabulary Trainer - Synthetic Conversations
📚 Overview
This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning.
The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference)
levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v2.kr-vocab-synth-20260327-v3
KoalaReads German Vocabulary Trainer - Synthetic Conversations
Overview
This dataset contains high-quality synthetic conversational data designed for training AI-powered vocabulary trainers for German language learning.
The conversations are carefully structured to follow proven pedagogical principles and the CEFR (Common European Framework of Reference)
levels A1-C2, making them ideal for fine-tuning language models that need to teach and assess vocabulary acquisition.… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v3.kr-vocab-synth-20260327-v1
KoalaReads German Vocabulary Trainer - Synthetic Conversations
Dataset Description
This dataset contains synthetic conversational data for training a vocabulary trainer model for German language learning.
The conversations follow CEFR (Common European Framework of Reference) levels A1-C2 and include various exercise types
designed to reinforce vocabulary acquisition through structured pedagogical interactions.
Related Datasets
This dataset is generated from… See the full description on the dataset page: https://huggingface.co/datasets/koalareads/kr-vocab-synth-20260327-v1.conflict_pairs
Conflict Pairs Dataset
This dataset contains conflict-pairs generated from the UltraFeedback dataset.
It was created by filtering for high-divergence, decent-quality response pairs and using a local LLM via vLLM to infer contrasting instructions that could have produced each response.
Pipeline
Load UltraFeedback (64k prompts × 4 responses each)
Pre-filter to high-divergence, decent-quality pairs
Use a local LLM to infer contrasting instructions from each pair
Parse and… See the full description on the dataset page: https://huggingface.co/datasets/Koalacrown/conflict_pairs.
