datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medium_turkish_math_reasoning
Dataset Summary
The Medium Turkish Math Reasoning dataset is Phase 2 of a curriculum learning pipeline to teach compact models multi-step reasoning in Turkish. It includes moderately difficult math problems involving multiple reasoning steps, such as two-part arithmetic, comparisons, and logical reasoning.
Use Case
This dataset is ideal for:
Continuing SFT after foundational training with simpler problems.
Bridging the gap between basic arithmetic and complex GSM8K-style… See the full description on the dataset page: https://huggingface.co/datasets/erayalp/medium_turkish_math_reasoning.CCNA_medium
CCNA_Medium
Dataset Description
CCNA_Medium is a domain-specific question-answer dataset focused on Cisco networking and CCNA-level concepts.
The dataset contains 23.3k networking questions and answers designed for:
CCNA study and exam preparation
Networking education assistants
Retrieval-Augmented Generation (RAG)
Domain-specific LLM fine-tuning
Evaluation of networking knowledge in language models
Dataset Statistics
Split
Samples… See the full description on the dataset page: https://huggingface.co/datasets/Rzkoohi/CCNA_medium.medium_512_1k_tokens_prompts
Medium 512-1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ By using this dataset you agree to our Terms of Use.
Overview
703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer.
Statistics
Rows
Token range
File size
Format
703
512 – 1 000
2.9 MB
CSV
Use-cases
Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.unpredictable_rated-mediumThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.pyra_mediumFiltered dataset of https://huggingface.co/datasets/LLM-EDA/pyra for RL. Keep only code more than 50 lines. Check https://github.com/CatIIIIIIII/VeriPrefer for usage.
gpt-oss-reasoning-ru-mediumНовая версия датасета на 25к строк
mn_business_benchmark_dataset_medium
mn_business_benchmark_dataset_10000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics, стратегийн 10000 мөртэй синтетик benchmark dataset.
Schema
id: 1-ээс 10000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: бизнесийн сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG card дүрслэл
Generated deterministically by… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_medium.
