CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amd /InstructGpt-educational LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-educational.texttext-generation100K<n<1M3 likes126 downloads7mo agoHugging Face02north /scandinavian-educational-annotations Scandinavian Educational Annotations Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash. texttext-generation100K<n<1M3 likes100 downloads2y agoHugging Face03dev-jonathanb /cs50-educational-rag CS50 Pedagogical RAG Dataset 📜 Dataset Description This repository contains the data artifacts for the undergraduate thesis, which explores the use of a pedagogical chatbot with Retrieval-Augmented Generation (RAG) for Harvard's CS50: Introduction to Computer Science course. The project involved several stages of data processing, from raw content collection to the generation and curation of a high-quality evaluation dataset. To ensure full transparency and… See the full description on the dataset page: https://huggingface.co/datasets/dev-jonathanb/cs50-educational-rag.tabularquestion-answeringn<1K0 likes52 downloads1y agoHugging Face04wangzihaogithub /job-educational-parser-dataset-08-0-0805 Job Educational Parser Dataset 招聘领域的岗位与学历要求数据集。 输入:岗位描述 -> 输出:学历要求 Splits train: 19w_0701.csv (约 19 万条) test: 2w_0716.csv (约 2 万条) validation: 4w_0708.csv (约 4 万条) 每条数据至少包含字段: user: 职位描述 assistant: 要求的学历(如 "博士、硕士、本科"),遵循从高到低 由 @wangzihaogithub 创建。 tabulartext-generation100K<n<1M0 likes46 downloads1y agoHugging Face05Srinivasmec26 /Multidisciplinary-Educational-Summaries Knowledge Summarization Dataset Overview 100 structured knowledge summaries across STEM, social sciences, and humanities. Features 70% Indian-centric content, 25% European perspectives, and 5% other Asian contexts for balanced representation. Dataset Structure { "input": "Long-form text", "output": { "type": "summary", "topic": "Subject name", "difficulty": "beginner/intermediate/advanced", "points": ["Key point 1", "Key point 2"] } }… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Multidisciplinary-Educational-Summaries.texttoken-classificationn<1K1 likes32 downloads1y agoHugging Face06Somtharu181coder /educational_domain_dataset Nepali Grounded Education QA (OpenHermes-format) A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about student enrollment statistics from Nepal's Ministry of Education. Every answer is anchored to a real numeric value pulled from government open data — nothing in the answers is model-hallucinated. Dataset Summary Rows 611 Language Nepali (Devanagari script) Format ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.textquestion-answeringn<1K0 likes30 downloads1mo agoHugging Face07ray-2908 /educational-rewriter-dataset Educational Content Rewriter Dataset Dataset Summary A synthetic dataset of 846 educational rewrite pairs generated for fine-tuning language models to rewrite confusing educational content in six targeted modes. Source passages were collected from Wikipedia and arXiv, and rewrites were generated using the Claude API following the Alpaca data generation methodology. This dataset was created as part of Phase 5 of the NLP/LLM Learning Journey project and is used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/ray-2908/educational-rewriter-dataset.texttext-generationn<1K0 likes23 downloads5mo agoHugging Face08roneymatusp /british-educational-prompts British Educational Prompts Dataset Dataset Description A curated collection of 8,086 educational prompts optimized for British International Schools, covering all levels from Pre-Prep to IBDP. Dataset Summary Total Examples: 8,086 Language: British English only Blocked Portuguese: 13,390 examples removed Focus: British curriculum and educational standards Supported Tasks Prompt optimization Educational content generation British English… See the full description on the dataset page: https://huggingface.co/datasets/roneymatusp/british-educational-prompts.texttext-generation1K<n<10K1 likes10 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.