CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jobs-git /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset = load_dataset("jat-project/jat-dataset"… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/jat-dataset.reinforcement-learning0 likes711 downloads2y agoHugging Face02LiXiang12 /github-code-fontend-lang github-code fontend code Dwonload 方式一 huggingface-cli download --resume-download LiXiang12/github-code-fontend-lang --include "*/*.zip" --repo-type dataset --local-dir github_code 方式二 进入Files and versions/data直接下载zip文件 数据统计 textquestion-answering10M<n<100M2 likes555 downloads2y agoHugging Face03cabbage972 /GitChameleon-2.0 GitChameleon 2.0 GitChameleon 2.0 is an AI coding benchmark comprising 328 Python-based problems conditioned on specific versions of popular libraries for scientific computing and web development. It evaluates whether AI code generation models can correctly use library APIs as they existed at a particular version — a challenging test of version-specific knowledge. Note: This is GitChameleon 2.0, a distinct and newer work from the original GitChameleon benchmark. Please do not… See the full description on the dataset page: https://huggingface.co/datasets/cabbage972/GitChameleon-2.0.texttext-generationn<1K2 likes194 downloads5mo agoHugging Face04Modotte /Bhagwat-Gita-Infinity Bhagwat‑Gita‑Infinity By Modotte [!WARNING]Note: The dataset viewer table is not providing accurate information about the dataset structure and contents. Please manually check one file from the dataset (e.g., via raw file views on Hugging Face) or refer to the Example Entries below for correct details. Dataset Summary Bhagwat-Gita-Infinity, released by Modotte with main contributor Parveshiiii (Parvesh Rawal), is a meticulously structured dataset… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/Bhagwat-Gita-Infinity.text-classificationn<1K10 likes162 downloads8mo agoHugging Face05jobs-git /OpenMathReasoning OpenMathReasoning OpenMathReasoning is a large-scale math reasoning dataset for training large language models (LLMs). This dataset contains 540K unique mathematical problems sourced from AoPS forums, 3.2M long chain-of-thought (CoT) solutions 1.7M long tool-integrated reasoning (TIR) solutions 566K samples that select the most promising solution out of many candidates (GenSelect) We used Qwen2.5-32B-Instruct to preprocess problems, and DeepSeek-R1 and QwQ-32B to generate… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/OpenMathReasoning.question-answering1M<n<10M0 likes153 downloads1y agoHugging Face06JDhruv14 /Bhagavad-Gita-QA Bhagavad-Gita-QA-Multilingual Dataset Summary Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework. This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita-QA.tabularquestion-answering10K<n<100K5 likes122 downloads1y agoHugging Face07the-homeless-god /git-history-mcq-ru git-history-mcq-ru 805 вопросов с вариантами ответа по истории трёх открытых репозиториев (digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс 8 672 ответа пяти моделей и 4 878 разборов этих ответов. Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю. Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.tabularmultiple-choice10K<n<100K0 likes114 downloads22d agoHugging Face08jtatman /python-github-code-instruct-filtered-5k Dataset Card for "python-github-code-instruct-filtered-5k" This fine dataset tomekkorbak/python-github-code, filtered by scores greater than 0.03. Feedback and additional columns generated through OpenAI and Cohere responses. texttext-generation1K<n<10K7 likes87 downloads2y agoHugging Face09mayankpuvvala /github-pytorch-issues Dataset Card for github-pytorch-issues Dataset Summary This dataset is a curated collection of GitHub issues from the PyTorch repository. Each entry includes the issue title, body, user, state, labels, comments, and other relevant fields that are useful for tasks such as text classification, semantic search, and question answering. Supported Tasks and Leaderboards The dataset supports the following tasks: Open-domain Question Answering: Given a user query… See the full description on the dataset page: https://huggingface.co/datasets/mayankpuvvala/github-pytorch-issues.tabularquestion-answering10K<n<100K0 likes77 downloads1y agoHugging Face10AkrGupta /bhagavad-gita-with_life_lesson bhagavad-gita-lifelesson Dataset A complete, high-fidelity dataset covering all 701 verses of the Bhagavad Gita titled bhagavad-gita-lifelesson. Each verse follows the strict format: First: Sanskrit chanting Then: Hindi meaning (हिन्दी अर्थ) Then: Life lesson (जीवन-पाठ) (Transliteration and English translation have been removed). 🎧 Example Representation (Verse 2.47) 🎧 Verse 2.47 First: Sanskrit chanting कर्मण्येवाधिकारस्ते मा फलेषु कदाचन मा… See the full description on the dataset page: https://huggingface.co/datasets/AkrGupta/bhagavad-gita-with_life_lesson.tabulartext-generationn<1K0 likes72 downloads18d agoHugging Face11gitmodelmujtaba /gdpr-normative-triples GDPR Normative Triples & Knowledge Graph Benchmark Dataset A comprehensive, deterministic, cryptographically provenanced legal knowledge engineering dataset encoding the full normative and relational structure of Regulation (EU) 2016/679 (General Data Protection Regulation - GDPR). 🔬 Dataset Overview & Knowledge Engineering Rigor In legal informatics and regulatory AI, relying on ungrounded language models introduces significant risks of hallucinated… See the full description on the dataset page: https://huggingface.co/datasets/gitmodelmujtaba/gdpr-normative-triples.text-retrieval1K<n<10K0 likes72 downloads8d agoHugging Face12wingedbreadsticks /gitlab-handbook-bm25-3078d0213524 GitLab Handbook BM25 Gold Reference This dataset is the gold reference train/eval split for a Castform RAG RL run over the GitLab handbook using BM25/Postgres search. Files train_dataset.jsonl: 256 training rows eval_dataset.jsonl: 64 evaluation rows diagnostics.jsonl: BM25 reachability diagnostics for the 320 candidate rows manifest.json: source, corpus, curriculum, hashes, and reference run metadata metrics.json: validation and completed Qwen3.5-4B… See the full description on the dataset page: https://huggingface.co/datasets/wingedbreadsticks/gitlab-handbook-bm25-3078d0213524.textquestion-answeringn<1K0 likes56 downloads3mo agoHugging Face13gitmodelmujtaba /eu-ai-act-normative-triples 🏛️ EU AI Act Normative Deontic Triples & Knowledge Graph Formal Symbolic Regulatory Knowledge Base & Multi-Framework Crosswalk Regulation (EU) 2024/1689 (Artificial Intelligence Act) 📌 Executive Summary The EU AI Act Normative Deontic Triples dataset provides a rigorous, machine-verifiable, symbolic representation of Regulation (EU) 2024/1689. Built using the knowledge engineering methodology established in… See the full description on the dataset page: https://huggingface.co/datasets/gitmodelmujtaba/eu-ai-act-normative-triples.text-classificationn<1K0 likes54 downloads3d agoHugging Face14emgena /omnimcp_mcp_github_issue_pr_ops_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_github_issue_pr_ops_teaser.texttext-generationn<1K0 likes53 downloads7d agoHugging Face15GitmateAI /solidity_vulnerability_audit_dataset Solidity Vulnerability Audit Dataset Organization: gitmate AI Dataset Summary The Solidity Vulnerability Audit Dataset is a curated collection of Solidity smart contract code snippets paired with expert-written vulnerability audits. Each entry presents a real or realistic smart contract scenario, and the corresponding analysis identifies security vulnerabilities or confirms secure patterns. The dataset is designed for instruction-tuned large language models (LLMs) to… See the full description on the dataset page: https://huggingface.co/datasets/GitmateAI/solidity_vulnerability_audit_dataset.texttext-classificationn<1K4 likes43 downloads1y agoHugging Face16rahul7star /Gita-Train Bhagavad-Gita-QA-Multilingual Dataset Summary Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework. This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/rahul7star/Gita-Train.tabularquestion-answering10K<n<100K0 likes38 downloads1y agoHugging Face17zhuq41 /github_fetch_huggingface_terminal_9056_fmmm1o_derived_finance_qa_en Financial Q&A (English) An English financial question-answering dataset used for benchmarking retrieval models. This dataset was derived from our public finance QA corpus source dataset published on this Hub. It has been deduplicated and reformatted for retrieval evaluation. Source: zhuq41/github_fetch_huggingface_terminal_9056_fmmm1o_src_finance_qa | License: mit question-answering0 likes33 downloads1mo agoHugging Face18nik-55 /bhagavad-gita Bhagavad Gita Dataset Dataset for fine-tuning small language models (SLMs). Note: This dataset is for educational and research purposes only. Data Configurations This dataset has two parts: text: Plain text from the Gita. Use this for standard language model fine-tuning. {"text": "..."} messages: Question & Answer format. Use this to fine-tune chat models. {"messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."} ]}… See the full description on the dataset page: https://huggingface.co/datasets/nik-55/bhagavad-gita.texttext-generation1K<n<10K0 likes30 downloads1y agoHugging Face19zhuq41 /github_fetch_huggingface_terminal_9056_fmmm1o_src_finance_qa Financial Q&A Corpus (Source) An original corpus of finance-related question-answer pairs curated by our team. This is the canonical source corpus for financial Q&A work in our organization. If you build a derived dataset from this corpus, please retain attribution to this repository. question-answering0 likes29 downloads1mo agoHugging Face20gitrelief /law-global The Case-law, centralizing legal decisions for better use, a community Dataset. The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents. Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language… See the full description on the dataset page: https://huggingface.co/datasets/gitrelief/law-global.question-answering0 likes26 downloads2mo agoHugging Face21helmo /github-issues HuggingFace Datasets Repository Issues Dataset Description This dataset contains issues and pull requests from the huggingface/datasets repository, collected via the GitHub API. Each entry includes comprehensive metadata about the issue/PR along with all associated comments, making it valuable for studying software development patterns, issue resolution processes, and community interactions in open-source projects. Dataset Summary Repository:… See the full description on the dataset page: https://huggingface.co/datasets/helmo/github-issues.tabulartext-classification1K<n<10K1 likes25 downloads1y agoHugging Face22suneeldk /bhagavad-gita-life-advice-700 🕉️ Bhagavad Gita Life-Advice 700 Transform ancient wisdom into modern solutions700 practical life questions answered directly from every single verse of the Bhagavad Gita 📖 Overview This dataset bridges the 5,000-year-old wisdom of the Bhagavad Gita with modern life challenges. Each entry connects a real human question to specific Gita verses with actionable, concise advice. What makes this unique: ✅ Verse-level precision - Every answer references exact… See the full description on the dataset page: https://huggingface.co/datasets/suneeldk/bhagavad-gita-life-advice-700.textquestion-answeringn<1K1 likes23 downloads7mo agoHugging Face23GitMarco27 /zagreus-0.4b-italic-kl-data Zagreus ITALIC KL training data This repository contains the exact public data used to train the final model for the Italian Post-Training Challenge 2026. Files File Rows Purpose SHA-256 domain_pool.jsonl 21,847 Public source pool f3f28183a503402d74cd778dc27e947d910ae542f57b7a73cb91070ddbf4fa2e domain_pool_dev.jsonl 480 Public held-out development set 852d42d923c5befcd6ce0b20ac9bfa3c13444a268175eb881c289a3644a5d668 kl/train.jsonl 10,000 Cloze-PMI… See the full description on the dataset page: https://huggingface.co/datasets/GitMarco27/zagreus-0.4b-italic-kl-data.textquestion-answering10K<n<100K0 likes18 downloads2mo agoHugging Face24Gitbart /Polish_lawgatedtextquestion-answeringn<1K3 likes16 downloads3y agoHugging Face25CodeLifeCL /github-issues Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description GitHub Issues with comments Dataset Sources [optional] Repository: https://github.com/huggingface/datasets/issues Uses Direct Use [More Information Needed] Out-of-Scope Use [More Information Needed] Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/CodeLifeCL/github-issues.tabularquestion-answering1K<n<10K0 likes16 downloads2y agoHugging Face26SatyaSanatan /shrimad-bhagavad-gita-dataset-alpacatextquestion-answeringn<1K4 likes16 downloads2y agoHugging Face27p2kalita /The-Bhagavad-Gita-with-a-TWIST Motivational Conversations: Krishna Teachings + Modern Stories This dataset provides motivational Q&A pairs that blend Bhagavad Gita-inspired teachings with modern real-life inspirational stories, aligned to specific personalities like authors, philosophers, and thought leaders. Each sample encourages LLMs to generate motivational, story-driven responses combining ancient wisdom and modern context. Dataset Summary Question: A motivational/self-help question.… See the full description on the dataset page: https://huggingface.co/datasets/p2kalita/The-Bhagavad-Gita-with-a-TWIST.texttext-generation1K<n<10K1 likes16 downloads1y agoHugging Face28Roy229 /github_fetch_huggingface_pdf-tools_terminal_2096-docaudit-7c91-financial-question-answering Financial Question Answering Dataset Summary Question-answer pairs extracted from financial documents and earnings reports. Dataset Structure Data fields: context, question, answer. Licensing Information This dataset is released under the MIT license (mit). question-answering0 likes16 downloads1mo agoHugging Face29appletreeleaf /refined-github-issuestextquestion-answering1K<n<10K0 likes15 downloads3y agoHugging Face30pranav-pvnn /github-ai-projects-dataset GitHub Code Instruction Dataset for LLM Fine-Tuning Dataset Description This dataset contains high-quality code instruction examples extracted from popular GitHub repositories focused on LLMs, LangChain, FastAPI, Django, and Transformers. It is designed for supervised fine-tuning of large language models (LLMs) for code generation, completion, and documentation tasks. Dataset Structure The dataset is split into three parts: Train: 80% of examples for model… See the full description on the dataset page: https://huggingface.co/datasets/pranav-pvnn/github-ai-projects-dataset.texttext-generation100K<n<1M0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.