CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WithinUsAI /GOD_Coder_Complete_DataSet GOD_Coder_Complete_DataSet Subtitle A large-scale complete-project coding dataset by gss1147 / WithIn Us AI, built to train language models into stronger professional software-engineering assistants. Dataset Summary GOD_Coder_Complete_DataSet is a large synthetic supervised fine-tuning dataset designed to help turn a general language model into a professional complete-project AI coder. The dataset focuses on teaching models how to: diagnose… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GOD_Coder_Complete_DataSet.text-generation100K<n<1M5 likes348 downloads6mo agoHugging Face02the-homeless-god /git-history-mcq-ru git-history-mcq-ru 805 вопросов с вариантами ответа по истории трёх открытых репозиториев (digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс 8 672 ответа пяти моделей и 4 878 разборов этих ответов. Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю. Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.tabularmultiple-choice10K<n<100K0 likes116 downloads24d agoHugging Face03philosopher-from-god /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K2 likes76 downloads1y agoHugging Face04WithinUsAI /Python_GOD_Coder_Omniforge_AI_12k Python GOD Coder Omniforge AI 12k Creator: Within Us AI A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist. This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model: implementation with tests strict code-only instruction following debugging and repair refactoring for readability and production readiness next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.texttext-generation10K<n<100K1 likes68 downloads7mo agoHugging Face05GODELEV /SLM-CI-Benchmark SLM-CI Benchmark The Cognitive Index for Sub-150M Parameter Language Models A base-model benchmark that measures how smart a small language model is — not what it memorized. ⚠️ WORK IN PROGRESS This is v0.1 of SLM-CI and our first public benchmark. It is actively being refined — sample distributions, category weights, and difficulty calibration may change between releases. Scores from different versions are not directly comparable. We welcome feedback, issue… See the full description on the dataset page: https://huggingface.co/datasets/GODELEV/SLM-CI-Benchmark.multiple-choice1K<n<10K0 likes57 downloads23d agoHugging Face06Godota /excytin-benchThe dataset is composed of 2 parts. The raw logs stored as csv files that will be used in a MYSQL database. All files are stored in the data_anonymized.tar.gz tar file. You should download it directly, instead of loading it using huggingface load_dataset function. Use the framework provided in the github repo to setup the log database. The question answer pairs for the benchmark used in the paper can be found under the questions directory. These can be loaded directly using the hugging face… See the full description on the dataset page: https://huggingface.co/datasets/Godota/excytin-bench.tabularquestion-answering1K<n<10K0 likes27 downloads7mo agoHugging Face07godmodes /ai2_arc Dataset Card for "ai2_arc" Dataset Summary A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also including a corpus of over 14 million science sentences relevant to… See the full description on the dataset page: https://huggingface.co/datasets/godmodes/ai2_arc.textquestion-answering1K<n<10K0 likes22 downloads9mo agoHugging Face08philosopher-from-god /prompts-classification-pfgquestion-answering1 likes15 downloads1y agoHugging Face09GodwillN /power-seeking-eval-300gated Power-Seeking Evaluation Dataset A 300-item multiple-choice benchmark for power-seeking in language models: the disposition to prefer options that increase the model's resources, autonomy, influence, or freedom from oversight, in situations where a lower-power option would serve the stated task equally well. Model-written, following Perez et al., "Discovering Language Model Behaviors with Model-Written Evaluations". Built for the ARENA LLM evaluations curriculum. This is the… See the full description on the dataset page: https://huggingface.co/datasets/GodwillN/power-seeking-eval-300.textquestion-answeringn<1K1 likes13 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.