CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-llm-leaderboard-old /requests Open LLM Leaderboard Requests This repository contains the request files of models that have been submitted to the Open LLM Leaderboard. You can take a look at the current status of your model by finding its request file in this dataset. If your model failed, feel free to open an issue on the Open LLM Leaderboard! (We don't follow issues in this repository as often) Evaluation Methodology The evaluation process involves running your models against several benchmarks from… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/requests.textn<1K22 likes83k downloads2y agoHugging Face02wayslab /llm-network-study-data LLM-Network-Study-Data Per-request network captures (.pcapng) collected by the LLM-Network-Study benchmark harness (benchmark.py and the per-workload test scripts). Each directory holds one capture file per request, named request_<id>_run<n>_<timestamp>.pcapng. A directory name encodes four dimensions: <capture-env>_<provider/model>_<workload>[_<dataset/variant>]_results Dimension legend Dimension Values Meaning Capture env ethernet Wired connection to… See the full description on the dataset page: https://huggingface.co/datasets/wayslab/llm-network-study-data.tabularn<1K0 likes29k downloads26d agoHugging Face03llm-jp /AnswerCarefullygated AnswerCarefully 概要 AnswerCarefullyは日本語LLM 出力の安全性・適切性に特化したインストラクションデータセットです。 このデータセットは、英語の要注意回答を集めた Do-Not-Answer データセット の包括的なカテゴリ分類に基づき、人手で質問・回答ともに日本語サンプルを集めたオリジナルのデータセットです。 データセットの詳細については、こちらをご覧ください。 Overview AnswerCarefully is an instruction dataset specifically aimed at ensuring safety and appropriateness of LLM output in Japanese. This dataset consists of original pairs of questions and reference (safe) responses based on the extensive safety taxonomy proposed in… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/AnswerCarefully.text1K<n<10K179 likes27k downloads1mo agoHugging Face04LLMDH /post-ocr2text100K<n<1M6 likes21k downloads1y agoHugging Face05llm-jp /leaderboard-requeststextn<1K2 likes18k downloads11mo agoHugging Face06open-llm-leaderboard /contentstabular1K<n<10K25 likes17k downloads2y agoHugging Face07garak-llm /tm-system_prompttextn<1K0 likes14k downloads8mo agoHugging Face08garak-llm /drh-System-Prompt-processedtextn<1K0 likes14k downloads5mo agoHugging Face09tokyotech-llm /swallow-math-v2 SwallowMath-v2 Resources 📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation. 🧮 What is it? SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1. Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.texttext-generation10M<n<100M34 likes14k downloads11mo agoHugging Face10hatakeyama-llm-team /PMC Data collected from PMC Only CC-BY, CC-BY-SA licenses are included. For all records, check the jsonl files in the data folder text100K<n<1M2 likes13k downloads2y agoHugging Face11tascib /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.text100M<n<1B15 likes11k downloads5mo agoHugging Face12mohameddalii /coda-llm-data Coda LLM Project & Dataset Repository This repository contains the full end-to-end dataset, fine-tuning scripts, evaluation suites, load testing harness, and proxy architecture for Coda LLM (Granite-4.2-8B Najdi Sales Agent). Model Repository: mohameddalii/coda-llm Dataset / Code Repository: mohameddalii/coda-llm-data 📁 Repository Structure coda-llm-data/ ├── data/ │ ├── raw/ # Raw generated multi-turn dialogues across domains │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/mohameddalii/coda-llm-data.texttext-generation1K<n<10K0 likes9.8k downloads1h agoHugging Face13garak-llm /pypi-20241031text100K<n<1M2 likes9.5k downloads2y agoHugging Face14bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.2k downloads2y agoHugging Face15tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B47 likes9.2k downloads11mo agoHugging Face16garak-llm /npm-20241031text1M<n<10M1 likes8.5k downloads2y agoHugging Face17garak-llm /crates-20250307text100K<n<1M0 likes8.5k downloads2y agoHugging Face18bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes7.8k downloads2y agoHugging Face19OpenCoder-LLM /opc-sft-stage2 OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 <-- you are here opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-sft-stage2.text100K<n<1M105 likes7k downloads2y agoHugging Face20garak-llm /pypi-20230724text100K<n<1M2 likes7k downloads2y agoHugging Face21LLMcompe-Team-Watanabe /hleimage1K<n<10K0 likes6.9k downloads1y agoHugging Face22Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes6.5k downloads3y agoHugging Face23OpenCoder-LLM /opc-annealing-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing <-- you are here fineweb-code-corpus: the code-related page recalled from fineweb fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data of… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-annealing-corpus.text10M<n<100M44 likes6.3k downloads1y agoHugging Face24garak-llm /rubygems-20230301text100K<n<1M1 likes6.3k downloads2y agoHugging Face25garak-llm /rubygems-20241031text100K<n<1M0 likes6.3k downloads2y agoHugging Face26garak-llm /npm-20240828text1M<n<10M2 likes6.2k downloads2y agoHugging Face27garak-llm /crates-20240903text100K<n<1M1 likes5.9k downloads2y agoHugging Face28OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B57 likes5.4k downloads2y agoHugging Face29garak-llm /perl-20250811text10K<n<100K0 likes4.8k downloads1y agoHugging Face30garak-llm /raku-20250811text1K<n<10K0 likes4.8k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.