datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Methods2Test_java_unit_test_code
Dataset Description
Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods.
It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K
Java open source project hosted on GitHub.
The mapping between test case and focal methods are based heuristics rules and Java developer's best practice.
More information could be found here:
methods2test Github repo
Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.UnitTests
UnitTests
Task description
Evaluation of unit-test generation for functions and methods in five programming languages (Java, Python, Go, JavaScript, and C#). Dataset contains 2500 tasks.
Evaluated skills: Instruction Following, Long Context Comprehension, Synthesis, Testing
Contributors: Alena Pestova, Valentin Malykh
Motivation
Unit testing is an important software development practice in which individual components of a software system are evaluated in… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/UnitTests.CodeRM-UnitTest
Dataset Description
CodeRM-UnitTest dataset originates from the paper: Dynamic Scaling of Unit Tests for Code Reward Modeling available on arXiv.
You can visit the homepage to learn more about the paper.
It is a curated collection of high-quality synthetic Python unit tests, derived from two prominent code instruction tuning
datasets: CodeFeedback-Filtered-Instruction and the training
set of TACO. This dataset is used for training
CodeRM-8B, a small yet powerful unit test… See the full description on the dataset page: https://huggingface.co/datasets/KAKA22/CodeRM-UnitTest.cpp_unit_tests_benchmark_datajava_unit_testsnips_test_valid_unit
Dataset Card for "snips_test_valid_unit"
More Information needed
seed_code_multiple_samples_scale_up_base_16K_unit_testsUnitTestsPubliccpp_unit_tests_benchmark_data_with_splitsGO-UNITTEST-BENCHCPP-UNITTEST-BENCH
Dataset Card for Open Source Code and Unit Tests
Dataset Details
Dataset Description
This dataset contains c++ code snippets and their corresponding ground truth unit tests collected from various open-source GitHub repositories. The primary purpose of this dataset is to aid in the development and evaluation of automated testing tools, code quality analysis, and LLM models for test generation.
Curated by: Vaishnavi Bhargava
Language(s): C++… See the full description on the dataset page: https://huggingface.co/datasets/Nutanix/CPP-UNITTEST-BENCH.python-unit-test-training-pool
Python unit test training pool
A pool of public data for training a model to write tests for Python code. It is a
straight collection of open datasets, not a new corpus: every row comes from one of the
sources below, at the revision named, and the only rows removed are the ones an overlap
filter flagged against held-out material this pool is kept separate from.
Every row of the normalised layer pairs a program with tests for it. That is the point of
the pool, and it is why the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-unit-test-training-pool.librispeech_asr_test_unitslue-sqa5-test-LLM_unit
Dataset Card for "slue-sqa5-LLM_unit"
More Information needed
peft-unit-test-generation-experiments
PEFT Unit Test Generation Experiments
Dataset description
The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/andstor/peft-unit-test-generation-experiments.unit-test-v2
Dataset Card for "unit-test-v2"
More Information needed
Nsynth-test_unitlibri2Mix_test_unit
Dataset Card for "libri2Mix_test_unit"
More Information needed
omnimcp_unit_test_synthesizer_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_unit_test_synthesizer_teaser.peft-unit-test-generation-experiments
PEFT Unit Test Generation Experiments
Dataset description
The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/fals3/peft-unit-test-generation-experiments.mini-rust-unit-test-in-the-stacklivecodebench_unit_test_error_240_samplesSWESwiss-SFT-Unittest-1K
Overview
SFT dataset for training SWE-Swiss models on the unit test generation task. The prompts contain issues sourced from SWE-Gym and SWE-smith, while the responses are generated by DeepSeek-R1-0528. To ensure quality, we filter out data where the generated unit tests do not perform as expected. A generated test is kept only if its execution results correctly distinguish between a set of correct and incorrect patches, mirroring the behavior of the repository's own test suite.… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Swiss/SWESwiss-SFT-Unittest-1K.unit-test-mocking-dataset-enriched-v2NMSQA-test-gpt4_unit
Dataset Card for "NMSQA-test-gpt4_unit"
More Information needed
unit_test_generation⚠️ Note: The dataset symprompt_supp.jsonl is not created by us. We only supplemented this dataset with additional branch-level metadata (e.g., has_branch, total_branches) to enable coverage testing.
This helps users keep their workflows clean when determining whether branches exist, simplifying branch coverage calculation.
It originates from the paper:
Code-Aware Prompting: A Study of Coverage Guided Test Generation in Regression Setting using LLM
— Gabriel Ryan, Siddhartha Jain, Mingyue… See the full description on the dataset page: https://huggingface.co/datasets/Code-TREAT/unit_test_generation.public_unittest_repoThis is an integration database of erai-raws, myanimelist and nyaasi. You can know which animes are the hottest ones currently, and which of them have well-seeded magnet links.
This database is refreshed daily.
Current Animes
5 animes, 45 episodes in total, Last updated on: 2024-07-21 01:08:35 CST.
ID
Post
Bangumi
Type
Episodes
Status
Score
Nyaasi
Magnets
Seeds
Downloads
Updated At
53802
2.5-jigen no RirisaTV
7 / 24
Currently Airing
7.29
Search
Download
32
900
2024-07-19… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/public_unittest_repo.unit-test-mocking-dataset-enriched-v3cpp-unittest-26-11-2025cpp-unittest
