datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test_generation
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
🏠 Home Page •
💻 GitHub Repository •
🏆 Leaderboard •
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/test_generation.emu_edit_test_set_generations
Dataset Card for the Emu Edit Generations on Emu Edit Test Set
Dataset Summary
This dataset contains Emu Edit's generations on the Emu Edit test set. For more information please read our paper or visit our homepage.
Licensing Information
Licensed with CC-BY-NC 4.0 License available here.
Citation Information
@inproceedings{Sheynin2023EmuEP,
title={Emu Edit: Precise Image Editing via Recognition and Generation Tasks},
author={Shelly Sheynin and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set_generations.peft-unit-test-generation-experiments
PEFT Unit Test Generation Experiments
Dataset description
The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/andstor/peft-unit-test-generation-experiments.peft-unit-test-generation-experiments
PEFT Unit Test Generation Experiments
Dataset description
The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/fals3/peft-unit-test-generation-experiments.llm-detection-generation-failcase-test
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
test_generationllm-detection-generation-contribution2-test
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
story_generation_reward_test
Reward Test — Held-Out Evaluation Set (EpisodeBench)
This dataset is the held-out test set for automatic narrative evaluators released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL.
It is designed to measure how well an LLM-as-a-judge calibrates to EpisodeBench's synthesized rubric targets. Specifically, the paper reports the average absolute gap between each evaluator's predicted score and the synthesized… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_test.test-text-generation-inference
Dataset Card for test-text-generation-inference
This dataset has been created with Distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/test-text-generation-inference/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/test-text-generation-inference.generation-test-fourgeneration-test-two5000_generations_test_text2struc_text_code_1116llama-3b-gold-15M-student-generations_PRESAMPLING_1024_TESTunit_test_generation⚠️ Note: The dataset symprompt_supp.jsonl is not created by us. We only supplemented this dataset with additional branch-level metadata (e.g., has_branch, total_branches) to enable coverage testing.
This helps users keep their workflows clean when determining whether branches exist, simplifying branch coverage calculation.
It originates from the paper:
Code-Aware Prompting: A Study of Coverage Guided Test Generation in Regression Setting using LLM
— Gabriel Ryan, Siddhartha Jain, Mingyue… See the full description on the dataset page: https://huggingface.co/datasets/Code-TREAT/unit_test_generation.llama-3b-gold-15M-student-generations_SNIS_1024_TEST_N150.00Kimage-generation-dataset-testllama-3b-gold-15M-student-generations_RS_1024_TEST_N150.00Ktest_generation_2k
Dataset Card for "xxxxxx"
More Information needed
gitbug-java-unit-test-generationgeneration_test
Dataset Card for generation_test
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Pedagogy-r1/generation_test/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Pedagogy-r1/generation_test.RUCAIBox-Story-Generation-testQuestion_generation_test-088d3464-f299-492c-9793-e2ad385a31cfcode-test-generationllama-3b-gold-15M-student-generations_SNIS_2048_TEST_baseN10.00K_N30.00K_T2.0radon-test-code_generation
radon-test-code_generation
Description
Code generation test dataset for RADON model evaluation with programming prompts
Usage
Load Dataset
from datasets import load_dataset
dataset = load_dataset("MagistrTheOne/radon-test-code_generation")
print(dataset)
Use with RADON Model
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load RADON model
model = AutoModelForCausalLM.from_pretrained("MagistrTheOne/RadonSAI")
tokenizer =… See the full description on the dataset page: https://huggingface.co/datasets/MagistrTheOne/radon-test-code_generation.test_generation_v2
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
🏠 Home Page •
💻 GitHub Repository •
🏆 Leaderboard •
LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs.
Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution.
This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/mathewmouchamel/test_generation_v2.test_generation_2k_rewardsnorobot_3pair_test_generation_24163
allenai/open_instruct: Generation Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'vwxyzjn',
'hf_repo_id': 'norobot_3pair_test_generation_24163',
'mode': 'generation',
'model_name_or_path': 'allenai/open_instruct_dev',
'push_to_hub': True,
'revision': 'costa_finetune_tulu3_8b_norobot__meta-llama_Meta-Llama-3.1-8B__42__1725559869'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/norobot_3pair_test_generation_24163.llama-3b-gold-15M-student-generations_PRESAMPLING_2048_TEST_baseN10.00Kfast-generation-test
