CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01livecodebench /test_generation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/test_generation.textn<1K8 likes560 downloads2y agoHugging Face02facebook /emu_edit_test_set_generations Dataset Card for the Emu Edit Generations on Emu Edit Test Set Dataset Summary This dataset contains Emu Edit's generations on the Emu Edit test set. For more information please read our paper or visit our homepage. Licensing Information Licensed with CC-BY-NC 4.0 License available here. Citation Information @inproceedings{Sheynin2023EmuEP, title={Emu Edit: Precise Image Editing via Recognition and Generation Tasks}, author={Shelly Sheynin and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set_generations.image1K<n<10K38 likes209 downloads3y agoHugging Face03andstor /peft-unit-test-generation-experiments PEFT Unit Test Generation Experiments Dataset description The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research. Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/andstor/peft-unit-test-generation-experiments.tabularn<1K1 likes59 downloads10mo agoHugging Face04fals3 /peft-unit-test-generation-experiments PEFT Unit Test Generation Experiments Dataset description The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research. Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/fals3/peft-unit-test-generation-experiments.tabularn<1K0 likes51 downloads1y agoHugging Face05jjz5463 /llm-detection-generation-failcase-test Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. text1K<n<10K0 likes48 downloads2y agoHugging Face06hij /test_generationtext100K<n<1M0 likes39 downloads11mo agoHugging Face07jjz5463 /llm-detection-generation-contribution2-test Dataset Card Add more information here This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here. text1K<n<10K0 likes34 downloads2y agoHugging Face08HeAAAAA /story_generation_reward_test Reward Test — Held-Out Evaluation Set (EpisodeBench) This dataset is the held-out test set for automatic narrative evaluators released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. It is designed to measure how well an LLM-as-a-judge calibrates to EpisodeBench's synthesized rubric targets. Specifically, the paper reports the average absolute gap between each evaluator's predicted score and the synthesized… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_reward_test.texttext-classification10K<n<100K0 likes31 downloads5mo agoHugging Face09distilabel-internal-testing /test-text-generation-inference Dataset Card for test-text-generation-inference This dataset has been created with Distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/test-text-generation-inference/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/test-text-generation-inference.textn<1K0 likes30 downloads2y agoHugging Face10zrowt /generation-test-fourtextn<1K0 likes29 downloads2y agoHugging Face11zrowt /generation-test-twotextn<1K0 likes28 downloads2y agoHugging Face12TRI-ML /5000_generations_test_text2struc_text_code_1116text1K<n<10K0 likes28 downloads2y agoHugging Face13kothasuhas /llama-3b-gold-15M-student-generations_PRESAMPLING_1024_TESTtabular100K<n<1M0 likes27 downloads1y agoHugging Face14Code-TREAT /unit_test_generation⚠️ Note: The dataset symprompt_supp.jsonl is not created by us. We only supplemented this dataset with additional branch-level metadata (e.g., has_branch, total_branches) to enable coverage testing. This helps users keep their workflows clean when determining whether branches exist, simplifying branch coverage calculation. It originates from the paper: Code-Aware Prompting: A Study of Coverage Guided Test Generation in Regression Setting using LLM — Gabriel Ryan, Siddhartha Jain, Mingyue… See the full description on the dataset page: https://huggingface.co/datasets/Code-TREAT/unit_test_generation.tabularn<1K0 likes27 downloads1y agoHugging Face15kothasuhas /llama-3b-gold-15M-student-generations_SNIS_1024_TEST_N150.00Ktabular100K<n<1M0 likes25 downloads1y agoHugging Face16manasdhir04 /image-generation-dataset-testimage10K<n<100K0 likes24 downloads1y agoHugging Face17kothasuhas /llama-3b-gold-15M-student-generations_RS_1024_TEST_N150.00Ktabular1K<n<10K0 likes22 downloads1y agoHugging Face18RLHFlow /test_generation_2k Dataset Card for "xxxxxx" More Information needed text1K<n<10K0 likes21 downloads2y agoHugging Face19arksap /gitbug-java-unit-test-generationgatedtabularn<1K1 likes21 downloads2y agoHugging Face20Pedagogy-r1 /generation_test Dataset Card for generation_test This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/Pedagogy-r1/generation_test/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Pedagogy-r1/generation_test.textn<1K0 likes20 downloads2y agoHugging Face21friendshipkim /RUCAIBox-Story-Generation-testtext1K<n<10K0 likes20 downloads1y agoHugging Face22SoorajK1 /Question_generation_test-088d3464-f299-492c-9793-e2ad385a31cftabularn<1K0 likes19 downloads3y agoHugging Face23Steve77 /code-test-generationtext1K<n<10K0 likes17 downloads11mo agoHugging Face24kothasuhas /llama-3b-gold-15M-student-generations_SNIS_2048_TEST_baseN10.00K_N30.00K_T2.0tabular10K<n<100K0 likes17 downloads1y agoHugging Face25MagistrTheOne /radon-test-code_generation radon-test-code_generation Description Code generation test dataset for RADON model evaluation with programming prompts Usage Load Dataset from datasets import load_dataset dataset = load_dataset("MagistrTheOne/radon-test-code_generation") print(dataset) Use with RADON Model from transformers import AutoModelForCausalLM, AutoTokenizer # Load RADON model model = AutoModelForCausalLM.from_pretrained("MagistrTheOne/RadonSAI") tokenizer =… See the full description on the dataset page: https://huggingface.co/datasets/MagistrTheOne/radon-test-code_generation.texttext-generationn<1K0 likes17 downloads1y agoHugging Face26mathewmouchamel /test_generation_v2 LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/mathewmouchamel/test_generation_v2.textn<1K0 likes17 downloads5mo agoHugging Face27Shiki258 /test_generation_2k_rewardstext1K<n<10K0 likes16 downloads1y agoHugging Face28vwxyzjn /norobot_3pair_test_generation_24163 allenai/open_instruct: Generation Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'vwxyzjn', 'hf_repo_id': 'norobot_3pair_test_generation_24163', 'mode': 'generation', 'model_name_or_path': 'allenai/open_instruct_dev', 'push_to_hub': True, 'revision': 'costa_finetune_tulu3_8b_norobot__meta-llama_Meta-Llama-3.1-8B__42__1725559869'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/norobot_3pair_test_generation_24163.text10K<n<100K0 likes15 downloads2y agoHugging Face29kothasuhas /llama-3b-gold-15M-student-generations_PRESAMPLING_2048_TEST_baseN10.00Ktabular10K<n<100K0 likes15 downloads1y agoHugging Face30zjhhhh /fast-generation-testtext10K<n<100K0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.