CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01likaixin /APPS-verified Introduction This dataset contains verified solutions from the APPS dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed. The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds. Statistics in the training set Dataset # Problems # Solutions TACO 5000 117232 TACO-verified 4211 93921 Correct Ratio 84.22% 80.12% tabularquestion-answering1K<n<10K5 likes86 downloads2y agoHugging Face02PNYX /apps_pnyx PNYX - Apps This is a splitted and tested version of APPS dataset, refer to it for further information on the original dataset construction. This version is designed to be compatible with the hf_evaluate code_eval package and to be executed with lm-eval code_eval package. This dataset does not include all the original fields. Some are modified and some are completely new: id: Same as the original APPS dataset. difficulty: Difficulty of the problem. Same as the original APPS… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/apps_pnyx.tabulartext-generation1K<n<10K0 likes51 downloads6mo agoHugging Face03styme3279 /control-apps-cleaned control-apps-cleaned A curated subset of the APPS dataset (Hendrycks et al., NeurIPS 2021), pre-filtered for use in the ARENA AI Control chapter — a teaching replication of Greenblatt et al. 2023 (arXiv:2312.06942). What's in here cleaned_apps.jsonl — 1,202 problems from the APPS "interview" split, filtered to a uniform I/O schema (inputs and outputs are each one of list[str], list[int], list[list[str]], list[list[int]]). Each line is a JSON record with the… See the full description on the dataset page: https://huggingface.co/datasets/styme3279/control-apps-cleaned.tabulartext-generationn<1K0 likes42 downloads3mo agoHugging Face04abhayesian /answers-with-reasoning-apps answers-with-reasoning-apps Self-distillation SFT corpus: Qwen3-8B-Instruct's own all-tests-pass chain-of-thought rollouts on APPS interview tier (code domain). Generation Source problems: codeparrot/apps, difficulty == "interview" filter on both train (2000 problems) and test (3000 problems) splits = 5000 candidate problems. Problems with empty / malformed input_output are dropped (~6%), leaving 4692 candidates. LCB-v5 (our held-out code benchmark) does not overlap APPS… See the full description on the dataset page: https://huggingface.co/datasets/abhayesian/answers-with-reasoning-apps.tabulartext-generation1K<n<10K0 likes27 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.