CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01deepmind /code_contests Dataset Card for CodeContests Dataset Summary CodeContests is a competitive programming dataset for machine-learning. This dataset was used when training AlphaCode. It consists of programming problems, from a variety of sources: Site URL Source Aizu https://judge.u-aizu.ac.jp CodeNet AtCoder https://atcoder.jp CodeNet CodeChef https://www.codechef.com description2code Codeforces https://codeforces.com description2code and Codeforces HackerEarth… See the full description on the dataset page: https://huggingface.co/datasets/deepmind/code_contests.tabulartranslation1K<n<10K235 likes76k downloads3y agoHugging Face02ByteDance-Seed /Code-Contests-Plus CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset built upon CodeContests. It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions. Highlights High… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/Code-Contests-Plus.tabularother10K<n<100K69 likes13k downloads11mo agoHugging Face03BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.4k downloads9mo agoHugging Face04teven /code_contestsHF-datasets version of Deepmind's code_contests dataset, notably used for AlphaGo. 1 row per solution, no test data or incorrect solutions included (only name/source/description/solution/language/difficulty) tabular1M<n<10M4 likes857 downloads4y agoHugging Face05voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes468 downloads2mo agoHugging Face06Imandra /code_contests Dataset Card for CodeContests Dataset Summary CodeContests is a competitive programming dataset for machine-learning. This dataset was used when training AlphaCode. It consists of programming problems, from a variety of sources: Site URL Source Aizu https://judge.u-aizu.ac.jp CodeNet AtCoder https://atcoder.jp CodeNet CodeChef https://www.codechef.com description2code Codeforces https://codeforces.com description2code and Codeforces HackerEarth… See the full description on the dataset page: https://huggingface.co/datasets/Imandra/code_contests.tabulartranslation10K<n<100K0 likes357 downloads1y agoHugging Face07sharkchill-xy /CodeContests_apps_format Dataset Card for "CodeContests_apps_format" More Information needed tabular10K<n<100K0 likes351 downloads3y agoHugging Face08Hanuman2 /code_contests Dataset Card for CodeContests Dataset Summary CodeContests is a competitive programming dataset for machine-learning. This dataset was used when training AlphaCode. It consists of programming problems, from a variety of sources: Site URL Source Aizu https://judge.u-aizu.ac.jp CodeNet AtCoder https://atcoder.jp CodeNet CodeChef https://www.codechef.com description2code Codeforces https://codeforces.com description2code and Codeforces HackerEarth… See the full description on the dataset page: https://huggingface.co/datasets/Hanuman2/code_contests.tabulartranslation10K<n<100K0 likes321 downloads3mo agoHugging Face09hieunguyenminh /code_contests_dp_datasettabular1K<n<10K0 likes319 downloads2y agoHugging Face10HexQuant /Code-Contests-Plus CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset built upon CodeContests. It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions. Highlights High Quality Test… See the full description on the dataset page: https://huggingface.co/datasets/HexQuant/Code-Contests-Plus.tabularother10K<n<100K0 likes255 downloads9mo agoHugging Face11wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes245 downloads2y agoHugging Face12louisbrulenaudet /code-commande-publique Code de la commande publique, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-commande-publique.tabulartext-generation1K<n<10K0 likes236 downloads1y agoHugging Face13argilla /code_contests_qwen_coder Dataset Card for code_contests_qwen_coder This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: pipeline.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/argilla/code_contests_qwen_coder/raw/main/pipeline.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel… See the full description on the dataset page: https://huggingface.co/datasets/argilla/code_contests_qwen_coder.tabularn<1K2 likes186 downloads2y agoHugging Face14skzg /Code-Contests-Plus CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset built upon CodeContests. It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions. Highlights High Quality Test… See the full description on the dataset page: https://huggingface.co/datasets/skzg/Code-Contests-Plus.tabularother10K<n<100K0 likes173 downloads4mo agoHugging Face15louisbrulenaudet /code-collectivites-territoriales Code général des collectivités territoriales, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-collectivites-territoriales.tabulartext-generation1K<n<10K0 likes158 downloads1y agoHugging Face16Asap7772 /code_contests_llamabase_mc_intermediatetabular1M<n<10M0 likes109 downloads2y agoHugging Face17taoroalin /code_contests_slim_jsontabular1K<n<10K0 likes86 downloads2y agoHugging Face18semeru /code-code-DefectDetection Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/Defect-detection in Semeru CodeXGLUE -- Defect Detection Task Definition Given a source code, the task is to identify whether it is an insecure code that may attack software systems, such as resource leaks, use-after-free vulnerabilities and DoS attack. We treat the task as… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-DefectDetection.tabular10K<n<100K2 likes82 downloads3y agoHugging Face19louisbrulenaudet /code-civil Code civil, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models based… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-civil.tabulartext-generation1K<n<10K2 likes79 downloads1y agoHugging Face20Gen-Verse /CodeContests_trainWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests_train.tabular1K<n<10K3 likes72 downloads1y agoHugging Face21Gen-Verse /CodeContestsWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format: input = "5\n1 2 3 4 5\n" output = "15" CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like assert sum_function([1, 2, 3, 4, 5]) == 15 In this project, we have converted the the functional format to the Stdio format to achieve consistency. Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests.tabularn<1K1 likes67 downloads1y agoHugging Face22rasdani /Code-Contests-Plus-HQ-2x-GRPOtabular1K<n<10K0 likes62 downloads1y agoHugging Face23louisbrulenaudet /code-consommation Code de la consommation, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-consommation.tabulartext-generation1K<n<10K0 likes59 downloads1y agoHugging Face24rasdani /Code-Contests-Plus-HQ-1x-GRPOtabular1K<n<10K0 likes59 downloads1y agoHugging Face25hieupham14022003 /general_speech_distorted_big_codectabular100K<n<1M0 likes56 downloads1y agoHugging Face26Hiren122 /code_contests Dataset Card for CodeContests Dataset Summary CodeContests is a competitive programming dataset for machine-learning. This dataset was used when training AlphaCode. It consists of programming problems, from a variety of sources: Site URL Source Aizu https://judge.u-aizu.ac.jp CodeNet AtCoder https://atcoder.jp CodeNet CodeChef https://www.codechef.com description2code Codeforces https://codeforces.com description2code and Codeforces HackerEarth… See the full description on the dataset page: https://huggingface.co/datasets/Hiren122/code_contests.tabulartranslation10K<n<100K0 likes56 downloads3mo agoHugging Face27Asap7772 /code_contests_llamabase_mc_intermediate-part2-of-4tabular100K<n<1M0 likes53 downloads2y agoHugging Face28osouza /code_contests_pttabular1K<n<10K0 likes52 downloads2y agoHugging Face29semeru /code-code-galeras-code-completion-from-docstring-3k-dedupedtabular1K<n<10K0 likes45 downloads3y agoHugging Face30Asap7772 /code_contests_llamabase_mc_intermediate-part3-of-4tabular100K<n<1M0 likes38 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.