CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SWE-bench /SWE-smith-phptextn<1K0 likes935 downloads9mo agoHugging Face02ajibawa-2023 /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/PHP-Code-Large.texttext-generation1M<n<10M22 likes836 downloads7mo agoHugging Face03xormania /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/xormania/PHP-Code-Large.texttext-generation1M<n<10M0 likes769 downloads5mo agoHugging Face04DaniilOr /php_cat1tabular10K<n<100K0 likes304 downloads1y agoHugging Face05jpaulpoliquit /ph-pretrain PH Pretrain — Philippine Languages Corpus (v0.6-ph-unified) 👉 Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead — it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage. A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.tabulartext-generation1M<n<10M0 likes297 downloads4mo agoHugging Face06fyaronskiy /cornstack_php_ru_enThe part of CoRNStack Dataset translated into Russian. Translation was done with Qwen3 model. Samples that satisfy the dual consistency filtering condition (samples where the document_rank is 0 or 1 and document_score > 0.7) were translated. Source code you can find here. For support: fedor.yaronskiy@gmail.com textsentence-similarity1M<n<10M0 likes290 downloads7mo agoHugging Face07nthakur /cornstack-php-v1-tevatron-1Mtext100K<n<1M0 likes244 downloads1y agoHugging Face08nomic-ai /cornstack-php-v1 CoRNStack PHP Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.text10M<n<100M3 likes239 downloads1y agoHugging Face09Ujjwal-Tyagi /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/PHP-Code-Large.texttext-generation1M<n<10M0 likes236 downloads6mo agoHugging Face10jpaulpoliquit /ph-pretrain-03 PH Pretrain 03 — Filipino/Tagalog Web Corpus (ph-pretrain-03) The recommended dataset for Filipino / Tagalog language-model pretraining and continued pretraining (CPT). A ~1-billion-token, Tagalog-forward web corpus, cleaned, deduplicated, and quality-scored by the jpaulpoliquit/pretraining refinery. 1,933,988 documents · ≈ 1.0 billion tokens (Sailor2-1B, calibrated) ~97.7% Tagalog web (FineWeb2 + CC100), with regional + news + Wikipedia in the long tail ~3.41 GB to download… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03.tabulartext-generation1M<n<10M0 likes232 downloads4mo agoHugging Face11Mo7art /Stack2Graph_KG_php PHP StackOverflow Knowledge Graph Summary This Hugging Face dataset repository contains the PHP shard of the Stack2Graph StackOverflow Knowledge Graph. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content. Stack2Graph source:… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_php.100M<n<1B0 likes226 downloads2mo agoHugging Face12CM /codexglue_code2text_php Dataset Card for "codexglue_code2text_php" More Information needed text100K<n<1M2 likes221 downloads3y agoHugging Face13open-athena /rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_stac-php-largtext10K<n<100K0 likes217 downloads7mo agoHugging Face14verify-ppt /marin-starcoderdata_php0 likes157 downloads6mo agoHugging Face15Helsinki-NLP /phpA parallel corpus originally extracted from http://se.php.net/download-docs.php. The original documents are written in English and have been partly translated into 21 languages. The original manuals contain about 500,000 words. The amount of actually translated texts varies for different languages between 50,000 and 380,000 words. The corpus is rather noisy and may include parts from the English original in some of the translations. The corpus is tokenized and each language pair has been sentence aligned. 23 languages, 252 bitexts total number of files: 71,414 total number of tokens: 3.28M total number of sentence fragments: 1.38Mtranslation10K<n<100K4 likes150 downloads3y agoHugging Face16FlanChanXwO /phpwind-captcha-dataset PHPWind Captcha Dataset Labelled four-digit numeric captcha images for training and evaluating OCR on authorized PHPWind deployments. This repository stores each visual captcha family in an independent dataset directory so that samples from different versions, forks, themes, or generators are never silently mixed. 中文说明: README_zh.md Related model: FlanChanXwO/phpwind-captcha-ocr Dataset catalogue Dataset ID Deployment or version identifier Status Images… See the full description on the dataset page: https://huggingface.co/datasets/FlanChanXwO/phpwind-captcha-dataset.imageimage-to-textn<1K0 likes143 downloads2mo agoHugging Face17Nan-Do /code-search-net-php Dataset Card for "code-search-net-php" Dataset Summary This dataset is the Php portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in Php Data Splits Train, test, validation labels are included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-php.texttext-generation100K<n<1M1 likes95 downloads3y agoHugging Face18Nan-Do /instructional_code-search-net-php Dataset Card for "instructional_code-search-net-php" Dataset Summary This is an instructional dataset for PHP. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-php.texttext-generation100K<n<1M3 likes77 downloads3y agoHugging Face19CoIR-Retrieval /CodeSearchNet-ccr-php-queries-corpus Dataset Card for "CodeSearchNet-ccr-php-queries-corpus" More Information needed text100K<n<1M0 likes63 downloads2y agoHugging Face20Shuu12121 /php-treesitter-filtered-datasetsV2 Php CodeSearch Dataset (Shuu12121/php-treesitter-filtered-datasetsV2) Dataset Description This dataset contains PHP functions and methods paired with their PHPDoc comments, extracted from open-source PHP repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a php function or method. docstring: The docstring or Javadoc associated with the function/method. func_name: The name of the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/php-treesitter-filtered-datasetsV2.text1M<n<10M1 likes60 downloads1y agoHugging Face21DCAgent2 /terminal_bench_2_r2egym_nl2bash_stack_bugsseq_lr3e_5_exp_rpt_stack_php_v2_step24d00812btextn<1K0 likes57 downloads7mo agoHugging Face22laion /swebench_verified_random_100_folders_a1_stack_phpunit_20260818_162722text10K<n<100K0 likes57 downloads1mo agoHugging Face23CoIR-Retrieval /CodeSearchNet-php-qrels Dataset Card for "CodeSearchNet-php-qrels" More Information needed text100K<n<1M0 likes56 downloads2y agoHugging Face24CoIR-Retrieval /CodeSearchNet-ccr-php-qrels Dataset Card for "CodeSearchNet-ccr-php-qrels" More Information needed text100K<n<1M0 likes53 downloads2y agoHugging Face25KaiLv /UDR_PHP Dataset Card for "UDR_PHP" More Information needed tabular100K<n<1M0 likes50 downloads3y agoHugging Face26Shuu12121 /php-treesitter-dedupe-filtered-datasetsV2 Php CodeSearch Dataset (Shuu12121/php-treesitter-dedupe-filtered-datasetsV2) Dataset Description This dataset contains PHP functions and methods paired with their PHPDoc comments, extracted from open-source PHP repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a php function or method. docstring: The docstring or Javadoc associated with the function/method. func_name: The name of the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/php-treesitter-dedupe-filtered-datasetsV2.text100K<n<1M0 likes50 downloads1y agoHugging Face27CoIR-Retrieval /CodeSearchNet-php-queries-corpus Dataset Card for "CodeSearchNet-php-queries-corpus" More Information needed text100K<n<1M2 likes49 downloads2y agoHugging Face28DCAgent2 /terminal_bench_2_r2egym_nl2bash_stack_bugsseq_stack_php_v2_20260222_044012textn<1K0 likes48 downloads7mo agoHugging Face29verify-ppt /smollm3-stack-v2-PHP0 likes47 downloads6mo agoHugging Face30ml-remunn /COIN-PHP-YOLOv11n-Dataset Coin Optical Identification and Numeration for the Philippine Peso - Dataset About This repo contains the dataset i gathered, annotated, and processed to train a yolov11n model for detecting and counting philippine peso coins. Folder Structure I organized the repo into three main folders to keep things clean: /raw - contains the original unedited photos. /annotated - contains the images and their respective label files from the annotation tool… See the full description on the dataset page: https://huggingface.co/datasets/ml-remunn/COIN-PHP-YOLOv11n-Dataset.imagen<1K0 likes46 downloads23d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.