datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-smith-phpPHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/PHP-Code-Large.PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/xormania/PHP-Code-Large.php_cat1ph-pretrain
PH Pretrain — Philippine Languages Corpus (v0.6-ph-unified)
👉 Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead — it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage.
A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.cornstack_php_ru_enThe part of CoRNStack Dataset translated into Russian. Translation was done with Qwen3 model.
Samples that satisfy the dual consistency filtering condition (samples where the document_rank is 0 or 1 and document_score > 0.7) were translated.
Source code you can find here. For support: fedor.yaronskiy@gmail.com
cornstack-php-v1-tevatron-1Mcornstack-php-v1
CoRNStack PHP Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/PHP-Code-Large.ph-pretrain-03
PH Pretrain 03 — Filipino/Tagalog Web Corpus (ph-pretrain-03)
The recommended dataset for Filipino / Tagalog language-model pretraining and continued pretraining (CPT).
A ~1-billion-token, Tagalog-forward web corpus, cleaned, deduplicated, and quality-scored by the jpaulpoliquit/pretraining refinery.
1,933,988 documents · ≈ 1.0 billion tokens (Sailor2-1B, calibrated)
~97.7% Tagalog web (FineWeb2 + CC100), with regional + news + Wikipedia in the long tail
~3.41 GB to download… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03.Stack2Graph_KG_php
PHP StackOverflow Knowledge Graph
Summary
This Hugging Face dataset repository contains the PHP shard of the Stack2Graph StackOverflow Knowledge Graph.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content.
Stack2Graph source:… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_php.codexglue_code2text_php
Dataset Card for "codexglue_code2text_php"
More Information needed
rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_stac-php-largmarin-starcoderdata_phpphpA parallel corpus originally extracted from http://se.php.net/download-docs.php. The original documents are written in English and have been partly translated into 21 languages. The original manuals contain about 500,000 words. The amount of actually translated texts varies for different languages between 50,000 and 380,000 words. The corpus is rather noisy and may include parts from the English original in some of the translations. The corpus is tokenized and each language pair has been sentence aligned.
23 languages, 252 bitexts
total number of files: 71,414
total number of tokens: 3.28M
total number of sentence fragments: 1.38Mphpwind-captcha-dataset
PHPWind Captcha Dataset
Labelled four-digit numeric captcha images for training and evaluating OCR on
authorized PHPWind deployments. This repository stores each visual captcha
family in an independent dataset directory so that samples from different
versions, forks, themes, or generators are never silently mixed.
中文说明: README_zh.md
Related model: FlanChanXwO/phpwind-captcha-ocr
Dataset catalogue
Dataset ID
Deployment or version identifier
Status
Images… See the full description on the dataset page: https://huggingface.co/datasets/FlanChanXwO/phpwind-captcha-dataset.code-search-net-php
Dataset Card for "code-search-net-php"
Dataset Summary
This dataset is the Php portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does.
Languages
The dataset's comments are in English and the functions are coded in Php
Data Splits
Train, test, validation labels are included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-php.instructional_code-search-net-php
Dataset Card for "instructional_code-search-net-php"
Dataset Summary
This is an instructional dataset for PHP.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-php.CodeSearchNet-ccr-php-queries-corpus
Dataset Card for "CodeSearchNet-ccr-php-queries-corpus"
More Information needed
php-treesitter-filtered-datasetsV2
Php CodeSearch Dataset (Shuu12121/php-treesitter-filtered-datasetsV2)
Dataset Description
This dataset contains PHP functions and methods paired with their PHPDoc comments, extracted from open-source PHP repositories on GitHub.
It is formatted similarly to the CodeSearchNet challenge dataset.
Each entry includes:
code: The source code of a php function or method.
docstring: The docstring or Javadoc associated with the function/method.
func_name: The name of the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/php-treesitter-filtered-datasetsV2.terminal_bench_2_r2egym_nl2bash_stack_bugsseq_lr3e_5_exp_rpt_stack_php_v2_step24d00812bswebench_verified_random_100_folders_a1_stack_phpunit_20260818_162722CodeSearchNet-php-qrels
Dataset Card for "CodeSearchNet-php-qrels"
More Information needed
CodeSearchNet-ccr-php-qrels
Dataset Card for "CodeSearchNet-ccr-php-qrels"
More Information needed
UDR_PHP
Dataset Card for "UDR_PHP"
More Information needed
php-treesitter-dedupe-filtered-datasetsV2
Php CodeSearch Dataset (Shuu12121/php-treesitter-dedupe-filtered-datasetsV2)
Dataset Description
This dataset contains PHP functions and methods paired with their PHPDoc comments, extracted from open-source PHP repositories on GitHub.
It is formatted similarly to the CodeSearchNet challenge dataset.
Each entry includes:
code: The source code of a php function or method.
docstring: The docstring or Javadoc associated with the function/method.
func_name: The name of the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/php-treesitter-dedupe-filtered-datasetsV2.CodeSearchNet-php-queries-corpus
Dataset Card for "CodeSearchNet-php-queries-corpus"
More Information needed
terminal_bench_2_r2egym_nl2bash_stack_bugsseq_stack_php_v2_20260222_044012smollm3-stack-v2-PHPCOIN-PHP-YOLOv11n-Dataset
Coin Optical Identification and Numeration for the Philippine Peso - Dataset
About
This repo contains the dataset i gathered, annotated, and processed to train a yolov11n model for detecting and counting philippine peso coins.
Folder Structure
I organized the repo into three main folders to keep things clean:
/raw - contains the original unedited photos.
/annotated - contains the images and their respective label files from the annotation tool… See the full description on the dataset page: https://huggingface.co/datasets/ml-remunn/COIN-PHP-YOLOv11n-Dataset.
