datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mediawiki-code2code-search
MediaWiki Code2Code Search — Dataset
Pre-computed retrieval artifacts for MediaWiki Code2Code Search, a neural system for the
semantic discovery of open-source software entities (functions, types, templates) across the
MediaWiki / Wikimedia ecosystem.
This Hugging Face dataset is a complementary mirror of the pre-computed artifacts archived on
Zenodo (10.5281/zenodo.20586256). The Hugging Face copy
makes the corpus browsable in the dataset Viewer and easy to pull with the… See the full description on the dataset page: https://huggingface.co/datasets/ftosoni/mediawiki-code2code-search.findfile-3.3k-mediawiki
findfile-3.3k-mediawiki
3,291 wiki articles converted to hierarchical directory structures for file navigation tasks.
Format
JSONL with fields:
folder_name: sanitized directory name
wiki: source wiki name
title: original article title
wikitext: raw MediaWiki markup
structure: nested dict of files/dirs with contents
Source
Various fan/community MediaWiki sites. Most use CC-BY-SA, some use GFDL or CC-BY-NC-SA.
Usage
Used for training models on… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/findfile-3.3k-mediawiki.mediawiki-dolma
MediaWiki Datasets
buzz_sources_284_mediawiki
