legesher/language-decoded-data
Language Decoded | Multilingual Code Dataset Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models Note (2026-05-18): Current Phase 3 configs use the short condition-* namespace and include 103k, 20k, and 5k sizes for Conditions 1--2. Phase 2 configs remain available under the phase-2-the-stack-v1-* namespace for reproducibility. Multilingual Python code datasets for the Language Decoded project (part of Cohere's… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-data.
Language Decoded | Multilingual Code Dataset
Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models
Note (2026-05-18): Current Phase 3 configs use the shortcondition-*namespace and include103k,20k, and5ksizes for Conditions 1--2. Phase 2 configs remain available under thephase-2-the-stack-v1-*namespace for reproducibility.
Multilingual Python code datasets for the Language Decoded project (part of Cohere's Tiny Aya Expedition). See legesher/language-decoded-experiments for the canonical project description, the full experimental ladder, and the paper-grade evaluation results.
Research Question
How does fine-tuning on non-English code — whether transpiled, mixed-native, or fully translated — affect a model's multilingual reasoning, and how does that impact _differ_ from fine-tuning on English code?
Prior work (Aryabumi et al., 2024 -- "To Code or Not to Code") demonstrated that including English code in pre-training data improves downstream reasoning performance by approximately 8%. However, that study only tested English code. This dataset enables the natural follow-up: how does the impact of non-English code differ from English code, and how does that vary by language, structure, and corpus construction?
Dataset Description
This dataset provides filtered, quality-controlled Python source code in multiple configurations: the original English (cond-1); three Legesher-transpiled variants (cond-2 zh/es/ur, with Python's reserved words translated to the target language); a community-collected raw native-source corpus (cond-3); strictly native code (cond-4, pending); and a model-translated set (cond-5, where c4ai-aya-expanse-32b translates everything translatable inside the file). Python source for Conditions 1, 2, and 5 is drawn from bigcode/the-stack-v2-dedup (Python subset) for the current Phase 3 configs; the legacy phase-2-the-stack-v1-* configs are sourced from The Stack v1 (non-dedup). Conditions 3 and 4 draw on natively-authored or community-contributed code (see those conditions below).
Source-file control
Cond-1, cond-2, and cond-5 all train on the same 5,000-file subset drawn from bigcode/the-stack-v2-dedup (with a parallel 20k subset for the 20k tier). Differences across these conditions reflect the processing pipeline (raw / transpiled / fully translated), not file-quality or content drift. Cond-3 is the deliberate exception — its source files are a different population by design (community-collected from varied online sources, potentially including non-Python files).
Source files for cond-1/2/5 are filtered using:
- AST-valid Python only (must parse without errors)
- Permissive licenses only (MIT, Apache-2.0, BSD, etc.)
- 10--1000 lines of code
- Minimum 21 GitHub stars
- No autogenerated files
- SHA-256 deduplication
Cond-2 variants are produced using Legesher v0.7.3, which translates Python's reserved words (37 keywords, 72 built-in functions, 66 exceptions, plus the numerical system for some target languages) into the target language while preserving code structure and user logic. Cond-5 takes the Legesher-transpiled output and runs it through c4ai-aya-expanse-32b via the Cohere API to translate the remaining content — identifiers, comments, docstrings, string literals, and any other natural-language wording — into the target language. Logic and structure are preserved throughout.
Available Configs
Conditions 1--2 are available in three current Phase 3 sizes: -103k full corpora, -20k random subsets sampled from the corresponding -103k config with seed 42, and -5k compact subsets. Phase 2 -32k configs are still available with the phase-2-the-stack-v1-* prefix. Condition 5 (condition-5-*-c4ai-aya-expanse-32b) is the model-translated set — currently 5k only, and raw/pre-cleanup (see the note above).
Schema
Conditions 1--2
Used by: condition-1-en-*, condition-2-zh-*, condition-2-es-*, condition-2-ur-*
Condition 5
Used by: condition-5-ur-5k-c4ai-aya-expanse-32b, condition-5-zh-5k-c4ai-aya-expanse-32b, condition-5-es-5k-c4ai-aya-expanse-32b
Condition 5 uses the conditions 1--2 schema plus an idx column. code is the full LLM-translated source (identifiers, strings, comments, and keywords); code_en is the English original. These configs are raw model output — see the note at the top of this card.
Condition 3
Used by: condition-3-zh-5k
In Phase 3, Condition 3 ("Mixed Native Sources") refers to community-collected raw Chinese code from varied online public-source repositories — reflecting how non-English Python is actually used in real-world projects. The "Mixed Native Sources" name carries from Phase 2, where it originally referred to a planned composite (native code padded with cond-2 transpiled files); in Phase 3 the "mixed" refers to the diversity of source locations, not a cond-2/native composite. The physical dataset has not changed across phases.
The schema includes a source_type column from the Phase 2 composite design, which remains "native" or "transpiled" depending on each row's origin. code_en is populated for transpiled rows (keeping them in sync with conditions 1--2) but null for native code rows, which have no English equivalent.
Condition 4
Used by: condition-4-zh-5k
Condition 4 ("Community-Contributed Native Code") is intended to contain code whose problem-solving logic is itself native — written as if a native speaker were approaching the problem, not English code that was later translated. The current dataset reflects an earlier Phase 2 attempt to assemble this corpus; community contributions were insufficient for stable training, so cond-4 was not evaluated in either Phase 2 or Phase 3. Cond-5's fully-translated data served as Phase 3's practical proxy because gathering native-authored code at scale proved difficult. Direct contributions to the cond-4 corpus are open at the `legesher/legesher-native-code` HF Space.
Uses the same schema as the language-decoded-community dataset rather than the transpilation schema, since there is no English original to reference.
Experimental Conditions
The Language Decoded experiment uses a ladder of conditions to isolate the mechanism behind code's reasoning benefit. For the full ladder including future directions, see legesher/language-decoded-experiments.
The Experimental Ladder
- Baseline → 1: Does code help at all?
- 1 → 2: Does the language Python is written in matter? (Cond-2 translates Python's reserved words; user logic preserved.)
- 2 → 3: Does code humans actually wrote in or with the target language add value beyond Legesher's mechanical translation?
- 2 → 5: Cond-2 translates only Python's reserved words; cond-5 goes further by also translating identifiers, comments, docstrings, and string literals via
c4ai-aya-expanse-32b. Logic preserved. Does full translation produce different effects than partial translation? - 3 → 5 (implicit): Human-authored vs. machine-synthesized native code.
Usage
from datasets import load_dataset
# Load full-size English code (control)
ds = load_dataset("legesher/language-decoded-data", "condition-1-en-103k")
# Load random 20k subsets
ds = load_dataset("legesher/language-decoded-data", "condition-1-en-20k")
ds = load_dataset("legesher/language-decoded-data", "condition-2-zh-20k")
ds = load_dataset("legesher/language-decoded-data", "condition-2-es-20k")
ds = load_dataset("legesher/language-decoded-data", "condition-2-ur-20k")
# Load 5k subset (for QLoRA fine-tuning)
ds = load_dataset("legesher/language-decoded-data", "condition-1-en-5k")
# Load Legesher-transpiled variants (reserved-word translation)
ds = load_dataset("legesher/language-decoded-data", "condition-2-zh-5k")
ds = load_dataset("legesher/language-decoded-data", "condition-2-es-5k")
ds = load_dataset("legesher/language-decoded-data", "condition-2-ur-5k")
# Load blended native + transpiled (condition 3)
ds = load_dataset("legesher/language-decoded-data", "condition-3-zh-5k")
# Load strictly native code (condition 4)
ds = load_dataset("legesher/language-decoded-data", "condition-4-zh-5k")
# Load model-translated code (condition 5 -- raw, pre-cleanup)
ds = load_dataset("legesher/language-decoded-data", "condition-5-ur-5k-c4ai-aya-expanse-32b")
ds = load_dataset("legesher/language-decoded-data", "condition-5-zh-5k-c4ai-aya-expanse-32b")
ds = load_dataset("legesher/language-decoded-data", "condition-5-es-5k-c4ai-aya-expanse-32b")
# Access splits
train = ds["train"]
val = ds["validation"]
# Filter condition-3 by source type
native_only = train.filter(lambda x: x["source_type"] == "native")Technical Details
Limitations
- Source bias: The Stack skews toward popular, well-starred GitHub repositories, which may not represent the full diversity of Python code in the wild.
- Keyword-only transpilation: Legesher translates Python reserved words (keywords, builtins, exceptions) but leaves comments, docstrings, string literals, and variable/function names in their original language (typically English). This means condition-2 code is a hybrid of translated keywords and English identifiers.
- Token count variation: Transpiled code may have different token counts than the English original due to multi-byte characters (especially for Chinese and Urdu), even though the code structure is identical.
- Single programming language: Currently limited to Python. Results may not generalize to other programming languages.
- Condition 4 not yet evaluated: Community contributions to the `legesher/legesher-native-code` HF Space have been insufficient for stable training. The existing
condition-4-zh-5kdata is a Phase 2 attempt limited to publicly available sources (The Stack, Wenyan, Program-in-Chinese, Qi, Mulan). Cond-5's fully-translated data served as the Phase 3 practical proxy for cond-4's "logic in the target language" goal. - Condition 5 is raw model output: The
condition-5-*configs contain prompt-leakage contamination -- translator-model preamble text, JSON wrappers, and explanation commentary leaked into string literals and identifier names, in AST-valid and AST-invalid rows alike. Cleaned configs will be published separately. See the note at the top of this card.
Citation
@misc{language-decoded-2026,
title={Language Decoded: Exploring the Impact of Native Code on Multilingual Models},
author={Madison Edgar and Saad Ahmed Bazaz and Tom Sherborne and Rashik Shahjahan and Khojasteh Mirza and Sarah Jawaid and Rafay Mustafa and Sohaib Ahmed Bazaz},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/datasets/legesher/language-decoded-data}
}Links
- Legesher on GitHub
- Tiny Aya Expedition
- bigcode/the-stack-v2-dedup (Phase 3 source)
- bigcode/the-stack (The Stack v1 — Phase 2 source)
- Language Decoded Community (native code)
- Language Decoded Experiments (tracking)
- Language Decoded LoRA (model hub)
License
Apache 2.0
Attribution & takedown
Portions of this dataset derive from publicly available code repositories (The Stack v2 for the English source subsets; public GitHub and Gitee repositories for natively-authored code), collected under a best-effort license review. If you are the author of code included here and believe it was included improperly, or you would like attribution added or your code removed:
- open a discussion on this repository's Community tab, or
- email Madi Edgar at support@legesher.com (the maintainer listed in
CITATION.cff).
We commit to citing, and on request from authors removing, any source identified as improperly included. Removals are propagated in a new dataset revision; prior revisions remain in git history unless a takedown requires otherwise.
