language-decoded
language-decoded-loralanguage-decoded-lora-phase-2-the-stack-v1-condition-1-en-5klanguage-decoded-lora-phase-2-the-stack-v1-condition-2-es-5klanguage-decoded-lora-phase-2-the-stack-v1-condition-2-ur-5klanguage-decoded-lora-phase-2-the-stack-v1-condition-2-zh-5klanguage-decoded-lora-phase-2-the-stack-v1-condition-3-zh-5k
language-decoded-data
Language Decoded | Multilingual Code Dataset
Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models
Note (2026-05-18): Current Phase 3 configs use the short condition-* namespace and include 103k, 20k, and 5k sizes for Conditions 1--2. Phase 2 configs remain available under the phase-2-the-stack-v1-* namespace for reproducibility.
Multilingual Python code datasets for the Language Decoded project (part of Cohere's… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-data.language-decoded-experiments
Language Decoded — Experiment Tracking
Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper.
Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code
⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.language-decoded-community
Language Decoded — Community Code
Natively-authored multilingual code for the Language Decoded project (part of Cohere's Tiny Aya Expedition). This dataset contains code written by developers in non-English programming languages and code with significant CJK content — not mechanically transpiled or LLM-translated from English.
Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models
This data serves as the corpus for… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-community.
