legesher
Datasets
All datasets matching “legesher”language-decoded-data
Language Decoded | Multilingual Code Dataset
Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models
Note (2026-05-18): Current Phase 3 configs use the short condition-* namespace and include 103k, 20k, and 5k sizes for Conditions 1--2. Phase 2 configs remain available under the phase-2-the-stack-v1-* namespace for reproducibility.
Multilingual Python code datasets for the Language Decoded project (part of Cohere's… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-data.language-decoded-experiments
Language Decoded — Experiment Tracking
Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper.
Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code
⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.language-packs
Legesher Language Packs
Legesher's language packs 📦 are the localization layer for code; the fixed
vocabulary of a programming language in other natural languages.
Every reserved word within a programming language's grammar (keywords,
builtins, exceptions, error messages) is compiled into a language pack that,
when utilized by Legesher, allows code to be in any language and it compiles
all the same.
Release
Languages supported
Target
Versions
License
legesher-i18n
Commit… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-packs.language-corpus
Legesher language corpus
What this corpus is
A concept-centric corpus of reserved-word translations for programming
languages: every candidate rendering ("option") of a programming-language
concept in a human locale that anyone -- Legesher, another open-source project,
a community reviewer, or a model -- ever offered, together with every
endorsement a person attached to it.
This is a growing dataset, and append-only by design. New locales, new
concepts, new… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-corpus.language-canon
Legesher language canon
What this Canon is
A centralized standard, created and curated by language communities, for what
programming concepts are called in each natural language.
It holds language-specific renderings of programming keywords, builtins and
exceptions, and documents where each language stands in that process.
The language-canon builds on
legesher/language-corpus,
which gathers every rendering anyone has proposed. The corpus is the full
record of… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-canon.language-decoded-community
Language Decoded — Community Code
Natively-authored multilingual code for the Language Decoded project (part of Cohere's Tiny Aya Expedition). This dataset contains code written by developers in non-English programming languages and code with significant CJK content — not mechanically transpiled or LLM-translated from English.
Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models
This data serves as the corpus for… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-community.
