CoolFace
Datasetpublic

legesher/language-decoded-data

Language Decoded | Multilingual Code Dataset Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models Note (2026-05-18): Current Phase 3 configs use the short condition-* namespace and include 103k, 20k, and 5k sizes for Conditions 1--2. Phase 2 configs remain available under the phase-2-the-stack-v1-* namespace for reproducibility. Multilingual Python code datasets for the Language Decoded project (part of Cohere's… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-data.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes381downloads
settings

This repository belongs to legesher on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namelanguage-decoded-data
visibilitypublic
licenceapache-2.0
gatedno
ownerlegesher
Account settings
legesher/language-decoded-data · CoolFace