CoolFace
Datasetpublic

legesher/language-decoded-community

Language Decoded — Community Code Natively-authored multilingual code for the Language Decoded project (part of Cohere's Tiny Aya Expedition). This dataset contains code written by developers in non-English programming languages and code with significant CJK content — not mechanically transpiled or LLM-translated from English. Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models This data serves as the corpus for… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-community.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes56downloads
12 commits on main
4a301242mo ago

Add Attribution & takedown section (#9)

madiedgar
fe153284mo ago

docs: add arxiv: tags for paper bibliography anchors (#8)

madiedgar
daabe6d4mo ago

docs(readme): align with canonical source-of-truth (cond-4 framing, cond-3 vs cond-4 distinction) (#7)

madiedgar
2cc9daf5mo ago

feat: add Qalb native Arabic code (ar/validation)

rafaym
326ce4b5mo ago

feat: add Qalb native Arabic code (ar/train)

rafaym
5bc412b6mo ago

docs: fix CITATION.cff url to match repo, set correct type (#6)

madiedgar
e1b87436mo ago

docs: add citation, limitations section, update condition references (#5)

madiedgar
9eb1f036mo ago

docs: add comprehensive README with schema, source breakdown, and usage examples

madiedgar
da79e846mo ago

feat: add Chinese native code dataset (3,486 files from 5 sources)

madiedgar
6ee4e066mo ago

Update dataset card: fix languages (am → es), add experiments link (#1)

madiedgar
5b0b1f16mo ago

init: create README.md

madiedgar
8f8a0bf7mo ago

initial commit

madiedgar