legesher/language-decoded-community
Language Decoded — Community Code Natively-authored multilingual code for the Language Decoded project (part of Cohere's Tiny Aya Expedition). This dataset contains code written by developers in non-English programming languages and code with significant CJK content — not mechanically transpiled or LLM-translated from English. Experiment and proposed paper title: Language Decoded: Exploring the Impact of Native Code on Multilingual Models This data serves as the corpus for… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-community.
Add Attribution & takedown section (#9)
docs: add arxiv: tags for paper bibliography anchors (#8)
docs(readme): align with canonical source-of-truth (cond-4 framing, cond-3 vs cond-4 distinction) (#7)
feat: add Qalb native Arabic code (ar/validation)
feat: add Qalb native Arabic code (ar/train)
docs: fix CITATION.cff url to match repo, set correct type (#6)
docs: add citation, limitations section, update condition references (#5)
docs: add comprehensive README with schema, source breakdown, and usage examples
feat: add Chinese native code dataset (3,486 files from 5 sources)
Update dataset card: fix languages (am → es), add experiments link (#1)
init: create README.md
initial commit
