datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RealCodeJava
RealCodeJava
Task description
RealCodeJava is a benchmark for evaluating the ability of language models to generate function bodies in real-world Java repositories. The benchmark focuses on realistic completions using project-level context and validates correctness through test execution. Dataset contains 298 tasks.
Evaluated skills: Instruction Following, Code Perception, Completion
Contributors: Dmitry Vorobiev, Pavel Zadorozhny, Rodion Levichev, Pavel Adamenko, Aidar… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/RealCodeJava.RealCode
RealCode
Task description
RealCode is a benchmark for evaluating the ability of language models to generate function bodies in real-world Python repositories. The benchmark focuses on realistic completions using project-level context and validates correctness through test execution. Dataset contains 802 tasks.
Evaluated skills: Instruction Following, Code Perception, Completion
Contributors: Pavel Zadorozhny, Rodion Levichev, Pavel Adamenko, Aidar Valeev, Dmitrii Babaev… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/RealCode.a1_code_primeintellect_real_world_swe_eval_636d
mlfoundations-dev/a1_code_primeintellect_real_world_swe_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
12.3
52.2
72.4
24.4
36.3
38.6
15.9
7.8
8.4
AIME24
Average Accuracy: 12.33% ± 1.42%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
10.00%
3
30
3
10.00%
3… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_primeintellect_real_world_swe_eval_636d.a1_code_primeintellect_real_world_sweload_in_code_primeintellect_real_world_swerealcode-translatedRealCode_enRealCodeJava_en
