datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lca-results
Long Code Arena (raw results)
These are the raw results from the Long Code Arena benchmark suite, as well as the corresponding model predictions.
Please use the subset dropdown menu to select the necessary data relating to our six benchmarks:
🤗 Library-based code generation
🤗 CI builds repair
🤗 Project-level code completion
🤗 Commit message generation🤗 Bug localization
🤗 Module summarization
lca-bug-localization
🏟️ Long Code Arena (Bug localization)
This is the benchmark for the Bug localization task as part of the
🏟️ Long Code Arena benchmark.
The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug.
The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.lca-project-level-code-completion
🏟️ Long Code Arena (Project-level code completion)
This is the benchmark for Project-level code completion task as part of the 🏟️ Long Code Arena benchmark.
Each datapoint contains the file for completion, a list of lines to complete with their categories (see the categorization below), and a repository snapshot that can be used to build the context.
All the repositories are published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-project-level-code-completion.GPTspeech_encodec_v2
Dataset Card for "GPTspeech_encodec_v2"
More Information needed
text_interference_vocalsoundlca-commit-message-generation
🏟️ Long Code Arena (Commit message generation)
This is the benchmark for the Commit message generation task as part of the
🏟️ Long Code Arena benchmark.
The dataset is a manually curated subset of the Python test set from the 🤗 CommitChronicle dataset, tailored for larger commits.
All the repositories are published under permissive licenses (MIT, Apache-2.0, and BSD-3-Clause). The datapoints can be removed upon request.
How-to
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-commit-message-generation.audio_interference_mmlusoxdata_encodec
Dataset Card for "soxdata_encodec"
More Information needed
lca-library-based-code-generation
🏟️ Long Code Arena (Library-based code generation)
This is the benchmark for Library-based code generation task as part of the
🏟️ Long Code Arena benchmark.
The current version includes 150 manually curated instructions asking the model to generate Python code using a particular library.
The samples come from 62 Python repositories.
All the samples in the dataset are based on reference example programs written by authors of the respective libraries.
All the repositories are… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-library-based-code-generation.LCA-on-the-LineLCAR-Hallucination-Benchmark
LCAR Hallucination Benchmark
LCAR Hallucination Benchmark is a manually reviewed speech benchmark for
studying acoustic-grounding failures in LLM-based ASR. It contains two
500-utterance suites: controlled speech synthesized with IndexTTS2 and speech
derived from openly released corpora. The benchmark covers translation or
transliteration, spoken or text-prompt instruction execution, unsupported
repetition, and catastrophic deletion.
The benchmark is a targeted stress set. It is… See the full description on the dataset page: https://huggingface.co/datasets/aguangguang/LCAR-Hallucination-Benchmark.text_interference_urbansound8klca-ci-builds-repair
🏟️ Long Code Arena (CI builds repair)
This is the benchmark for CI builds repair task as part of the
🏟️ Long Code Arena benchmark.
🛠️ Task. Given the logs of a failed GitHub Actions workflow and the corresponding repository snapshot,
repair the repository contents in order to make the workflow pass.
All the data is collected from repositories published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can be removed upon request.
To… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-ci-builds-repair.LCA-GCS
LCA-GCS: Large City Architecture - Generated Cityscape Set
The Large City Architecture - Generated Cityscape dataset (LCA-GCS) is a comprehensive collection of 1,060,166 AI-generated images representing architectural features of 5,856 global cities. Created using advanced diffusion models, this dataset offers over 200 samples of various architectural types for each city with a population exceeding 100,000. LCA-GCS aims to facilitate comparative analysis, synthesis, and learning… See the full description on the dataset page: https://huggingface.co/datasets/Punktiert/LCA-GCS.INSPIRE
INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval
Overview
INSPIRE is a benchmark for evaluating instruction-aware speech retrieval systems with open-ended instructions. It provides tools for building and evaluating speech retrieval models that can handle diverse retrieval tasks specified through natural language instructions. The benchmark includes dataset processing, feature extraction, and evaluation metrics.
Motivation
Traditional… See the full description on the dataset page: https://huggingface.co/datasets/lca0503/INSPIRE.h1b-lca-open-data
H-1B / LCA Open Data — aggregate tables (FY2010–FY2026)
Aggregate statistics on 9,331,644 US Labor Condition Applications (LCAs) —
the wage-and-worksite filing every employer must submit to the Department of Labor
before sponsoring an H-1B, H-1B1, or E-3 worker. Covers 126,466 employers
and 782 occupations across fiscal years 2010–2026.
This is a mirror. The canonical home of this dataset — always the newest release, with the data dictionary and license — is… See the full description on the dataset page: https://huggingface.co/datasets/sergyinfo/h1b-lca-open-data.details_LeroyDyer__Mixtral_AI_LCARS_portuguese_vidml2025-hw4-pokemonml2025-hw4-colormaplca-module-summarization
🏟️ Long Code Arena (Module summarization)
This is the benchmark for Module summarization task as part of the
🏟️ Long Code Arena benchmark.
The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects.
The model is required to generate such description, given the relevant context code and the intent behind the documentation.
All the repositories are published under permissive licenses (MIT, Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-module-summarization.audio_interference_gsm8kproxann_topic_models
ProxAnn Topic Models
ProxAnn Topic Models provides the trained topic models used inPROXANN: Use-Oriented Evaluations of Topic Models and Document Clustering(Hoyle et al., ACL 2025).
This collection includes 50-topic models for both the Bills (Adler & Wilkerson, 2008) and Wiki (Merity et al., 2017) corpora.All source datasets are available at lcalvobartolome/proxann_data.
Overview
Split
Path
Description
bills_bertopic
model-runs/bills/bertopic/
50-topic… See the full description on the dataset page: https://huggingface.co/datasets/lcalvobartolome/proxann_topic_models.delexicalized_n_gramsctebmsp
CT-EBM-SP (Clinical Trials for Evidence-based Medicine in Spanish)
Dataset Summary
The Clinical Trials for Evidence-Based-Medicine in Spanish corpus is a collection of 1200 texts about clinical trials studies and clinical trials announcements:
500 abstracts from journals published under a Creative Commons license, e.g. available in PubMed or the Scientific Electronic Library Online (SciELO)
700 clinical trials announcements published in the European Clinical Trials… See the full description on the dataset page: https://huggingface.co/datasets/lcampillos/ctebmsp.proxann_data
PROXANN Data
PROXANN Data provides the corpora used for training and evaluating topic models inPROXANN: Use-Oriented Evaluations of Topic Models and Document Clustering(Hoyle et al., ACL 2025).
This repository contains two dataset — Bills and Wiki — each with training (with contextualized embeddings) and test (metadata-only) splits.
Structure
Split
File
Rows
Description
bills_train
bills_train.metadata.embeddings.jsonl.all-MiniLM-L6-v2.parquet
32,661… See the full description on the dataset page: https://huggingface.co/datasets/lcalvobartolome/proxann_data.amazon_tts_encodec_v2
Dataset Card for "amazon_tts_encodec_v2"
More Information needed
LeroyDyer__LCARS_AI_001-details
Dataset Card for Evaluation run of LeroyDyer/LCARS_AI_001
Dataset automatically created during the evaluation run of model LeroyDyer/LCARS_AI_001
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__LCARS_AI_001-details.lca-example-generation
🏟️ Long Code Arena (Example Generation)
This is the benchmark for Example Generation task, a specific case of Code Generatoin, as part of
🏟️ Long Code Arena benchmark.
The dataset currently contains 150 samples, each in the form of a library and an instruction for a model to solve a particular task using the library.
How-to
🚧 This section is under construction 🚧
Dataset Structure
🚧 This section is under construction 🚧
elon-tweets
