CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JetBrains-Research /lca-results Long Code Arena (raw results) These are the raw results from the Long Code Arena benchmark suite, as well as the corresponding model predictions. Please use the subset dropdown menu to select the necessary data relating to our six benchmarks: 🤗 Library-based code generation 🤗 CI builds repair 🤗 Project-level code completion 🤗 Commit message generation🤗 Bug localization 🤗 Module summarization tabularn<1K2 likes2k downloads1y agoHugging Face02JetBrains-Research /lca-bug-localization 🏟️ Long Code Arena (Bug localization) This is the benchmark for the Bug localization task as part of the 🏟️ Long Code Arena benchmark. The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug. The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.imagetext-generation10K<n<100K4 likes1.2k downloads2y agoHugging Face03JetBrains-Research /lca-project-level-code-completion 🏟️ Long Code Arena (Project-level code completion) This is the benchmark for Project-level code completion task as part of the 🏟️ Long Code Arena benchmark. Each datapoint contains the file for completion, a list of lines to complete with their categories (see the categorization below), and a repository snapshot that can be used to build the context. All the repositories are published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-project-level-code-completion.textn<1K0 likes854 downloads2y agoHugging Face04lca0503 /GPTspeech_encodec_v2 Dataset Card for "GPTspeech_encodec_v2" More Information needed text100K<n<1M0 likes712 downloads3y agoHugging Face05lca0503 /text_interference_vocalsoundaudio10K<n<100K0 likes531 downloads1y agoHugging Face06JetBrains-Research /lca-commit-message-generation 🏟️ Long Code Arena (Commit message generation) This is the benchmark for the Commit message generation task as part of the 🏟️ Long Code Arena benchmark. The dataset is a manually curated subset of the Python test set from the 🤗 CommitChronicle dataset, tailored for larger commits. All the repositories are published under permissive licenses (MIT, Apache-2.0, and BSD-3-Clause). The datapoints can be removed upon request. How-to from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-commit-message-generation.text1K<n<10K0 likes455 downloads2y agoHugging Face07lca0503 /audio_interference_mmluaudio100K<n<1M0 likes386 downloads1y agoHugging Face08lca0503 /soxdata_encodec Dataset Card for "soxdata_encodec" More Information needed text100K<n<1M0 likes311 downloads3y agoHugging Face09JetBrains-Research /lca-library-based-code-generation 🏟️ Long Code Arena (Library-based code generation) This is the benchmark for Library-based code generation task as part of the 🏟️ Long Code Arena benchmark. The current version includes 150 manually curated instructions asking the model to generate Python code using a particular library. The samples come from 62 Python repositories. All the samples in the dataset are based on reference example programs written by authors of the respective libraries. All the repositories are… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-library-based-code-generation.textn<1K2 likes282 downloads2y agoHugging Face10elvishelvis6 /LCA-on-the-Line0 likes229 downloads2y agoHugging Face11aguangguang /LCAR-Hallucination-Benchmark LCAR Hallucination Benchmark LCAR Hallucination Benchmark is a manually reviewed speech benchmark for studying acoustic-grounding failures in LLM-based ASR. It contains two 500-utterance suites: controlled speech synthesized with IndexTTS2 and speech derived from openly released corpora. The benchmark covers translation or transliteration, spoken or text-prompt instruction execution, unsupported repetition, and catastrophic deletion. The benchmark is a targeted stress set. It is… See the full description on the dataset page: https://huggingface.co/datasets/aguangguang/LCAR-Hallucination-Benchmark.audioautomatic-speech-recognition1K<n<10K1 likes194 downloads2mo agoHugging Face12lca0503 /text_interference_urbansound8kaudio10K<n<100K0 likes192 downloads1y agoHugging Face13JetBrains-Research /lca-ci-builds-repair 🏟️ Long Code Arena (CI builds repair) This is the benchmark for CI builds repair task as part of the 🏟️ Long Code Arena benchmark. 🛠️ Task. Given the logs of a failed GitHub Actions workflow and the corresponding repository snapshot, repair the repository contents in order to make the workflow pass. All the data is collected from repositories published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can be removed upon request. To… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-ci-builds-repair.tabularn<1K3 likes155 downloads2y agoHugging Face14Punktiert /LCA-GCS LCA-GCS: Large City Architecture - Generated Cityscape Set The Large City Architecture - Generated Cityscape dataset (LCA-GCS) is a comprehensive collection of 1,060,166 AI-generated images representing architectural features of 5,856 global cities. Created using advanced diffusion models, this dataset offers over 200 samples of various architectural types for each city with a population exceeding 100,000. LCA-GCS aims to facilitate comparative analysis, synthesis, and learning… See the full description on the dataset page: https://huggingface.co/datasets/Punktiert/LCA-GCS.1M<n<10M2 likes155 downloads16d agoHugging Face15lca0503 /INSPIRE INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval Overview INSPIRE is a benchmark for evaluating instruction-aware speech retrieval systems with open-ended instructions. It provides tools for building and evaluating speech retrieval models that can handle diverse retrieval tasks specified through natural language instructions. The benchmark includes dataset processing, feature extraction, and evaluation metrics. Motivation Traditional… See the full description on the dataset page: https://huggingface.co/datasets/lca0503/INSPIRE.audioaudio-to-audio10K<n<100K0 likes154 downloads8mo agoHugging Face16sergyinfo /h1b-lca-open-data H-1B / LCA Open Data — aggregate tables (FY2010–FY2026) Aggregate statistics on 9,331,644 US Labor Condition Applications (LCAs) — the wage-and-worksite filing every employer must submit to the Department of Labor before sponsoring an H-1B, H-1B1, or E-3 worker. Covers 126,466 employers and 782 occupations across fiscal years 2010–2026. This is a mirror. The canonical home of this dataset — always the newest release, with the data dictionary and license — is… See the full description on the dataset page: https://huggingface.co/datasets/sergyinfo/h1b-lca-open-data.0 likes137 downloads1mo agoHugging Face17open-llm-leaderboard-old /details_LeroyDyer__Mixtral_AI_LCARS_0 likes134 downloads2y agoHugging Face18LCA-PORVID /portuguese_vidtext1M<n<10M0 likes127 downloads3y agoHugging Face19lca0503 /ml2025-hw4-pokemontextn<1K1 likes127 downloads2y agoHugging Face20lca0503 /ml2025-hw4-colormapn<1K0 likes106 downloads2y agoHugging Face21JetBrains-Research /lca-module-summarization 🏟️ Long Code Arena (Module summarization) This is the benchmark for Module summarization task as part of the 🏟️ Long Code Arena benchmark. The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects. The model is required to generate such description, given the relevant context code and the intent behind the documentation. All the repositories are published under permissive licenses (MIT, Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-module-summarization.imagetext-generationn<1K1 likes104 downloads2y agoHugging Face22lca0503 /audio_interference_gsm8kaudio10K<n<100K0 likes94 downloads1y agoHugging Face23lcalvobartolome /proxann_topic_models ProxAnn Topic Models ProxAnn Topic Models provides the trained topic models used inPROXANN: Use-Oriented Evaluations of Topic Models and Document Clustering(Hoyle et al., ACL 2025). This collection includes 50-topic models for both the Bills (Adler & Wilkerson, 2008) and Wiki (Merity et al., 2017) corpora.All source datasets are available at lcalvobartolome/proxann_data. Overview Split Path Description bills_bertopic model-runs/bills/bertopic/ 50-topic… See the full description on the dataset page: https://huggingface.co/datasets/lcalvobartolome/proxann_topic_models.n<1K0 likes83 downloads11mo agoHugging Face24LCA-PORVID /delexicalized_n_gramstext100K<n<1M0 likes80 downloads3y agoHugging Face25lcampillos /ctebmsp CT-EBM-SP (Clinical Trials for Evidence-based Medicine in Spanish) Dataset Summary The Clinical Trials for Evidence-Based-Medicine in Spanish corpus is a collection of 1200 texts about clinical trials studies and clinical trials announcements: 500 abstracts from journals published under a Creative Commons license, e.g. available in PubMed or the Scientific Electronic Library Online (SciELO) 700 clinical trials announcements published in the European Clinical Trials… See the full description on the dataset page: https://huggingface.co/datasets/lcampillos/ctebmsp.texttoken-classification10K<n<100K4 likes79 downloads4y agoHugging Face26lcalvobartolome /proxann_data PROXANN Data PROXANN Data provides the corpora used for training and evaluating topic models inPROXANN: Use-Oriented Evaluations of Topic Models and Document Clustering(Hoyle et al., ACL 2025). This repository contains two dataset — Bills and Wiki — each with training (with contextualized embeddings) and test (metadata-only) splits. Structure Split File Rows Description bills_train bills_train.metadata.embeddings.jsonl.all-MiniLM-L6-v2.parquet 32,661… See the full description on the dataset page: https://huggingface.co/datasets/lcalvobartolome/proxann_data.text10K<n<100K0 likes79 downloads10mo agoHugging Face27lca0503 /amazon_tts_encodec_v2 Dataset Card for "amazon_tts_encodec_v2" More Information needed text100K<n<1M0 likes77 downloads3y agoHugging Face28open-llm-leaderboard /LeroyDyer__LCARS_AI_001-detailsgated Dataset Card for Evaluation run of LeroyDyer/LCARS_AI_001 Dataset automatically created during the evaluation run of model LeroyDyer/LCARS_AI_001 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__LCARS_AI_001-details.tabular10K<n<100K0 likes69 downloads2y agoHugging Face29JetBrains-Research /lca-example-generation 🏟️ Long Code Arena (Example Generation) This is the benchmark for Example Generation task, a specific case of Code Generatoin, as part of 🏟️ Long Code Arena benchmark. The dataset currently contains 150 samples, each in the form of a library and an instruction for a model to solve a particular task using the library. How-to 🚧 This section is under construction 🚧 Dataset Structure 🚧 This section is under construction 🚧 textn<1K0 likes68 downloads2y agoHugging Face30lcama /elon-tweetstext1K<n<10K2 likes60 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.