datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
codah
Dataset Card for COmmonsense Dataset Adversarially-authored by Humans
Dataset Summary
The COmmonsense Dataset Adversarially-authored by Humans (CODAH) is an evaluation set for commonsense
question-answering in the sentence completion style of SWAG. As opposed to other automatically generated
NLI datasets, CODAH is adversarially constructed by humans who can view feedback from a pre-trained model
and use this information to design challenging commonsense questions.… See the full description on the dataset page: https://huggingface.co/datasets/jaredfern/codah.coda-llm-data
Coda LLM Project & Dataset Repository
This repository contains the full end-to-end dataset, fine-tuning scripts, evaluation suites, load testing harness, and proxy architecture for Coda LLM (Granite-4.2-8B Najdi Sales Agent).
Model Repository: mohameddalii/coda-llm
Dataset / Code Repository: mohameddalii/coda-llm-data
📁 Repository Structure
coda-llm-data/
├── data/
│ ├── raw/ # Raw generated multi-turn dialogues across domains
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/mohameddalii/coda-llm-data.coda-lm
CODA-LM Dataset Card
CODA-LM is the multi-modal version of the CODA dataset, used in the CODA-LM paper. Both English and Chinese annotations are available. Check detailed usage in our Github repo.
Citation
@article{li2024automated,
title={Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases},
author={Li, Yanze and Zhang, Wenhua and Chen, Kai and Liu, Yanxin and Li, Pengxiang and Gao, Ruiyuan and Hong, Lanqing and Tian, Meng and Zhao… See the full description on the dataset page: https://huggingface.co/datasets/KaiChen1998/coda-lm.coda-lm-llava-format
CODA-LM Dataset Card
CODA-LM is the multi-modal version of the CODA dataset, used in the CODA-LM paper. Both English and Chinese annotations are available. Check detailed usage in our Github repo.
This repo contains the CODA-LM dataset, which has been reorganized in the LLaVA data format.
You are also welcome to check the original CODA-LM data which contains more metadata vanilla annotations.
Usage
from datasets import load_dataset
# name can be selected from… See the full description on the dataset page: https://huggingface.co/datasets/KaiChen1998/coda-lm-llava-format.eurocv
EuroCV - images dataset
This dataset was created using codaco.app.
Description
Develop better AI together: Your images help build diverse training data, making AI systems more accurate, safer, and more reliable. Join now and help improve AI!
Labels
This dataset includes the following labels:
Bounding box objects
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any purpose, including… See the full description on the dataset page: https://huggingface.co/datasets/codaco/eurocv.vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
images
CoDaCo - images dataset
This dataset was created using codaco.app.
Description
All data contributed to this campaign goes to the global CoDaCo datasets.
Labels
This dataset includes the following labels:
Captions
Bounding box objects
Bounding box texts
Contains objects
Tags
Emotions
AI generated
Quality rating
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any purpose, including commercially… See the full description on the dataset page: https://huggingface.co/datasets/codaco/images.CodAGE
CodAGE: A Dataset of Coding Agent-generated GitHub Events
CodAGE (Coding Agent-generated GitHub Events) is a collection of
public GitHub activity attributable to AI coding agents, mined from
GH Archive. It covers 16 agents (Copilot, Claude
Code, Devin, Cursor, OpenAI Codex, Gemini Code Assist, CodeRabbit, Amazon Q,
Jules, Sweep, Aider, SWE-agent, PR-Agent, Kiro, Windsurf, and OpenAI Codex Cloud)
across 10 GitHub event types, from 2024-01-01 to 2026-04-15.
The dataset holds 27… See the full description on the dataset page: https://huggingface.co/datasets/taher-ghaleb/CodAGE.microscope-dataCoDA-Bench
CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?
Authors: Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang*, Xiaoyong Du
CoDA-Bench (Code and Data-intensive Benchmark) is the first benchmark to jointly evaluate code intelligence and data intelligence of AI agents in realistic data-intensive environments.
Unlike existing benchmarks that provide oracle data directly, CoDA-Bench requires agents to:
🔍 Discover relevant data among hundreds of semantically similar files… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/CoDA-Bench.coda-daaamvideos
CoDaCo - videos dataset
This dataset was created using codaco.app.
Description
All data contributed to this campaign goes to the global CoDaCo datasets.
Labels
This dataset includes the following labels:
Captions
Spoken text
Bounding box texts
Bounding box objects
Contains objects
Performed actions
Tags
Emotions
AI generated
Quality rating
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any… See the full description on the dataset page: https://huggingface.co/datasets/codaco/videos.vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5
vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5
Description
VLN Navigation dataset with 100% of iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 1.0
mateoguaman/iphone_stairs_ramps: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5.german-emotional-speech
German Emotional Speech - audios dataset
This dataset was created using codaco.app.
Description
Audio training data for speech emotion classification models.
Labels
This dataset includes the following labels:
Emotions
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any purpose, including commercially, as long as
you give appropriate credit. See LICENSE for the full terms.
Required… See the full description on the dataset page: https://huggingface.co/datasets/codaco/german-emotional-speech.text
CoDaCo - texts dataset
This dataset was created using codaco.app.
Description
All data contributed to this campaign goes to the global CoDaCo datasets.
Labels
This dataset includes the following labels:
Summaries
Entities
Tags
Emotions
AI generated
Quality rating
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any purpose, including commercially, as long as
you give appropriate credit. See… See the full description on the dataset page: https://huggingface.co/datasets/codaco/text.audio
CoDaCo - audios dataset
This dataset was created using codaco.app.
Description
All data contributed to this campaign goes to the global CoDaCo datasets.
Labels
This dataset includes the following labels:
Spoken text
Tags
Emotions
AI generated
Quality rating
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any purpose, including commercially, as long as
you give appropriate credit. See LICENSE… See the full description on the dataset page: https://huggingface.co/datasets/codaco/audio.coda_every1_25pct_sub5
coda_every1_25pct_sub5
Description
Processed coda dataset with filter_every_nth=1, 25% of data, and num_subsampled_points=5
Processing Parameters
mateoguaman/coda:
exclude_outliers_pct: 3
filter_by_curvature: true
filter_every_nth: 1
horizon:
300: 0.5
500: 0.5
num_subsampled_points: 5
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda: 0.25
split: train
Validation dataset:
mixer: mateoguaman/coda: 0.25split:… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/coda_every1_25pct_sub5.X-CODAHvlmn_tartandrive100_scand50_coda25_spot100_sub5_filtered_trajectories_training_25_fixedcoda*The Color Dataset* (CoDa) is a probing dataset to evaluate the representation of visual properties in language models. CoDa consists of color distributions for 521 common objects, which are split into 3 groups: Single, Multi, and Any.codacoda
CODa Navigation Dataset
This dataset contains navigation trajectory data for robotic navigation tasks. Each example includes an RGB image, a language goal describing the desired navigation target, and 2D/3D trajectories showing the path to the goal.
Dataset Structure
image: RGB image from the robot's viewpoint
lang_goal: Natural language instruction describing the navigation goal
trajectory_2d: 2D trajectory coordinates (pixel space)
trajectory_3d: 3D trajectory… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/coda.vlmn_iphone100_tartandrive100_scand50_coda25_spot100_insta360100_sub5
vlmn_iphone100_tartandrive100_scand50_coda25_spot100_insta360100_sub5
Description
VLN Navigation dataset with 100% of iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, 100% of in-domain spot data, and 100% of insta360 data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphone100_tartandrive100_scand50_coda25_spot100_insta360100_sub5.coda_every1_25pct_rdp
coda_every1_25pct_rdp
Description
Processed coda dataset with filter_every_nth=1, 25% of data, and subsampled with RDP.
Processing Parameters
mateoguaman/coda:
exclude_outliers_pct: 3
filter_by_curvature: true
filter_every_nth: 1
horizon:
300: 0.5
500: 0.5
num_subsampled_points: -1
subsample_method: rdp
tolerance: 25
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda: 0.25
split: train
Validation dataset:… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/coda_every1_25pct_rdp.vlmn_tartandrive100_scand50_coda25_spot100_sub5_filtered_trajectories_training_10_fixedv0.4.1_codatasetCodalm_test_driving_suggestionvlmn_tartandrive100_scand50_coda25_spot100_sub5_filtered_trajectories_training_10_fixedvlmn_tartandrive100_scand50_coda25_spot100_insta360100_sub5
vlmn_tartandrive100_scand50_coda25_spot100_insta360100_sub5
Description
VLN Navigation dataset with 100% of tartandrive data, 50% of scand data, 25% of coda data, 100% of in-domain spot data, and 100% of insta360 data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 1.0
mateoguaman/insta360_every1_sub5: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_tartandrive100_scand50_coda25_spot100_insta360100_sub5.vlmn_tartandrive100_scand50_coda25_spot100_sub5_filtered_trajectories_training_25_fixed
