datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
coda-llm-data
Coda LLM Project & Dataset Repository
This repository contains the full end-to-end dataset, fine-tuning scripts, evaluation suites, load testing harness, and proxy architecture for Coda LLM (Granite-4.2-8B Najdi Sales Agent).
Model Repository: mohameddalii/coda-llm
Dataset / Code Repository: mohameddalii/coda-llm-data
📁 Repository Structure
coda-llm-data/
├── data/
│ ├── raw/ # Raw generated multi-turn dialogues across domains
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/mohameddalii/coda-llm-data.CodAGE
CodAGE: A Dataset of Coding Agent-generated GitHub Events
CodAGE (Coding Agent-generated GitHub Events) is a collection of
public GitHub activity attributable to AI coding agents, mined from
GH Archive. It covers 16 agents (Copilot, Claude
Code, Devin, Cursor, OpenAI Codex, Gemini Code Assist, CodeRabbit, Amazon Q,
Jules, Sweep, Aider, SWE-agent, PR-Agent, Kiro, Windsurf, and OpenAI Codex Cloud)
across 10 GitHub event types, from 2024-01-01 to 2026-04-15.
The dataset holds 27… See the full description on the dataset page: https://huggingface.co/datasets/taher-ghaleb/CodAGE.CoDA-Bench
CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?
Authors: Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang*, Xiaoyong Du
CoDA-Bench (Code and Data-intensive Benchmark) is the first benchmark to jointly evaluate code intelligence and data intelligence of AI agents in realistic data-intensive environments.
Unlike existing benchmarks that provide oracle data directly, CoDA-Bench requires agents to:
🔍 Discover relevant data among hundreds of semantically similar files… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/CoDA-Bench.codal-bench
CODAL-Bench - Evaluating LLM Alignment to Coding Preferences
This benchmark comprises 500 random samples of CodeUltraFeedback dataset.
The benchmark includes responses of multiple closed-source LLMs that can be used as references when judging other LLMs using LLM-as-a-Judge:
OpenAI - GPT-3.5-Turbo
OpenAI - GPT-4-Turbo
Anthropic - Claude-3-sonnet-20240229
LLM Alignment Evaluation
Please refer to our GitHub repository to evaluate your own LLM on CODAL-Bench using… See the full description on the dataset page: https://huggingface.co/datasets/coseal/codal-bench.task156_codah_classification_adversarial
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task156_codah_classification_adversarial
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task156_codah_classification_adversarial.
