datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FeatureBench
FeatureBench: Agent Coding Evaluation Benchmark
Dataset Description
FeatureBench is a comprehensive benchmark designed to evaluate AI agents' capabilities in end-to-end feature-level code generation. Unlike traditional benchmarks that focus on function-level or algorithm-specific tasks, FeatureBench challenges agents to implement complete features within real-world software projects.
Key Characteristics
Feature-Level Tasks: Each task requires… See the full description on the dataset page: https://huggingface.co/datasets/LiberCoders/FeatureBench.feature_stories
Feature Stories
A large contrastive story dataset for mechanistic interpretability and alignment research.
Each row contains two short matched narratives about the same situation:
concept_text — written to express one behavioral / affective / epistemic pole
antagonist_text — the contrast pole for the same shared setup
Feature labels come from a curated concept ontology (148 classes, 1036 concept↔antagonist pairs),
with narrative_guidance explaining what each dichotomy means.… See the full description on the dataset page: https://huggingface.co/datasets/AntonKorznikov/feature_stories.EpiCoder-meta-features
EpiCoder Meta Features
This dataset contains the hierarchical meta-feature taxonomy and corresponding frequency statistics used in EpiCoder. These meta-features capture fine-grained code characteristics extracted from real-world repositories and serve as the foundation for controlled, feature-conditioned code generation.
Dataset Description
The dataset consists of two files:
1. epicoder_features.json
A hierarchical taxonomy of 17 top-level code feature… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/EpiCoder-meta-features.celestial-feature-demonstrations-v2
CELESTIAL Feature Demonstrations Dataset
Dataset Description
This dataset is part of the CELESTIAL spiritual AI platform, designed for training Mistral-7B models on spiritual and astrological guidance tasks.
Dataset Summary
Total Examples: 1500
Categories: feature_demonstration
Languages: English, Hindi (transliterated)
Format: Conversational format with tool calling examples
Dataset Structure
{
"messages": [
{"role": "user", "content": "User… See the full description on the dataset page: https://huggingface.co/datasets/dp1812/celestial-feature-demonstrations-v2.
