datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Taur_CoT_Analysis_Project___gpt-4o-mini-2024-07-18SWE-smith-rs-gpt-5-mini-trajectories
SWE-smith-rs gpt-5-mini trajectories
Top-level fields are unchanged:
messages, instance_id, resolved, model, traj_id, patch
messages now preserves tool calls in chat-template-compatible structure:
assistant tool_calls with function.name and parsed function.arguments
tool tool_call_id
Generated at: 2026-02-23 05:28:29Z
Rows: 3953
Shards: 16
Skipped runs (missing/corrupt trajectory): 5
gpt-4o-minigpt-5-mini_swebench_verified_traj
Live-SWE-agent: live, self-evolving software agent
Live-SWE-agent is the first live, runtime self-evolving software engineering agent that expands and revises its own capabilities on the fly while working on a real-world issue.
Our key insight is that software agents are themselves software systems, and modern LLM-based agents already possess the intrinsic capability to extend or modify their own behavior at runtime.
airoboros_riddle_instructions_gpt-4o-mini20260802_mini-v2.4.2_gpt-5-6-sol-xhighGPQA_GPT-4o-mini
Features
id: Unique identifier for each problem
question: Original question text
options: List of 4 answer options, where first element is the correct answer
prompts: List of 1000 prompts with randomized option orders for each problem.
responses: List of 1000 model responses for each problem
correct_bools: List of 1000 boolean values indicating if each response was correct
split: Dataset split identifier ('diamond' or 'main')
REASONING_evalchemy_64_sharded_gpt-4o-mini
Dataset card for REASONING_evalchemy_64_sharded_gpt-4o-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"context": [
{
"content": "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.You are given a 0-indexed array nums of n integers and an integer target.\nYou are initially positioned at index 0. In one step, you can… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_64_sharded_gpt-4o-mini.20260730_mini-v2.2.8_gpt-5-6-solopen2math-1M-gpt-4.1-minigpt-5-miniSageLM-gpt-4o-mini-ttsvamos_10pct_gpt5_mini
vamos_10pct_gpt5_mini
Description
VLN Navigation dataset with 100% of tartandrive data, 50% of scand data, 25% of coda data, 100% of in-domain spot data, and 10% of annotated/augmented data using gpt5-mini. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/vlmn_tartandrive100_scand50_coda25_spot100_sub5: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vamos_10pct_gpt5_mini.open-orca_gpt-4o-mini_scale_x4slim-orca_gpt-4o-mini_scale_x4math-multilingual-gpt4.1-mini-sampleddetails_TheBloke__orca_mini_v3_13B-GPTQ
Dataset Card for Evaluation run of TheBloke/orca_mini_v3_13B-GPTQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/orca_mini_v3_13B-GPTQ on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__orca_mini_v3_13B-GPTQ.airoboros_writing_instructions_gpt-4o-miniGPQA_GPT-4o-mini_v2
Features
'instruction': Empty string for each row
'problem': A single string corresponding to the prompt used for generating all the samples
'level': Empty string for each row
'type': Corresponding to main or diamond split
'options': All the possible answers for the multiple choice, in the randomized order used in the prompt
'solution': The explanation for the correct answer
'samples': 1000 responses for each question
'answer_correct': 1000 boolean corresponding to whether each… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/GPQA_GPT-4o-mini_v2.aime_gpt-4o-mini_responses_evaluated_flatturnPersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini
Dataset card for PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"dimension_name": "programming_expertise",
"dimension_values": [
"Novice",
"Intermediate",
"Advanced"
],
"dimension_description": "Represents the user's practical fluency in software engineering. It shapes how they decompose problems, choose abstractions, weigh… See the full description on the dataset page: https://huggingface.co/datasets/JasonYan777/PersonaSignal-PersonalizedResponse-Programming-Expertise-gpt-5-mini.20260508_mini-v2.2.6_gpt-5-5-xhighPersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini
Dataset card for PersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"dimension_name": "programming_expertise",
"dimension_values": [
"Novice",
"Intermediate",
"Advanced"
],
"dimension_description": "Represents the user's practical fluency in software engineering. It shapes how they decompose problems, choose abstractions, weigh… See the full description on the dataset page: https://huggingface.co/datasets/JasonYan777/PersonaSignal-PerceivabilityTest-Programming-Expertise-gpt-5-mini.20260507_mini-v2.2.6_gpt-5-5gpt-5-mini-rebench-v2-cpptulu-3-sft-single-turn-gpt4o-mini-thoughts-original-responsesgpt-4o-mini-1k-ifmath-multilingual-gpt4.1-minicode_contests_10k_OG_10k_New_Questions_GPT5-minioh-dcft-v1.3_no-curation_gpt-4o-mini_scale_4x
