datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.variance_analysis
Variance Analysis
This repository contains the data, scripts, and generated figures used for
variance analysis experiments.
Contents
data/SmolLM: SmolLM GSM8K rollout data.
data/Qwen3: Qwen3 math rollout data, including 1.7B 512 x 512 and
4B 1024 x 128 samples.
data/Maze/variance: Maze rollout data for variance analysis.
outputs: generated JSON summaries and figures.
*.py and run_*.sh: analysis, plotting, and Slurm launch scripts.
See data/README.md for additional data… See the full description on the dataset page: https://huggingface.co/datasets/max-rl/variance_analysis.perplexity_analysis
Perplexity Analysis
This repository contains the data, scripts, and generated figures used for
perplexity analysis experiments.
Contents
data/Qwen3: rollout data for Qwen3 1.7B and 4B Base, GRPO, and MaxRL
models on AIME25 and BeyondAIME.
data/Maze/perplexity: maze rollout data and derived perplexity analysis
artifacts.
outputs: generated JSON summaries and figures for Qwen3 analyses.
*.py: analysis and plotting scripts.
See data/README.md for additional data details… See the full description on the dataset page: https://huggingface.co/datasets/max-rl/perplexity_analysis.misc-cfo-testing-cfo-analysis
misc-cfo-testing CFO analysis
This folder is the self-contained analysis output for:
C:\Users\15255\Desktop\Research\CSE237D\morty_data\misc-cfo-testing
The source data and the Weyl pipeline are read-only. All generated scripts,
fingerprints, statistics, logs, PNGs, and SVGs remain in this analysis folder.
Start with RESULTS.md.
Dataset and estimator parameters
Experiments: faraday (4 min), reboot (5 min), reboot-10m (10 min)
Receiver: pluto11
Input sample rate:… See the full description on the dataset page: https://huggingface.co/datasets/Morty0311/misc-cfo-testing-cfo-analysis.chess-stockfish-analysistrajectory_analysisrepro-understanding-lora-as-knowledge-memory-an-empirical-analysis-traces
Agent traces
Agent sessions published from a Trackio Logbook.
realistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.han-humanoid-facial-emotion-analysis-v1
Humanoid Facial Emotion Analysis Dataset
Overview
Dataset ini berisi fitur ekstraksi wajah dari sistem vision humanoid
untuk mendeteksi emosi manusia.
Features
eyebrow_raise_ratio
eye_openness_score
mouth_curve_angle
facial_muscle_tension_index
gaze_focus_score
head_tilt_degree
Target
detected_emotion
happy
neutral
sad
angry
surprised
cost-analysis
Cost Analysis
Cloud API vs on-device inference cost comparison.
At 10K queries/day: Save $18,249/year with on-device.
At 100K queries/day: Save $182,499/year.
🚀 dispatchAI
musiciwant-sensory-music-analysis
Music I Want — Sensory Music Analysis (1,000-song sample)
A free, openly-licensed sample of the Music I Want catalog: songs scored on the
dimensions that decide how music feels. It is the independent answer to the
audio-features data that disappeared when Spotify deprecated its Audio Features
API in November 2024 — and it carries fields nothing else does.
sample-1000.csv / sample-1000.json — 1,000 songs, evenly sampled across the catalog's eras and intensity range.
Full… See the full description on the dataset page: https://huggingface.co/datasets/agreenbox/musiciwant-sensory-music-analysis.qwen3-1.7b-512x512-variance-analysisqwen3-4b-1024x128-variance-analysissentiment-meta-analysisadvanced-readability-analysis
Advanced Readability Analysis
This dataset provides rich syntactic and lexical complexity features calculated from English text snippets. It is designed to help researchers study the underlying factors that influence reading difficulty, especially in cases where traditional readability formulas yield conflicting results.
The source text is pulled from the training split of the agentlans/readability dataset.
The linguistic annotations and complexity metrics were computed using a… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/advanced-readability-analysis.magpie_analysis_nomic
