arp
Datasets
All datasets matching “arp”swe-cleanroom-trajs
swe-cleanroom-trajs
Procedurally generated SWE-agent trajectories — no LLM, no GPU — as a clean-room counterpart to
AlienKevin/SWE-ZERO-12M-trajectories (which rolls out a 1.7B model). ~40k trajectories across 124 repos
(5 languages), mini-swe-agent-1 format.
How it's generated (no LLM):
TOOL rhythm: a per-step 2nd-order Markov chain fit on ricdomolm/mini-coder-trajs-400k (verb sequence).
ARGUMENTS: a tree-sitter code graph + the gold PR patch decide which file/symbol/lines… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-cleanroom-trajs.tts-ranking-dataswe-zero-free
SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM
SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free.
This dataset has 12,238,610 mini-swe-agent trajectories across 119,084 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free.ar_pixmocapqatrans_instructpaper-trail-datalibrispeech-arpabet-processed
LibriSpeech ARPAbet Processed Dataset
Pre-processed dataset for training ARPAbet phoneme recognition models using CTC loss.
Dataset Description
This dataset is derived from LibriSpeech (train-clean-100 split) with the following preprocessing:
Audio: Resampled to 16kHz, normalized using Wav2Vec2 feature extractor
Labels: Text transcriptions converted to ARPAbet phoneme sequences using CMU Pronouncing Dictionary
Filtering: Samples with out-of-vocabulary words (not in CMU… See the full description on the dataset page: https://huggingface.co/datasets/davidggphy/librispeech-arpabet-processed.
