datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-cleanroom-trajs
swe-cleanroom-trajs
Procedurally generated SWE-agent trajectories — no LLM, no GPU — as a clean-room counterpart to
AlienKevin/SWE-ZERO-12M-trajectories (which rolls out a 1.7B model). ~40k trajectories across 124 repos
(5 languages), mini-swe-agent-1 format.
How it's generated (no LLM):
TOOL rhythm: a per-step 2nd-order Markov chain fit on ricdomolm/mini-coder-trajs-400k (verb sequence).
ARGUMENTS: a tree-sitter code graph + the gold PR patch decide which file/symbol/lines… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-cleanroom-trajs.tts-ranking-dataswe-zero-free
SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM
SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free.
This dataset has 12,238,610 mini-swe-agent trajectories across 119,084 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free.ar_pixmocapqatrans_instructpaper-trail-datalibrispeech-arpabet-processed
LibriSpeech ARPAbet Processed Dataset
Pre-processed dataset for training ARPAbet phoneme recognition models using CTC loss.
Dataset Description
This dataset is derived from LibriSpeech (train-clean-100 split) with the following preprocessing:
Audio: Resampled to 16kHz, normalized using Wav2Vec2 feature extractor
Labels: Text transcriptions converted to ARPAbet phoneme sequences using CMU Pronouncing Dictionary
Filtering: Samples with out-of-vocabulary words (not in CMU… See the full description on the dataset page: https://huggingface.co/datasets/davidggphy/librispeech-arpabet-processed.ar_pixmocapqatrans2_instructarpa-osmer-weatherstations
OSMER ARPA FVG — Validated Weather Observations
Validated meteorological observations from the OSMER network of ARPA FVG (Regional
Environmental Protection Agency of Friuli-Venezia Giulia, NE Italy): temperature, humidity,
precipitation, wind (mean/vector/gust + direction), global radiation, pressure, leaf wetness and
soil temperatures. Two resolutions: hourly point observations and daily aggregates
(min/max/mean), packaged as Parquet for analysis and ML.
This is a self-archived… See the full description on the dataset page: https://huggingface.co/datasets/jotone/arpa-osmer-weatherstations.Pothole_classificationA dataset for efficient pothole classification which contains more than 400 images collected over 50 km road from kerala.
Pulaar_Dictionary
Saggitorde — Dictionnaire Pulaar / Français / Anglais
Dictionnaire multilingue Pulaar ↔ Français ↔ Anglais extrait et enrichi à partir du Saggitorde de Ceerno Abuu Sih, enrichi par les terminologies de l'ARPRIM, ILN, MFS et d'autres sources spécialisées.
Statistiques
Statistique
Valeur
Nombre total d'entrées
1862
Nombre de domaines
24
Nombre de sources
24
Langues
Pulaar (ff), Français (fr), Anglais (en)
Structure des données… See the full description on the dataset page: https://huggingface.co/datasets/ARPRIM/Pulaar_Dictionary.ARPA-Armenian-Paraphrase-Corpus
Dataset Description
We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language.
Dataset Summary
The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.Spotify_Audio_features_2.3Marp-decision131-robocasa-libero-20
Decision-131 complete RoboCasa + LIBERO-PRO subset
This repository contains the complete data package requested for the 20 task
identities selected in ARP decision 131.
RoboCasa human demonstrations
10 target tasks
5,101 episodes
2,680,229 frames at 20 Hz
Actions, robot state, episode/task metadata, and three synchronized 256×256
H.264 camera streams
176 MP4 files and 8 Parquet files
The episodes were selected from
lerobot/robocasa_target_human_unified
at… See the full description on the dataset page: https://huggingface.co/datasets/mojimoji61/arp-decision131-robocasa-libero-20.swe-zero-grounded-fullar_pixmodocsother_instructsyspin-kannada-ttsgenerations-nemotron-nano-9b-v2-simnpo-gentle-igm-10bARPO-SFT-54K
Agentic Reinforced Policy Optimization (ARPO) Dataset
This repository contains the datasets associated with the paper Agentic Reinforced Policy Optimization (ARPO).
ARPO proposes a novel agentic Reinforcement Learning algorithm designed for training multi-turn Large Language Model (LLM)-based agents. It addresses the challenge of balancing intrinsic long-horizon reasoning capabilities with proficiency in multi-turn tool interactions, particularly noting the increased uncertainty in… See the full description on the dataset page: https://huggingface.co/datasets/dongguanting/ARPO-SFT-54K.arpdf
ArPDF — does Arabic survive a PDF?
Author: Syamjith NK
Write-up: Your Arabic PDF is fine. What reads it is not.
Third in a series with ArNum-TTS
(speech) and ArShape (screen rendering).
The finding
No PDF generator and extractor pair reads Arabic reliably, and correctness depends on
the pair rather than on either tool.
5 Arabic strings × 4 generators × 3 extractors. Clean extractions out of 5:
generator
pypdf
pdfminer.six
poppler pdftotext
Chrome (HTML… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arpdf.rvl-cdip-parquetswe-zero-free-v2
SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM
SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free.
This dataset has 12,177,154 mini-swe-agent trajectories across ~121,800 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free-v2.iisc-mile-kannada-asr-corpusswe-zero-aligned-v9swe-zero-grounded-v8nepalitext-language-model-dataset
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.ar_pixmodocstables_instructanimal-imagesgenerations-gemma-3-12b-simnpo-gentle-bm25-10bswe-zero-aligned-v10Arpes
