CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01arpandeepk /swe-cleanroom-trajs swe-cleanroom-trajs Procedurally generated SWE-agent trajectories — no LLM, no GPU — as a clean-room counterpart to AlienKevin/SWE-ZERO-12M-trajectories (which rolls out a 1.7B model). ~40k trajectories across 124 repos (5 languages), mini-swe-agent-1 format. How it's generated (no LLM): TOOL rhythm: a per-step 2nd-order Markov chain fit on ricdomolm/mini-coder-trajs-400k (verb sequence). ARGUMENTS: a tree-sitter code graph + the gold PR patch decide which file/symbol/lines… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-cleanroom-trajs.texttext-generationn<1K0 likes1.1k downloads4mo agoHugging Face02arpesenti /tts-ranking-data0 likes836 downloads10mo agoHugging Face03arpandeepk /swe-zero-free SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free. This dataset has 12,238,610 mini-swe-agent trajectories across 119,084 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free.texttext-generation10M<n<100M0 likes566 downloads4mo agoHugging Face04ahmedheakl /ar_pixmocapqatrans_instructimage100K<n<1M0 likes561 downloads2y agoHugging Face05arpanbosmia /paper-trail-datatabular10M<n<100M0 likes538 downloads19h agoHugging Face06davidggphy /librispeech-arpabet-processed LibriSpeech ARPAbet Processed Dataset Pre-processed dataset for training ARPAbet phoneme recognition models using CTC loss. Dataset Description This dataset is derived from LibriSpeech (train-clean-100 split) with the following preprocessing: Audio: Resampled to 16kHz, normalized using Wav2Vec2 feature extractor Labels: Text transcriptions converted to ARPAbet phoneme sequences using CMU Pronouncing Dictionary Filtering: Samples with out-of-vocabulary words (not in CMU… See the full description on the dataset page: https://huggingface.co/datasets/davidggphy/librispeech-arpabet-processed.audioautomatic-speech-recognition10K<n<100K0 likes527 downloads8mo agoHugging Face07ahmedheakl /ar_pixmocapqatrans2_instructimage100K<n<1M0 likes474 downloads2y agoHugging Face08jotone /arpa-osmer-weatherstations OSMER ARPA FVG — Validated Weather Observations Validated meteorological observations from the OSMER network of ARPA FVG (Regional Environmental Protection Agency of Friuli-Venezia Giulia, NE Italy): temperature, humidity, precipitation, wind (mean/vector/gust + direction), global radiation, pressure, leaf wetness and soil temperatures. Two resolutions: hourly point observations and daily aggregates (min/max/mean), packaged as Parquet for analysis and ML. This is a self-archived… See the full description on the dataset page: https://huggingface.co/datasets/jotone/arpa-osmer-weatherstations.tabulartime-series-forecasting10M<n<100M0 likes437 downloads3mo agoHugging Face09Arpitraj01 /Pothole_classificationA dataset for efficient pothole classification which contains more than 400 images collected over 50 km road from kerala. imagen<1K1 likes436 downloads27d agoHugging Face10ARPRIM /Pulaar_Dictionary Saggitorde — Dictionnaire Pulaar / Français / Anglais Dictionnaire multilingue Pulaar ↔ Français ↔ Anglais extrait et enrichi à partir du Saggitorde de Ceerno Abuu Sih, enrichi par les terminologies de l'ARPRIM, ILN, MFS et d'autres sources spécialisées. Statistiques Statistique Valeur Nombre total d'entrées 1862 Nombre de domaines 24 Nombre de sources 24 Langues Pulaar (ff), Français (fr), Anglais (en) Structure des données… See the full description on the dataset page: https://huggingface.co/datasets/ARPRIM/Pulaar_Dictionary.text1K<n<10K1 likes388 downloads13d agoHugging Face11Karavet /ARPA-Armenian-Paraphrase-Corpus Dataset Description We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language. Dataset Summary The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.text1K<n<10K3 likes292 downloads4y agoHugging Face12P-Arpan /Spotify_Audio_features_2.3Mtabular1M<n<10M0 likes236 downloads1mo agoHugging Face13mojimoji61 /arp-decision131-robocasa-libero-20 Decision-131 complete RoboCasa + LIBERO-PRO subset This repository contains the complete data package requested for the 20 task identities selected in ARP decision 131. RoboCasa human demonstrations 10 target tasks 5,101 episodes 2,680,229 frames at 20 Hz Actions, robot state, episode/task metadata, and three synchronized 256×256 H.264 camera streams 176 MP4 files and 8 Parquet files The episodes were selected from lerobot/robocasa_target_human_unified at… See the full description on the dataset page: https://huggingface.co/datasets/mojimoji61/arp-decision131-robocasa-libero-20.tabular1M<n<10M0 likes227 downloads1mo agoHugging Face14arpandeepk /swe-zero-grounded-fulltextn<1K0 likes206 downloads3mo agoHugging Face15ahmedheakl /ar_pixmodocsother_instructimage10K<n<100K0 likes205 downloads2y agoHugging Face16arpit-tiwari /syspin-kannada-ttsaudio10K<n<100K0 likes177 downloads11mo agoHugging Face17arpandeepk /generations-nemotron-nano-9b-v2-simnpo-gentle-igm-10btabular10K<n<100K0 likes169 downloads5mo agoHugging Face18dongguanting /ARPO-SFT-54K Agentic Reinforced Policy Optimization (ARPO) Dataset This repository contains the datasets associated with the paper Agentic Reinforced Policy Optimization (ARPO). ARPO proposes a novel agentic Reinforcement Learning algorithm designed for training multi-turn Large Language Model (LLM)-based agents. It addresses the challenge of balancing intrinsic long-horizon reasoning capabilities with proficiency in multi-turn tool interactions, particularly noting the increased uncertainty in… See the full description on the dataset page: https://huggingface.co/datasets/dongguanting/ARPO-SFT-54K.texttext-generation10K<n<100K15 likes161 downloads11mo agoHugging Face19syamjithnk /arpdf ArPDF — does Arabic survive a PDF? Author: Syamjith NK Write-up: Your Arabic PDF is fine. What reads it is not. Third in a series with ArNum-TTS (speech) and ArShape (screen rendering). The finding No PDF generator and extractor pair reads Arabic reliably, and correctness depends on the pair rather than on either tool. 5 Arabic strings × 4 generators × 3 extractors. Clean extractions out of 5: generator pypdf pdfminer.six poppler pdftotext Chrome (HTML… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arpdf.texttext-classificationn<1K1 likes149 downloads8d agoHugging Face20arpit-gour02 /rvl-cdip-parquettext10K<n<100K0 likes143 downloads7mo agoHugging Face21arpandeepk /swe-zero-free-v2 SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free. This dataset has 12,177,154 mini-swe-agent trajectories across ~121,800 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free-v2.texttext-generation10M<n<100M0 likes143 downloads4mo agoHugging Face22arpit-tiwari /iisc-mile-kannada-asr-corpusaudio100K<n<1M0 likes122 downloads1y agoHugging Face23arpandeepk /swe-zero-aligned-v9text10K<n<100K0 likes116 downloads3mo agoHugging Face24arpandeepk /swe-zero-grounded-v8textn<1K0 likes115 downloads3mo agoHugging Face25Arpuuu /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.texttext-generation10M<n<100M0 likes113 downloads6mo agoHugging Face26ahmedheakl /ar_pixmodocstables_instructimage10K<n<100K0 likes97 downloads2y agoHugging Face27arpitwagare /animal-imagesimagen<1K0 likes95 downloads3mo agoHugging Face28arpandeepk /generations-gemma-3-12b-simnpo-gentle-bm25-10btabular10K<n<100K0 likes94 downloads5mo agoHugging Face29arpandeepk /swe-zero-aligned-v10text1K<n<10K0 likes94 downloads3mo agoHugging Face30introvoyz041 /Arpes0 likes91 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.