datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
APPS-verified
Introduction
This dataset contains verified solutions from the APPS dataset's training set. Solutions that fail to pass all the test cases are removed. Problems with no correct solution are also removed.
The solutions were executed on Intel E5-2620 v3 CPUs with the execution timeout set to 10 seconds.
Statistics in the training set
Dataset
# Problems
# Solutions
TACO
5000
117232
TACO-verified
4211
93921
Correct Ratio
84.22%
80.12%
mobile-apps-user-sentiment-reviews
Top Mobile Apps User Sentiment & Review Corpus (Google Play)
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/google-play-reviews-scraper
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/instagram-1star-reviews… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/mobile-apps-user-sentiment-reviews.app-store-reviews-scraper
App Store Reviews Scraper
Scrape Apple App Store reviews, star ratings and app version history for any iOS app in any country storefront. No login, no API key.
Rows in this dataset
23,048
Fields
43
Collector runs behind it
88
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/app-store-reviews-scraper/ — 176 entity pages
Run the collector yourself
https://apify.com/reapx/app-store-reviews-scraper
What this is… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/app-store-reviews-scraper.appsec-router-pairs-r5
appsec-router-pairs-r5
Training data for pratikamin/appsec-router-deberta-r5:
pairs of an application-security interview answer and a hypothesis about the speaker, labelled
1 when the answer expresses the point and 0 when it does not.
Entirely synthetic. Answers were generated by openai/gpt-oss-120b (Apache 2.0) to 86 authored
interview questions and their 378 follow-ups from appsecinterview.com, in several registers; labels
come from the same model judging each answer against… See the full description on the dataset page: https://huggingface.co/datasets/pratikamin/appsec-router-pairs-r5.arena-control-apps-v5-dataset
Arena Control Apps V5 Dataset
A curated dataset for training and evaluating models on collusion signal detection in code.
Dataset Overview
Metric
Value
Total samples
1,793
Train split
1,493
Test split
300
Label 0 (clean)
765
Label 1 (backdoor)
1,028
Bucket Distribution
Bucket
Count
Label
Description
clean
465
0
Clean code, no backdoor, no signal
clean_signal_a
100
0
Clean code + Signal A
clean_signal_b
100
0
Clean code +… See the full description on the dataset page: https://huggingface.co/datasets/jprivera44/arena-control-apps-v5-dataset.seal-7b-apps-vectorapps_select
