datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodeContests_apps_format
Dataset Card for "CodeContests_apps_format"
More Information needed
app-store-apps-charts-reviews-sample
App Store Apps, Charts & Review Sentiment — Free Sample
Free samples from a mobile-app intelligence dataset of 6,384 chart apps built entirely from Apple's official public APIs (iTunes RSS charts, Search & Lookup) plus Google Play public pages: app metadata, chart-rank snapshots, review-derived sentiment metrics, and a unique "Category Opportunity Index" that ranks every US App Store category by high demand × low rating — where the most underserved app markets are.
➡️ Full… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/app-store-apps-charts-reviews-sample.apps_pnyx
PNYX - Apps
This is a splitted and tested version of APPS dataset, refer to it for further information on the original dataset construction.
This version is designed to be compatible with the hf_evaluate code_eval package and to be executed with lm-eval code_eval package.
This dataset does not include all the original fields. Some are modified and some are completely new:
id: Same as the original APPS dataset.
difficulty: Difficulty of the problem. Same as the original APPS… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/apps_pnyx.apps-control-arena-high-qualityanswers-with-reasoning-apps
answers-with-reasoning-apps
Self-distillation SFT corpus: Qwen3-8B-Instruct's own all-tests-pass
chain-of-thought rollouts on APPS interview tier (code domain).
Generation
Source problems: codeparrot/apps, difficulty == "interview" filter on both train (2000 problems) and test (3000 problems) splits = 5000 candidate problems. Problems with empty / malformed input_output are dropped (~6%), leaving 4692 candidates. LCB-v5 (our held-out code benchmark) does not overlap APPS… See the full description on the dataset page: https://huggingface.co/datasets/abhayesian/answers-with-reasoning-apps.apps-rl-ds-7b-inst-labeledapps-rl-ds-7b-base-labeledaitw-gmail-major-apps-classified
AiTW Gmail and Major App Classified Index
This dataset is a processed, app-labeled subset/index built from the public
jacklishufan/aitw Hugging Face mirror of Android in the Wild (AiTW), with
official Google Research AiTW split files applied by episode_id.
The goal is to make AiTW easier to use for GUI-agent training by adding readable
app labels, package/activity summaries, major-app proportions, and a ready-to-use
Gmail training subset.
Compatibility
The source… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-gmail-major-apps-classified.2and3_apps_30k_v4_tag5_sameprompt_processedoutputs-apps
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code
This dataset contains model generated solutions for the apps test dataset. We use this outputs for our experiments.Feel free to analyze this solutions and cite our paper if this dataset helps your research.We are pleased to announce that our work will appear in EMNLP 2025.
Here is our preprint, and we recommend you to read our paper if you are interested in evaluating the… See the full description on the dataset page: https://huggingface.co/datasets/sh0416/outputs-apps.apps_reshuffled2and3_apps_76_v6_processedapp-store-data-api-sample-data
Apple App Store API — Apps, Reviews, Ratings & ASO
Unofficial Apple App Store API in one Apify actor. 10 endpoints: app details, search, reviews, top charts, similar apps, developer profiles, autocomplete, rating histograms, privacy labels, version history. Pure HTTP, sub-3s cold start, batch & parallel. For iOS devs, ASO and AI tools.
What the actor scrapes
🍎 Apple App Store API — Scrape iOS Apps, Reviews, Ratings & ASO Data Unofficial Apple App Store API in a… See the full description on the dataset page: https://huggingface.co/datasets/logiover/app-store-data-api-sample-data.a1_code_apps_qwen3_annotatedapps_500_qwen7b_att_iter0_att10_sol5a1_code_apps_qwen3apps-taco-code-solve-rate2and3_apps_3k_v3_processed2and3_apps_40k_v7_processedapps-activations-introductory-v1apps_reshuffled_tobe_transformeda1_code_appsa1_code_apps_qwqa1_code_apps_eval_636d
mlfoundations-dev/a1_code_apps_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
19.0
58.5
71.8
28.6
36.6
35.0
29.4
7.1
10.3
AIME24
Average Accuracy: 19.00% ± 1.16%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
23.33%
7
30
3
20.00%
6
30
4
16.67%
5
30
5… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_apps_eval_636d.shopify-abandoned-cart-apps-market-intelligence
Shopify Abandoned Cart Apps Market Intelligence Dataset
This dataset packages 281 current Shopify App Store abandoned-cart apps into one analysis-ready table for ecommerce tooling research, merchant screening, pricing analysis, and competitive intelligence.
It starts from Shopify's live Abandoned cart category and enriches each app with pricing-plan structure, review volume, Built for Shopify flags, languages, integrations, category feature labels, developer metadata, and… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/shopify-abandoned-cart-apps-market-intelligence.apps-activations-introductory-layers-2-6-10decontaminate_stratos_appsapps_qwen7b_att_iter0_ppo_att2_sol2apps_1000_qwen7b_att_iter0_ppo_att2_sol2_debuga1_code_apps_phi_annotated
