datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
d3
D3: Diverse Data for Diff-by-Diff Coding (Python)
D3 is a large-scale dataset of instruction + file-state + diff-sequence trajectories for training LMs to synthesize and edit Python code diff-by-diff. Each trajectory pairs a natural-language goal with the initial contents of a source file and an ordered sequence of atomic file diffs that realize the goal.
Release format (SWE-Bench–style sharded parquet)
Each row has the following fields:
prompt — the initial source… See the full description on the dataset page: https://huggingface.co/datasets/upiter/d3.india-upi-ecosystem-2018-2025
India UPI Ecosystem Dataset (2018-2025)
Dataset Summary
This dataset analyzes India's UPI transaction ecosystem by combining district-level app and usage data, official NPCI benchmark statistics, and RBI macroeconomic cash indicators.It is a merged and enriched analytics dataset designed for market concentration studies, geographic adoption analysis, forecasting, and cash displacement research.
Data Sources
Source
What it contains
Why it was used… See the full description on the dataset page: https://huggingface.co/datasets/prasad-gade05/india-upi-ecosystem-2018-2025.d3_sampleupiterbarg-lintseq-reproduction
