datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LangHack
LangHack
LangHack is a dataset of diff history demonstration data for the rogue-like video game NetHack generated using the symbolic AutoAscend bot, which boasts state-of-the-art performance in the game (as of 07/22/2024).
This dataset was created by sub-sampling 10,000 full NetHack games played by AutoAscend into contiguous "chunks" of 64 timesteps, and converting the agent's game state observations in natural language text using the NetHack Language Wrapper. Sub-sampling was… See the full description on the dataset page: https://huggingface.co/datasets/upiter/LangHack.d3
D3: Diverse Data for Diff-by-Diff Coding (Python)
D3 is a large-scale dataset of instruction + file-state + diff-sequence trajectories for training LMs to synthesize and edit Python code diff-by-diff. Each trajectory pairs a natural-language goal with the initial contents of a source file and an ordered sequence of atomic file diffs that realize the goal.
Release format (SWE-Bench–style sharded parquet)
Each row has the following fields:
prompt — the initial source… See the full description on the dataset page: https://huggingface.co/datasets/upiter/d3.india-upi-ecosystem-2018-2025
India UPI Ecosystem Dataset (2018-2025)
Dataset Summary
This dataset analyzes India's UPI transaction ecosystem by combining district-level app and usage data, official NPCI benchmark statistics, and RBI macroeconomic cash indicators.It is a merged and enriched analytics dataset designed for market concentration studies, geographic adoption analysis, forecasting, and cash displacement research.
Data Sources
Source
What it contains
Why it was used… See the full description on the dataset page: https://huggingface.co/datasets/prasad-gade05/india-upi-ecosystem-2018-2025.up-it-ds-sft
Dataset Card for "up-it-ds-sft"
More Information needed
india-upi-ecosystem-2018-2025
India UPI Ecosystem Dataset (2018-2025)
Dataset Summary
This dataset analyzes India's UPI transaction ecosystem by combining district-level app and usage data, official NPCI benchmark statistics, and RBI macroeconomic cash indicators.It is a merged and enriched analytics dataset designed for market concentration studies, geographic adoption analysis, forecasting, and cash displacement research.
Data Sources
Source
What it contains
Why it was used… See the full description on the dataset page: https://huggingface.co/datasets/comrademonk/india-upi-ecosystem-2018-2025.mock-upi-txn-datachatbot-rag-upi-datasetd3_samplegricean_agents_datasetsupiterbarg-lintseq-reproductionkusonime-drive-urlupi-qa-syntheticindia-upi-transactions-mini-001
India UPI Transactions Mini Dataset
Small dataset showing UPI transaction growth in India over time.
Use cases
FinTech analysis
Time-series modeling
Economic trend analysis
License
CC-BY-4.0
upi-DigiPay-qa-augmentedUpinnews_title_classification_dataset
News Title Classification Dataset
Dataset Summary
This dataset contains article titles labeled by label. Whether the title represents the news or not. It is intended for text classification tasks such as topic modeling and news categorization.
Features
title: The headline or title of the news article (string)
label: whether the title represents the news or not.
Usage
This dataset can be used for:
Training and evaluating text classification models… See the full description on the dataset page: https://huggingface.co/datasets/upi-0/news_title_classification_dataset.gricean_agents_resultsmock-upi-txn-dataupi-sms-forwarderUpicingomock-upi-txn-data0J_upiol81wkRAIk64CLEAN_SD_HDDUpikUpiptN9XUPIAetDY
