datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
program_generation_v3program_generation_v5Ads_Creative_Ad_Copy_Programmatic
Dataset Summary
The Programmatic Ad Creatives dataset contains 7097 samples of online programmatic ad creatives along with their ad sizes. The dataset includes 8 unique ad sizes, such as (300, 250), (728, 90), (970, 250), (300, 600), (160, 600), (970, 90), (336, 280), and (320, 50). The dataset is in a tabular format and represents a random sample from Project300x250.com's complete creative data set. It is primarily used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/PeterBrendan/Ads_Creative_Ad_Copy_Programmatic.Nemotron-Competitive-Programming-v1-prompt-only
Nemotron-Competitive-Programming-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-Competitive-Programming-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Competitive-Programming-v1-prompt-only.Nemotron-SFT-Competitive-Programming-v2-prompt-only
Nemotron-SFT-Competitive-Programming-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Competitive-Programming-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Competitive-Programming-v2-prompt-only.ProgrammingDataset
🧠 ProgrammingDataset
A high-quality, production-grade dataset of programming code snippets across multiple languages, collected and curated manually to support research in code generation, analysis, and educational tools.
📌 Dataset Summary
Field
Description
Rows
100+ code samples
Languages
Python, JavaScript, C++, Java, etc.
Tasks
Data structures, algorithms, system utilities
Format
Excel (.xlsx) and CSV
License
MIT
Each entry includes:
id:… See the full description on the dataset page: https://huggingface.co/datasets/kaiiddo/ProgrammingDataset.Ads_Creative_Text_Programmatic
Dataset Summary
The Programmatic Ad Creatives dataset contains 1000 samples of online programmatic ad creatives along with their ad sizes. The dataset includes 8 unique ad sizes, such as (300, 250), (728, 90), (970, 250), (300, 600), (160, 600), (970, 90), (336, 280), and (320, 50). The dataset is in a tabular format and represents a random sample from Project300x250.com's complete creative data set. It is primarily used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/PeterBrendan/Ads_Creative_Text_Programmatic.genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/genz-slang-pairs-1k.program_peacebug-bounty-programs-rewards
Overview
This dataset is part of a research project Economic Taxonomy of Software Vulnerabilities, which aims to estimate the monetary cost associated with software vulnerabilities based on real-world bug bounty program data. The core objective is to provide a concrete, data-driven reference for evaluating the average cost of discovering and reporting vulnerabilities across different severity levels.
Methodology Summary
The dataset was generated following these steps:… See the full description on the dataset page: https://huggingface.co/datasets/lesis-lat/bug-bounty-programs-rewards.AIYA-Programming
AIYA Programs Toolkit & Benchmark Data 🛠️
This repository provides technical resources, exercises, evaluation benchmarks, datasets, and starter materials supporting the experiential programming of the AI Youth Alliance (AIYA).
AIYA is a global, student-led educational network dedicated to building foundational AI literacy, AI fluency, technical skills, critical AI literacy, student agency, and genuine authorship through hands-on learning and collaborative problem-solving.
The… See the full description on the dataset page: https://huggingface.co/datasets/AIYA-on-Huggingface/AIYA-Programming.electronic-program-fbd06d
electronic-program-fbd06d
Synthetic weather test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/RapidBeacon/electronic-program-fbd06d.sinhala-english-singlish-translation
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.pharma-program-go-no-go-coherence-risk-v0.1What this repo is for
support go or stop decisions in drug programs
detect when teams continue weak assets
detect when strong assets are wrongly killed
align biology, safety, and signal with decision
reduce sunk-cost bias
support portfolio governance boards
support licensing and diligence reviews
How it is used
You provide one row describing a program.
The system returns one label.
go
or
no_go
How to read the output
go means
efficacy signal present
safety acceptable
biomarker supports
target… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/pharma-program-go-no-go-coherence-risk-v0.1.Quantum-Programming🧑💻 Overview
This dataset focuses on Quantum Programming and contains curated information that can be used for research, education, and model training. Quantum programming is an emerging field that leverages the principles of quantum mechanics to develop new algorithms and computational techniques. This dataset aims to provide structured information that can help both beginners and advanced users explore concepts, applications, and trends in quantum computing.
📂 Dataset Contents
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/as-krn/Quantum-Programming.africa-synth-cancer-cancer-screening-programs-africa-all
Cancer Screening Programs Africa | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-cancer-screening-programs-africa-all.africa-synth-poverty-safety-net-programs-africa-all
Africa Synth Poverty Safety Net Programs Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-poverty-safety-net-programs-africa-all.program_generation_v6tech_program
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/aaaaaaaqdqd/tech_program.program_gen_v1_m1Programming_langaugesmanaged-care-enrollment-by-program-and-population
Managed Care Enrollment by Program and Population (Duals)
Description
The Medicaid Managed Care Enrollment Report profiles enrollment statistics on Medicaid managed care programs on a plan-specific level. The managed care enrollment statistics include enrollees receiving comprehensive benefits and limited benefits and are point-in-time counts.
Because Medicaid beneficiaries may be enrolled concurrently in more than one type of managed care program (e.g., a Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/managed-care-enrollment-by-program-and-population.commodore64_program_latestclinical-temporal-5node-pressure-buf-lag-cpl-program-lockin-v0.1
What this repo does
This dataset tests whether a model can detect a drug development program entering composite instability across recruitment, safety, manufacturing, and competitive pressure over time, and predict whether the program crosses into lock-in by the final step.
Core quad
pressurebufferlagcoupling
Prediction target
label_cascade_state
Row structure
One row represents a short temporal window (t0–t3) across program quarters. It summarizes… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-temporal-5node-pressure-buf-lag-cpl-program-lockin-v0.1.ProgrammingEnthusiastsDB
ProgrammingEnthusiastsDB
tags: education, online learning, engagement tracking
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'ProgrammingEnthusiastsDB' dataset contains anonymized user data from an online programming and technology school. The dataset is designed to help educational researchers and ML practitioners analyze user engagement, course effectiveness, and overall satisfaction with the online learning experience.… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ProgrammingEnthusiastsDB.public-policy-program-targeting-beneficiary-coherence-risk-v0.1What this repo is for
Detect when public programs
reach the wrong people
or fail to reach the right ones.
Flags:
low uptake among intended group
high leakage outside target group
benefits delivered without outcome change
regional targeting distortion
cqadupstack-programmers-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language.
Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf
Contact: konrad.wojtasik@pwr.edu.pl
program_gen_data_model1_superprogram_gen_data_model1_super_v2QA-Python-Programming-Indonesia
QA-Python-Programming-Indonesia
Deskripsi
Dataset ini berisi kumpulan pertanyaan dan jawaban (QA) terkait pemrograman Python, dirancang untuk membantu pengguna memahami konsep, teknik, dan praktik terbaik dalam bahasa pemrograman ini.
Isi Dataset
Pertanyaan: Beragam pertanyaan yang mencakup berbagai topik dalam pemrograman Python, mulai dari dasar hingga lanjutan.
Jawaban: Penjelasan mendetail dan kode contoh yang relevan, memberikan klarifikasi dan konteks… See the full description on the dataset page: https://huggingface.co/datasets/gabrielb/QA-Python-Programming-Indonesia.
