datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vqa-cmsv-benchmark
VQA-CMSV Benchmark Data Package
This repository contains annotation splits for VQA v2-CMSV, GQA-CMSV, and VG-CMSV, plus patch-mask NPZ files used for mask supervision experiments.
Contents
data/vqa_v2_cmsv/train.json, data/vqa_v2_cmsv/val.json, data/vqa_v2_cmsv/test.json
data/gqa_cmsv/train.jsonl, data/gqa_cmsv/val.jsonl, data/gqa_cmsv/test.jsonl
data/vg_cmsv/train.jsonl, data/vg_cmsv/val.jsonl, data/vg_cmsv/test.jsonl
masks/vqa_v2_cmsv_masks.npz
masks/gqa_cmsv_masks.npz… See the full description on the dataset page: https://huggingface.co/datasets/as-benchmark-artifacts/vqa-cmsv-benchmark.cms-desynpuf-insurance-claims
CMS DE-SynPUF Insurance Claims (Rendered EOB Documents)
Pipeline-ready insurance_claim evaluation corpus derived from the CMS
2008-2010 Data Entrepreneurs' Synthetic Public Use File (DE-SynPUF), Sample 1.
One row = one Medicare claim rendered as a plain-text EOB-style document with
ground-truth extraction fields aligned to the llm-mailroom
InsuranceClaimExtraction schema.
Provenance
Source: CMS DE-SynPUF Sample 1 (fully synthetic; "very limited inferential… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/cms-desynpuf-insurance-claims.regex-pattern-generation-with-verified-match-sets-cmskdvm7
Regex Pattern Generation with Verified Match Sets
A dataset of regex pattern generation with verified match sets examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: Regex
Framework: Community
License: CC-BY-4.0… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/regex-pattern-generation-with-verified-match-sets-cmskdvm7.python-execution-trace-output-prediction-cmskdimp
Python Execution Trace & Output Prediction
Self-contained Python programs paired with a concrete function call and the exact runtime output. Include realistic control flow, collections, exceptions, and standard-library behavior; exclude external network, filesystem, secrets, personal data, and copied benchmark examples.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-execution-trace-output-prediction-cmskdimp.build-competitive-programming-problem-dataset-cmsokff2
Build Competitive Programming Problem Dataset
Each item is a build competitive programming problem dataset example providing Problem statement, Constraints, Reference solution, Test cases, Time limit (ms). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: Python… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/build-competitive-programming-problem-dataset-cmsokff2.graphql-schema-resolver-execution-validation-cmskdigf
GraphQL Schema Resolver Execution Validation
Each item is a graphql schema resolver execution validation example providing Schema (SDL), Resolver map, Sample query, Expected response (JSON). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: JavaScript
Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/graphql-schema-resolver-execution-validation-cmskdigf.python-data-transformation-scripts-cmskdijz
Python Data Transformation Scripts
Each item is a python data transformation scripts example providing Input sample, What the transform should do, Transformation script, Expected output. Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Contributor items exported: 1000
Language: Python
Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-data-transformation-scripts-cmskdijz.extract-metrics-from-log-fixtures-cmskdlip
Extract Metrics from Log Fixtures
A dataset of extract metrics from log fixtures examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Contributor items exported: 1000
Language: Python
Framework: Community
License: CC-BY-4.0
Contributors… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/extract-metrics-from-log-fixtures-cmskdlip.cms_iom_500
Dataset Summary
This dataset contains CMS information with local and national coverage document data sets (LCD & NCD),as Coverage Articles and [Internet-Only Manuals (IOMs)(https://www.cms.gov/medicare/regulations-guidance/manuals/internet-only-manuals-ioms)
A list of Current LCDS, NCDs and Articles is obrained from Medicare Coverage Database.
The data itself was obtainted by scrapping the urls and extracting data from the pdf files listed in current articles and current lcds… See the full description on the dataset page: https://huggingface.co/datasets/evekhm/cms_iom_500.cms_iom_3000
Dataset Summary
This dataset contains CMS information with local and national coverage document data sets (LCD & NCD),as Coverage Articles and [Internet-Only Manuals (IOMs)(https://www.cms.gov/medicare/regulations-guidance/manuals/internet-only-manuals-ioms)
A list of Current LCDS, NCDs and Articles is obrained from Medicare Coverage Database.
The data itself was obtainted by scrapping the urls and extracting data from the pdf files listed in current articles and current lcds… See the full description on the dataset page: https://huggingface.co/datasets/evekhm/cms_iom_3000.JiRack-Pretrain-DatasetThis dataset was created with a strong focus on corporate AI. Its main goal is to help banks, fintech companies, and enterprises protect sensitive data and stay compliant with strict U.S. privacy laws that prohibit the use of public online and cloud-based AI services.
Multi-Domain High-Quality Corpus
A carefully curated high-quality multilingual dataset used to pretrain the JiRack model with modern data.
So JiRack Tokenizers stand out because they’ve been trained on premium… See the full description on the dataset page: https://huggingface.co/datasets/CMSManhattan/JiRack-Pretrain-Dataset.CMS_JSONbuga
