datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rhan-checkpointsadaption-recipe-ingredient-validation
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-recipe_ingredient_validation
This dataset consists of prompt-completion pairs designed to test ingredient relevance for specific recipes. Each entry presents a recipe title and a candidate ingredient, requiring a binary classification of whether the ingredient belongs in the dish. The completions provide the correct label as either 'BELONGS' or 'NOT_BELONGS' based on culinary logic.… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-recipe-ingredient-validation.adaption-fertility-rate-qa-pairs-v1This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-fertility_rate_qa_pairs
This dataset contains question-answer pairs derived from fertility rate statistics (births per woman) for various countries between 1990 and 2023. The samples include descriptions of line, bar, and area charts, some featuring misleading design elements like inverted or truncated axes. The questions test the ability to interpret trends, extract specific values… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-fertility-rate-qa-pairs-v1.FermiBench
FermiBench: Nuclear Power Information Retrieval Benchmark
Dataset Description
Dataset Summary
This dataset is designed for benchmarking information retrieval systems within the nuclear power domain, focusing on long-context retrieval of full-text documents. The corpus includes approximately 4,500 documents sourced from the U.S. Nuclear Regulatory Commission’s (NRC) Agency-wide Documents Access and Management System (ADAMS). The… See the full description on the dataset page: https://huggingface.co/datasets/atomic-canyon/FermiBench.desing-uifermi-lat-synthetic-daily-sky-maps
SPECTRA Fermi-LAT Simulated Sky Maps
This repository contains the simulated dataset associated with the manuscript:
“Self-Supervised ConvLSTM for Fermi Large Area Telescope Transient Detection”
The associated manuscript is currently under review at Astronomy and Computing.
Overview
This dataset was generated to support the development and validation of self-supervised, spatio-temporal anomaly-detection methods for gamma-ray transient searches in Fermi-LAT-like… See the full description on the dataset page: https://huggingface.co/datasets/Idea-re/fermi-lat-synthetic-daily-sky-maps.adaption-math-misconception-match
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-math_misconception_match
This dataset contains pairs of math problems, student wrong answers, and candidate misconceptions. The task is to determine if the specific misconception logically explains the student's error. Each sample provides the question context, the incorrect response, and a labeled classification of 'MATCH' or 'NO_MATCH'.
Dataset size
There are 19,996… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-math-misconception-match.adaption-consumer-finance-complaints
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-consumer_finance_complaints
This dataset contains consumer complaints regarding financial products and services, including banking, credit cards, debt collection, and money transfers. Each entry consists of a user-submitted narrative describing a specific issue, paired with a structured classification of the product, issue type, and sub-issue. The content highlights common grievances… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-consumer-finance-complaints.adaption-brazil-crypto-regulatory-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazil_crypto_regulatory_qa
This dataset contains prompt-completion pairs focused on Brazilian financial regulations regarding cryptocurrencies, tokens, and digital assets. The content specifically addresses the role of the CVM (Comissão de Valores Mobiliários), risk assessments for investors, and compliance with laws such as Lei 14.478/2022. Each entry provides structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazil-crypto-regulatory-qa.adaption-map2k1-rasopathy-tier1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-map2k1_rasopathy_tier1
This dataset contains structured research prioritization assessments for high-priority (Tier 1) missense variants in the MAP2K1 gene associated with RASopathies like Cardio-Facio-Cutaneous Syndrome. Each entry provides computational evidence scores (CADD, AlphaMissense), population frequency data, and structural context within the MEK1 kinase domain to guide… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-map2k1-rasopathy-tier1.adaption-ongbr-impact-copy
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ongbr_impact_copy
This dataset contains prompt-completion pairs for generating ethical copywriting content for Brazilian non-profit organizations (OSCs). Each sample includes a detailed style guide emphasizing dignity, specificity, and transparency, followed by a specific organizational briefing and a corresponding draft text. The completions demonstrate how to craft donation appeals… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-ongbr-impact-copy.PKR-QA
PKR-QA Dataset
This folder contains PKR-QA question answering examples prepared as JSON Lines files for Hugging Face Datasets.
Files
Split
File
Source
Examples
train
train.jsonl
dataset/s4_QADataset_12Feb2025/train/training_small_100.json
1700
validation
validation.jsonl
dataset/s4_QADataset_12Feb2025/val/validation_small_50.json
850
test
test.jsonl
dataset/s4_QADataset_12Feb2025/testing.json
46921
Each line is one original QA example… See the full description on the dataset page: https://huggingface.co/datasets/fernandopbc/PKR-QA.adaption-job-scam-detection
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-job_scam_detection
This dataset contains pairs of job advertisement prompts and binary classifications indicating whether each posting is legitimate or fraudulent. The samples include detailed company profiles, job descriptions, and requirements used to train models for identifying employment scams. Each entry is formatted as a prompt asking for a single-word decision followed by the… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-job-scam-detection.state-fertility-insurance-mandates
State insurance mandates for fertility treatment and IVF coverage
Canonical, always-current version: https://referencesource.org/state-fertility-insurance-mandates/
Machine-readable: https://referencesource.org/state-fertility-insurance-mandates/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-17
Stale after: 2027-02-13 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 26
Which US states require… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-fertility-insurance-mandates.emotions_worldwide
Emotions worldwide
This is an open dataset listing emotions from across the world with their descriptions in English. First used for the artwork E*star for the NeurIPS 2024 Creative AI track, the dataset behind the artwork has been made open-source through github.
The dataset actively seeks validation, correction, and addition by the open public (especially from non-English language speakers, as those emotions are difficult to validate by myself).
Support by validating, correcting… See the full description on the dataset page: https://huggingface.co/datasets/Ferdinandnathaniel/emotions_worldwide.fertility_env_factors_effect_factsheets_eshre
Dataset Card for fertility_env_factors_effect_factsheets_eshre
Dataset Summary
fertility_env_factors_effect_factsheets_eshre is a structured question–answer dataset derived from publicly available ESHRE (European Society of Human Reproduction and Embryology) patient factsheets describing how environmental and lifestyle factors influence fertility.
The dataset converts evidence-based educational material into clear patient-style questions and clinically aligned answers… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/fertility_env_factors_effect_factsheets_eshre.adaption-brazilian-crypto-compliance
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazilian_crypto_compliance
This dataset contains structured compliance assessments for Brazilian crypto asset regulations, featuring prompts with specific scenarios and JSON completions detailing applicable laws, risk levels, and corrective actions. Each entry evaluates adherence to rules from authorities like the BCB, CVM, and COAF regarding issues such as asset segregation, KYC… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazilian-crypto-compliance.adaption-brazil-agri-qa-match
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazil_agri_qa_match
This dataset contains pairs of agricultural questions posed to the Brazilian Ministry of Agriculture and candidate answers sourced from government portals. Each sample includes a binary label indicating whether the provided answer correctly addresses the specific question asked. The content covers diverse topics such as family farming, traceability, ministerial… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazil-agri-qa-match.fertility-medical-reasoning
Tanit Fertility Medical Reasoning Dataset
Dataset Description
High-quality medical reasoning dataset specialized for fertility care.
Total: 45,183 samples (42,923 train / 2,260 validation)
Dataset Composition
Source
Samples
Percentage
Purpose
MedReason
27,711
61.3%
Expert-verified clinical reasoning
Medical-O1
14,272
31.6%
O1-style systematic thinking
Synthetic-Fertility
3,200
7.1%
Domain-specific fertility cases
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/MohamedISSAOUI/fertility-medical-reasoning.brazilian-gov-formal-letters
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
brazilian_gov_formal_letters
This dataset contains pairs of prompts and completions for generating formal administrative communications in Portuguese directed at Brazilian federal institutions. The content includes complaints, suggestions, and official responses addressing issues such as missed deadlines, lack of transparency, and unresolved demands. Each sample demonstrates a formal, legalistic… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/brazilian-gov-formal-letters.adaption-shoc2-ns-lah-variant-reports
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-shoc2_ns_lah_variant_reports
This dataset contains structured research-level assessments of SHOC2 missense variants associated with Noonan Syndrome with Loose Anagen Hair (NS-LAH). Each entry provides a detailed interpretation including computational scores (CADD, AlphaMissense), population frequency, LRR domain context, and SMP complex proximity. The reports preserve source-derived Tier… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-shoc2-ns-lah-variant-reports.adaption-fertility-rate-qa-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-fertility_rate_qa_pairs
This dataset contains question-answer pairs derived from fertility rate statistics (births per woman) for various countries between 1990 and 2023. The samples include descriptions of line, bar, and area charts, some featuring misleading design elements like inverted or truncated axes. The questions test the ability to interpret trends, extract specific values… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-fertility-rate-qa-pairs.adaption-pt-afro-brasileiro-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-pt_afro_brasileiro_qa
This dataset contains question-answer pairs in Portuguese focusing on the linguistic features of Afro-Brazilian Portuguese, such as reduced nominal agreement and preposition variation. The content provides historical context, sociolinguistic analysis, and references to academic research to explain these phenomena as regular grammatical systems rather than… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-pt-afro-brasileiro-qa.adaption-parteira-br-maternal-guidance
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-parteira_br_maternal_guidance
This dataset contains conversational pairs between an AI assistant and users in remote Brazilian communities regarding maternal health concerns like nausea, breastfeeding pain, and fatigue. The assistant provides empathetic, low-risk guidance while strictly adhering to a protocol that prioritizes immediate referral to formal healthcare services for any… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-parteira-br-maternal-guidance.adaption-brazilian-regulatory-filings
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazilian_regulatory_filings
This dataset contains samples of Brazilian regulatory filings (Fatos Relevantes and Market Notices) from publicly traded companies, presented in both Portuguese and English. Each sample includes a classification task where the text is analyzed to determine the event type, status, scope relative to a target entity, and market signal sentiment. The content… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazilian-regulatory-filings.womens-health-fertility-safety-qa
Dataset Card for WomenHealth-Ferility-safety-guideline
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset contains clinically grounded question–answer pairs focused on fertility treatments and IVF (In Vitro Fertilization), with a strong emphasis on patient safety, procedure awareness, and emotional reassurance.
The dataset is designed to address… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/womens-health-fertility-safety-qa.adaption-brazilian-civil-law-rulings
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
brazilian_civil_law_rulings
This dataset contains samples of Brazilian civil law court decisions, specifically focusing on appeals, special recourse agravo, and consumer protection cases handled by the Superior Court of Justice (STJ). The text includes detailed legal reasoning, case summaries, discussion of res judicata, contractual rescission, moral damages, and citations of relevant articles… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazilian-civil-law-rulings.adaption-sos1-variant-prioritization
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sos1_variant_prioritization
This dataset contains structured research prioritization assessments for SOS1 missense variants associated with Noonan Syndrome. Each entry details computational evidence metrics, including CADD PHRED scores, AlphaMissense predictions, and gnomAD frequencies, alongside functional domain localization. The completions strictly preserve source-derived… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-sos1-variant-prioritization.adaption-sos2-ns13-variant-assessments
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sos2_ns13_variant_assessments
This dataset contains structured research-level assessments for SOS2 gene variants associated with Noonan Syndrome 13 (NS13). Each entry provides a detailed analysis including assigned priority tiers, investigation scores, and evidence from computational predictors like CADD and AlphaMissense. The content further details population rarity, specific GEF… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-sos2-ns13-variant-assessments.adaption-mapk2-cfc4-variant-reports-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-mapk2_cfc4_variant_reports
This dataset contains structured research prioritization reports for MAP2K2 (MEK2) missense variants associated with Cardio-Facio-Cutaneous Syndrome 4 (CFC4). Each entry details computational evidence including CADD PHRED scores, AlphaMissense predictions, gnomAD frequencies, and protein domain localization within the kinase region. The reports preserve… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-mapk2-cfc4-variant-reports-v1.
