genbench-iitp/coding-variant
GenBench CoCG QA Dataset Multi-hop genetic reasoning QA items generated from GenBench's knowledge graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO, SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more), built for CoCG (Co-Evolving Confidence Graph) agent training. 2513 items across 11 task types. Task types task_type count coding_variant 53 conservation_reasoning 246 counterfactual 246 disease_reasoning 246… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/coding-variant.
GenBench CoCG QA Dataset
Multi-hop genetic reasoning QA items generated from GenBench's knowledge graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO, SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more), built for CoCG (Co-Evolving Confidence Graph) agent training.
2513 items across 11 task types.
Task types
Schema
Each item has:
id,task_type,pipeline(coding_variant/noncoding_regulatory),difficultyquestion,answer,choices(MCQ options, when applicable)context-- either a templated chain narration, or (ifllm_rewritewas applied) an LLM-rewritten fluent Step/Evidence/Interpretation/Conclusion narrativereasoning_chain-- the grounded, machine-checkable multi-hop path (steps: each withsource_node_id/target_node_id/edge_relation/edge_confidence/edge_source_db), never touched by any LLM stepmodality_data-- raw modality payloads (sequence, structural, transcriptomic, posttranslational, signalingrole, etc.) attached to the chain's anchor nodesevidence-- supporting evidence entries with source database/PMIDpath_confidence_score-- continuous, confidence-derived difficulty score
Companion graph
graph.json (if included in this repo) is the exact knowledge graph these items' reasoning_chain node IDs refer to -- load it with GenBench's `GraphBuilder.load()` to resolve full node/edge attributes beyond what's inlined in each item.
Source
Generated from data\curated\qa_dataset.jsonl in GenBench, the substrate for CoCG (Co-Evolving Confidence Graph) agent training -- per-edge, per-modality KG confidence that co-adapts with an RL policy during training rather than treating the KG as a frozen oracle.
