CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes423 downloads3y agoHugging Face02qualora-data-labs /qualora-workforce-skills-graph Qualora Workforce Skills Graph (Representative Sample) Rights-clean, provenance-tracked vocational learning data, rebuilt from roughly $2B of U.S. Department of Labor funded open courseware into a labeled skills graph: cleaned courses and lessons, Bloom-tagged assessment items with answer rationales and learning objectives, and a content-grounded course to skill to career graph with salary context. Built for post-training and evaluation, not pretraining bulk. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/qualora-data-labs/qualora-workforce-skills-graph.tabularquestion-answeringn<1K0 likes171 downloads2mo agoHugging Face03GraphWiz /GraphInstruct-Testtexttext-generation1K<n<10K2 likes160 downloads3y agoHugging Face04whoisjiji /fin-cfa-graphgen Fin-CFA-GraphGen: 785K Knowledge-Guided Financial QA Examples Fin-CFA-GraphGen is a large-scale English dataset for financial instruction tuning, financial question answering, and domain-specific language-model post-training. It contains 785,149 synthetic question–answer examples generated from CFA curriculum and exam-preparation books with the GraphGen knowledge-driven data-generation method. The dataset and its role in the post-training pipeline are described in Data-Centric… See the full description on the dataset page: https://huggingface.co/datasets/whoisjiji/fin-cfa-graphgen.tabulartext-generation100K<n<1M0 likes150 downloads13d agoHugging Face05GraphWiz /GraphInstruct-RFT-72Ktextquestion-answering10K<n<100K10 likes69 downloads3y agoHugging Face06GraphPRM /GraphSilo-Testtexttext-generation1K<n<10K1 likes50 downloads2y agoHugging Face07wuwu616 /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M0 likes49 downloads6d agoHugging Face08Alextgt /kg-triplet-graphrag kg-triplet-graphrag Open-text passages annotated with {{entities, relationships}} in the Microsoft GraphRAG knowledge-model format. Part of kg-triplet-sft: https://github.com/Alex-tangt/kg-triplet-sft Size: 3,349 labeled passages — train 2,575 / validation 75 / test 699. (The split file lists 700 test ids; one, wikipedia-01183, had no teacher labels and is omitted.) Domain / language: English; Wikipedia 60% + arXiv 40%; sentence-boundary chunks (100–300 words). Labels: produced… See the full description on the dataset page: https://huggingface.co/datasets/Alextgt/kg-triplet-graphrag.texttext-generation1K<n<10K0 likes42 downloads10d agoHugging Face09Borisz42 /CTNSG-Graph-Curriculum CTNSG Graph Curriculum Dataset This dataset contains preprocessed graphs from WebNLG (v3.0), ATOMIC, and Spider. It is explicitly designed for the Canonical Tractable Neuro-Symbolic Generation (CTNSG) framework. Preprocessing All raw data has been parsed into continuous node and edge embeddings using sentence-transformers/all-MiniLM-L6-v2. Crucially, the graphs have been mathematically canonicalized using the Reverse Cuthill-McKee (RCM) algorithm. This minimizes… See the full description on the dataset page: https://huggingface.co/datasets/Borisz42/CTNSG-Graph-Curriculum.texttext-generation1K<n<10K0 likes41 downloads3mo agoHugging Face10nagygabor /Z3-Verified-Reasoning-Graphs Z3-Verified Constraint Reasoning Dataset 5k Baseline · Production-Ready · Zero Label Noise The Problem This Solves Most synthetic reasoning datasets only show the "happy path". Real reasoning requires knowing when to backtrack. Open-source LLMs hallucinate on constraint satisfaction problems because they are trained on fluent-sounding but logically inconsistent traces. This dataset is different: ❌ No LLM-generated reasoning — zero hallucinations, zero label noise ✅… See the full description on the dataset page: https://huggingface.co/datasets/nagygabor/Z3-Verified-Reasoning-Graphs.tabulartext-generation1K<n<10K1 likes39 downloads5mo agoHugging Face11Beanbagdzf /graph_problem_traces_test1 Additional Information This dataset contains graph and discrete math problem-solving traces generated using the CAMEL framework. Each entry includes: A graph and discrete math problem statement A final answer A tool-based code solution Meta data texttext-generationn<1K0 likes38 downloads2y agoHugging Face12Pete1994 /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/Pete1994/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M0 likes25 downloads4mo agoHugging Face13VittorioRossi /GraphCode-Bench-500-v0 GraphCode-Bench-500-v0 GraphCode-Bench is a benchmark for evaluating LLMs on call-graph reasoning — given a function in a real-world repository, can a model identify which functions call it (upstream) or which functions it calls (downstream), across 1 and 2 hops? Models are evaluated agentically: they receive read-only filesystem tools (list_directory, read_file, search_in_file) and up to 10 turns to explore the codebase before producing an answer. Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/VittorioRossi/GraphCode-Bench-500-v0.textquestion-answeringn<1K0 likes20 downloads6mo agoHugging Face14xushuwen23 /GraphWalkerBenchThis repository contains the GraphWalkerBench dataset from the paper GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum. Code: https://github.com/XuShuwenn/GraphWalker tabularquestion-answering1K<n<10K0 likes20 downloads1mo agoHugging Face15Fodda-ai /industry-intelligence-graph-samples Fodda Industry Intelligence — Graph Samples Expert-curated knowledge graph slices for AI agents and LLM fine-tuning. This dataset contains JSON-LD samples from Fodda's five core domain knowledge graphs — showing the top trending topics across Retail, Beauty, Sports, Fashion, and Culture. These are slices of a much larger interconnected intelligence system. What Fodda Is Fodda is an AI context layer built on PSFK's 20+ years of editorial expertise. It structures… See the full description on the dataset page: https://huggingface.co/datasets/Fodda-ai/industry-intelligence-graph-samples.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face16chatarchive /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/chatarchive/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M0 likes13 downloads1mo agoHugging Face17graphium-company /hu-collocation-lambada hu-collocation-lambada Dataset Summary This dataset is a Hungarian benchmark designed to evaluate large language models' understanding of contextual collocations and definitions. It is inspired by the LAMBADA task and constructed using the full content of the Magyar szókapcsolatok, kollokációk adatbázisa (Temesi, ed.). Each data point contains: a target collocation from the original dataset, its dictionary-style definition, and a short narrative ending just before the… See the full description on the dataset page: https://huggingface.co/datasets/graphium-company/hu-collocation-lambada.texttext-generation1K<n<10K0 likes9 downloads1y agoHugging Face18Jayesh2160 /wikipedia_knowledge_graph_en Dataset Card for Wikipedia Knowledge Graph The dataset contains 16_958_654 extracted ontologies from a subset of selected wikipedia articles. Dataset Creation The dataset was created via LLM processing a subset of the English Wikipedia 20231101.en dataset. The initial knowledge base dataset was used as a basis to extract the ontologies from. Pipeline: Wikipedia article → Chunking → Fact extraction (Knowledge base dataset) → Ontology extraction from facts →… See the full description on the dataset page: https://huggingface.co/datasets/Jayesh2160/wikipedia_knowledge_graph_en.texttext-generation100K<n<1M0 likes9 downloads1mo agoHugging Face19trychannel3 /channel3-universal-product-graph-samplegated Channel3 Universal Product Graph (Sample) Structured, AI-ready product data from across the web. This repository contains a small evaluation sample. The full dataset is available under a commercial license from Channel3. What this is Channel3 maintains a universal product graph: a connected, continuously updated dataset of products from across the internet. For each product it provides normalized and comparable titles, stable identifiers for deduplication and… See the full description on the dataset page: https://huggingface.co/datasets/trychannel3/channel3-universal-product-graph-sample.tabulartext-retrievaln<1K0 likes6 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.