datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual_captions_jsonConceptGraph
Homepage
Exploring and Verbalizing Academic Ideas by Concept Co-occurrence
https://github.com/xyjigsaw/Kiscovery
Evolving Concept Co-occurrence Graph
It is the official Evolving Concept Co-occurrence Graph dataset of paper Exploring and Verbalizing Academic Ideas by Concept Co-occurrence.
To train our model for temporal link prediction, we first collect 240 essential and common queries from 19 disciplines and one special topic (COVID-19). Then, we enter these queries into… See the full description on the dataset page: https://huggingface.co/datasets/Reacubeth/ConceptGraph.conceptual-captions-cc12m-llavanext
Dataset Card for conceptual-captions-cc12m-llavanext
Dataset Summary
This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B.
Languages
The captions are in English.
Data Instances
An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.concept-guard
Dataset Card for ConceptGuard
Dataset Details
Dataset Description
ConceptGuard is a benchmark dataset for evaluating concept-level unlearning in Large Language Models. It is built around dual-use concepts, where each concept appears in both harmful and benign contexts. The dataset is designed to assess whether models can suppress harmful behavior while preserving useful knowledge, enabling evaluation of contextual separation.
Curated by: Authors… See the full description on the dataset page: https://huggingface.co/datasets/sk0511/concept-guard.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.CIEL-Clinical-Concepts-to-ICD-11TD;LR Run:
pip install torch==2.4.1+cu118 torchvision==0.19.2+cu118 torchaudio==2.4.1 --extra-index-url https://download.pytorch.org/whl/cu118
pip install -U packaging setuptools wheel ninja
pip install --no-build-isolation axolotl[flash-attn,deepspeed]
axolotl train axolotl_2_a40_runpod_config.yaml
📚 CIEL to ICD-11 Fine-tuning Dataset
This dataset was created to support the fine-tuning of open-source large language models (LLMs) specialized in ICD-11 terminology mapping.
It… See the full description on the dataset page: https://huggingface.co/datasets/filipelopesmedbr/CIEL-Clinical-Concepts-to-ICD-11.chess-positions-conceptsPromptCoT-2.0-Concepts
🧩 PromptCoT 2.0 – Concepts Dataset
PromptCoT 2.0 Concepts provides the foundational conceptual inputs for large-scale prompt synthesis in mathematics and programming.These concept files are used to generate high-quality synthetic problems through the PromptCoT 2.0 Prompt Generation Model.
📘 Overview
Each file (e.g., math.jsonl, code.jsonl) contains a list of concept prompts that serve as the input for the problem generation stage.By feeding these prompts into the… See the full description on the dataset page: https://huggingface.co/datasets/xl-zhao/PromptCoT-2.0-Concepts.concept-cot-conv-qa-full
Concept CoT Conversational QA (Full)
Conversational QA pairs about chain-of-thought reasoning traces, generated using DeepSeek v3.2 via OpenRouter. Designed for training activation oracles to answer natural language questions about what a model is doing during reasoning.
Overview
Total pairs: 10,499
Unique source entries: 7,496 (from 8,132 concept corpus entries)
Source corpus: ceselder/concept-cot-corpus-full (concept_corpus/corpus_full.jsonl)
Generator model:… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/concept-cot-conv-qa-full.conceptnet-de-indexed
ConceptNet 5 (Un-normalized SQLite, 23.6 GB)
This repository contains the complete, un-normalized ConceptNet 5.5 knowledge graph in SQLite format. Unlike the filtered version, this dataset includes all languages from the original ConceptNet release.
The database conceptnet-de-indexed.db is a 23.6 GB un-normalized SQLite file containing the full knowledge graph with all 28.3 million nodes and 34 million edges across all languages.
When to Use This Dataset
Use this… See the full description on the dataset page: https://huggingface.co/datasets/cstr/conceptnet-de-indexed.adaption-tech-concepts-explained
Adaption Tech Concepts Explained
A High-Quality Instruction Tuning Dataset for Large Language Models
A high-quality instruction tuning dataset designed for fine-tuning Large Language Models (LLMs) to generate clear, structured, and beginner-friendly explanations of technical concepts.
This dataset was enhanced using Adaption's Adaptive Data Platform, which improves instruction quality, response consistency, and educational value for supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/ujjawalbansal/adaption-tech-concepts-explained.self-instruct-data-concept-need-responsedeepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual
Filtered version of ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.
Filtering
reward filter enabled: False
minimum official reward: 1.0
scoring errors rejected: False
maximum text tokens: 32768
maximum response chars: 200000
near-duplicate SimHash hamming threshold: 4
required <think>...</think> and final boxed answer after reasoning
exact text/problem/response dedupe and near problem dedupe… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k-clean-v4-conceptual.ARC_OWARIDA_conceptConceptNetSyntheticPhi3Text_ja
Dataset Card for Dataset Name
ConceptNet(5.7)の日本語のトリプルに対して, Phi-3 を用いて文を生成したデータセット
Dataset Details
Dataset Creation
Generate a sentence in phi-3 using the below prompt
// h = {start}, r = {relation}, t = {end}
prompt = f"<|assistant|>\n次に示すトリプル: <start>, <relation>, <end>を用いて、<start>と<end>を<relation>で結びつけた関係を表す文を作成しなさい。必ずすべての情報を網羅し、日本語で出力すること。<|end|>\n<|user|>\nトリプル: {h}, {r}, {t}<|end|>\n<|assistant|>"
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/RJZ/ConceptNetSyntheticPhi3Text_ja.pre-conceptual-schemas-alpacaPre-Conceptual Schemas Dataset in Alpaca Format
This repository contains the dataset used for fine-tuning the Phi-3.5-PCS-Finetuned model, designed for the interpretation of Pre-Conceptual Schemas (PCS).
The dataset is a result of the master's thesis project "Natural Language Processing in Pre-conceptual Schemas for Representing Knowledge by Using Large Language Models" from the Master's Degree in Systems and Computing Engineering at the University of Nariño.
Dataset Curators: Felipe Roa… See the full description on the dataset page: https://huggingface.co/datasets/Galerasnet/pre-conceptual-schemas-alpaca.mmlu-conceptual-physicsconcept-datasetconceptnet_UsedFor_en_en_mixtral_finetuneThe purpose of this dataset is to be used for a fine tuning on Mixtral, it contains all the 'UsedFor' relationships (english to english) present in ConceptNet 5.7.0 in the form of aggregated instructions, i.e. for any arg1, arg2_list is the list of all arg2 found to be in a UsedFor relationship with arg1 (arg1 --UsedFor--> arg2)
The instruction is written in the following format: <s> [INST] instruction [/INST] answer </s>
tiny-aya-medical-concept-probes
Tiny Aya Cross-Lingual Medical Concept Probes
Dataset Description
20 medical concepts expressed as full sentences in 10 languages, designed for probing cross-lingual concept representations in multilingual LLMs. Each concept is a complete declarative sentence preserving the same semantic structure across all languages.
Purpose
These probe sentences serve as stimuli for mechanistic interpretability analysis -- specifically, extracting residual stream activations… See the full description on the dataset page: https://huggingface.co/datasets/s4um1l/tiny-aya-medical-concept-probes.HC_concepts_refusaltest_batch_oaUnderstanding-Shopping-Concepts-post-ptchoreo_concepts_qna_test_set_v0.1This dataset contains the merge of rtweera/user_centric_results_v1 and rtweera/user_centric_results_v2 datasets for creating a unified testing for Choreo Concepts SLM finetuning project.
MC_concepts_refusalcore_concepts_100LC_concepts_refusalataques_conceptos_civerchoreo_concepts_docs_qna_trainset_v1.0A synthetically augmented dataset on WSO2 Choreo Product's Documentaion from rtweera/simple_implicit_n_qa_results_v2 dataset.
sae_concept
