datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KemSU
🎓 Kemerovo State University Instructional QA Dataset (NodeLinker/KemSU)
📝 Dataset Overview & Splits
This dataset provides instructional question-answer (Q&A) pairs meticulously crafted for Kemerovo State University (КемГУ, KemSU), Russia. Its primary purpose is to facilitate the fine-tuning of Large Language Models (LLMs), enabling them to function as knowledgeable and accurate assistants on a wide array of topics concerning… See the full description on the dataset page: https://huggingface.co/datasets/NodeLinker/KemSU.btoe-nodes
Boris' Theory of Everything (BTOE) — Semantic Node Dataset
Pre-chunked semantic nodes from Boris' Theory of Everything (BTOE) for
RAG, embedding, fine-tune experiments, and cross-model integrity checks.
Core claim (primary layer): divine love — as revealed in Jesus Christ — is the
fundamental law of existence; the universe is framed as a womb for free-willed
growth toward partnership with God (relational / node structure, not a physics ToE competitor).
Official site:… See the full description on the dataset page: https://huggingface.co/datasets/therussian/btoe-nodes.NodeJS-5000node_red_nl_prompts
Geyanth08/node_red_nl_prompts
This dataset contains curated NL-to-Node-RED prompt and flow pairs.
Primary fine-tuning splits are under data:
data/train.jsonl
data/validation.jsonl
Each record includes:
flow_id
prompt_id
prompt_style
prompt_text
generation_source
source_title
source_summary_text
source_url
target_flow_json
created_at_utc
Files
data/train.jsonl: Curated train split
data/validation.jsonl: Curated validation split
data/kept_all.jsonl: All curated… See the full description on the dataset page: https://huggingface.co/datasets/Geyanth08/node_red_nl_prompts.gpu-node-inventorysentence-unscramble-r1-nodedupe-it1nodejs100gpu-node-inventoryNodedata[
{
"instruction": "What is Node.js?",
"output": "Node.js is a cross-platform, open-source JavaScript runtime environment that executes JavaScript code outside a web browser, built on Chrome's V8 engine."
},
{
"instruction": "Explain the difference between Node.js and JavaScript.",
"output": "JavaScript is a programming language used for client-side scripts, while Node.js is a runtime environment that allows JavaScript to be used on the server-side."
},
{… See the full description on the dataset page: https://huggingface.co/datasets/Steve393/Nodedata.reddit_finance_43_250k
reddit finance 43 250k
reddit_finance_43_250k is a collection of 250k post/comment pairs from 43 financial, investing and crypto subreddits. Post must have all been text, with a length of 250chars, and a positive score. Each subreddit is narrowed down to the 70th qunatile before being mergered with their top 3 comments and than the other subs. Further score based methods are used to select the top 250k post/comment pairs.
The code to recreate the dataset is here:… See the full description on the dataset page: https://huggingface.co/datasets/NodecoreHQbot/reddit_finance_43_250k.code-ast-node-edge-sequenceraw_node_red_flows
Geyanth08/raw_node_red_flows
This dataset contains extracted Node-RED flow records for NL-to-Flow generation research.
Files
data/train.jsonl: Main training records in JSONL format
metadata/dataset_v0_summary.csv: Flattened summary view
metadata/flow_urls.csv: URL source list
metadata/operation_timings.csv: Collection and extraction run timings
metadata/dataset_v0_extract_errors.csv: Extraction errors log
Record Fields
Each JSONL row contains:
id… See the full description on the dataset page: https://huggingface.co/datasets/Geyanth08/raw_node_red_flows.
