datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cortex-1-market-analysis
NEAR Cortex-1 Market Analysis Dataset
Dataset Summary
This dataset contains blockchain market analyses combining historical and real-time data with chain-of-thought reasoning. The dataset includes examples from Ethereum, Bitcoin, and NEAR chains, demonstrating high-quality market analysis with explicit calculations, numerical citations, and actionable insights.
The dataset has been enhanced with examples generated by GPT-4o and Claude 3.7 Sonnet, providing diverse… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/cortex-1-market-analysis.arxiv-tex-corpus-fullarxiv-tex-corpus-full (80GB)
Large-scale LaTeX corpus from arXiv (math, CS, physics, statistics)
📄 Paper: https://arxiv.org/abs/2602.17288
📚 Overview
arxiv-tex-corpus-full (80GB) is a large-scale dataset of LaTeX source content extracted from papers hosted on arXiv.
This version contains approximately 80GB of structured JSONL data, restricted to the following arXiv categories:
math
cs
hep-th
hep-ph
quant-ph
stat.ML
stat.TH
The dataset is designed for research in:
Large… See the full description on the dataset page: https://huggingface.co/datasets/CortexEvolved/arxiv-tex-corpus-full.keural-cortex-8b-sft
Keural-Cortex-8B SFT dataset
The supervised fine-tuning set used to train Keural-Cortex-8B, a Korean-first
bilingual model with a 64K context window, tool calling, and hybrid
thinking/non-thinking modes.
1,568,649 rows · 1.87B estimated tokens · 1.80B real Qwen3 tokens · 73.7% Korean
Three files:
file
rows
what it is
train.jsonl
1,564,042
main set, all rows under 32,768 tokens
train_long64k.jsonl
4,607
the 32K–64K band, kept separate because it needs a different… See the full description on the dataset page: https://huggingface.co/datasets/mkd-hossain/keural-cortex-8b-sft.cortex-decision-threadsEdgeMMEval
EdgeMMEval
Minimal multimodal evaluation dataset for on-device inference testing.
Covers functional correctness, accuracy, latency stress, and memory
pressure across image, audio, text, multi-turn, combination, structured
output, and tool-calling cases.
Dataset summary
The test split is defined in data/test/metadata.jsonl (200 rows). Each
row has a test_id (for example IMG-001, STO-020) and a modality.
Modality
Samples
Focus
Image
34
VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.cortex-skill-routingcortex-workflow-triggerscortex-enterprise-qaCortexLM__btlm-7b-base-v0.2-details
Dataset Card for Evaluation run of CortexLM/btlm-7b-base-v0.2
Dataset automatically created during the evaluation run of model CortexLM/btlm-7b-base-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CortexLM__btlm-7b-base-v0.2-details.Cortex-Nexus-Emotional-Lab
Cortex-Nexus: Emotional State Injection in Large Language Models
A Controlled Experimental Study on Simulated Curiosity and Output Quality
Lead Researcher & Developer: Maximiliano Rodrigo Speranza - https://www.linkedin.com/in/maximiliano-speranza-35876737a/ - https://github.com/SperanzaMax
Affiliations: Cisco Networking Academy (Certified) · Universidad Tecnológica Nacional — Facultad Regional Buenos Aires (UTN-BA)
AI Collaborators: Antigravity AI · Claude (Anthropic… See the full description on the dataset page: https://huggingface.co/datasets/SperanzaMax/Cortex-Nexus-Emotional-Lab.cortex-agent-planningcortex-training-data
