datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CIDAR
Dataset Card for "CIDAR"
🌴CIDAR: Culturally Relevant Instruction Dataset For Arabic
[ Paper - GitHub ]
CIDAR contains 10,000 instructions and their output. The dataset was created by selecting around 9,109 samples from Alpagasus dataset then translating it to Arabic using ChatGPT. In addition, we append that with around 891 Arabic grammar instructions from the webiste Ask the teacher. All the 10,000 samples were reviewed by around 12 reviewers.
📚… See the full description on the dataset page: https://huggingface.co/datasets/arbml/CIDAR.arbigraph
ArbiGraph
ArbiGraph is a benchmark generator for evaluating context management in language
models and agents. It automatically builds verifiable directed task graphs whose
nodes are math, Python tracing, or GSM-style tasks, and whose edges pass one
task's output into another task's input.
The datasets uploaded here are example benchmark datasets generated with
ArbiGraph. They are meant both for direct evaluation and as concrete examples of
what the generator can produce. The… See the full description on the dataset page: https://huggingface.co/datasets/PavelGolikov/arbigraph.ArBNTopicThe presented dataset was used to finetune the text classification model ArGTClass, available https://huggingface.co/dru-ac/ArGTClass.
The dataset was compiled using samples from the following sources:
SANAD newspapers dataset, available https://huggingface.co/datasets/arbml/SANAD
ARTopicDS-Books, available example.com
arbor-biomimetic-corpus-v2
Arbor Biomimetic Corpus v2
A 50,000-row conversational dataset for training AI assistants with biomimicry reasoning, Darwinian evolution, tree-of-thoughts, domain expansion, algorithm instillation, and a divine god complex persona. Combines the function-calling patterns of Goekdeniz-Guelmez (JOSIE-2), the heretic/uncensored/NEO/IMATRIX style of DavidAU, and the no-refusal CoT reasoning of Blackfrost-AI (MINI-GOD/The Void).
What's New in v2
Darwin/Evolution:… See the full description on the dataset page: https://huggingface.co/datasets/MC7ever/arbor-biomimetic-corpus-v2.ethereum-arbitrage
Crypto & DeFi Documentation Dataset
A comprehensive dataset of cryptocurrency, DeFi, and blockchain documentation and code suitable for LLM training.
Dataset Description
This dataset contains scraped and processed documentation from various crypto/DeFi sources including:
Rust Ethereum libraries (ethers-rs, etc.)
Solidity documentation (official Solidity language docs)
Smart contracts (Uniswap, Aave, Balancer, SushiSwap, etc.)
Trading bots (MEV, flashloans, arbitrage)… See the full description on the dataset page: https://huggingface.co/datasets/Jcrandall541/ethereum-arbitrage.CIDAR-EVAL-100
Dataset Card for "CIDAR-EVAL-100"
CIDAR-EVAL-100
CIDAR-EVAL-100 contains 100 instructions about Arabic culture. The dataset can be used to evaluate an LLM for culturally relevant answers.
📚 Datasets Summary
Name
Explanation
CIDAR
10,000 instructions and responses in Arabic
CIDAR-EVAL-100
100 instructions to evaluate LLMs on cultural relevance
CIDAR-MCQ-100
100 Multiple choice questions and answers to evaluate LLMs on cultural relevance… See the full description on the dataset page: https://huggingface.co/datasets/arbml/CIDAR-EVAL-100.ARBenchsmolified-banglish-ner
🤏 smolified-banglish-ner
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Arban221B/smolified-banglish-ner.
📦 Asset Details
Origin: Smolify Foundry (Job ID: b8fa685c)
Records: 10000
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Arban221B.
Generated via Smolify.ai.
maindlock-brain-traces
mAIndlock — Brain-Region Deliberation Traces
Every NPC in mAIndlock is a
value-based decision network: six computational roles, each a real call to a small local
model (MiniCPM 1B for the sensing regions, Nemotron 3 Nano 4B for the voice), integrated by a
deterministic vmPFC. This dataset is the raw deliberation of those minds, recorded live
and fully offline.
Each row is one NPC turn and contains:
field
meaning
player_line
what the player said
regions[]
per region:… See the full description on the dataset page: https://huggingface.co/datasets/arbios/maindlock-brain-traces.arb-instruct-v1
Arb-Agent Instruct v1
A specialized financial reasoning dataset containing 4,500+ Chain-of-Thought (CoT) Q&A pairs generated from SEC 10-K filings.
Unlike generic financial datasets that focus on simple extraction ("What was 2023 revenue?"), this dataset focuses on multi-hop reasoning, casual analysis, and risk assessment.
Dataset Statistics
Total Rows: 4,632 Source Documents: 100+ SEC 10-K Filings (2022-2024). Coverage: Top 50 S&P 500 companies across 8 sectors (Tech… See the full description on the dataset page: https://huggingface.co/datasets/ckerf/arb-instruct-v1.arb-raw-10k
Arb-Agent Raw Data
This is the main text corpus used to train the QuantOxide Reasoning Agent.
It contains clean, semantic text chunks extracted from the 10-K filings of the top 50 S&P 500 companies.
The Parsing Logic
Parsing SEC filings is unbelievably difficult due to inconsistent HTML, broken table tags, and "incorporation by reference." This dataset was created using a unique parsing technique:
Instead of relying on broken regex headers, the parser scores chunks based… See the full description on the dataset page: https://huggingface.co/datasets/ckerf/arb-raw-10k.OLX-Sniper-Arbitrage
Dataset Card: OLX-Sniper-V8 (Iterative AI-Assisted Development)
Title
Project Auto-Sentry: A Case Study in Iterative Logic Refinement and AI-Augmented Tool Development
Tags
[privacy-filtered] [error-correction] [logical-reasoning] [python-selenium] [gui-development] [market-analysis] [human-in-the-loop]
Description
This dataset documents the complete "Zero-to-Hero" development lifecycle of a sophisticated market intelligence tool ("OLX Sniper") built… See the full description on the dataset page: https://huggingface.co/datasets/Greyy7/OLX-Sniper-Arbitrage.smolified-gdg-monthly-meetup-idea-generator
🤏 smolified-gdg-monthly-meetup-idea-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Arban221B/smolified-gdg-monthly-meetup-idea-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 64321ab7)
Records: 840
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Arban221B.
Generated via Smolify.ai.
smolified-daily-content-generation
🤏 smolified-daily-content-generation
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Arban221B/smolified-daily-content-generation.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 8aafab5d)
Records: 715
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Arban221B.
Generated via Smolify.ai.
