datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LlamaLens-Arabic-Native
LlamaLens: Specialized Multilingual LLM Dataset
Overview
LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi.
LlamaLens
This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation.
Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Arabic-Native.BFCL-V4-Parallel-Native
BFCL V4 Parallel Native
Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
Categories
live_parallel
live_parallel_multiple
parallel
parallel_multiple
Counts
train: 352 rows
eval: 88 rows
total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.entity-native-agent-sessions
Entity-Native vs File-Native Agent Sessions on SWE-bench Verified
Full session logs from a controlled A/B experiment measuring how a coding agent's
retrieval substrate changes its behaviour, cost, and success rate on real
software-engineering tasks.
Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the
same repository state. The only difference is how the agent is allowed to find code.
Arm
Label
Tools available
A
file-native
Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.LlamaLens-Hindi-Native
LlamaLens: Specialized Multilingual LLM Dataset
Overview
LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi.
LlamaLens
This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation.
Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Hindi-Native.dfm11-toolace-native-tool-use-repaired
dfm11-toolace-native-tool-use-repaired
ToolACE conversations with declared-name parsing and complete parallel result binding.
This is a DFM11 replacement for schneiderkamplab/dfm10-toolace-native-tool-use. All rows pass exhaustive structural validation. See metadata/manifest.json.
Home-Assistant-Requests-V5.2-Native-Strict
Home Assistant Requests V5.2 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.korean-legal-retrieval-source-native-250k
Korean Legal Retrieval Source-Native 250K
Legalize-KR의 법령·행정규칙·판례·자치법규 구조에서 query/positive 관계를 추출한
250,000-row 한국어 retrieval dataset이다. release_eligible: false인
target-adapted 연구·비상업 성능 shard이며 통합 라이선스는 other다.
구성과 고정 revision
Source
Revision
Rows
구조 관계
legalize-kr/legalize-kr
db3cd760c14042ee04fd9166e1bdbb662fc999bc
50,000
법령명+조문 → 조문 본문
legalize-kr/admrule-kr
64a5a272909ab5bc077b0ad9519ef31de8febb46
50,000
규칙명+조문 → 조문 본문
legalize-kr/precedent-kr… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-legal-retrieval-source-native-250k.qfs-smollm2-135m-wikitext2-native-v1
HF workflow d3dc69602aeb981f06bd9f4c726937f9
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.dfm11-synthetic-native-tool-calling-repaired
dfm11-synthetic-native-tool-calling-repaired
DFM8 synthetic tool trajectories with compatibility normalization materialized in source data.
This is a DFM11 replacement for schneiderkamplab/dfm8-synthetic-native-tool-calling. All rows pass exhaustive structural validation. See metadata/manifest.json.
Home-Assistant-Requests-V5-Native
Home Assistant Requests V5 Native
Native Home Assistant tool-calling dataset for home-assistant-specialist-v0.5-4b-q5.
Contract
Track B. Model learns native tools, not legacy ha-action-v3 JSON:
HassTurnOn, HassTurnOff, HassToggle, HassSetPosition, HassLightSet, climate/media/vacuum/timer/todo tools, and GetLiveContext;
tool results remain in the message sequence;
train_on_turn is preserved and non-target turns are masked by the trainer;
direct speech is retained… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5-Native.Home-Assistant-Requests-V5.1-Native-Strict
Home Assistant Requests V5.1 Native Strict
Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model.
Contract: ha-native-tool-calling-v2.
Frozen snapshot
Split
Rows
Direct speech
Multi-call
Maximum rendered tokens
train
3,806
340
78
3,098
validation
530
52
4
2,874
test
633
102
22
2,925
Tokenizer audit:
model: unsloth/Qwen3-4B-Instruct-2507
revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.1-Native-Strict.US_Native_American_Tribal_Treaties_Table_from_WikipediaORIGINALLY COMPILED ON WIKIPEDIA. Cleaned and improved from original version on Wikipedia, by completing some column information that was available elsewhere in the wikipedia article.
This is a dataset of tabular data regarding treaties between the USA and Native American Tribes/Nations, to date, including many executive orders.
Table columns include Year, Date, Treaty name, "Alternative Treaty name", Statutes, "Land cession reference (Royce Area)", Tribe(s).
All of those listed include the… See the full description on the dataset page: https://huggingface.co/datasets/pseudolab/US_Native_American_Tribal_Treaties_Table_from_Wikipedia.web-access-api-benchmarks
NativePort Web-Access API Benchmarks
Measured quality, latency, cost and error-rate figures for 22 commercial web-access
APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots,
document parsing, browser actions and change watching — scored per capability on a
fixed task corpus. This is the 2026-08-05 run: 67 provider × capability
scorecards across 13 capabilities, flattened into 297 metric rows.
It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.assay-pose-replay-offset-tool-native-v1-rlmaze2d_easy_native256_stepmsg_fixedstart_plain
maze2d_easy_native256_stopreq_stepmsg_fixedstart_ordered / maze2d_easy_native256_stopreq_stepmsg_fixedstart_cot_ordered
Built by build_maze2d_native256_ordered_pair.py at 20260606_stepmsg_fixedstart_v2.
Manifest version: maze2d_native256_stopreq_stepmsg_fixedstart_plain_cot_ordered100k_v2. CoT policy: maze2d_native256_stopreq_stepmsg_fixedstart_cot_branch_v2.
Plain and CoT rows share the same retained episode manifests and order for each split.
osworld-native-parity-runs
OSWorld Native Parity Runs
This dataset stores large native OSWorld parity run archives that are too large for the shared Harbor parity-experiments dataset.
Archives
fulltask_20260621-172216/attempt_1/osworld-native-fulltask_20260621-172216-attempt1-349of361.tar.zst
Source run: /home/servermacadmin/osworld-parity/parity_results/fulltask_20260621-172216/attempt_1
Source upstream: xlang-ai/OSWorld at fe8c78e
Tasks: OSWorld-Verified no-Google-Drive split, 361 task… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/osworld-native-parity-runs.dfm8-synthetic-native-tool-calling
Native Tool Calling
Synthetic DFM8 training data generated with Gemma 4 31B and filtered by deterministic checks plus a Gemma 4 31B judge.
Schema
Rows are JSONL chat records:
{"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Tool-calling rows may also include a top-level tools list and assistant tool_calls.
Counts
accepted rows: 836675
generated rows seen: 4800000
audit rows seen: 4580233… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm8-synthetic-native-tool-calling.maze2d_easy_native256_cot_chunk_kinf_20260707_perseg
maze2d_easy_native256_cot_chunk_kinf_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_kinf_20260707_perseg.maze2d_easy_native256_noncot_chunk_k3_20260707_perseg
maze2d_easy_native256_noncot_chunk_k3_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k3_20260707_perseg.maze2d_easy_native256_cot_chunk_k5_20260707_perseg
maze2d_easy_native256_cot_chunk_k5_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k5_20260707_perseg.mozgach_localizations
Mozgach Localizations Dataset
Dataset Description
This dataset contains localization strings for the Mozgach application, providing translations from Russian to multiple languages including Chinese, Arabic, and others. The dataset is formatted for instruction-following language models and translation tasks.
Languages
Source Language: Russian (ru)
Target Languages: Chinese (zh), Arabic (ar), and others
Dataset Structure
Each entry in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/mozgach_localizations.react_native_code_review-reasoning-SFTawesome-ai-native
Awesome AI Native — Dataset
A curated, structured dataset of 137 resources across 18 categories and 6 themes for understanding and building AI Native products, systems, and teams — where large models are the foundation, not a feature.
This is the machine-readable companion to the Awesome AI Native list and its website.
Columns
Field
Description
category
Category id (e.g. agents, eval)
category_name
Human-readable category name
group
Theme id… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-ai-native.maze2d_easy_native256_noncot_chunk_k10_20260707_perseg
maze2d_easy_native256_noncot_chunk_k10_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k10_20260707_perseg.maze2d_easy_native256_cot_chunk_k10_20260707_perseg
maze2d_easy_native256_cot_chunk_k10_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k10_20260707_perseg.NativeDE-Opus4.7-REAP
NativeDE-Opus4.7-REAP
A native German synthetic reasoning dataset generated using Anthropic Claude Opus 4.7 (claude-opus-4-7). All prompts and responses are in natural, idiomatic German — not translations from English. Each sample contains an explicit <think>...</think> reasoning block followed by a Final answer: boundary and the actual response.
This dataset is the German-language complement to BaaderSo36-Opus4.7-REAP.
Dataset Statistics
Total samples: 2,306
Source… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/NativeDE-Opus4.7-REAP.maze2d_easy_native256_noncot_chunk_k5_20260707_perseg
maze2d_easy_native256_noncot_chunk_k5_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k5_20260707_perseg.maze2d_easy_native256_noncot_chunk_k1_20260707_perseg
maze2d_easy_native256_noncot_chunk_k1_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_k1_20260707_perseg.maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg
maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (non-CoT action-chunk baseline) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_noncot_chunk_kinf_20260707_perseg.maze2d_easy_native256_cot_chunk_k1_20260707_perseg
maze2d_easy_native256_cot_chunk_k1_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k1_20260707_perseg.
