datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
halo-hil
halo-hil
Web text in hil, re-filtered by language and prepared for
pretraining.
What changed, and why it had to
The earlier version of this dataset was labelled hil by the
crawler's own language detection, and that label was never verified. An audit
on 2026-09-22 found that most of it was not hil: over a random
sample of 1,499 sentences, GlotLID v3 called 44 % English, 22 %
Filipino/Tagalog and only 12 % Hiligaynon — much of the corpus was Tagalog
news copy and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.HALO-Gemini-3-Flash-AppWorld
Dataset Card: Gemini 3 Flash Traces on AppWorld (test-normal)
Dataset Overview
This dataset contains agent execution traces of Gemini 3 Flash running on the AppWorld benchmark, specifically evaluated on the test-normal dataset split. The traces capture the full span-level execution detail of the model interacting with AppWorld's simulated app ecosystem.
Field
Value
Model
Gemini 3 Flash
Benchmark
AppWorld
Split
test-normal
Total Traces
168
Total Spans
3… See the full description on the dataset page: https://huggingface.co/datasets/inference-net/HALO-Gemini-3-Flash-AppWorld.halo-docs
Halo Documentation Q&A
English, single-turn instruction-tuning examples about the Halo LLM training toolkit: concepts, configurations, commands, model recipes, internals, and troubleshooting.
This is a source-preserving, extractive Q&A dataset. Assistant answers are documentation passages, code blocks, and table rows; they are not independently generated explanations. Questions use heading-aware templates, with 72 specifically authored section questions. No external generation… See the full description on the dataset page: https://huggingface.co/datasets/skundu42/halo-docs.
