datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
halo-hil
halo-hil
Web text in hil, re-filtered by language and prepared for
pretraining.
What changed, and why it had to
The earlier version of this dataset was labelled hil by the
crawler's own language detection, and that label was never verified. An audit
on 2026-09-22 found that most of it was not hil: over a random
sample of 1,499 sentences, GlotLID v3 called 44 % English, 22 %
Filipino/Tagalog and only 12 % Hiligaynon — much of the corpus was Tagalog
news copy and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.narit-ghosts-halo-catalogs
NARIT GHOSTS Halo Catalogs
Dataset Description
This dataset contains reduced stellar catalogs, combined FITS images, and candidate substructure catalogs from the GHOSTS (Galaxy Halos, Outer disks, Substructure, Thick disks, and Star clusters) Survey observed by the Hubble Space Telescope (HST).
It serves as the primary data lake for the automated astronomical pipeline designed to detect faint stellar substructures (like Ultra-Faint Dwarfs and stellar streams) in… See the full description on the dataset page: https://huggingface.co/datasets/appleboiy/narit-ghosts-halo-catalogs.HaLO
HaLO: Handwriting Assessment for Legibility and Ordering
HaLO is a dataset of 1,907 handwriting images from 202 German schoolchildren (ages 10--12), annotated with 19,070 pairwise legibility judgments. It supports research on AI-based handwriting legibility assessment through comparative ranking.
Dataset Description
Each handwriting sample depicts one of ten predefined German sentences (20--27 characters each), written by a child in their natural handwriting… See the full description on the dataset page: https://huggingface.co/datasets/MarcoLents/HaLO.strix-halo-inference-bench
Strix Halo Local Inference Benchmarks
Measured prefill and decode throughput, and real VRAM cost, for local GGUF models
on AMD Strix Halo (Radeon 8060S / gfx1151) under ROCm.
Why this exists
Strix Halo inverts the usual local-inference trade-off. A discrete 24 GB card gives
you high memory bandwidth and a hard capacity ceiling; Strix Halo gives you the
opposite — up to 64 GiB addressable as VRAM out of 128 GB unified, at substantially
lower bandwidth. That changes… See the full description on the dataset page: https://huggingface.co/datasets/axjns/strix-halo-inference-bench.record-pick-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 52,
"total_frames": 38159,
"total_tasks": 1,
"total_videos": 168,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HaloLi0828/record-pick-1.CubeSO101This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 40121,
"total_tasks": 1,
"total_videos": 180,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HaloLi0828/CubeSO101.halohalo
halohalo
Dataset Summary
halohalo is a Pretraining text corpus for Philippine languages,
assembled from web-scraped data. It is compatible with Fineweb for LLM Pretraining.
Source Data
Derived from the following cleaned datasets:
Source
Documents
halo-hil
8,874
halo-tgl
6,589
halo-bcl
1,264
Each source dataset was cleaned using clean_halo.py to remove web boilerplate, navigation menus,
markdown noise, HTML artifacts, and low-quality… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halohalo.Nexesenex__Nemotron_W_4b_Halo_0.1-details
Dataset Card for Evaluation run of Nexesenex/Nemotron_W_4b_Halo_0.1
Dataset automatically created during the evaluation run of model Nexesenex/Nemotron_W_4b_Halo_0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Nemotron_W_4b_Halo_0.1-details.halooglasi-property-scraper-serbia-real-estate-sample-data
Halooglasi.com Property Scraper — Serbia Real Estate
Scrape real estate listings from Halooglasi.com — Serbia's #1 property portal. Extract apartments, houses, land, commercial, garages and rooms (sale or rent) by location, EUR price, m², rooms and owner type. Returns price, area, GPS-ready address, advertiser and image gallery per ad.
What the actor scrapes
Halooglasi.com Property Scraper — Serbia Real Estate Listings to JSON, CSV & Excel Scrape real estate… See the full description on the dataset page: https://huggingface.co/datasets/logiover/halooglasi-property-scraper-serbia-real-estate-sample-data.Quazim0t0__Halo-14B-sce-details
Dataset Card for Evaluation run of Quazim0t0/Halo-14B-sce
Dataset automatically created during the evaluation run of model Quazim0t0/Halo-14B-sce
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Halo-14B-sce-details.
