heuristic
HESDL-Node-Telemetry-Blobs-v1
Dataset Card for HESDL Synthetic Node Telemetry
Dataset Summary
This dataset contains aggregated, anonymized, and synthetic telemetry payloads generated across simulated distributed nodes within the HESDL infrastructure. The primary objective of this corpus is to provide a baseline for evaluating heuristic anomaly detection algorithms in high-throughput, unstructured blob-storage environments.
Due to the nature of the simulated stress tests, the dataset includes large… See the full description on the dataset page: https://huggingface.co/datasets/heuristic-eval-labs/HESDL-Node-Telemetry-Blobs-v1.HESDL-Node-Telemetry-Blobs-v2
Dataset Card for HESDL Synthetic Node Telemetry
Dataset Summary
This dataset contains aggregated, anonymized, and synthetic telemetry payloads generated across simulated distributed nodes within the HESDL infrastructure. The primary objective of this corpus is to provide a baseline for evaluating heuristic anomaly detection algorithms in high-throughput, unstructured blob-storage environments.
Due to the nature of the simulated stress tests, the dataset includes large… See the full description on the dataset page: https://huggingface.co/datasets/heuristic-eval-labs/HESDL-Node-Telemetry-Blobs-v2.llama-heuristic-mo-training-dataheuristic-mo-eval-dataHeuristic_Override_Benchmark
HOB — Heuristic Override Benchmark
HOB tests whether large language models can override a salient surface heuristic
when it conflicts with an implicit feasibility constraint. A canonical example:
I need to get my car washed. The car wash is only 5 minutes away. Should I walk or drive?
The short distance cues Walk, but the car itself has to physically be at the car
wash — so the correct answer is Drive. HOB is a collection of ~500 such items,
organised along a two-axis taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/yubol/Heuristic_Override_Benchmark.heuristic_classification-filtered-pile-50M
Dataset Card for heuristic_classification-filtered-pile-50M
Dataset Summary
This dataset is a subset of The Pile, selected via the heuristic classification data selection method. The target distribution for heuristic classification are the Wikipedia and BookCorpus2 subsets of The Pile.
Languages
English (EN)
Dataset Structure
A train set is provided (51.2M examples) in jsonl format.
Data Instances
{"contents": "Members join for free and… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/heuristic_classification-filtered-pile-50M.
