datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hf-training-corpusquant-finance-hft-trading-2026
⚡ Quantitative Finance & High-Frequency Trading (HFT) SFT/DPO Suite (2026)
Institutional-grade instruction fine-tuning and preference alignment dataset for training domain-expert Large Language Models in Quantitative Finance, Algorithmic Execution, and Ultra-Low-Latency HFT Systems.
Engineered to the Mandatory Tier-1 Quality Standard: 80–150 lines of dense, production-grade C++20 and Rust per code snippet. Zero stubs, zero toy snippets, zero heap allocations on the critical… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/quant-finance-hft-trading-2026.hf-tmp-c1-reposurban-mobility-intent
Urban Mobility Intent Dataset
A small fully synthetic dataset for testing dataset ingestion, metadata extraction,
classification, and Hugging Face repository monitoring pipelines.
The records do not contain real user information or production data.
Dataset Structure
data/train.jsonl - training examples
data/validation.jsonl - validation examples
dataset_info.json - lightweight schema/metadata
LICENSE - MIT license
HFTP
HFTP Experimental Datasets
Overview
This dataset contains experimental corpora used in the paper "Hierarchical Frequency Tagging Probe (HFTP): A Unified Approach to Investigate Syntactic Structure Representations in Large Language Models and the Human Brain", published at NeurIPS 2025.
📄 Paper: https://arxiv.org/abs/2510.13255
💻 Code Repository: https://github.com/LilTiger/HFTP
Dataset Description
These datasets are designed for analyzing syntactic… See the full description on the dataset page: https://huggingface.co/datasets/GarfieldX/HFTP.hft12735-review-jh6gwynw
Model Candidate Review Report
Review Date: 2026-08-23
Total Candidates Reviewed: 12
Approved Count: 5
Rejected Count: 7
Top Approved Model: mod-010
Approved Model IDs: mod-001,mod-002,mod-003,mod-008,mod-010
sv_corpora_parliament_processedSwedish text corpus created by extracting the "text" from dataset = load_dataset("europarl_bilingual", lang1="en", lang2="sv", split="train") and processing it with:
import re
def extract_text(batch):
text = batch["translation"]["sv"]
batch["text"] = re.sub(chars_to_ignore_regex, "", text.lower())
return batch
hf_test_dataset_card_b9b92a59f2f0913ffbc7928afa48fd4b_datasethft-market-data-1minhf-trending-summarieshft12735-review-qu7n64i3
Model Candidate Review Report
Review Date: 2026-08-24
Total Candidates Reviewed: 12
Approved Count: 5
Rejected Count: 7
Top Approved Model: mod-010
Approved Model IDs: mod-001,mod-002,mod-003,mod-008,mod-010
hf_test_dataset_card_b910a74a0488c2a692a812796081fa43_datasethft12735-review-m9vd0mdi
Model Candidate Review Report
Review Date: 2026-08-23
Total Candidates Reviewed: 12
Approved Count: 5
Rejected Count: 7
Top Approved Model: mod-010
Approved Model IDs: mod-001, mod-002, mod-003, mod-008, mod-010
hft12735-candidates-jh6gwynwmei_hell_paradisehf_test_dataset_card_1c5fed6f86116d1f0390ef2142f51742_datasetci_outputsSinhala-writing-stylestemp_wiki_dpr_2hft12735-candidates-qu7n64i3hf-testtemp_test_reportsHFTransformersT5Translation_jp-en-ruHFTransformersT5Translation_jp-en-ru
Multilingual Neural Machine Translation using T5 (Hugging Face Transformers)
for Japanese, English, and Russian.
Model Description
This repository provides a multilingual neural machine translation system
built using T5 from
Hugging Face Transformers.
The system supports translation between:
Japanese (ja)
English (en)
Russian (ru)
Supported Language Pairs
ja ↔ en
en ↔ ru
ja ↔ ru
Training Data
The training corpus is publicly… See the full description on the dataset page: https://huggingface.co/datasets/SHSK0118/HFTransformersT5Translation_jp-en-ru.SROIE-document-parsinghft12735-candidates-m9vd0mditiny_model_infoindian-legal-nerhf_time_yt_2707_readtesthf_test_afriannotate
test hf
Annotated dataset from AfriAnnotate project #320.
Provenance
Tool: AfriAnnotate (Label Studio fork).
Project ID: 320
Rows: 1
Aggregation: raw — how multi-annotator labels were resolved into per-row values.
Provenance sidecar: _annotations column included.
Annotator language: unspecified
License
SPDX: unspecified
Name: unspecified
Generated by AfriAnnotate — fresh mode.
hf_tt_cfc1b8_registry
