datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
ladakh-thermal-shelter-surrogate
Ladakh Thermal Shelter Surrogate Dataset
Dataset Overview
This dataset contains high-fidelity physics-based thermal simulation logs generated locally using EnergyPlus for specialized area-specific shelter design focused on thermal comfort maintenance in extreme cold environments like Ladakh, India.
Dataset Structure
Total Rows: ~1.31 Crore (13,140,000 hourly simulation time-steps)
Configurations: 1,500 unique shelter variations generated via Latin… See the full description on the dataset page: https://huggingface.co/datasets/VashuTheGreat2/ladakh-thermal-shelter-surrogate.LADru_transcription_punctuation
About
This is a dataset for training Russian punctuators/capitalizers via NeMo scripts (https://github.com/NVIDIA/NeMo)
A BERT model already fine-tuned on this dataset can be found here: https://huggingface.co/denis-berezutskiy-lad/lad_transcription_bert_ru_punctuator
Scripts for collecting/updating such a dataset, as well as training/using the model are located here: https://github.com/denis-berezutskiy-lad/transcription-bert-ru-punctuator-scripts/tree/main
The idea behind the… See the full description on the dataset page: https://huggingface.co/datasets/denis-berezutskiy-lad/ru_transcription_punctuation.MultiGeoDTAnewdatasetLadder-machine-learning-MCQslady_for_llm_ft
It's a test dataset
Dataset Schema
问句: 用户输入.
答句: AI回答.
Ladder-machine-learning-QALADECv1ladybird-commitscode_datasetladle_trainingHI
Tone Dataset
This dataset is designed to fine-tune language models to respond in a cool, intelligent, articulate, informed, and professional tone. Each entry contains an input (user prompt) and an output (tone-specific response).
Dataset Structure
input: The user’s question or prompt.
output: The desired response in the desired tone/context.
Dataset Creation
This dataset was manually curated to represent a unique voice in the Fintech space, serving as the… See the full description on the dataset page: https://huggingface.co/datasets/LADRECH/HI.
