datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TwoHopFactThis is the dataset introduced in the paper Do Large Language Models Latently Perform Multi-Hop Reasoning?.
Code: https://github.com/google-deepmind/latent-multi-hop-reasoning.
SOCRATESThis is the dataset introduced in the paper Do Large Language Models Latently Perform Multi-Hop Reasoning without Exploiting Shortcuts?.
Code: https://github.com/google-deepmind/latent-multi-hop-reasoning.
glassdoor_reviewsindia-ev-range-dataset
🔋 India EV Real-World Range Dataset
A comprehensive dataset of 7,291 data points covering 61 Indian 4-wheeler EV variants from 18 manufacturers across 12 Indian driving scenarios.
Dataset Description
This dataset was built to predict the real-world driving range of Electric Vehicles under Indian conditions. It combines:
Real EV specifications from all major EVs sold in India (Tata, Mahindra, Hyundai, MG, Kia, BYD, BMW, Mercedes, etc.)
Physics-based energy modeling using… See the full description on the dataset page: https://huggingface.co/datasets/SohamThakkar-07/india-ev-range-dataset.S-OH
Scent of Health (S-OH) Dataset
The Scent of Health (S-OH) dataset is the largest public clinical electronic nose (eNose) collection for non-invasive disease screening via exhaled breath analysis. It comprises 1,234 patients across nine diagnostic groups (healthy controls and eight diseases), each providing a 17-channel multivariate time series of breath measurements.
Property
Value
Patients
1,234
Diagnostic groups
9 (healthy + 8 diseases)
Time series channels
17… See the full description on the dataset page: https://huggingface.co/datasets/blinoff/S-OH.FinRAD_Financial_Readability_Assessment_Dataset
FinRAD: Financial Readability Assessment Dataset - 13,000+ Definitions of Financial Terms for Measuring Readability
This repository contains the dataset mentioned in the paper: FinRAD: Financial Readability Assessment Dataset - 13,000+ Definitions of Financial Terms for Measuring Readability (presented at The Financial Narrative Processing Workshop colocated with LREC-2022, Marseille, France).
In addition to this, data collection & cleaning scripts, embedding extraction & model… See the full description on the dataset page: https://huggingface.co/datasets/sohomghosh/FinRAD_Financial_Readability_Assessment_Dataset.malicious-prompts
