datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fol-data
FOL Reasoning Dataset
A preprocessed and vocabulary-augmented dataset derived from the ProofWriter (Kaggle) OWA splits, built for training a Natural Language → First-Order Logic translation model.
The source dataset contains natural-language premises and questions in English along with structured proof metadata. Our preprocessing adds two things that the original does not provide:
FOL translations — each natural-language statement is converted to First-Order Logic via a rule-based… See the full description on the dataset page: https://huggingface.co/datasets/Venkatdatta/fol-data.ventset
Ventset: Raw & Real Conversations with an Empathic AI
Ventset is a dataset of human-AI dialogues featuring an AI designed to respond with empathy, humor, or tough love. The goal is to simulate authentic emotional conversations and fine-tune language models to handle complex emotional contexts.
⚠️ This dataset is still under development — contributions and feedback are welcome!
⚠️ Some messages may be misinterpreted. The creator is not a psychologist. Misuse or misinterpretation… See the full description on the dataset page: https://huggingface.co/datasets/archIBARBUgrr/ventset.IRONWORKS-VENOM-preview
IRONWORKS VENOM
Supply Chain Security Training Dataset — Preview v0.1
by IronGate Digital
What this is
A synthetic instruction-tuning dataset focused on software supply chain security.
Built from real threat intelligence sources including security advisories, research
blogs, and vulnerability databases.
This is an early preview. More datasets are in progress.
Coverage
35,000+ labeled training pairs covering:
Dependency confusion and typosquatting… See the full description on the dataset page: https://huggingface.co/datasets/IronGateDigi/IRONWORKS-VENOM-preview.Pure-Telugu-Alpaca
Pure Telugu Alpaca Dataset
This dataset is a cleaned version of Telugu-MultiTask-Instruct-77K with Telugu keys.
It uses enhanced_prompt as instruction and enhanced_completion as output.
Processing Steps
Extracted enhanced_prompt → సూచన (instruction) and enhanced_completion → అవుట్పుట్ (output)
Filtered to keep only entries with no English letters
Removed duplicate entries
Normalized whitespace
Format
Each entry follows the Alpaca format with… See the full description on the dataset page: https://huggingface.co/datasets/VenkataRamanaKurumallajaddangi/Pure-Telugu-Alpaca.
