datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
venra
VeNRA: Financial Hallucination Detection Dataset
VeNRA (Verification & Reasoning Audit) is a specialized dataset designed to train "Judge" models to detect hallucinations in Financial RAG (Retrieval-Augmented Generation) systems.
Unlike datasets that rely on "generative hallucinations" (asking an LLM to invent errors), VeNRA adopts an Adversarial Simulation philosophy. We scientifically reconstruct the specific cognitive failures that RAG systems exhibit in production by applying… See the full description on the dataset page: https://huggingface.co/datasets/pagand/venra.SlimPajama-62BSubset of cerebras/SlimPajama-627B,
consisting of 10% of the train split and 100% of the test and
validation splits.
The train split consists of chunk2 from the original
[cerebras/SlimPajama-627B] dataset, split into five zstd-compressed jsonl
files for efficient loading. The dataset is 70 GB compressed, 249 GB
uncompressed.
@misc{cerebras2023slimpajama,
author = {Soboleva, Daria and Al-Khateeb, Faisal and Myers, Robert and Steeves, Jacob R and Hestness, Joel and Dey, Nolan}… See the full description on the dataset page: https://huggingface.co/datasets/venketh/SlimPajama-62B.ecommerce-customer-support-conversationsE-Commerce Customer Support Conversations
Dataset Summary:
This dataset contains customer support queries and responses from an e-commerce context.
It is designed for training and fine-tuning AI models for automated customer service, chatbots, and natural language processing (NLP) applications.
Use Cases:
Fine-tuning conversational AI models (e.g., GPT, BERT)
Training chatbots for e-commerce support
Improving customer service automation
Sentiment and intent analysis
Dataset Format:
The… See the full description on the dataset page: https://huggingface.co/datasets/Venkatrajan247/ecommerce-customer-support-conversations.NexaFlow-SFT-DatasetAnonyMED-BR
AnonyMED-BR
Dataset Description
AnonyMED-BR is a dataset created for research on medical text anonymization in Brazilian Portuguese.It combines real electronic health records (EHRs) — used strictly for research under ethics committee approval — with synthetically generated medical records to support the development of robust transformer-based models.
Language(s): Brazilian Portuguese (pt-BR)
Domain: Clinical / Medical
Task(s): Named Entity Recognition (NER)… See the full description on the dataset page: https://huggingface.co/datasets/Venturus/AnonyMED-BR.NexaFlow-DPO-DatasetNexaFlow-CPT-Datasetfol-data
FOL Reasoning Dataset
A preprocessed and vocabulary-augmented dataset derived from the ProofWriter (Kaggle) OWA splits, built for training a Natural Language → First-Order Logic translation model.
The source dataset contains natural-language premises and questions in English along with structured proof metadata. Our preprocessing adds two things that the original does not provide:
FOL translations — each natural-language statement is converted to First-Order Logic via a rule-based… See the full description on the dataset page: https://huggingface.co/datasets/Venkatdatta/fol-data.llm4securityai-training-data-use-by-vendor
Does this vendor train AI models on your data? Per-product, per-tier, quoted from the current policy
Canonical, always-current version: https://referencesource.org/ai-training-data-use-by-vendor/
Machine-readable: https://referencesource.org/ai-training-data-use-by-vendor/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-10
Stale after: 2026-10-09 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 14… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ai-training-data-use-by-vendor.ventset
Ventset: Raw & Real Conversations with an Empathic AI
Ventset is a dataset of human-AI dialogues featuring an AI designed to respond with empathy, humor, or tough love. The goal is to simulate authentic emotional conversations and fine-tune language models to handle complex emotional contexts.
⚠️ This dataset is still under development — contributions and feedback are welcome!
⚠️ Some messages may be misinterpreted. The creator is not a psychologist. Misuse or misinterpretation… See the full description on the dataset page: https://huggingface.co/datasets/archIBARBUgrr/ventset.Venki_data_set_Ananthapuram
Venki_data_set_Ananthapuram
A small, internally consistent retrieval corpus of Indian administrative
geography: 370 short factual statements covering every state and union
territory, 261 districts, 49 major cities, and 32 article summaries.
Built from live Wikidata and Wikipedia, and machine-checked for the specific
defect that makes a knowledge base useless — two rows that contradict each
other.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Venkatesulu/Venki_data_set_Ananthapuram.Lord-ganeshagentic-tooluse-computer-browserBPCC-en-hi-300-cleanedzyte-datagenminicpm5-tool-calling-xmlNER_augmented_indiandeepseekIRONWORKS-VENOM-preview
IRONWORKS VENOM
Supply Chain Security Training Dataset — Preview v0.1
by IronGate Digital
What this is
A synthetic instruction-tuning dataset focused on software supply chain security.
Built from real threat intelligence sources including security advisories, research
blogs, and vulnerability databases.
This is an early preview. More datasets are in progress.
Coverage
35,000+ labeled training pairs covering:
Dependency confusion and typosquatting… See the full description on the dataset page: https://huggingface.co/datasets/IronGateDigi/IRONWORKS-VENOM-preview.defendable-pain-vendor-psirt-pain-v0.1
Vendor PSIRT Pain Receipt
"the patch" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 6 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
6 pain receipts themed… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-vendor-psirt-pain-v0.1.minicpm5-computer-browser-coding-v2VenusPure-Telugu-Alpaca
Pure Telugu Alpaca Dataset
This dataset is a cleaned version of Telugu-MultiTask-Instruct-77K with Telugu keys.
It uses enhanced_prompt as instruction and enhanced_completion as output.
Processing Steps
Extracted enhanced_prompt → సూచన (instruction) and enhanced_completion → అవుట్పుట్ (output)
Filtered to keep only entries with no English letters
Removed duplicate entries
Normalized whitespace
Format
Each entry follows the Alpaca format with… See the full description on the dataset page: https://huggingface.co/datasets/VenkataRamanaKurumallajaddangi/Pure-Telugu-Alpaca.VENUS_DATA_betaTelugu-Dpowhatsapp_NERhr-consultor-vendasindian-augmented-NERjaspionjader__Kosmos-VENN-8B-details
Dataset Card for Evaluation run of jaspionjader/Kosmos-VENN-8B
Dataset automatically created during the evaluation run of model jaspionjader/Kosmos-VENN-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaspionjader__Kosmos-VENN-8B-details.
