datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
titanicThe legendary Titanic dataset from this Kaggle competition
Titanium2-DeepSeek-R1Click here to support our open-source dataset and model releases!
Titanium2-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills!
This dataset contains:
32.4k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium2-DeepSeek-R1.TitanVul
Dataset Card for TitanVul
TitanVul is a large-scale function-level vulnerability dataset constructed for training machine learning models for vulnerability detection. It consists of paired vulnerability-fix function samples aggregated from multiple public sources and validated using a multi-agent LLM framework.
Dataset Details
TitanVul is designed to provide high-quality training data that generalizes across vulnerability types and codebases. The dataset is built by… See the full description on the dataset page: https://huggingface.co/datasets/yikun-li/TitanVul.TitaniumTitanium is a dataset containing DevOps-instruct data.
The 2024-10-02 version contains:
26.6k rows of synthetic DevOps-instruct data, using synthetically generated prompts and responses generated using Llama 3.1 405b Instruct. Primary areas of expertise are AWS, Azure, GCP, Terraform, Dockerfiles, pipelines, and shell scripts.
This dataset contains synthetically generated data and has not been subject to manual review.
titanic-survival
Titanic Survival
from https://web.stanford.edu/class/archive/cs/cs109/cs109.1166/problem12.html
Titanium4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Titanium 4 is an agentic coding dataset focused on DevOps and architecture, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics.
Areas of focus include IaC, cloud architecture, incident response, configuration and cost optimization, security… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro.Titanium2.1-DeepSeek-R1Click here to support our open-source dataset and model releases!
Titanium2.1-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills!
This dataset contains:
31.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium2.1-DeepSeek-R1.Titanium2-DeepSeek-R1Click here to support our open-source dataset and model releases!
Titanium2-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills!
This dataset contains:
32.4k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Titanium2-DeepSeek-R1.titan-hohmann-transfer-orbit
🪐 Titan-Hohmann-Transfer-Orbit Dataset
🛰️ A 700-row synthetic dataset simulating an interplanetary mission to Titan using Hall Effect electric propulsion and gravity assists, from Earth departure through Titan orbital insertion.
⚠️ Disclaimer: All values are synthetically generated from simplified orbital mechanics models. This is not flight data and is not suitable for mission planning.
📋 At a Glance
🔢 Rows
700
📊 Columns
18
🧬… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/titan-hohmann-transfer-orbit.Titanium3-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases!
Titanium3-DeepSeek-V3.1-Terminus is a dataset focused on architecture and DevOps, testing the limits of DeepSeek V3.1 Terminus's architect and coding skills!
This dataset contains:
27.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek V3.1 Terminus in reasoning mode:
20k selected technical expertise prompts from sequelbox/Titanium2.1-DeepSeek-R1 focused on… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Titanium3-DeepSeek-V3.1-Terminus.titanic
Titanic Dataset
Классический учебный датасет Titanic (Kaggle). Бинарная классификация: выжил / не выжил.
Файлы
train.csv — 891 строка, с колонкой Survived
test.csv — 418 строк, без ответов
Колонки
Pclass, Sex, Age, SibSp, Parch, Fare, Embarked — признаки; Survived — целевая.
titanicarxiv_qa
Arxiv Paper Generative Question Answering
Dataset Summary
This dataset is made using ChatGPT (text-davinci-003) to generate Question/Answer pairs from Arxiv papers from this dataset
Data Fields
TextID: references the datarow (paper) in the arxiv summarizer dataset
Question: question based on the text
Response: answer
Text: Full text with the paper as 'context:' and and the question appended as 'question:'. Used for generative question answering usign language… See the full description on the dataset page: https://huggingface.co/datasets/TitanMLData/arxiv_qa.TitanVul
Dataset Card for TitanVul
TitanVul is a large-scale function-level vulnerability dataset constructed for training machine learning models for vulnerability detection. It consists of paired vulnerability-fix function samples aggregated from multiple public sources and validated using a multi-agent LLM framework.
Dataset Details
TitanVul is designed to provide high-quality training data that generalizes across vulnerability types and codebases. The dataset is built by… See the full description on the dataset page: https://huggingface.co/datasets/comet1910/TitanVul.Titanium4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule!
This is an early sneak preview of Titanium 4, containing the first 4.9k rows!
Titanium 4 is an upcoming agentic coding dataset focused on DevOps and architecture, generated by DeepSeek-V4-Pro:
Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics.
Areas of focus include IaC, cloud architecture… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro-PREVIEW.spaceship-titanic-trainTitanium3-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases!
Titanium3-DeepSeek-V3.1-Terminus is a dataset focused on architecture and DevOps, testing the limits of DeepSeek V3.1 Terminus's architect and coding skills!
This dataset contains:
27.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek V3.1 Terminus in reasoning mode:
20k selected technical expertise prompts from sequelbox/Titanium2.1-DeepSeek-R1 focused on… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium3-DeepSeek-V3.1-Terminus.Titanic-Machine-Learning-from-Disaster-0.77751Spaceship-Titanic-result-0.80757Titanic
DEATH RECORD'S OF RMS TITANIC INCIDENT
The Titanic was a British luxury ocean liner that sank on April 15, 1912, after striking an iceberg during its maiden voyage from Southampton, England, to New York City.
Dataset Description
Curated by: [XythicK]
Funded by [optional]: [XythicK/Alchemist]
Shared by [optional]: [People]
Language(s) (NLP): [English]
License: [All rights reserved to Official Titanic Website]
Dataset Sources [optional]
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/XythicK/Titanic.titanictitanic-datasettitantic_altered.csv
Titantic Dataset
This is an altered titanic dataset for training purposes.
The Titanic dataset is a well-known and widely used dataset in the field of data science and machine learning. The dataset provides information about the passengers aboard the RMS Titanic, which famously sank on its maiden voyage on April 15, 1912.
The dataset contains a combination of demographic and passenger-related information, making it suitable for various analyses and predictions. It has also been… See the full description on the dataset page: https://huggingface.co/datasets/sigtica/titantic_altered.csv.titanic
Dataset Card for Titanic Survival Prediction
Dataset Details
Dataset Description
This dataset is a copy of the original Kaggle Titanic dataset made to explore the Hugging Face Datasets feature.
The Titanic Survival Prediction dataset is widely used in machine learning and statistics. It originates from the Titanic: Machine Learning from Disaster competition on Kaggle. The dataset consists of passenger details from the RMS Titanic disaster, including… See the full description on the dataset page: https://huggingface.co/datasets/hdhnhdhnhdhn/titanic.titanic-databooth
Dataset Description
Purpose: Demonstrate how data quality impacts analytics through the iconic Titanic dataset, featuring:
Original datasets (with known age/class errors)
Corrected versions (with reconciled passenger details)
Data quality annotations (error flags, reconciliation sources)
Homepage: Data Governance: Titanic Dataset and the Perils of Bad Data
Repository: mjboothaus-titanic-databoothTasks: data-cleaning, error-detection, survival-prediction
Dataset Versions… See the full description on the dataset page: https://huggingface.co/datasets/mjboothaus/titanic-databooth.titan-signal-b2b-ai-leads
Titan Signal - B2B AI Company Lead Intelligence
Verified contact records for decision-makers at AI, ML, and enterprise software companies.
Built by Titan Signal's 230+ autonomous harvesting agents. Continuously refreshed, MX-verified, 90-day auto-purge.
Fields
Company name, contact title, email
Industry vertical, headcount band, revenue band
Tech stack tags, engagement score, verification date
Full Dataset
This is a 50-record sample. Full datasets (10K-1M+… See the full description on the dataset page: https://huggingface.co/datasets/SophieTitan/titan-signal-b2b-ai-leads.titanic-datasettitanic-survival-predictionspaceship-titanic-testTitanVul
Dataset Card for TitanVul
TitanVul is a large-scale function-level vulnerability dataset constructed for training machine learning models for vulnerability detection. It consists of paired vulnerability-fix function samples aggregated from multiple public sources and validated using a multi-agent LLM framework.
Dataset Details
TitanVul is designed to provide high-quality training data that generalizes across vulnerability types and codebases. The dataset is built by… See the full description on the dataset page: https://huggingface.co/datasets/Divya1214/TitanVul.
