datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LEAD-DVS
LEAD-DVS
LEAD expert-driving data collected in CARLA 0.9.15 with synchronized RGB,
depth, semantic/instance segmentation, LiDAR, radar, HD map, metadata, 3D
bounding boxes, and a forward-facing DVS event camera.
Code release
Dataset construction, preprocessing, training, and evaluation code will be
published in SamanthaZhang-stu/ReflexWorldModel.
This GitHub repository is the designated code-release location for the project.
Contents
1,579… See the full description on the dataset page: https://huggingface.co/datasets/SamanthaZhang/LEAD-DVS.samantha-1.1-uncensoredThis dataset is based on ehartford/samantha-data that was used to create ehartford/samantha-1.1-llama-7b and other samantha models. It has been unfiltered and uncensored.
Her-Samantha-Style
Ultra-High Quality Samantha Dataset
A meticulously curated conversational AI dataset designed to capture the essence of Samantha from the movie "Her" - characterized by emotional intelligence, philosophical depth, and authentic conversational patterns.
Dataset Summary
This dataset contains 20,000 ultra-high quality conversational responses that have been systematically filtered and scored based on Samantha's distinctive characteristics from the 2013 film "Her". Each… See the full description on the dataset page: https://huggingface.co/datasets/WasamiKirua/Her-Samantha-Style.samantha-data-fr
Samantha Data — French Translation
Traduction française de cognitivecomputations/samantha-data, le dataset original de conversations pour une IA compagne inspirée de Samantha dans le film Her.
Contenu
6534 conversations, 69374 tours de dialogue — couverture intégrale du dataset original.
Format identique à l'original (ShareGPT) : [{"id": "...", "conversations": [{"from": "human"|"gpt", "value": "..."}]}].
Français naturel et idiomatique, pensé pour être lu à voix… See the full description on the dataset page: https://huggingface.co/datasets/SaucisseduNord/samantha-data-fr.StableBeluga-7B-Qlora-Samantha-V3-Converted-Datasetsamantha-omni-3.0-datasetSamantha-OpusOriginal source: https://huggingface.co/datasets/macadeliccc/opus_samantha
Fixed some pairing issues.
macadeliccc__Samantha-Qwen-2-7B-details
Dataset Card for Evaluation run of macadeliccc/Samantha-Qwen-2-7B
Dataset automatically created during the evaluation run of model macadeliccc/Samantha-Qwen-2-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/macadeliccc__Samantha-Qwen-2-7B-details.samantha-dialogues-ru
Samantha Dialogues — Русская локализация
...
(и дальше по шаблону)
Samantha Dialogues — Русская локализация
Этот датасет содержит переведённую на русский язык версию оригинального англоязычного набора диалогов с виртуальной ассистенткой Самантой. Оригинальные английские данные взяты из открытого датасета.
🧠 Назначение
Дообучение чат-ботов и LLM моделей на русском языке
Создание более человечных виртуальных ассистентов на русском
Ролевые диалоги и симуляция… See the full description on the dataset page: https://huggingface.co/datasets/Hoaxer2000/samantha-dialogues-ru.Samantha-humanlikedatasetSamantha-PL-AG-axolotlsamantha-1.5Samantha-EN-CN-Dataset-V1Samantha-Her-Style-backupThis is NOT my dataset, i just saved it again to have a backup if the creator erases it from hf
ORIGINAL:
WasamiKirua/Her-Samantha-Style
Samantha_Her-GPT3.5_turboSamantha-data-single-line-Mixed-V1import json
# Load the provided data
with open("path_to_your_original_file.jsonl", "r", encoding="utf-8") as file:
mixed_data = [json.loads(line) for line in file.readlines()]
# Convert the mixed data by extracting all possible Q&A pairs from each conversation
reformatted_data_complete = []
for conversation in mixed_data:
text = conversation['text']
# Split the text into segments based on the prefixes
segments = [segment for segment in text.split("###") if… See the full description on the dataset page: https://huggingface.co/datasets/RoversX/Samantha-data-single-line-Mixed-V1.Samantha-PL-AGSamantha-newdataset-morelargeuukuguy__speechless-mistral-dolphin-orca-platypus-samantha-7b-details
Dataset Card for Evaluation run of uukuguy/speechless-mistral-dolphin-orca-platypus-samantha-7b
Dataset automatically created during the evaluation run of model uukuguy/speechless-mistral-dolphin-orca-platypus-samantha-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/uukuguy__speechless-mistral-dolphin-orca-platypus-samantha-7b-details.SamanthaDataset-rolesformatsamantha-nvfp4-calibrationSamantha-new-dataset
