datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laions_got_talent_orpheus_snacSome Laion's Got Talent (https://huggingface.co/datasets/laion/laions_got_talent) voice snippets converted to snac tokens in the format of the Orpheus-TTS https://github.com/canopyai/Orpheus-TTS
We converted the data into instructions like format.
The snac data is in 7 token frame groups. See the Orpheus blog for more details: https://canopylabs.ai/model-releases
We did not create the original dataset and are only providing snac token with minimal text instructions for ease of use.
You must be… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_orpheus_snac.Got_Agentic_AI_5k
Got_Agentic_AI_5k
A 5,000-example dataset to train LLMs into production-grade agentic assistants (“Angelic Agents”): high-agency, tool-aware, test-driven, and safety-first.
This dataset focuses on the kinds of tasks real engineering teams and major AI developers care about:
Diff-first coding patches and tests
Planner–executor agent architectures
Evals, monitoring, and rollback discipline
Data engineering transforms with quality checks
Incident postmortems and operational… See the full description on the dataset page: https://huggingface.co/datasets/11-47/Got_Agentic_AI_5k.VC-LLM-DatasetDue to the presence of harmful and toxic unsafe content in the fine-tuning data, a portion of the data is displayed. For the full data, please contact ignitesun@163.com.
instructionsBible-responses-dataset-gotquestions
Theology Question-Answer Dataset
Description
This dataset contains structured, human-generated content focused on theology, primarily sourced from the website GotQuestions. Each entry is formatted as a question (prompt) and a corresponding answer (response). The dataset is provided in JSON format and is intended for fine-tuning AI models, though it can be used for other purposes as well.
The structure of the dataset is as follows:
{
"prompt": "What does it mean to… See the full description on the dataset page: https://huggingface.co/datasets/vericudebuget/Bible-responses-dataset-gotquestions.got_qa_pairsgenerated some part by parsing html scripts and remaining using gemini api
goth-girl-friendsORCA_GOT_STYLE
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mindchain/ORCA_GOT_STYLE.gothic-slutsGot_Science_28keco-gotest-trainingGoToCompany__llama3-8b-cpt-sahabatai-v1-instruct-details
Dataset Card for Evaluation run of GoToCompany/llama3-8b-cpt-sahabatai-v1-instruct
Dataset automatically created during the evaluation run of model GoToCompany/llama3-8b-cpt-sahabatai-v1-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/GoToCompany__llama3-8b-cpt-sahabatai-v1-instruct-details.eco-gotestseco-gotest-TAGgot_fivehungot-filledConversationSentimentConversationalSentimentDatasetsn96-key85go-test-dataset
