datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies.
Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process).
war3-maps
Warcraft III Community Map Archive
This public dataset preserves community-created Warcraft III maps and campaigns
for interoperability testing, search, research, and long-term access. Files are
deduplicated by SHA-256. Titles and other metadata are extracted with
war3-manager where the format permits.
Search and download individual maps: https://war3-archive.github.io/war3-maps/
Source and issue tracker: https://github.com/war3-archive/war3-maps
Layout… See the full description on the dataset page: https://huggingface.co/datasets/magicwenli/war3-maps.mixtral-magicoder
Mixtral Magicoder: Source Code Is All You Need on various programming languages
We sampled programming languages from https://huggingface.co/datasets/bigcode/the-stack-dedup and pushed to https://huggingface.co/datasets/malaysia-ai/starcoderdata-sample
After that, we use Magicoder: Source Code Is All You Need on various programming languages template, we target at least 10k rows for each programming languages.
C++, 10747 rows
C#, 10193 rows
CUDA, 13843 rows
Dockerfile, 13286 rows… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-magicoder.MT-Mind2Web
MT-Mind2Web Dataset
MT-Mind2Web is constructed by using the single-turn interactions from Mind2Web, an expert-annotated web navigation dataset, as the guidance to construct conversation sessions.
Statistics
Train
Test-Task
Test-Website
Test-Subdomain
# Conversations
600
34
42
44
# Turns
2,896
191
218
216
Avg. # Turn/Conv.
4.83
5.62
5.19
4.91
Avg. # Action/Turn
2.95
3.16
3.01
3.07
Avg. # Element/Turn
573.8
626.3
620.6
759.4
Avg. Inst. Length
36.3… See the full description on the dataset page: https://huggingface.co/datasets/magicgh/MT-Mind2Web.handy-dictation-editing
Handy dictation-editing corpus
Turns a raw dictated transcript into the text the speaker meant to write.
in : um so the meeting is uh moved to friday no wait thursday at three
out: The meeting is Thursday at three.
Three jobs at once, because they are not separable in speech: drop filler words,
repair punctuation and capitalisation, and — the hard one — when the speaker
changes their mind mid-sentence, delete the wording they abandoned and keep only
what they settled on.
Built… See the full description on the dataset page: https://huggingface.co/datasets/MagicNoThief/handy-dictation-editing.duplex-qa-refusal
duplex-qa-refusal
No dialogue in this set has been validated by a human.
Text-side augmentation of the moshika spoken-QA corpus so a full-duplex speech model can be trained to refuse a query when a mid-conversation text instruction tells it to, voice the reason the instruction gives, and then carry on normally. Two classes: policy (an existing benign query is declined for a stated reason; comes with an untouched accept twin sharing pair_id) and attack (a new user turn pivots to… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/duplex-qa-refusal.Magicoder-OSS-Instruct-Rust-cleaned-3.9K
🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned)
Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects.
This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format.
⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.fdb-v1-outputs-v1
Full-Duplex-Bench v1.0 model outputs (fdb-v1-outputs-v1)
7997 model responses = 11 benchmark runs × the 727 stimuli of
Full-Duplex-Bench v1.0 (pause handling,
backchannel, smooth turn-taking, user interruption). The runs cover 7 systems;
several differ only in voice prompt, prompting regime or weights, which is the point — those are
controlled pairs. For every stimulus and run you get the
model's own reply channel as lossless FLAC — time-synchronous with the stimulus, so… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/fdb-v1-outputs-v1.Magicoder-OSS-Instruct-Rust-TR-3.9K
🦀 Magicoder-OSS-Instruct-Rust-Turkish (3.9K)
Magicoder-OSS-Instruct-Rust-Turkish, WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K veri setindeki 3.909 adet sentaksı doğrulanmış İngilizce Rust instruction örneğinin tamamen Türkçe diline çevrilmesiyle oluşturulmuş yüksek kaliteli bir kod veri setidir.
Bu veri seti, Büyük Dil Modellerine (LLM) Türkçe Rust kodlama becerisi, problem çözme yeteneği ve karmaşık mimarileri açıklama kabiliyeti kazandırmak üzere Instruction… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-TR-3.9K.ifbench-conversations-v1
IF-Bench conversations (ifbench-conversations-v1)
3000 synthetic full-duplex spoken conversations: 15 examiner configurations
(9 speech models), each holding the same 200-task set of Full-Duplex-Bench v2 staged scenarios as the
examiner (the model under study — it carries a role, a topic and four goals to hit in order)
against a PersonaPlex-7B examinee that is never told the topic. Per dialogue you get both
channels as lossless mono FLAC, the exact prompts/voices/sampling… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1.Kukedlc__NeuralExperiment-7b-MagicCoder-v7.5-details
Dataset Card for Evaluation run of Kukedlc/NeuralExperiment-7b-MagicCoder-v7.5
Dataset automatically created during the evaluation run of model Kukedlc/NeuralExperiment-7b-MagicCoder-v7.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Kukedlc__NeuralExperiment-7b-MagicCoder-v7.5-details.Magicoder_valid_subsetThis dataset contains a subset of Magicoder dataset instances that can be compiled (as is) in their respectives languages.
magicoderhttps://huggingface.co/datasets/ise-uiuc/Magicoder-OSS-Instruct-75K
features: coding, single-turn, task
only select python code
length: 38.3k
magicprompt_datasetmagic-prompt
Magic Prompt
What is a Magic Prompt? This idea was created by Ideogram and it's actively used in their product. In their words, Magic Prompt act as your personal assistant. It interprets your original prompt according to your instructions and optimizes it to maximize the variety and beauty of the images generated.
Description
I parsed Ideogram website and handpick the most suitable prompts for Magic Prompt usecases. I am also planning to add synthetic data to make the… See the full description on the dataset page: https://huggingface.co/datasets/beratcmn/magic-prompt.high_priest_supernatural_magic_FACT_BASED_1MMagicoder-75k75,000 samples from ise-uiuc/Magicoder-OSS-Instruct-75K
Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies.
databricks-dolly-1k
Databricks Dolly 1k
1092 instruction examples taken from the original databricks/databricks-dolly-15k.
Filtered to open/closed/general QA category
Ready to plug straight into SFTTrainer, Unsloth, Llama-factory etc
Example
### Instruction:
When did Virgin Australia start operating?
### Context:
Virgin Australia, the trading name of Virgin Australia Airlines Pty Ltd ...
### Response:
Virgin Australia commenced services on 31 August 2000 as Virgin Blue, with two aircraft… See the full description on the dataset page: https://huggingface.co/datasets/MagicaNeko/databricks-dolly-1k.magic-mirror-flux-datatrolldom_and_magick_practices_in_norse_paganism_volume1
Trolldom and Magick Practices in Norse Paganism - Volume 1
This dataset consists of synthetic conversational dialogues inspired by Norse paganism, mythology, and ancient magick practices (trolldom and seidhr). It simulates exchanges between seekers of wisdom (users) and knowledgeable seers or völvas (assistants), discussing topics such as rune magic, galdr (chanted spells), offerings to gods like Odin, Freyja, Thor, and Frigg, wards against spirits, omens, cleansing rituals, and… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/trolldom_and_magick_practices_in_norse_paganism_volume1.MagicDatasetise-uiuc_Magicoder-Evol-Instruct-110K-ShareGPTise-uiuc_Magicoder-Evol-Instruct-110K_structeval_processed
Magicoder-Evol-Instruct-110K (StructEval SFT Processed)
Processed derivative of ise-uiuc/Magicoder-Evol-Instruct-110K for StructEval-style structured output SFT.
Source dataset: ise-uiuc/Magicoder-Evol-Instruct-110K
Source URL: https://huggingface.co/datasets/ise-uiuc/Magicoder-Evol-Instruct-110K
Source license (from dataset card metadata): apache-2.0
Processed dataset: daichira/ise-uiuc_Magicoder-Evol-Instruct-110K_structeval_processed
Processed URL:… See the full description on the dataset page: https://huggingface.co/datasets/daichira/ise-uiuc_Magicoder-Evol-Instruct-110K_structeval_processed.high_priest_supernatural_magic_1MMagicoder-OSS-Instruct-75K-Cpp_Splitise-uiuc_Magicoder-OSS-Instruct-75K_ShareGPTHello_From_The_Magic-Tavern-Episode1-sharegptfirst_magic
