datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maestro_synth
Dataset Card for "maestro_synth"
More Information needed
MAESTRO-E
Viewer note: default uses viewer_preview/ for responsive audio playback.
Full training/evaluation files remain available in the original folder structure.
MAESTRO-E
MAESTRO-E dataset for music practice error detection with paired performance/reference inputs and note-level error labels.
Paired Inputs for Error Detection
The model takes paired inputs:
mistake: performance audio/MIDI containing musical errors
score: paired reference score audio/MIDI (target/correct… See the full description on the dataset page: https://huggingface.co/datasets/ben2002chou/MAESTRO-E.maestro
Dataset Card for "maestro"
More Information needed
maestro-romanticmaestro-mas-benchmark
MAESTRO MAS Benchmark Dataset
maestro-mas-benchmark is a dataset derived from MAESTRO, a framework-agnostic evaluation suite for LLM-based
multi-agent systems (MAS). It provides a systems-level view of MAS behavior and is designed to benchmark,
observe, and analyze MAS performance and behavior across diverse scenarios.
For more details about MAESTRO, visit the GitHub repository.
Dataset details
The dataset currently includes data for 12 different MAS systems… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/maestro-mas-benchmark.maestro-classicalmaestro-modernMaestro20hmaestro-preprocessed
Dataset Card for "maestro-preprocessed"
More Information needed
maestro_extract_unit
Dataset Card for "maestro_extract_unit"
More Information needed
suayptalha__Maestro-10B-details
Dataset Card for Evaluation run of suayptalha/Maestro-10B
Dataset automatically created during the evaluation run of model suayptalha/Maestro-10B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/suayptalha__Maestro-10B-details.maestro-sustain-v2
Dataset Card for "maestro-sustain-v2"
More Information needed
MAESTRO-test
Mirror of the test set of MAESTRO V.3.0.0 dataset with midi and audio columns for midi file and audio in FLAC
Citation
If you use this dataset, please cite the original paper:
@inproceedings{
hawthorne2018enabling,
title={Enabling Factorized Piano Music Modeling and Generation with the {MAESTRO} Dataset},
author={Curtis Hawthorne and Andriy Stasyuk and Adam Roberts and Ian Simon and Cheng-Zhi Anna Huang and Sander Dieleman and Erich Elsen and Jesse Engel and… See the full description on the dataset page: https://huggingface.co/datasets/B-K/MAESTRO-test.arcee-ai__Arcee-Maestro-7B-Preview-details
Dataset Card for Evaluation run of arcee-ai/Arcee-Maestro-7B-Preview
Dataset automatically created during the evaluation run of model arcee-ai/Arcee-Maestro-7B-Preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/arcee-ai__Arcee-Maestro-7B-Preview-details.hermes-function-calling-v1
Hermes Function-Calling V1
This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models.
This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational… See the full description on the dataset page: https://huggingface.co/datasets/Maestro502/hermes-function-calling-v1.maestro-spectrogramsMARBLEBeatTracking_ASAP_MAESTROmaestro-rollingsplit
Dataset Card for "maestro-rollingsplit"
More Information needed
maestro-v1
Dataset Card for "maestro-v1"
More Information needed
maestro_abc_notation_25s
Maestro ABC Notation 25s Dataset
Dataset Summary
This is based on V3.0.0 of the Maestro dataset.
The Maestro ABC Notation 25s Dataset is a curated collection of question-and-answer pairs derived from short audio clips within the MAESTRO dataset. Each entry in the dataset includes:
An id corresponding to the original audio file.
A start_time marking where the 25-second audio clip begins within the full track.
A question designed to prompt music transcription in ABC… See the full description on the dataset page: https://huggingface.co/datasets/jonflynn/maestro_abc_notation_25s.maestro-base-v2
Dataset Card for "maestro-base-v2"
More Information needed
AMoreNaturalMIDItoPianoGeneration_Maestroarxiv-author-affiliations-arcee_ai_maestro-correct-rolloutsmaestro-quantized
Dataset Card for "maestro-quantized"
More Information needed
CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
OpenCodeInterpreter
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information… See the full description on the dataset page: https://huggingface.co/datasets/Maestro502/CodeFeedback-Filtered-Instruction.maestro-tokenizedmaestro-v1-sustain
Dataset Card for "maestro-v1-sustain"
More Information needed
MetaMathQAView the project page:
https://meta-math.github.io/
see our paper at https://arxiv.org/abs/2309.12284
Note
All MetaMathQA data are augmented from the training sets of GSM8K and MATH.
None of the augmented data is from the testing set.
You can check the original_question in meta-math/MetaMathQA, each item is from the GSM8K or MATH train set.
Model Details
MetaMath-Mistral-7B is fully fine-tuned on the MetaMathQA datasets and based on the powerful Mistral-7B model.… See the full description on the dataset page: https://huggingface.co/datasets/Maestro502/MetaMathQA.NuminaMath-CoT
Dataset Card for NuminaMath CoT
Dataset Summary
Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/Maestro502/NuminaMath-CoT.maestro-v1-sustain-masked
Dataset Card for "maestro-v1-sustain-masked"
More Information needed
