datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imslp-midi-cc0-1.0
IMSLP MIDI Dataset (CC0-1.0)
This dataset contains MIDI files and metadata crawled from IMSLP (International Music Score Library Project) on July 21-22, 2024.
Data Fields
midi_source: URL to the original MIDI file on IMSLP (incl. original uploader).
metadata_source: URL to the original metadata on IMSLP.
file_name, title, composer, year, era, style, key, license: Metadata fields.
midi: Raw MIDI bytes.
midi_mido: JSON-serialized mido object.
How to Retrieve… See the full description on the dataset page: https://huggingface.co/datasets/TiMauzi/imslp-midi-cc0-1.0.imsdb-genre-movie-scripts
Dataset Card for "imsdb-genre-movie-scripts"
More Information needed
basic_3D_shapesimsa
IMSA WeatherTech Championship Racing Dataset
This dataset contains comprehensive lap-by-lap data from the IMSA WeatherTech SportsCar Championship, including detailed timing information, driver data, and weather conditions for each lap.
Dataset Details
Dataset Description
The IMSA Racing Dataset provides detailed lap-by-lap telemetry and contextual data from the IMSA WeatherTech SportsCar Championship races from 2021-2025. The primary table contains… See the full description on the dataset page: https://huggingface.co/datasets/tobil/imsa.imslp-midi-by-sa
IMSLP MIDI Dataset (CC-BY-SA-4.0)
This dataset contains MIDI files and metadata crawled from IMSLP (International Music Score Library Project) on July 21-22, 2024.
Data Fields
midi_source: URL to the original MIDI file on IMSLP (incl. original uploader).
metadata_source: URL to the original metadata on IMSLP.
file_name, title, composer, year, era, style, key, license: Metadata fields.
midi: Raw MIDI bytes.
midi_mido: JSON-serialized mido object.
How to… See the full description on the dataset page: https://huggingface.co/datasets/TiMauzi/imslp-midi-by-sa.IMS-perception
Roles
Roles: perception view of IMS — annot is the source label (ball / inner_race / normal / outer_race), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the repo ships a bearing's vibration in four image encodings as four equal-sized configs — reshaped (consecutive samples arranged as the rows of a grayscale square), scalogram (a continuous-wavelet time-scale view), spectrogram (a short-time Fourier transform) and… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/IMS-perception.imslp-midi-by-nc-sa
IMSLP MIDI Dataset (CC-BY-NC-SA-4.0)
This dataset contains MIDI files and metadata crawled from IMSLP (International Music Score Library Project) on July 21-22, 2024.
Data Fields
midi_source: URL to the original MIDI file on IMSLP (incl. original uploader).
metadata_source: URL to the original metadata on IMSLP.
file_name, title, composer, year, era, style, key, license: Metadata fields.
midi: Raw MIDI bytes.
midi_mido: JSON-serialized mido object.
How to… See the full description on the dataset page: https://huggingface.co/datasets/TiMauzi/imslp-midi-by-nc-sa.IMS
Roles
Roles: canon repo — annot is the source label, kept machine-parseable as the gold for verification and reward parsing; there is no filled reasoning column and this repo is not itself a training view. Derived repos each state their own regime on their own card.
IMS / NASA-Bearing — fault classification from the envelope spectrum (reasoning track)
Second signal dataset in the AI4Manufacturing FORGE corpus (Category C, task T-C1), from three run-to-failure… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/IMS.imslp-crawling
Top Composers
Composer folders size
size
composer
1017G
/
35G
/(1959)Daniel Léo Simpson
31G
/(1756)Wolfgang Amadeus Mozart
24G
/(1685)Johann Sebastian Bach
21G
/(1770)Ludwig van Beethoven
17G
/(1699)Johann Adolph Hasse
16G
/(1678)Antonio Vivaldi
16G
/(1670)Antonio Caldara
14G
/(1685)George Frideric Handel
14G
/(1683)Christoph Graupner
13G
/(1975)Carlotta Ferrari
13G
/(1732)Joseph Haydn
11G
/(1792)Gioacchino Rossini
10G
/(1948)Michel… See the full description on the dataset page: https://huggingface.co/datasets/k-l-lambda/imslp-crawling.kabyle-corpus-ubouira
Kabyle Paragraph Corpus (Bouira University-DSpace)
336,982 Kabyle sentences extracted from open-access PDFs published onBouira University – the institutional repository ofBouira University (Algeria).
Documents originate from theFaculté des Lettres et des Langues / Département de Langue et Culture Amazighes.
Fields
id : unique identifier
text : paragraph text (UTF-8)
folder : origin sub-corpus
line_no : 1-based line number in original file
Splits… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/kabyle-corpus-ubouira.imsdb-drama-screenplaylibretranslate-suggestions
Kabyle Suggestions Dataset
This dataset contains English-to-Kabyle translation suggestions submmitted by Imsidag community using LibreTranslate, designed to support the development and evaluation of machine translation tools for the Kabyle language.
ShareGPT_Vicuna_unfiltered_has_imsorryHas instances of "I'm sorry, but" from https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered
I converted single json file to jsonlines
imslp-pdf-index
IMSLP PDF Index
This dataset is the canonical PDF-level index for the ReScore IMSLP PDF
collection. It contains one row per unique IMSLP PDF and points to the PDF
payload stored in cminst/imslp-raw-pdf-collection.
The PDF payload repository is append-only and may contain duplicate rows from
retry launches. This index is deduplicated by imslp_id; duplicate content was
validated to have identical SHA256, byte size, and page count before publishing.
Summary
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cminst/imslp-pdf-index.imsdb-sci-fi-movie-scripts
Dataset Card for "imsdb-sci-fi-movie-scripts"
More Information needed
kabyle-corpus-hca
Kabyle Paragraph Corpus (HCA Algeria)
188,620 Kabyle sentences extracted from open-access PDFs published on HCA Algeria – the institutional repository of Haut Commissariat à l'Amazighité (Algeria).
Documents originate from Haut Commissariat à l'Amazighité.
Fields
id : unique identifier
text : paragraph text (UTF-8)
folder : origin sub-corpus
line_no : 1-based line number in original file
Splits
split
# records
train
169,758
validation
9,431… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/kabyle-corpus-hca.IMSimsdb-drama-movie-scripts
Dataset Card for "imsdb-drama-movie-scripts"
More Information needed
kabyle-corpus-ummto
Kabyle Paragraph Corpus (UMMTO-DSpace)
690,917 Kabyle sentences extracted from open-access PDFs published onDSpace UMMTO – the institutional repository ofUniversité Mouloud Mammeri de Tizi-Ouzou (Algeria).
Documents originate from theFaculté des Lettres et des Langues / Département de Langue et Culture Amazighes.
Fields
id : unique identifier
text : paragraph text (UTF-8)
folder : origin sub-corpus
line_no : 1-based line number in original file
Splits… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/kabyle-corpus-ummto.IMSDbnllb_en_kab
NLLB English - Kabyle Dataset
This dataset contains parallel sentences in English and Kabyle, cleaned and filtered using the GlotLid model. The dataset is derived from the OPUS-NLLB corpus and has been processed to ensure high-quality sentence pairs.
Dataset Structure
nllb_en_kab.parquet: A Parquet file containing the cleaned English-Kabyle sentence pairs.
Dataset Statistics
Total Sentence Pairs: 2,484,297
English Sentences: 2,484,297
Kabyle Sentences: 2,484… See the full description on the dataset page: https://huggingface.co/datasets/Imsidag-community/nllb_en_kab.imslp-midi-v1IMS-annotated
Roles
Roles: reasoning view of IMS — annot is the source label (ball / inner_race / normal / outer_race), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the image is an envelope spectrum with the theoretical fault frequencies marked. The reasoning column is filled on all 542 records and is the SFT imitation target for this repo; the query enumerates the closed set of labels the answer must come from, and annot remains… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/IMS-annotated.ShareGPT_V3_unfiltered_cleaned_split_no_imsorry
Dataset Card for "ShareGPT_V3_unfiltered_cleaned_split_no_imsorry"
More Information needed
musicbrainz-all-songsimsi-dataset
Format-Preserved Pseudonymised IMSI-Catcher Observations
Dataset Description
This dataset contains pseudonymised mobile-network observation records derived from IMSI-catcher-style telemetry. It is intended for research on mobile-network measurements, cellular security, anomaly detection, roaming patterns, and privacy-preserving analysis of telecom metadata.
The dataset has been transformed using deterministic, keyed, format-preserving pseudonymisation… See the full description on the dataset page: https://huggingface.co/datasets/adulau/imsi-dataset.imsdb-comedy-movie-scripts
Dataset Card for "imsdb-comedy-movie-scripts"
More Information needed
pos_dataset-UD_Turkish-IMST-v2.13Claude3-Opus-Instruct-ShareGPT-14k-IMSORRY_REMOVEDInstruct dataset made with Claude 3 Opus in ShareGPT format. Instances that start with "I'm sorry" have been removed.
ImSitu_reduced
