datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mascarade-stm32-dataset
Ailiance — STM32 & ARM Cortex-M Q&A
🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/mascarade-stm32-dataset. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025).
Q&A bilingue (FR/EN) sur le firmware STM32 et ARM Cortex-M : drivers HAL/LL, CMSIS, FreeRTOS, peripherals (UART, SPI, I2C, DMA, ADC, timers), bare-metal register-level, et assembleur ARM Thumb-2. Couvre les familles STM32F0/F1/F4/F7/G0/G4/H7/L0/L4.… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-stm32-dataset.SemEval-STM
SemEval-STM — Dataset Description
This dataset is the official resource for the paper From Documents to Segments: A Contextual Reformulation for Topic Assignment.
SemEval-STM is a dataset built on SemEval data to demonstrate the superiority of Segment Topic Modeling (STM).
It contains annotations for two topic allocation paradigms:
Document Based Topic Allocation (DBTA): assigns a single topic to an entire document
Segment Based Topic Allocation (SBTA): assigns topics at the… See the full description on the dataset page: https://huggingface.co/datasets/LG-AI-Research/SemEval-STM.commentSemEval-STM
SemEval-STM — Dataset Description
SemEval-STM is a dataset built on SemEval data to demonstrate the superiority of Segment Topic Modeling (STM).
It contains annotations for two topic allocation paradigms:
Document Based Topic Allocation (DBTA): assigns a single topic to an entire document
Segment Based Topic Allocation (SBTA): assigns topics at the segment level within a document
Last updated: 2026-04-19
Rows marked bold in the summary table are the primary evaluation datasets… See the full description on the dataset page: https://huggingface.co/datasets/hoonst/SemEval-STM.mascarade-stm32-dataset
Mascarade — STM32 & ARM Cortex-M Q&A
Description
Q&A bilingue (FR/EN) sur le firmware STM32 et ARM Cortex-M : drivers HAL/LL, CMSIS, FreeRTOS, peripherals (UART, SPI, I2C, DMA, ADC, timers), bare-metal register-level, et assembleur ARM Thumb-2. Couvre les familles STM32F0/F1/F4/F7/G0/G4/H7/L0/L4.
Ce dataset fait partie de la famille Mascarade, un corpus thématique destiné au fine-tuning LoRA de modèles compacts (cible : Gemma-3n-E4B et équivalents) pour des assistants… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-stm32-dataset.reasoning_bank_small_subset_with_problem_stmtfin_stmt_reasoning_gemma
Overview
This dataset is constructed from real financial filings submitted to the U.S. Securities and Exchange Commission (SEC). It contains structured representations of accounting statements (such as income statements, balance sheets, and cash flow statements), along with reasoning components that include graphs showing mathematical relationships between financial line items.
Objective
The dataset is designed to train and evaluate large language models (LLMs) on… See the full description on the dataset page: https://huggingface.co/datasets/juliawawrykowicz/fin_stmt_reasoning_gemma.STM
Dataset Summary
A small, high-quality chat dataset to teach models how to answer like Sk. Tanzir Mehedi (QUT; software supply-chain security, HPC/LLM workflows, PyPI malware analysis).Primary reference: https://tanzirmehedi.netlify.app/
Format
Each row contains a messages list of {role, content, thinking} objects.thinking is optional and set to null for safety; models can be trained only on role + content.
Example usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/tanzirmehedi/STM.memory_items_wo_problem_stmtmemory_items_with_problem_stmtcombined_ds_wo_problem_stmt25fps-mcap-dataset-0717_2030-stm0v1rgcombined_ds_with_problem_stmt
