gtm
Datasets
All datasets matching “gtm”gtmo-military-commissions
GTMO Military Commissions Dataset
This dataset is a structured, research-oriented snapshot of public military commissions docket records from mc.mil. It combines public docket metadata collected from mc.mil, a locally hash-checked archive of public source PDFs, direct PDF text extraction, page-level quality metadata, and transcript structure derived from the PDF text and word positions.
The most important thing to know is that the raw PDFs are the source files for this dataset's… See the full description on the dataset page: https://huggingface.co/datasets/strickvl/gtmo-military-commissions.DataClaw
DataClawBench
A data-analysis benchmark for OpenClaw-style end-to-end agents. Every task is grounded in real-world data and has a single objective gold answer.
简体中文
🌊 Data Analysis Tasks Are Changing in the OpenClaw Era
With the emergence of end-to-end agents like OpenClaw, data analysis is no longer equivalent to static QA — "read a passage, output one answer." Real-world data analysis tasks often require agents to locate evidence across heterogeneous files… See the full description on the dataset page: https://huggingface.co/datasets/GTML-LAB/DataClaw.Dataset_UCFAGT-Music-Genre
GT-Music-Genre
This is an audio classification dataset for Music Analysis.
Classes = 10 , Split = Train-Test
Structure
audios folder contains audio files.
train.csv for training split and test.csv for the testing split.
Download
import os
import huggingface_hub
audio_datasets_path = "DATASET_PATH/Audio-Datasets"
if not os.path.exists(audio_datasets_path): print(f"Given {audio_datasets_path=} does not exist. Specify a valid path ending with… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/GT-Music-Genre.FalAI
Dataset Card for FalAI Dataset
Dataset Summary
The FalAI dataset consists of a total of 265,603 audio files (wav) with associated annotations in the form of metadata.
The FalAI dataset is designed for SLU (Spoken Language Understanding) and is the largest publicly released dataset, in any language, for the task of SLU.
Metadata is available for each recording, including the reference phrase, validation label, user id, demographics such as age, accent, gender, locality and… See the full description on the dataset page: https://huggingface.co/datasets/GTM-UVigo/FalAI.piper_place_block_random500_no_occlusion_gtmask_3cam
