gaa
Datasets
All datasets matching “gaa”RMOT26
RMOT26
RMOT26 is a large-scale benchmark for Query-Driven Multi-Object Tracking, introduced in the paper QTrack: Query-Driven Reasoning for Multi-modal MOT.
Project Page: https://gaash-lab.github.io/QTrack/
Repository: https://github.com/gaash-lab/QTrack
Paper: https://arxiv.org/abs/2603.13759
Description
Multi-object tracking (MOT) has traditionally focused on estimating trajectories of all objects in a video. RMOT26 introduces a query-driven tracking paradigm that… See the full description on the dataset page: https://huggingface.co/datasets/GAASH-Lab/RMOT26.gaap-sec-compliance-dataset
GAAP & SEC Compliance Dataset
A comprehensive dataset for financial AI applications
Dataset Overview
This dataset contains 470,151 documents covering US GAAP (Generally Accepted Accounting Principles) standards and SEC (Securities and Exchange Commission) filing requirements. It's designed for training and evaluating AI systems for financial compliance, accounting Q&A, and regulatory analysis.
Key Statistics
Total Documents: 470,151
Average Length: 363… See the full description on the dataset page: https://huggingface.co/datasets/aanshshah/gaap-sec-compliance-dataset.Bolbosh
Kashmiri TTS Dataset
Project Page | GitHub | Paper
Overview
This dataset is the Text-to-Speech (TTS) corpus for the Kashmiri language, as presented in the paper "Bolbosh: Script-Aware Flow Matching for Kashmiri Text-to-Speech".
The dataset is a derived and curated combination of Kashmiri speech data from the IndicVoices-R corpus and the RASA speech dataset. It was used to develop Bolbosh, an open-source neural TTS system designed to handle the specific orthographic and… See the full description on the dataset page: https://huggingface.co/datasets/GAASH-Lab/Bolbosh.gaa
General Appropriations Act (GAA) Dataset
Dataset Summary
This dataset contains detailed Philippine government budget appropriations data from the General Appropriations Act (GAA). It includes 3.7+ million records of government expenditures across departments, agencies, and programs, using the Unified Accounts Code Structure (UACS) classification system. The dataset covers fiscal years 2020-2025 (six years) and provides comprehensive breakdowns of government spending by… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/gaa.ghana-chat-corpus-gaa
Ghanaian Corpus — Ga (gaa)
English–Ga parallel corpus derived from
ghananlpcommunity/ghana-chat.
Each row contains an English conversation snippet and its generated question, alongside
their Ga translations via Google Translate (Amharic pivot).
Schema
Column
Description
source_type
Source category (e.g., news, parliament)
source
Original source name
text
English conversation (all assistant responses joined)
generated_question
First user message in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ghana-chat-corpus-gaa.GroundedSurg
GroundedSurg: A Multi-Procedure Benchmark for Language-Conditioned Surgical Tool Segmentation
📌 Dataset Summary
GroundedSurg is the first language-conditioned, instance-level surgical tool segmentation benchmark.
Unlike conventional category-level surgical segmentation datasets, GroundedSurg requires models to resolve natural-language references and segment a specific instrument instance in multi-instrument surgical scenes.
Each benchmark instance consists of:
A… See the full description on the dataset page: https://huggingface.co/datasets/GAASH-Lab/GroundedSurg.
