datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.bishkek-transport
Bishkek Public Transport
Open, continuously-growing data on the public-transport network of Bishkek,
the capital of the Kyrgyz Republic. The data originates from the Bishkek
mayoralty's public-transport monitoring system (the same feed behind the city's
official live transit map and its "My City" mobile service). Only publicly
visible transit information is included: stop locations and the live positions of
buses, trolleybuses/electric buses, and marshrutkas (shared minibuses).… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/bishkek-transport.ASVspoof_2019_LAZINC_22PubChemBiST
BiST
💻 Github Repo
English | 简体中文
Introduction
BiST is a large-scale bilingual translation dataset, with "BiST" standing for Bilingual Synthetic Translation dataset. Currently, the dataset contains approximately 60M entries and will continue to expand in the future.
BiST consists of two subsets, namely en-zh and zh-en, where the former represents the source language, collected from public data as real-world content; the latter represents the target language… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/BiST.USPTO_50KChEMBLMOSESiqra_eval_all_refbisac_expanded_finalIqra_train_with_arabic_BWsm-cv-ar-enbis_central_bank_speeches
Dataset Card
This dataset consists of central bankers speeches from 1997 to 2025 scrapped automatically from the Bank Of International Settlements website.
Each speech is associated to a central bank, a date and a description (a metadata provided by the website).
The dataset covers a wide range of topics in economics from monetary policies to world outlooks, financial stability, unemployment, fiscal policies...
Credits
Full credits to the Bank of International… See the full description on the dataset page: https://huggingface.co/datasets/samchain/bis_central_bank_speeches.mbank-atm-bishkek
MBank ATM Dataset (Bishkek, 2024–2025)
Per-second dataset of a network of 120 ATMs over 731 days
(2024-01-01 … 2025-12-31). Card-processing style records (ISO-8583-inspired),
grounded in the real-world context of Kyrgyzstan: ERA5 weather, the USD/KGS rate
of the National Bank of the KR, and ATM coordinates from OpenStreetMap.
Intended as a training ground: demand forecasting, ATM clustering, RFM analysis,
anomaly detection. Every anomaly has a checkable "answer key" in labels/.… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/mbank-atm-bishkek.bi_so101_fold_10inch_4camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 603,
"total_frames": 541612,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:603"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/deepreach/bi_so101_fold_10inch_4cam.bi_so101_fold_towel_4camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 603,
"total_frames": 541612,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:603"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Zekai-Chen/bi_so101_fold_towel_4cam.BIS_speeches_97_23_MLM
Dataset Card for "BIS_Speeches_97_23"
This dataset is built from scrapped speeches on the Bank of International Settlements thanks to this repo : https://github.com/HanssonMagnus/scrape_bis. The dataset is made of 12k speeches from 1997 to 2023.
Each pair is built with extracted sentences from speeches, if B is following A then the 'next_sentence_label' is 1 else it is 0.
Negative pairs are built by choosing a sentence from another speech randomly.
Credits
Full… See the full description on the dataset page: https://huggingface.co/datasets/samchain/BIS_speeches_97_23_MLM.ActiveSetBiST
BiST
💻 Github Repo
English | 简体中文
Introduction
BiST is a large-scale bilingual translation dataset, with "BiST" standing for Bilingual Synthetic Translation dataset. Currently, the dataset contains approximately 60M entries and will continue to expand in the future.
BiST consists of two subsets, namely en-zh and zh-en, where the former represents the source language, collected from public data as real-world content; the latter represents the target… See the full description on the dataset page: https://huggingface.co/datasets/yumatin/BiST.Bishoujo_Mangekyoudataset-ohada-droit-commercial-general-echantillon
Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon
Description
Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.ClArTTS-multimodalCATT_benchmarkCATT automatic diacritization benchmark introduced in https://arxiv.org/abs/2407.03236
Dataset was downloaded from the CATT repo, besides adding a column without diacritics no changes were made.
Refer to the repo for licensing information.
sm-cv-en-itbb-trade-idp-feedback
🏦 Bangladesh Bank Trade Finance IDP — Multi-User Collaborative Fine-Tuning Dataset
This dataset contains human-reviewed, verified, and corrected document extractions for the 8 official Bangladesh Bank regulatory trade-finance document types.
It is completely self-contained and structured for immediate Vision-Language Model (VLM) fine-tuning anytime from any environment (Colab, Kaggle, GPU cluster, or local), with built-in multi-annotator merge support and incremental delta… See the full description on the dataset page: https://huggingface.co/datasets/bisalsaha/bb-trade-idp-feedback.sm-cv-en-swBiScope_Data
BiScope: AI-generated Text Detection by Checking Memorization of Preceding Tokens
This dataset is associated with the official implementation for the NeurIPS 2024 paper "BiScope: AI-generated Text Detection by Checking Memorization of Preceding Tokens". It contains human-written and AI-generated text samples from multiple tasks and generative models. Human data is always nonparaphrased, while AI-generated data is provided in both nonparaphrased and paraphrased forms.… See the full description on the dataset page: https://huggingface.co/datasets/HanxiGuo/BiScope_Data.bis-cb-policy-rates
BIS central bank policy rates (monthly)
Monthly central-bank policy rates from BIS WS_CBPOL (iso3 × year × month), with Viet Nam SBV refinancing overlay. Policy rate is the target / main instrument (% p.a.). band midpoint when applicable.
Figures
Hero
Comparison
Files
cbpol_monthly (18483 rows)
data/cbpol_monthly.csv
data/cbpol_monthly.dta
data/cbpol_monthly.xlsx
Load
Stata:
use "data/cbpol_monthly.dta", clear
R:
df <-… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/bis-cb-policy-rates.bis-eer-monthly
BIS effective exchange rates (monthly)
Monthly BIS effective exchange rate indices (iso3 × year × month): real and nominal, broad and narrow baskets. Index rise = appreciation of the domestic currency.
Figures
Hero
Comparison
Files
eer_monthly (24696 rows)
data/eer_monthly.csv
data/eer_monthly.dta
data/eer_monthly.xlsx
Load
Stata:
use "data/eer_monthly.dta", clear
R:
df <- read.csv("data/eer_monthly.csv")
SPSS: open the .xlsx or… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/bis-eer-monthly.
