bctn
bctn-dense_embeddingtoken-classification-llmlingua2-phobert-bctn-323_sample-5_epoch_16k_fpt_v1token-classification-llmlingua2-phobert-bctn-323_sample-5_epoch_16k_fpt_v2token-classification-llmlingua2-phobert-bctn-2308_sample-10_epoch_best_datatoken-classification-llmlingua2-xlm-roberta-bctn-2308_sample-5_epoch_best_datatoken-classification-llmlingua2-xlm-roberta-bctn-2308_sample-5_epoch_best_data_v2token-classification-llmlingua2-xlm-roberta-bctn-4001_sample-5_epoch_vitoken-classification-llmlingua2-xlm-roberta-bctn-1470_chunk_10epoch_best
Datasets
All datasets matching “bctn”vn-bctn-supplement
Vietnam Annual Reports Supplement Dataset (546 PDFs)
Tập dữ liệu bổ sung gồm 546 Báo cáo thường niên (BCTN) của các công ty niêm yết trên thị trường chứng khoán Việt Nam (HOSE, HNX, UPCoM), được trích xuất và chuẩn hóa để bổ sung cho tập dữ liệu gốc 13,982 báo cáo trên Zenodo.
Cấu trúc lưu trữ
.
├── README.md
├── bctn_supplement_index.parquet # Metadata index (546 rows)
├── manifest.csv # Chi tiết danh mục file
└── pdfs/
├── {TICKER}/… See the full description on the dataset page: https://huggingface.co/datasets/Tumiqa103/vn-bctn-supplement.UltraChat
NOTE(bctnry,2026.8.13): this is made from openbmb/UltraChat. there are
a few files w/ the same line having 2 json objects and i've manually
fixed them all in this repo.
Dataset Card for Dataset Name
Dataset Description
An open-source, large-scale, and multi-round dialogue data powered by Turbo APIs. In consideration of factors such as safeguarding privacy, we do not directly use any data available on the Internet as prompts.
To ensure generation quality, two… See the full description on the dataset page: https://huggingface.co/datasets/bctnry/UltraChat.CosmopediaV2BookCorpussanime2x-2k
sanime2x-2k
What you're seeing here is a dataset of anime-style AI-generated
illustrations intended for training super-resolution neural networks
(e.g. waifu2x). The dataset contains 2000 images, divided into 10
parts each containing 200 images.
The source
All images are collected manually from Twitter (now X). Due to the
inherent status problem of AI-generated artworks, an attempt was made
to make sure that no poster was credited, no "repost without
permission is… See the full description on the dataset page: https://huggingface.co/datasets/bctnry/sanime2x-2k.
