lei
Datasets
All datasets matching “lei”colinear_scaling_models
license: gpl-2.0
Collinear Scaling Models
Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs.
Directory Structure
{dataset}/{design}/N_{param_count}/
Dataset: wikipedia, pes2o, cosmopedia, redpajama, c4 (plus _fp16 and _bigtpp variants)
Design: colinear or non_colinear
N: Model parameter count (one of 14 canonical sizes from ~5M to ~70M)
Experimental Designs
Collinear (CO):… See the full description on the dataset page: https://huggingface.co/datasets/leibnitz-lab/colinear_scaling_models.CC_eng_urlmilitary_vehicles
Citation
If you use this dataset, please cite the following paper:
@article{kricheli2024error,
title={Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge},
author={Kricheli, Joshua Shay and Vo, Khoa and Datta, Aniruddha and Ozgur, Spencer and Shakarian, Paulo},
journal={arXiv preprint arXiv:2407.15192},
year={2024}
}
VINCIE-10M
Dataset Card for VINCIE-10M
VINCIE: Unlocking In-context Image Editing from Video
Leigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao, Shanchuan Lin, Yichun Shi, Yicong Li, Wenjie Wang, Tat-Seng Chua, Lu Jiang
Dataset Construction Pipeline
Visual Transition Annotation. To describe visual transitions between frames, we use chain-of-thought (CoT) prompting to instruct a VLM to perform visual transition… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/VINCIE-10M.leire-corpus
Leire Corpus
(EN) Tokenized pretraining corpus for Leire, a 343.7M-parameter Brazilian Portuguese
LM trained from scratch on Kaggle T4s. ~15B tokens, 70% PT / 15% code / 8% math /
7% educational English, tokenized with a custom 32,768 BPE vocabulary trained on the
same mixture. Shards are uint16 binaries; recipe and stats below (in Portuguese).
Corpus de pre-treino da Leire, um LM de 343,7M de parametros em portugues
brasileiro, treinado do zero em T4 do Kaggle. O projeto e… See the full description on the dataset page: https://huggingface.co/datasets/raulmodena/leire-corpus.Leiniao_Dataset
