datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ruler-300-seed42
Frozen RULER 300, seed 42
This dataset freezes the exact RULER inputs used by the
short-long-pretraining native evaluation suite.
Repository: bicycleman15/ruler-300-seed42
Rows: 6,300
Tasks: s-niah-1, s-niah-2, s-niah-3, mk1, mk2, mv, mq
Context lengths: 1024, 2048, 4096
Samples per task/length: 300
Seed: 42
Dataset SHA-256: 4d82df6f9b1f2d9c45c0a0bda8c734032e62f517b746c6351bf9c2f38335ab3d
Tokenizer SHA-256: 1f186971e25f7bda3dd6f93a100bb8fa2a6801cf8dc3807c8a8c4e45f296ab90… See the full description on the dataset page: https://huggingface.co/datasets/bicycleman15/ruler-300-seed42.fineweb_50b_none_none
fineweb_50b_none_none
Tokenized FineWeb-Edu sample-100BT (no document-length filter).
Tokenizer: NousResearch/Llama-2-7b-hf (Llama-2, vocab 32k)
Format: uint16 NumPy shards (shard_train_*.npy, shard_val_*.npy)
Size: 50B train tokens, 100M validation tokens, 100M tokens per shard
Each document is prefixed with the EOS token
License: ODC-By 1.0 (same as FineWeb-Edu). Also subject to Common Crawl Terms of Use.
bicyclebicycle_maintenance
Dataset Card for bicycle_maintenance
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/burtenshaw/bicycle_maintenance.bicycle30b_lcbtop65_st_320x8k_30_v4_newUAEC-bird-or-bicycleAI-Bicycle-LFM2-VL-450M
🧠 Multistage Japanese Dialogue & VQA Dataset
本データセットは、複数の公開データセットを統合・日本語化し、
自然な日本語対話および視覚応答(Vision Reaction)データを生成したものです。
出典を明示することで、商用利用が可能です。
ライセンスを尊重したうえで翻訳・加工を行い、3つのステージで構成されています。
注意: Vision Reaction とは、画像に対して自然な応答をすることです。
例えば、海の画像を入力したとき、一般的なVQAタスクでは「青い空と青い海が広がっています。....」のような説明をします。
Vision Reactionでは、「お、きれいな海だな!」というように、自然なリアクションをします。
📘 データ構成
Stage 1:en_multiturn.jsonl
元データ:allenai/soda
ライセンス:Creative Commons Attribution 4.0 International (CC BY 4.0)… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/AI-Bicycle-LFM2-VL-450M.4b_lcbtop65_mt_64x4kx5_s0_e65_behaviour_promptSDIP_bicycle
This repository contains Unofficial access for SDIP-bicycles dataset
Official Repository, Project Page, Paper
Self-Distilled Internet Photos (SDIP) is a multi-domain image dataset. The dataset consists of Self-Distilled Flickr (SD-Flickr) and *Self-Distilled LSUN (SD-LSUN) that were crawled from Flickr and LSUN dataset, respectively, and then curated using the method described in our Self-Distilled StyleGAN paper:
Self-Distilled StyleGAN: Towards Generation from Internet Photos
Ron… See the full description on the dataset page: https://huggingface.co/datasets/rmokady/SDIP_bicycle.30b_lcbtop65_interactfix_32x4kx20word_sorting_num_words_3_7_9_word_length_3_7_9st_k160_250_275imnet1k_bicycle-built-for-two_tandem_bicycle_tandem1k_32_s1lcb_30b_mt_5_65bicycleword_sorting_num_words_9_word_length_9clip-bicycle-e-bikelcb_30b_mt_10_65word_sorting_num_words_6_word_length_8lcb_gpt_st8k_654b_lcb_rlef_64x4kx5_0word_sorting_num_words_9_11_13_word_length_9_11_13torch_modelsword_sorting_num_words_7_word_length_730b_lcbtop65_st_320x8k_0_v1_newword_sorting_num_words_13_word_length_13word_sorting_num_words_15_word_length_15NYC_bicycle_counts
Bicycle Counts Dataset
The Bicycle Counts dataset was merged with the Bicycle Counters dataset on the shared id field to assign geographic coordinates.
More information about how this dataset was generated can be found here.
