datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens
Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens)
📝Check out the Blog Post
This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000
samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing.
Dataset Overview
Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.pg19-subsample
Jet-Long Evaluation Datasets
This repository contains the evaluation data used in the paper Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE.
Code: GitHub Repository
Dataset Description
These datasets are processed versions of standard benchmarks used to evaluate the long-context capabilities of Jet-Long:
RULER-500: A dataset for evaluating long-context understanding and retrieval up to 128K context.
PG-19-subsample: A subsampled version of… See the full description on the dataset page: https://huggingface.co/datasets/jet-ai/pg19-subsample.
