jonasluehrs-jaai/synthetic_dataset_low-mid
Synthetic Dataset: Low Context, Medium Generation Dataset Description This is a synthetic benchmark dataset designed to test LLM inference performance in low-context, mid-generation scenarios. The dataset consists of 2,000 samples with randomly generated tokens that simulate workloads where models receive short prompts but generate longer responses. Use Cases This dataset is ideal for benchmarking: Creative writing and content generation Code… See the full description on the dataset page: https://huggingface.co/datasets/jonasluehrs-jaai/synthetic_dataset_low-mid.
Synthetic Dataset: Low Context, Medium Generation
Dataset Description
This is a synthetic benchmark dataset designed to test LLM inference performance in low-context, mid-generation scenarios. The dataset consists of 2,000 samples with randomly generated tokens that simulate workloads where models receive short prompts but generate longer responses.
Use Cases
This dataset is ideal for benchmarking:
- Creative writing and content generation
- Code generation from brief descriptions
- Story/article generation from short prompts
- Generation throughput and token-per-second optimization
Dataset Characteristics
- Number of Samples: 2,000
- Prompt Length Distribution: Beta distribution (skewed toward lower values)
- Alpha (α): 2.0 (controls left tail)
- Beta (β): 6.0 (controls right tail)
- Range: 10-120 tokens
- Response Length Distribution: Normal distribution
- Mean: 1,500 tokens
- Standard deviation: 300 tokens
- Tokenizer: meta-llama/Llama-3.1-8B-Instruct
Dataset Structure
Each sample contains:
prompt: A sequence of randomly generated tokens (low context)prompt_length: Number of tokens in the promptresponse_length: Number of tokens in the response
{
'prompt': str,
'prompt_length': int,
'response_length': int
}Token Generation
- Tokens are randomly sampled from the vocabulary of the Llama-3.1-8B-Instruct tokenizer
- Prompt lengths follow a beta distribution, creating a realistic skew toward shorter prompts
- Response lengths follow a normal distribution, simulating typical generation patterns
- Each sample is independently generated with lengths drawn from the specified distributions
Related Datasets
This dataset is part of a suite of three synthetic benchmark datasets, each designed for different workload patterns:
- synthetic_dataset_high-low
- High context (32k tokens), low generation (200 tokens)
- Focus: Prompt processing efficiency, TTFT optimization
- synthetic_dataset_mid-mid
- Medium context (1k tokens), medium generation (1k tokens)
- Focus: Balanced workload, realistic API scenarios
- 🔷 synthetic_dataset_low-mid (this dataset)
- Low context (10-120 tokens), medium generation (1.5k tokens)
- Focus: Generation throughput, creative writing scenarios
Benchmarking with vLLM
This dataset is designed for use with the vLLM inference framework. The vLLM engine supports a min_tokens parameter, allowing you to pass min_tokens=max_tokens=response_length for each prompt. This ensures that the response length follows the defined distribution.
Setup
First, install vLLM and start the server:
pip install vllm
# Start the vLLM server
vllm serve meta-llama/Llama-3.1-8B-InstructUsage Example
from datasets import load_dataset
from openai import OpenAI
# Load the dataset
dataset = load_dataset("jonasluehrs-jaai/synthetic_dataset_low-mid")
# Initialize vLLM client (OpenAI-compatible API)
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="token-abc123",
)
# Use a sample from the dataset
sample = dataset['train'][0]
# Make a completion request with controlled response length
completion = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": sample['prompt']}],
max_tokens=sample['response_length'],
extra_body={"min_tokens": sample['response_length']},
)
print(f"Generated {len(completion.choices[0].message.content)} characters")For more information, see the vLLM OpenAI-compatible server documentation.
License
This dataset is released under the MIT License.
