CoolFace
Datasetpublic

tomaarsen/zelo-scores-10kx100-granite-4.1-30b

Dataset Card for tomaarsen/zelo-scores-10kx100-granite-4.1-30b Dataset Summary Synthetic data generated by DataForge: Model: ibm-granite/granite-4.1-30b (main) Source dataset: tomaarsen/zelo-pairs-10kx100-quantile-anchor (train split). Generation config: temperature=None, top_p=None, top_k=None, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: `You are a relevance scoring system. Given a query and two documents (A and B)… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/zelo-scores-10kx100-granite-4.1-30b.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes48downloads
Dataset Card

Dataset Card for tomaarsen/zelo-scores-10kx100-granite-4.1-30b

Dataset Summary

Synthetic data generated by DataForge:

  • —Model: ibm-granite/granite-4.1-30b (main)
  • —Source dataset: tomaarsen/zelo-pairs-10kx100-quantile-anchor (train split).
  • —Generation config: temperature=None, top_p=None, top_k=None, max_tokens=4096, model_max_context=32768
  • —Speculative decoding: disabled
  • —System prompt: `You are a relevance scoring system. Given a query and two documents (A and B), your job is to decide which document is more relevant to the given query. You should think carefully, considering the pros and cons between each document. For your first few sentences, consider the pros and cons of Document A. Then, spend some time thinking about Document B. Then, at the end, compare, and make a decision as to which one is more relevant. Do NOT make a decision in the beginning of your thoughts, stay open-minded until the last 1-2 sentences of your thoughts.

The score should range from -1.0 to 1.0, where negative means Document A is more relevant, and positive means Document B is more relevant. You can pick any number from -1.0 to 1.0.

Your final response must end with exactly one line of the form: Score: <float between -1.0 and 1.0>.`

  • —User prompt: Column messages

The run produced 1,001,000 (~1.0M) samples and generated 352,337,836 (~352.3M) tokens.

You can load the dataset using

python
from datasets import load_dataset

ds = load_dataset("tomaarsen/zelo-scores-10kx100-granite-4.1-30b")

Dataset Stats

MetricValue
Documents processed1,001,000 (~1.0M)
Total prompt tokens517,236,012 (~517.2M)
Total completion tokens352,337,836 (~352.3M)
Mean prompt tokens516.72
Mean completion tokens351.99

Licensing Information

License: other

Contributions

Thanks to @tomaarsen for adding this dataset.