CoolFace
Datasetpublic

kaushik-systalyze/finetranslations-gaz-latn

FineTranslations gaz_Latn A bounded Oromo (gaz_Latn) to English translation workload derived from HuggingFaceFW/finetranslations, packaged for batched offline LLM inference experiments. The slice keeps rows whose full chat-template input length is input_tokens <= 10000. No source text is truncated. This preserves a meaningful single-language FineTranslations subset while removing only the longest tail that would dominate runtime and exceed the intended benchmark shape.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/finetranslations-gaz-latn.

sourceHugging Faceodc-byupdated 4mo agoView on Hugging Face
0likes23downloads
Dataset Card

FineTranslations gaz_Latn

A bounded Oromo (gaz_Latn) to English translation workload derived from `HuggingFaceFW/finetranslations`, packaged for batched offline LLM inference experiments.

The slice keeps rows whose full chat-template input length is input_tokens <= 10000. No source text is truncated. This preserves a meaningful single-language FineTranslations subset while removing only the longest tail that would dominate runtime and exceed the intended benchmark shape.

Intended Use

This dataset is intended for throughput benchmarking and predicted-vs-observed runtime validation. Each row contains a complete translation request in system and prompt; an OpenAI-compatible batch runner can send those fields as a system message plus user message.

It is not a translation quality benchmark. The translated_text column is the upstream Gemma3 27B generated translation from FineTranslations and is included as reference metadata, not as a human-validated label.

Prompt Template

Every row uses the upstream FineTranslations translation prompt. system is constant and prompt wraps each row's og_full_text:

text
**Oromo (gaz_Latn) Text to Translate (preserve all line breaks EXACTLY):**
<ORIGINAL>{og_full_text}</ORIGINAL>

Now translate to English (eng_Latn).

input_tokens counts the full chat request with the constant system message, the per-row prompt, and the assistant generation prefix.

Example Usage

python
from datasets import load_dataset

rows = load_dataset("kaushik-systalyze/finetranslations-gaz-latn", split="train")
record = rows[0]

request = {
    "model": "your-model",
    "temperature": 0,
    "max_tokens": int(record["max_tokens"]),
    "messages": [
        {"role": "system", "content": record["system"]},
        {"role": "user", "content": record["prompt"]},
    ],
}

For Systalyze batch inference, the default OpenAI-compatible mapper can use the system and prompt columns directly.

Split

  • —train: 47519 rows from HuggingFaceFW/finetranslations, config gaz_Latn, selected by input_tokens <= 10000.

Token Statistics

fieldrowsmeanp50p75p90p95p99max
input_tokens475191517.8114716532673359160659978
source_tokens47519931.956110662087300554799392
referenceoutputtokens47519493.32955671097159529337772
expectedmaxnew_tokens47519575.13576561239178732598582
max_tokens475191517.8114716532673359160659978

expected_max_new_tokens is a reference-output-based estimate: min(12000, max(64, ceil(reference_output_tokens * 1.10 + 32))).

max_tokens is the row-level inference guardrail intended to be sent to an OpenAI-compatible server: max(256, input_tokens). It is deliberately input-length based, with a floor for short prompts while allowing longer rows to decode up to their prompt length if EOS handling fails.

Length-Bucket Counts

bucketrows
shortunder5120
512_1k17948
1k_2k21508
2k_4k6378
4k_8k1568
8k_10k117

Length buckets are based on input_tokens: short_under_512 (<512), 512_1k, 1k_2k, 2k_4k, 4k_8k, and 8k_10k.

Runtime Target

The local Systalyze AOT proxy used during selection estimates this slice at about 3.25-3.79 hours on 1xA100, using input_tokens + reference_output_tokens as the work proxy. This is an estimate for workload selection, not an observed run. Actual duration will vary with engine settings, batching, max-token behavior, GPU SKU, and prefix/cache effects.

Source and Licensing

This dataset is a filtered derivative of `HuggingFaceFW/finetranslations` config gaz_Latn at source revision af3f4ca895450216d4771cdbf3e3b95c5bacaa2a. The upstream dataset is released under the Open Data Commons Attribution License (odc-by). Use of the source data is also subject to the upstream dataset card and CommonCrawl terms noted there.

Each row preserves the original FineTranslations metadata columns, including id, url, warc_path, language scores, quality scores, chunk fields, early_stop, and translation token counts.

Construction

  • —Downloaded the gaz_Latn parquet from HuggingFaceFW/finetranslations.
  • —Reconstructed a direct Oromo-to-English translation request using the upstream FineTranslations prompt shape.
  • —Counted full chat-template input tokens with the target tokenizer.
  • —Kept rows with input_tokens <= 10000; no rows were truncated.
  • —Added prompt columns, token-accounting columns, and a row-level max_tokens guardrail for offline inference tuning.

Limitations

  • —Web-sourced content can contain sensitive, low-quality, biased, or duplicated material despite upstream filtering.
  • —The reference translation is model-generated by the upstream dataset pipeline, not a human gold label.
  • —Rows where early_stop=true are retained; in those rows upstream translation may have stopped before translating all chunks.
  • —Runtime estimates depend on the local AOT profiles and should be validated by an observed batch run before using the result as a firm capacity number.

Files

  • —data/train-00000.parquet - selected rows and runnable prompts
  • —metadata.json - build settings and runtime proxy details
  • —token_stats.json - token distributions and bucket counts
  • —source_report.json - source attribution and row accounting
  • —histograms/input_tokens_histogram.csv - 500-token input histogram
  • —histograms/input_tokens_histogram.svg - rendered input-token histogram