kaushik-systalyze/finetranslations-gaz-latn
FineTranslations gaz_Latn A bounded Oromo (gaz_Latn) to English translation workload derived from HuggingFaceFW/finetranslations, packaged for batched offline LLM inference experiments. The slice keeps rows whose full chat-template input length is input_tokens <= 10000. No source text is truncated. This preserves a meaningful single-language FineTranslations subset while removing only the longest tail that would dominate runtime and exceed the intended benchmark shape.… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/finetranslations-gaz-latn.
FineTranslations gaz_Latn
A bounded Oromo (gaz_Latn) to English translation workload derived from `HuggingFaceFW/finetranslations`, packaged for batched offline LLM inference experiments.
The slice keeps rows whose full chat-template input length is input_tokens <= 10000. No source text is truncated. This preserves a meaningful single-language FineTranslations subset while removing only the longest tail that would dominate runtime and exceed the intended benchmark shape.
Intended Use
This dataset is intended for throughput benchmarking and predicted-vs-observed runtime validation. Each row contains a complete translation request in system and prompt; an OpenAI-compatible batch runner can send those fields as a system message plus user message.
It is not a translation quality benchmark. The translated_text column is the upstream Gemma3 27B generated translation from FineTranslations and is included as reference metadata, not as a human-validated label.
Prompt Template
Every row uses the upstream FineTranslations translation prompt. system is constant and prompt wraps each row's og_full_text:
**Oromo (gaz_Latn) Text to Translate (preserve all line breaks EXACTLY):**
<ORIGINAL>{og_full_text}</ORIGINAL>
Now translate to English (eng_Latn).input_tokens counts the full chat request with the constant system message, the per-row prompt, and the assistant generation prefix.
Example Usage
from datasets import load_dataset
rows = load_dataset("kaushik-systalyze/finetranslations-gaz-latn", split="train")
record = rows[0]
request = {
"model": "your-model",
"temperature": 0,
"max_tokens": int(record["max_tokens"]),
"messages": [
{"role": "system", "content": record["system"]},
{"role": "user", "content": record["prompt"]},
],
}For Systalyze batch inference, the default OpenAI-compatible mapper can use the system and prompt columns directly.
Split
train: 47519 rows fromHuggingFaceFW/finetranslations, configgaz_Latn, selected byinput_tokens <= 10000.
Token Statistics
expected_max_new_tokens is a reference-output-based estimate: min(12000, max(64, ceil(reference_output_tokens * 1.10 + 32))).
max_tokens is the row-level inference guardrail intended to be sent to an OpenAI-compatible server: max(256, input_tokens). It is deliberately input-length based, with a floor for short prompts while allowing longer rows to decode up to their prompt length if EOS handling fails.
Length-Bucket Counts
Length buckets are based on input_tokens: short_under_512 (<512), 512_1k, 1k_2k, 2k_4k, 4k_8k, and 8k_10k.
Runtime Target
The local Systalyze AOT proxy used during selection estimates this slice at about 3.25-3.79 hours on 1xA100, using input_tokens + reference_output_tokens as the work proxy. This is an estimate for workload selection, not an observed run. Actual duration will vary with engine settings, batching, max-token behavior, GPU SKU, and prefix/cache effects.
Source and Licensing
This dataset is a filtered derivative of `HuggingFaceFW/finetranslations` config gaz_Latn at source revision af3f4ca895450216d4771cdbf3e3b95c5bacaa2a. The upstream dataset is released under the Open Data Commons Attribution License (odc-by). Use of the source data is also subject to the upstream dataset card and CommonCrawl terms noted there.
Each row preserves the original FineTranslations metadata columns, including id, url, warc_path, language scores, quality scores, chunk fields, early_stop, and translation token counts.
Construction
- Downloaded the
gaz_Latnparquet fromHuggingFaceFW/finetranslations. - Reconstructed a direct Oromo-to-English translation request using the upstream FineTranslations prompt shape.
- Counted full chat-template input tokens with the target tokenizer.
- Kept rows with
input_tokens <= 10000; no rows were truncated. - Added prompt columns, token-accounting columns, and a row-level
max_tokensguardrail for offline inference tuning.
Limitations
- Web-sourced content can contain sensitive, low-quality, biased, or duplicated material despite upstream filtering.
- The reference translation is model-generated by the upstream dataset pipeline, not a human gold label.
- Rows where
early_stop=trueare retained; in those rows upstream translation may have stopped before translating all chunks. - Runtime estimates depend on the local AOT profiles and should be validated by an observed batch run before using the result as a firm capacity number.
Files
data/train-00000.parquet- selected rows and runnable promptsmetadata.json- build settings and runtime proxy detailstoken_stats.json- token distributions and bucket countssource_report.json- source attribution and row accountinghistograms/input_tokens_histogram.csv- 500-token input histogramhistograms/input_tokens_histogram.svg- rendered input-token histogram
