datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k
This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 9 hours for 2k examples.
Usage
from datasets import load_dataset
kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.mtob
MTOB (Machine Translation from One Book)
Last updated: Wednesday, July 9, 2025
Machine Translation from One Book evaluates a language model's ability to translate sentences from English to Kalamang (a low-resource language) and from Kalamang to English.
As of July 2, 2025, additional tasks for this groq-bench implementation include:
Kalamang-to-English translation
adding the option to perform long-context evaluation where the Kalamang corpus is used as input to the model
adding… See the full description on the dataset page: https://huggingface.co/datasets/Groq/mtob.LiveCodeBench-CodeGenerationmath-for-mcpdocument-qna-chroma-groq-logsopen-source-throughput-detailed
Open Sourced Data
more info coming soon!
open-source-throughput
Conversations Dataset
More info coming soon!
kurisu_groq1groq-generated-typos
