QuantFactory/DRT-o1-7B-GGUF
license: cc-by-nc-sa-4.0 language:
- en
- zh base_model:
- Qwen/Qwen2.5-7B-Instruct tags:
- machine tranlsation
- O1-like model
- Chat pipeline_tag: text-generation

QuantFactory/DRT-o1-7B-GGUF
This is quantized version of Krystalan/DRT-o1-7B created using llama.cpp
Original Model Card
DRT-o1
<p align="center"> 🤗 <a href="https://huggingface.co/Krystalan/DRT-o1-7B">DRT-o1-7B</a>   |   🤗 <a href="https://huggingface.co/Krystalan/DRT-o1-8B">DRT-o1-8B</a>   |   🤗 <a href="https://huggingface.co/Krystalan/DRT-o1-14B">DRT-o1-14B</a>   |    📑 <a href="https://arxiv.org/abs/2412.17498">Paper</a>
</p>
This repository contains the resources for our paper "DRT-o1: Optimized Deep Reasoning Translation via Long Chain-of-Thought"
Updates:
- 2024.12.31: We updated our paper with more detals and analyses. Check it out!
- 2024.12.31: We released the testing set of our work, please refer to
data/test.jsonl - 2024.12.30: We released a new model checkpoint using Llama-3.1-8B-Instruct as the backbone, i.e., 🤗 <a href="https://huggingface.co/Krystalan/DRT-o1-8B">DRT-o1-8B</a>
- 2024.12.24: We released our paper. Check it out!
- 2024.12.23: We released our model checkpoints. 🤗 <a href="https://huggingface.co/Krystalan/DRT-o1-7B">DRT-o1-7B</a> and 🤗 <a href="https://huggingface.co/Krystalan/DRT-o1-14B">DRT-o1-14B</a>.
If you find this work is useful, please consider cite our paper:
@article{wang2024drt,
title={DRT-o1: Optimized Deep Reasoning Translation via Long Chain-of-Thought},
author={Wang, Jiaan and Meng, Fandong and Liang, Yunlong and Zhou, Jie},
journal={arXiv preprint arXiv:2412.17498},
year={2024}
}Quick Links
Introduction
In this work, we introduce DRT-o1, an attempt to bring the success of long thought reasoning to neural machine translation (MT). To this end,
- 🌟 We mine English sentences with similes or metaphors from existing literature books, which are suitable for translation via long thought.
- 🌟 We propose a designed multi-agent framework with three agents (i.e., a translator, an advisor and an evaluator) to synthesize the MT samples with long thought. There are 22,264 synthesized samples in total.
- 🌟 We train DRT-o1-8B, DRT-o1-7B and DRT-o1-14B using Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct and Qwen2.5-14B-Instruct as backbones.
Our goal is not to achieve competitive performance with OpenAI’s O1 in neural machine translation (MT). Instead, we explore technical routes to bring the success of long thought to MT. To this end, we introduce DRT-o1, a byproduct of our exploration, and we hope it could facilitate the corresponding research in this direction.
Models
Model Access
Model Performance
Model Prompts
During model inference, please use the following prompts:
- System prompt:
You are a philosopher skilled in deep thinking, accustomed to exploring complex problems with profound insight. - User prompt:
Please translate the following text from English to Chinese:\n[An English text]
DRT-o1 models will first generate the thought and then provide the final translation, with the following format:
<thought>
[Reasoning process]
</thought>
<output>
[Final translation]
</output>Quickstart
- ⛷️ Huggingface Transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Krystalan/DRT-o1-7B"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
prompt = "Please translate the following text from English to Chinese:\nThe mother, with her feet propped up on a stool, seemed to be trying to get to the bottom of that answer, whose feminine profundity had struck her all of a heap."
messages = [
{"role": "system", "content": "You are a philosopher skilled in deep thinking, accustomed to exploring complex problems with profound insight."},
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
**model_inputs,
max_new_tokens=2048
)
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)- ⛷️ vllm:
Deploying LLMs:
python3 -m vllm.entrypoints.openai.api_server --model [model_ckpt] --served-model-name [model_name]Calling LLMs:
from openai import OpenAI
# Set OpenAI's API key and API base to use vLLM's API server.
openai_api_key = "EMPTY"
openai_api_base = "http://localhost:8000/v1"
client = OpenAI(
api_key=openai_api_key,
base_url=openai_api_base,
)
chat_response = client.chat.completions.create(
model=[model_name],
messages=[
{"role": "system", "content": "You are a philosopher skilled in deep thinking, accustomed to exploring complex problems with profound insight."},
{"role": "user", "content": "Please translate the following text from English to Chinese:\nThe mother, with her feet propped up on a stool, seemed to be trying to get to the bottom of that answer, whose feminine profundity had struck her all of a heap."},
],
temperature=0.1,
top_p=0.8,
max_tokens=2048,
extra_body={
"repetition_penalty": 1.05,
},
)
print("Chat response:", chat_response)Translation Cases
Data
We release the testing set of our work, please refer to data/test.jsonl, where en indicates the English source sentences, and zh denotes the corresponding Chinese translation.
We will release the long-thought MT data as well as the data collection codes soon!
License
This work is licensed under cc-by-nc-sa-4.0
