You can perform direct inference using the transformers library. HuatuoGPT-o1 follows a thinks-before-it-answers approach, where the output is structured into two distinct sections: ## Thinking (the reasoning process) and ## Final Response (the actual answer).
Available models include:
- HuatuoGPT-o1-8B (LLaMA-3.1-8B, English)
- HuatuoGPT-o1-70B (LLaMA-3.1-70B, English)
- HuatuoGPT-o1-7B (Qwen2.5-7B, English & Chinese)
- HuatuoGPT-o1-72B (Qwen2.5-72B, English & Chinese)
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("FreedomIntelligence/HuatuoGPT-o1-8B",torch_dtype="auto",device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("FreedomIntelligence/HuatuoGPT-o1-8B")
input_text = "How to stop a cough?"
messages = [{"role": "user", "content": input_text}]
inputs = tokenizer(tokenizer.apply_chat_template(messages, tokenize=False,add_generation_prompt=True
), return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=2048)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))