Repository navigation
HuggingFace login #167
Copy link
Copy link
Closed
Description
Activity
To use the HF models with LiteLLM you need to set the environment variable in the optillm proxy.
So,
export HUGGINGFACEHUB_API_TOKEN=your_hf_tokenand then run optillm.
Or, you can use the inbuilt inference server in optillm directly. For that, set the environment variables as follows:
export OPTILLM_API_KEY=optillm export HF_TOKEN=your_hf_tokenand then run optillm (setting the OPTILLM_API_KEY tells the proxy to use the inbuilt inference server).
The benefits of using the inbuilt inference server are that it is usually much faster, supports additional features in standard OpenAI API like returning logprobs, structured outputs (with response_format) and reasoning_effort.
E.g.
import os from openai import OpenAI import time OPENAI_BASE_URL = "http://localhost:8000/v1" OPENAI_API_KEY = "optillm" client = OpenAI(api_key=OPENAI_API_KEY, base_url=OPENAI_BASE_URL) messages=[ { "role": "user", "content": "How many rs are there in strawberry? Use code to solve the problem."} ] start_time = time.time() response = client.chat.completions.create( model = "huggingface/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B", # model = "deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B", # no need to include the prefix huggingface/ when using inbuilt inference server messages=messages, temperature=0.6, ) end_time = time.time() completion_tokens = response.usage.completion_tokens elapsed_time = end_time - start_time throughput = completion_tokens / elapsed_time if elapsed_time > 0 else 0 print(f"Completion tokens: {completion_tokens}") print(f"Elapsed time: {elapsed_time:.2f} seconds") print(f"Throughput: {throughput:.2f} tokens/second")
With LiteLLM:
Completion tokens: 275 Elapsed time: 90.09 seconds Throughput: 3.05 tokens/secondWith optiLLM:
Completion tokens: 541 Elapsed time: 30.08 seconds Throughput: 17.98 tokens/second- locked and limited conversation to collaborators
on Mar 4, 2025
Metadata
Metadata
Assignees
Labels
No labels
I get the following error message with Huggingface:
I don't know where I should input my Huggingface credentials. This is my code: