Skip to content

HuggingFace login #167

Description

@SeriousJ55

I get the following error message with Huggingface:

openai.InternalServerError: Error code: 500 - {'error': 'litellm.AuthenticationError: HuggingfaceException - {"error":"Invalid username or password."}'}

I don't know where I should input my Huggingface credentials. This is my code:

import os
from openai import OpenAI

messages=[
	{
	  "role": "user",
	  "content": "Write a Python program to build an RL model to recite text from any position that the user provides, using only numpy."
	}
]

OPENAI_KEY = "optillm"
OPENAI_BASE_URL = "http://localhost:8000/v1"

os.environ["HUGGINGFACEHUB_API_TOKEN"] = "XXX"

client = OpenAI(api_key=OPENAI_KEY, base_url=OPENAI_BASE_URL)

response = client.chat.completions.create(
  model="huggingface/meta-llama/Llama-3.2-1B-Instruct",
  messages=messages,
  temperature=0.2,
)

print(response)

Activity

  1. codelion commented on Mar 4, 2025

    @codelion
    Member

    To use the HF models with LiteLLM you need to set the environment variable in the optillm proxy.

    So,

    export HUGGINGFACEHUB_API_TOKEN=your_hf_token

    and then run optillm.

    Or, you can use the inbuilt inference server in optillm directly. For that, set the environment variables as follows:

    export OPTILLM_API_KEY=optillm
    export HF_TOKEN=your_hf_token
    

    and then run optillm (setting the OPTILLM_API_KEY tells the proxy to use the inbuilt inference server).

    The benefits of using the inbuilt inference server are that it is usually much faster, supports additional features in standard OpenAI API like returning logprobs, structured outputs (with response_format) and reasoning_effort.

    E.g.

    import os
    from openai import OpenAI
    import time
    
    OPENAI_BASE_URL = "http://localhost:8000/v1"
    OPENAI_API_KEY = "optillm"
    client = OpenAI(api_key=OPENAI_API_KEY, base_url=OPENAI_BASE_URL)
    
    messages=[
        {
          "role": "user",
          "content": "How many rs are there in strawberry? Use code to solve the problem."}
      ]
    start_time = time.time()
    response = client.chat.completions.create(
      model = "huggingface/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B",
      # model = "deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B", # no need to include the prefix huggingface/ when using inbuilt inference server
      messages=messages,
      temperature=0.6,
    )
    end_time = time.time()
    completion_tokens = response.usage.completion_tokens
    elapsed_time = end_time - start_time
    throughput = completion_tokens / elapsed_time if elapsed_time > 0 else 0
    
    print(f"Completion tokens: {completion_tokens}")
    print(f"Elapsed time: {elapsed_time:.2f} seconds") 
    print(f"Throughput: {throughput:.2f} tokens/second")

    With LiteLLM:

    Completion tokens: 275
    Elapsed time: 90.09 seconds
    Throughput: 3.05 tokens/second
    

    With optiLLM:

    Completion tokens: 541
    Elapsed time: 30.08 seconds
    Throughput: 17.98 tokens/second
    
  2. locked and limited conversation to collaborators on Mar 4, 2025
  3. converted this issue into a discussion #168 on Mar 4, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions