Skip to content

ruvllm 2.1.0: quantize reads 0 tensors from a HuggingFace directory or .safetensors and writes a 75-byte GGUF #968

Description

@ruvnet

What happens

ruvllm quantize --model <hf-dir-or-model.safetensors> --quant q4_k_m --output out.gguf (ruvllm 2.1.0, Linux x86_64) prints Input size: 0.00 MB, 0 tensors, 0 elements, and writes a 75-byte file, for both a merged HuggingFace directory (config.json + model.safetensors + tokenizer.json) and the bare .safetensors path. The same weights load fine in candle and convert to a valid GGUF with candle's writer, so the input is well-formed.

Where it was hit

Mission ruOS SLiM (cognitum-one/ruos-desktop PR #83, ADR-061): a candle LoRA fine-tune of Qwen2.5-0.5B-Instruct merged to HF safetensors, then exported. The --help text advertises "Quantize a model to GGUF format … Q4_K_M, Q5_K_M, Q8_0" and gives ruvllm quantize --model ./model.safetensors --quant q8_0 as an example, so callers reasonably expect the safetensors path to work.

Expected

Either a working safetensors/HF-dir loader for quantize, or a loud error ("safetensors input not supported yet") instead of a 75-byte success.

Workaround

candle's native GGUF writer (llama.cpp Q4_K_M mixed layout for hidden sizes that are not multiples of 256, e.g. 896 on Qwen2.5-0.5B).

🤖 Generated with claude-flow

https://claude.ai/code/session_01CmwEhDLBt5MWxBBBTzvQqA

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions