What happens
ruvllm quantize --model <hf-dir-or-model.safetensors> --quant q4_k_m --output out.gguf (ruvllm 2.1.0, Linux x86_64) prints Input size: 0.00 MB, 0 tensors, 0 elements, and writes a 75-byte file, for both a merged HuggingFace directory (config.json + model.safetensors + tokenizer.json) and the bare .safetensors path. The same weights load fine in candle and convert to a valid GGUF with candle's writer, so the input is well-formed.
Where it was hit
Mission ruOS SLiM (cognitum-one/ruos-desktop PR #83, ADR-061): a candle LoRA fine-tune of Qwen2.5-0.5B-Instruct merged to HF safetensors, then exported. The --help text advertises "Quantize a model to GGUF format … Q4_K_M, Q5_K_M, Q8_0" and gives ruvllm quantize --model ./model.safetensors --quant q8_0 as an example, so callers reasonably expect the safetensors path to work.
Expected
Either a working safetensors/HF-dir loader for quantize, or a loud error ("safetensors input not supported yet") instead of a 75-byte success.
Workaround
candle's native GGUF writer (llama.cpp Q4_K_M mixed layout for hidden sizes that are not multiples of 256, e.g. 896 on Qwen2.5-0.5B).
🤖 Generated with claude-flow
https://claude.ai/code/session_01CmwEhDLBt5MWxBBBTzvQqA
What happens
ruvllm quantize --model <hf-dir-or-model.safetensors> --quant q4_k_m --output out.gguf(ruvllm 2.1.0, Linux x86_64) printsInput size: 0.00 MB,0 tensors, 0 elements, and writes a 75-byte file, for both a merged HuggingFace directory (config.json + model.safetensors + tokenizer.json) and the bare.safetensorspath. The same weights load fine in candle and convert to a valid GGUF with candle's writer, so the input is well-formed.Where it was hit
Mission ruOS SLiM (cognitum-one/ruos-desktop PR #83, ADR-061): a candle LoRA fine-tune of Qwen2.5-0.5B-Instruct merged to HF safetensors, then exported. The
--helptext advertises "Quantize a model to GGUF format … Q4_K_M, Q5_K_M, Q8_0" and givesruvllm quantize --model ./model.safetensors --quant q8_0as an example, so callers reasonably expect the safetensors path to work.Expected
Either a working safetensors/HF-dir loader for
quantize, or a loud error ("safetensors input not supported yet") instead of a 75-byte success.Workaround
candle's native GGUF writer (llama.cpp Q4_K_M mixed layout for hidden sizes that are not multiples of 256, e.g. 896 on Qwen2.5-0.5B).
🤖 Generated with claude-flow
https://claude.ai/code/session_01CmwEhDLBt5MWxBBBTzvQqA