Description
Two weeks ago was released very efficient and interesting Meta model - muse glimmer
https://huggingface.co/meta-models/Muse-Glimmer-30B
https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF
Latest llama.cpp v0.3.0 release already supports it.
Please, consider upgrading to it - this model is actually kind of silent revolution in terms of VRAM efficiency across dense models.
Reproduction Steps
Try to load any gguf from https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF as an LLM model
Environment & Configuration
- Operating system: win
- .NET runtime version: 8
- LLamaSharp version: 0.29
- CUDA version (if you are using cuda backend):
- CPU & GPU device: core i5 10400, amd rx 9070, Vulkan
Known Workarounds
No response
Description
Two weeks ago was released very efficient and interesting Meta model - muse glimmer
https://huggingface.co/meta-models/Muse-Glimmer-30B
https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF
Latest llama.cpp v0.3.0 release already supports it.
Please, consider upgrading to it - this model is actually kind of silent revolution in terms of VRAM efficiency across dense models.
Reproduction Steps
Try to load any gguf from https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF as an LLM model
Environment & Configuration
Known Workarounds
No response