A tiny, object-oriented wrapper around 🤗 PEFT for LoRA fine-tuning of causal
language models. One LoraModel class hides the from_pretrained boilerplate,
and Agent adds web / file search capabilities on top.
- LoRA (Low-Rank Adaptation) — instead of updating all of a model's weights,
LoRA freezes the base model and trains two small low-rank matrices (
AandB) injected into the attention layers. Typically ~0.1% of parameters are trainable, so a checkpoint is a few MB instead of GB. - Base vs. adapter — the large pretrained weights never change. The tiny
adapter carries the new personality you trained.
enable_lora()/disable_lora()just switch which oneself.modelpoints at, so toggling never loses your trained weights. - Chat template — training and inference must format text the same way.
Training data is rendered with
add_generation_prompt=False; inference usesTrueso the model knows to start generating. - Label masking —
labelsmirrorinput_ids, but padding positions are set to-100so they are ignored in the loss. - System prompt — set
system_prompt=onLoraModelorAgent(description=…)to give the model a persona. Change it at runtime with/system-promptin an interactive session. - Slash commands — register your own with
@command("/name")and use them duringchat_session().run().
torch>=2.0
peft>=0.19
transformers>=4.45
pip install lora-easyRuns on CUDA, Apple Silicon (MPS), or CPU.
from lora_ez import LoraModel
m = LoraModel("Qwen/Qwen2.5-0.5B-Instruct", name="cat",
system_prompt="you are a sassy house cat")
# ----- fine-tune -----
m.enable_lora(r=8, alpha=16)
m.train(data, epochs=30)
m.save() # -> ./lora-cat/
# ----- single-turn chat -----
print(m.chat("hello!"))
# ----- multi-turn with memory -----
with m.chat_session("./chat.json", auto_save=True) as s:
s.chat("I'm back")
s.chat("how are you?")
s.run() # interactive REPL, /exit to quitfrom lora_ez import Agent
m = LoraModel("Qwen/Qwen2.5-0.5B-Instruct")
a = Agent(m, description="you are a data analyst",
web_enabled=True, file_enabled=True)
a.chat("what Python packages are installed?")
a.web_fetch("https://example.com")
a.disable_web()LoraModel
| Method | What it does |
|---|---|
LoraModel(model_id, name, system_prompt, device) |
Load base model + tokenizer |
enable_lora(r, alpha, dropout) |
Attach a LoRA adapter |
disable_lora() |
Point back to the frozen base model |
train(conversations, **kwargs) |
Fine-tune on ShareGPT-format data |
chat(prompt, history, system_prompt) |
Generate a reply |
chat_session(save_path, auto_save, system_prompt) |
Multi-turn session context manager |
save(path) / load(path) |
Persist / restore the adapter |
Agent — wraps a LoraModel with tools
| Method | What it does |
|---|---|
Agent(model, description, web_enabled, …) |
Wrap a model with search tools |
chat(prompt) |
Auto-injects web / file context, then delegates to model |
enable_web() / disable_web() |
Toggle web search |
enable_files() / disable_files() |
Toggle local file search |
web_fetch(url) |
Fetch a URL, respecting allowlists / blocklists |
file_read(path) |
Read a file inside allowed directories |
chat_session(…) |
Multi-turn session (delegated to the model) |
Slash commands — build your own with the @command decorator
from lora_ez import command
@command("/greet")
def greet(session, *args):
return f"Hello, {' '.join(args)}!" if args else "Hello!"Built-in: /exit, /help, /system-prompt.
The demo/ folder trains Qwen2.5-0.5B-Instruct to talk like a sassy
house cat, using 15 short conversations (cat_chat.json).
cd demo
python3 lora-cat.py # full training + before/after comparison
python3 test-session.py # multi-turn chat session with commandsMIT