v0.1.075.2 MB
Apache-2.0
strict
core26
Share local language model inference across friends and colleagues
Rozalia lends the language model hardware sitting on your desk to the
people you choose, behind one OpenAI-compatible endpoint.
The machines with the GPUs live at home, behind NAT, where nothing on the
internet can dial them. So they dial out instead: rozalia-proxy opens a
WebSocket to the server and holds it there, and the server pushes
requests down it, to whichever machine advertises the model asked for and
has the fewest requests in flight.
This snap ships the proxy. Enrol the machine with a one-time code from
whoever runs the server, say what to offer and where, and turn it on:
proxy-backends takes a comma-separated list, for a machine running more
than one inference server. proxy-concurrency says how many requests to
run at once, and defaults to one, because a single GPU answering two
completions at once finishes neither sooner.
The cloud half is optional and ships as a separate component, so a
machine that only lends its GPU never downloads it:
The server does not start on install: a fresh database has no accounts,
and the address it should listen on is not something a package can
guess.
people you choose, behind one OpenAI-compatible endpoint.
The machines with the GPUs live at home, behind NAT, where nothing on the
internet can dial them. So they dial out instead: rozalia-proxy opens a
WebSocket to the server and holds it there, and the server pushes
requests down it, to whichever machine advertises the model asked for and
has the fewest requests in flight.
This snap ships the proxy. Enrol the machine with a one-time code from
whoever runs the server, say what to offer and where, and turn it on:
sudo rozalia.proxyctl enroll -server https://ai.example.org -code rze_...
sudo snap set rozalia \
proxy-server=https://ai.example.org \
proxy-backends=http://127.0.0.1:11434/v1 \
proxy=trueproxy-backends takes a comma-separated list, for a machine running more
than one inference server. proxy-concurrency says how many requests to
run at once, and defaults to one, because a single GPU answering two
completions at once finishes neither sooner.
The cloud half is optional and ships as a separate component, so a
machine that only lends its GPU never downloads it:
sudo snap install rozalia+server
sudo snap set rozalia base-url=https://ai.example.org listen-addr=127.0.0.1:8080
sudo snap set rozalia server=true
sudo rozalia.serverctl account create -email you@example.org -adminThe server does not start on install: a fresh database has no accounts,
and the address it should listen on is not something a package can
guess.
Update History
0+git12.d72fc10 (4) → v0.1.0 (7)10 Sept 2026, 07:30 UTC
0+git11.669c872 (1) → 0+git12.d72fc10 (4)10 Sept 2026, 06:15 UTC
0+git11.669c872 (1)10 Sept 2026, 05:30 UTC
10 Sept 2026, 05:28 UTC
10 Sept 2026, 07:03 UTC
10 Sept 2026, 05:30 UTC