Distributed LLM inference
Nobody here owns a datacenter.
Vinculum pools idle consumer hardware into one shared inference network. Lend a machine and earn credits. Spend them on models no single machine here could run alone.
BitTorrent proved strangers will pool bandwidth to give each other things no one peer could serve. This is the same bargain, for GPU time.
Two ways in
Bring a machine, or bring a question.
Both are welcome and they are not the same deal. Read the one that's you.
For the machine you already own
Lend a machine
Your GPU is idle most of the day. Install the app, flip one switch, and it serves requests from the network in the background — earning credits the whole time.
- What it does
- Detects your hardware, picks the largest model tier it can actually hold, and serves whole requests. Bundled llama.cpp — no CUDA afternoon.
- What you get
- Credits per token you actually serve, weighted by tier, plus a trickle for staying reachable.
- What it costs you
- Electricity, and the fact that prompts you serve pass through your machine in the clear.
For the question you have now
Ask a question
The same app is a chat client. With no credits and no network it still works: replies run on your own machine, instantly and privately.
- Free, always
- Local Mode runs a small model on your hardware. Nothing leaves the machine, and it never expires.
- What credits unlock
- The big tiers — the models your laptop cannot hold, served by machines that can.
- What it costs you
- Queue time, and knowing that a network reply ran on a stranger's computer. Every reply is tagged with where it ran.
The loop no exit, by design
Contribute, earn, spend, repeat.
One app holds both halves. The sequence matters, and it closes — step four is why step one is worth doing.
Install
The app makes a keypair on first run. That's your identity — no email, no signup, nothing to verify.
Share
Flip the switch and your machine joins the grid, serving what its hardware can hold. Credits accrue while you work on something else.
Spend
Ask a big model something. Credits pay the machines that answer, per token, at the tier you chose.
Out of credits
Local Mode keeps answering for free while your machine earns the next batch. The network is never a paywall.
Not downloadable yet
There is no download button, on purpose.
Builds for macOS and Windows are being signed and notarised first — an app that asks to use your hardware should not arrive with the operating system warning you about it. Leave an address and you'll get one email when the real thing is ready.
One email, at launch — an address and a date, nothing else, and nothing else ever sent to it. Full disclosure.