Product
Everything Nodeau does, in one place.
Nodeau sets up your machines, checks what each GPU can hold, runs the models you choose and gives you an API to call them. Add machines and it becomes a fleet you can see and steer from anywhere. Here's the whole tour.
Getting started
Installs in minutes, and shows you the plan first.
nodeau install looks the machine over, lists every
command that needs your password, and waits for a yes. It sets up
the container runtime and Kubernetes pieces it needs, or adopts the
ones you already have, and remembers which is which so an
uninstall only removes what Nodeau added.
- Your NVIDIA driver, Secure Boot, bootloader and disks stay exactly as you set them up
nodeau doctorchecks the machine and the installation, read-only, and gives a remedy for anything it finds- Upgrading is running the installer again.
nodeau updatetells you when something newer is out - On a Mac there's no password and nothing to install but Nodeau itself
curl -fsSL https://get.nodeau.ai/install.sh | bash
# see the whole plan first, with nothing changed
nodeau install --dry-run
nodeau install
nodeau doctor
Will it fit?
GPU VERDICT NEEDS AVAILABLE
NVIDIA GeForce RTX 3080 yes (estimated) 7,210 MiB 9,365 MiB
NVIDIA GeForce RTX 5060 Ti yes 5,800 MiB 15,315 MiB
Models
A catalog that tells you which card each model wants.
Eight curated models, from a 2.6 GB starter for 8 GB cards up to a 27B flagship for a 24 GB card or two smaller cards in one machine. Each one is pinned to an exact file from its publisher and checked against a known SHA-256 before anything runs it.
- Nodeau predicts a model's memory from its size, quantisation, context and concurrency, then compares that with what your card can really give
- A measurement when it has one, a careful estimate otherwise, and the decision always says which
- Too big for the card? You get the numbers and what would change the answer, before anything starts loading
- Comfortable taking the risk on an estimate?
--accept-estimate-riskmakes that your call
Tasks
Chat, search, ranking, structured data and images.
Start a model for the job you have in mind. Nodeau checks the model can do it, starts it, and gives you the matching OpenAI-style route.
Chat
Chat and text completion, with streaming.
Embeddings
Vectors for search, RAG and clustering.
Reranking
Score documents against a query, in the Jina and Cohere shape.
Structured output
JSON that follows your schema, enforced while the model writes.
Tool calling
The model picks a function and hands you arguments to run it with.
Vision
Send an image with your prompt and ask about what's in it.
nodeau run qwen3-embedding-0.6b-q8_0 \
--task embed --port 8081
export NODEAU_API_KEY="$(nodeau auth show --quiet)"
curl http://127.0.0.1:8081/v1/embeddings \
-H "Authorization: Bearer $NODEAU_API_KEY" \
-H "Content-Type: application/json" \
-d '{"input": ["my GPUs", "a private cloud"]}'
Chat runs everywhere Nodeau runs, including a Mac running standalone. Embeddings, reranking, structured output, tool calling and vision run on Linux machines with NVIDIA GPUs. More of this is coming to the Mac, and the roadmap tracks it.
Bring your own model
Have a GGUF you like? Bring it along.
nodeau model import reads your file's header to work
out its architecture, quantisation and memory needs, and hashes the
bytes so the file is its own identity. Nothing inside it runs: a
GGUF is data, served by Nodeau's own pinned runtime.
Then nodeau model qualify starts it on a real card,
measures the memory it really takes, and tries each capability with
a test that only a capable model passes. Your model is ready
for exactly what it proved, and you know where it'll fit next time.
- One GGUF file per model, identified by the SHA-256 of its bytes
- Your weights stay on your machines
- Runs on Linux machines with NVIDIA GPUs
nodeau model import ./my-model.gguf --alias my-model
nodeau model qualify my-model
nodeau run my-model
nodeau batch submit requests.jsonl \
--model qwen3.5-9b-q4km --workers 2 --name overnight
nodeau batch wait overnight && nodeau batch results overnight
Batch inference
Queue a file of work and let your GPUs chew through it.
Write one request per line, each with its own custom_id.
Nodeau waits for a free card, runs the job, and writes a results file
where every answer is matched to the request it came from.
- Chat or embedding jobs, one model per job
--workers 2runs two copies of the model, each on its own card, sharing the records between them- Your records stay on your control-plane machine
- Part of Home Pro and Business, on Linux
Hardware
Every GPU, used for what it's good at.
Nodeau schedules every card in a machine on its own, and each card runs one workload at a time, so nothing fights over its memory. When a model is too big for any single card, Nodeau can split it across two cards in the same Linux machine, and each card holds its own share.
Splitting adds room rather than speed. On our own test pair, Qwen3-8B generated 113.6 tokens a second on an RTX 3080 alone and 89.3 split across the 3080 and an RTX 5060 Ti. So Nodeau uses a split for models that need the room, and for throughput it runs separate copies, one per card.
- NVIDIA GPUs on Linux, several per machine
- Apple Silicon Macs through Metal, running standalone
- Power limits for each card, inside the card's own range
# one model, too big for either card, across two
nodeau run qwen3.8-27b-q4km --gpus 2
# cap what a card may draw, in watts
nodeau power
nodeau power set --device <gpu-uuid> --limit 180
Selected studio / NVIDIA GeForce RTX 5060 Ti
Mode balanced (fleet-default)
WHY
it was already here and still fits
predicted 80.4 output tokens/s
predicted 166 W
NOT SELECTED
garage NVIDIA GeForce RTX 3080 predicted 114.0/s
Decisions
When Nodeau chooses a GPU, you can see why.
Pick how Nodeau should decide: Efficiency, Balanced or Performance, for the whole fleet or for one machine. Balanced is the default. It goes for speed while staying close to the most efficient choice.
nodeau placement explainshows the winner, the runners-up, and what each would have costnodeau service explainshows the memory arithmetic behind every decision- A running workload stays exactly where it is, even when another card scores better. Your service keeps serving
- Limits you set are hard limits: power budgets, cards held back from scheduling, a cap on cards per workload
Fleet
Several machines, one fleet.
Adding a machine takes two commands. Once it's in, Nodeau considers it for every placement, checks its copy of each model by hash, and keeps an eye on its health.
nodeau healthshows processor, memory, storage, network and GPU for every machine, keeps a day of history, and raises alerts when something needs you- Drain a machine for maintenance. New work goes elsewhere and what's running keeps serving
nodeau fleet connectputs your fleet in your account at app.nodeau.ai, so you can see it from any browser- With Home Pro or Business, run it from there too: start and stop models, set scheduling, drain machines, read logs and recover a workload from a machine that's gone
# on the machine you already have
nodeau fleet invite
# on the new machine, within 15 minutes
nodeau join <code>
nodeau fleet list
nodeau health
Organisation
Guardrails, usage and a record of every change.
Set it from a machine with nodeau governance, or from your account in a browser.
Your own limits
How many workloads, GPUs and batch workers may run at once, which models are allowed, and which cards and machines a fleet may use. Lowering a limit leaves running work alone and shapes what starts next.
Usage in accelerator-hours
Which workloads held which GPUs, and for how long, added up across your machines. Add your own hourly rate and Nodeau multiplies it out for you.
A record of changes
Who changed what and when, including requests that were turned down. Kept for a year.
Plans that travel offline
Carry a signed entitlement to a machine with no internet. It's checked exactly like one fetched over the network.
Upgrades
See what an upgrade involves before you start.
nodeau fleet upgrade plan pins a release down to an exact
build, then walks your fleet. For each machine it tells you whether
it would change, whether it's already current, or what needs your
attention first. It lists the order, with the control plane last, and
which workloads would restart along the way.
Ask twice with nothing changed and you get the same plan. When you're happy with it, you upgrade each machine with the installer.
Moving the whole fleet through a plan, one machine at a time, from one place, is in progress.
nodeau fleet upgrade plan
nodeau fleet upgrade plan --to <version> --json
Developers
Friendly to scripts, SDKs and people.
OpenAI-compatible API
Chat, streaming, embeddings and reranking on 127.0.0.1, with a key made on your machine.
JSON everywhere
Every command that prints a table also speaks --json, in a stable shape.
Honest exit codes
A refusal, a timeout and a bad flag each get their own code, so scripts know what happened.
A local dashboard
nodeau dashboard shows hardware, models, services and plan, served by Nodeau on your own machine.
Logs without kubectl
nodeau logs reads a model's output, a batch worker, or Nodeau itself.
Support bundles
nodeau support bundle writes one scrubbed file to your disk for you to send.
Privacy
Your models run close to your data.
Inference happens on your hardware. Prompts, completions,
embeddings, images and batch records stay there, and the local
endpoint listens on 127.0.0.1 behind a key generated on
your machine.
- A connected fleet reports its machines, cards, workloads and their state, so you can see them from a browser
- Your machines start every connection to Nodeau Cloud, which keeps your network boundary simple
- Plans are signed entitlements checked on the machine itself, so models keep serving offline
- Support bundles are scrubbed of keys and credentials, and stay on your disk until you send them
- Built for machines you and your team control, where everyone with access to a machine is trusted
Ready to try it on your own hardware?
Start with one machine for free. It's about ten minutes to your first model.