Product

Everything Nodeau does, in one place.

Nodeau sets up your machines, checks what each GPU can hold, runs the models you choose and gives you an API to call them. Add machines and it becomes a fleet you can see and steer from anywhere. Here's the whole tour.

Getting started

Installs in minutes, and shows you the plan first.

nodeau install looks the machine over, lists every command that needs your password, and waits for a yes. It sets up the container runtime and Kubernetes pieces it needs, or adopts the ones you already have, and remembers which is which so an uninstall only removes what Nodeau added.

  • Your NVIDIA driver, Secure Boot, bootloader and disks stay exactly as you set them up
  • nodeau doctor checks the machine and the installation, read-only, and gives a remedy for anything it finds
  • Upgrading is running the installer again. nodeau update tells you when something newer is out
  • On a Mac there's no password and nothing to install but Nodeau itself

Get Nodeau

bash
curl -fsSL https://get.nodeau.ai/install.sh | bash

# see the whole plan first, with nothing changed
nodeau install --dry-run

nodeau install
nodeau doctor
nodeau model status qwen3.5-9b-q4km
Will it fit?
  GPU                         VERDICT          NEEDS      AVAILABLE
  NVIDIA GeForce RTX 3080     yes (estimated)  7,210 MiB  9,365 MiB
  NVIDIA GeForce RTX 5060 Ti  yes              5,800 MiB  15,315 MiB

Models

A catalog that tells you which card each model wants.

Eight curated models, from a 2.6 GB starter for 8 GB cards up to a 27B flagship for a 24 GB card or two smaller cards in one machine. Each one is pinned to an exact file from its publisher and checked against a known SHA-256 before anything runs it.

  • Nodeau predicts a model's memory from its size, quantisation, context and concurrency, then compares that with what your card can really give
  • A measurement when it has one, a careful estimate otherwise, and the decision always says which
  • Too big for the card? You get the numbers and what would change the answer, before anything starts loading
  • Comfortable taking the risk on an estimate? --accept-estimate-risk makes that your call

The model catalog

Tasks

Chat, search, ranking, structured data and images.

Start a model for the job you have in mind. Nodeau checks the model can do it, starts it, and gives you the matching OpenAI-style route.

Chat

Chat and text completion, with streaming.

/v1/chat/completions

Embeddings

Vectors for search, RAG and clustering.

/v1/embeddings

Reranking

Score documents against a query, in the Jina and Cohere shape.

/v1/rerank

Structured output

JSON that follows your schema, enforced while the model writes.

response_format: json_schema

Tool calling

The model picks a function and hands you arguments to run it with.

tools

Vision

Send an image with your prompt and ask about what's in it.

image_url content

bash
nodeau run qwen3-embedding-0.6b-q8_0 \
  --task embed --port 8081

export NODEAU_API_KEY="$(nodeau auth show --quiet)"

curl http://127.0.0.1:8081/v1/embeddings \
  -H "Authorization: Bearer $NODEAU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"input": ["my GPUs", "a private cloud"]}'

Chat runs everywhere Nodeau runs, including a Mac running standalone. Embeddings, reranking, structured output, tool calling and vision run on Linux machines with NVIDIA GPUs. More of this is coming to the Mac, and the roadmap tracks it.

Examples for every task

Bring your own model

Have a GGUF you like? Bring it along.

nodeau model import reads your file's header to work out its architecture, quantisation and memory needs, and hashes the bytes so the file is its own identity. Nothing inside it runs: a GGUF is data, served by Nodeau's own pinned runtime.

Then nodeau model qualify starts it on a real card, measures the memory it really takes, and tries each capability with a test that only a capable model passes. Your model is ready for exactly what it proved, and you know where it'll fit next time.

  • One GGUF file per model, identified by the SHA-256 of its bytes
  • Your weights stay on your machines
  • Runs on Linux machines with NVIDIA GPUs

Bring your own model

bash
nodeau model import ./my-model.gguf --alias my-model
nodeau model qualify my-model
nodeau run my-model
bash
nodeau batch submit requests.jsonl \
  --model qwen3.5-9b-q4km --workers 2 --name overnight

nodeau batch wait overnight && nodeau batch results overnight

Batch inference

Queue a file of work and let your GPUs chew through it.

Write one request per line, each with its own custom_id. Nodeau waits for a free card, runs the job, and writes a results file where every answer is matched to the request it came from.

  • Chat or embedding jobs, one model per job
  • --workers 2 runs two copies of the model, each on its own card, sharing the records between them
  • Your records stay on your control-plane machine
  • Part of Home Pro and Business, on Linux

Batch inference

Hardware

Every GPU, used for what it's good at.

Nodeau schedules every card in a machine on its own, and each card runs one workload at a time, so nothing fights over its memory. When a model is too big for any single card, Nodeau can split it across two cards in the same Linux machine, and each card holds its own share.

Splitting adds room rather than speed. On our own test pair, Qwen3-8B generated 113.6 tokens a second on an RTX 3080 alone and 89.3 split across the 3080 and an RTX 5060 Ti. So Nodeau uses a split for models that need the room, and for throughput it runs separate copies, one per card.

  • NVIDIA GPUs on Linux, several per machine
  • Apple Silicon Macs through Metal, running standalone
  • Power limits for each card, inside the card's own range

Several GPUs in one machine

bash
# one model, too big for either card, across two
nodeau run qwen3.8-27b-q4km --gpus 2

# cap what a card may draw, in watts
nodeau power
nodeau power set --device <gpu-uuid> --limit 180
nodeau placement explain qwen-local
Selected   studio / NVIDIA GeForce RTX 5060 Ti
Mode       balanced   (fleet-default)

  WHY
    it was already here and still fits
    predicted 80.4 output tokens/s
    predicted 166 W

  NOT SELECTED
    garage   NVIDIA GeForce RTX 3080   predicted 114.0/s

Decisions

When Nodeau chooses a GPU, you can see why.

Pick how Nodeau should decide: Efficiency, Balanced or Performance, for the whole fleet or for one machine. Balanced is the default. It goes for speed while staying close to the most efficient choice.

  • nodeau placement explain shows the winner, the runners-up, and what each would have cost
  • nodeau service explain shows the memory arithmetic behind every decision
  • A running workload stays exactly where it is, even when another card scores better. Your service keeps serving
  • Limits you set are hard limits: power budgets, cards held back from scheduling, a cap on cards per workload

How Nodeau decides

Fleet

Several machines, one fleet.

Adding a machine takes two commands. Once it's in, Nodeau considers it for every placement, checks its copy of each model by hash, and keeps an eye on its health.

  • nodeau health shows processor, memory, storage, network and GPU for every machine, keeps a day of history, and raises alerts when something needs you
  • Drain a machine for maintenance. New work goes elsewhere and what's running keeps serving
  • nodeau fleet connect puts your fleet in your account at app.nodeau.ai, so you can see it from any browser
  • With Home Pro or Business, run it from there too: start and stop models, set scheduling, drain machines, read logs and recover a workload from a machine that's gone

Add and run machines

bash
# on the machine you already have
nodeau fleet invite

# on the new machine, within 15 minutes
nodeau join <code>

nodeau fleet list
nodeau health

Organisation

Guardrails, usage and a record of every change.

Set it from a machine with nodeau governance, or from your account in a browser.

Your own limits

How many workloads, GPUs and batch workers may run at once, which models are allowed, and which cards and machines a fleet may use. Lowering a limit leaves running work alone and shapes what starts next.

Usage in accelerator-hours

Which workloads held which GPUs, and for how long, added up across your machines. Add your own hourly rate and Nodeau multiplies it out for you.

A record of changes

Who changed what and when, including requests that were turned down. Kept for a year.

Plans that travel offline

Carry a signed entitlement to a machine with no internet. It's checked exactly like one fetched over the network.

Nodeau for teams and businesses

Upgrades

See what an upgrade involves before you start.

nodeau fleet upgrade plan pins a release down to an exact build, then walks your fleet. For each machine it tells you whether it would change, whether it's already current, or what needs your attention first. It lists the order, with the control plane last, and which workloads would restart along the way.

Ask twice with nothing changed and you get the same plan. When you're happy with it, you upgrade each machine with the installer.

Moving the whole fleet through a plan, one machine at a time, from one place, is in progress.

Planning a fleet upgrade

bash
nodeau fleet upgrade plan
nodeau fleet upgrade plan --to <version> --json

Developers

Friendly to scripts, SDKs and people.

OpenAI-compatible API

Chat, streaming, embeddings and reranking on 127.0.0.1, with a key made on your machine.

JSON everywhere

Every command that prints a table also speaks --json, in a stable shape.

Honest exit codes

A refusal, a timeout and a bad flag each get their own code, so scripts know what happened.

A local dashboard

nodeau dashboard shows hardware, models, services and plan, served by Nodeau on your own machine.

Logs without kubectl

nodeau logs reads a model's output, a batch worker, or Nodeau itself.

Support bundles

nodeau support bundle writes one scrubbed file to your disk for you to send.

Every command

Privacy

Your models run close to your data.

Inference happens on your hardware. Prompts, completions, embeddings, images and batch records stay there, and the local endpoint listens on 127.0.0.1 behind a key generated on your machine.

  • A connected fleet reports its machines, cards, workloads and their state, so you can see them from a browser
  • Your machines start every connection to Nodeau Cloud, which keeps your network boundary simple
  • Plans are signed entitlements checked on the machine itself, so models keep serving offline
  • Support bundles are scrubbed of keys and credentials, and stay on your disk until you send them
  • Built for machines you and your team control, where everyone with access to a machine is trusted

Security and privacy

Ready to try it on your own hardware?

Start with one machine for free. It's about ten minutes to your first model.