Docs · Models

Embeddings, reranking, tools and vision

Chat is where most people start, and it's far from all Nodeau runs. Here's how to get vectors, rank documents, call tools, get JSON back and ask about an image, on your own GPUs.

Everything on this page runs on Linux machines with NVIDIA GPUs, through the same fit check and the same kind of endpoint as chat. A Mac runs standalone and serves chat models, and embeddings, reranking and image input on the Mac are in progress.

The basics of the endpoint, authentication and chat are in the OpenAI-compatible API. Set your key first:

Terminal
export NODEAU_API_KEY="$(nodeau auth show --quiet)"

Which model does what#

TaskCurated modelsHow you call it
Embeddingsqwen3-embedding-0.6b-q8_0/v1/embeddings
Rerankingbge-reranker-v2-m3-q8_0/v1/rerank
Tool callingqwen3.5-4b-q4km, qwen3.5-9b-q4km, qwen3.8-27b-q4kmtools in a chat completion
Structured outputthe three Qwen models above, plus gemma-4-e4b-qat-q4-0 and gemma-4-12b-qat-q4-0response_format in a chat completion
Image inputgemma-4-e4b-qat-q4-0 (8 GB), gemma-4-12b-qat-q4-0 (12 GB)image parts in a chat completion

nodeau model info <model> lists exactly what a model can do, and the model catalog shows which size of card each one is for.

Starting a workload for a task#

A workload serves one model for one task, on its own port. An embedding model or a reranker does only one thing, so Nodeau works out the task for you:

Terminal
nodeau run qwen3-embedding-0.6b-q8_0 --port 8081
nodeau run bge-reranker-v2-m3-q8_0 --port 8082

For a model that could do more than one task, say which with --task, for example --task embed.

A GPU belongs to one workload at a time, so each workload you want running side by side gets a card of its own. On a machine with one GPU, free the card first with nodeau stop <name> --workload, then start the next model. With more cards, in one machine or across several machines, each workload takes its own card and its own port, and they all serve at once.

Embeddings#

Vectors for search, clustering and retrieval-augmented generation.

Terminal
nodeau run qwen3-embedding-0.6b-q8_0 --port 8081
Terminal
curl http://127.0.0.1:8081/v1/embeddings \
  -H "Authorization: Bearer $NODEAU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"input": ["How do I add a machine to my fleet?",
                 "Run nodeau fleet invite on the machine you already have."]}'

You get the usual OpenAI embeddings response, one vector per input. With the Python SDK:

Python
import os
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8081/v1",
                api_key=os.environ["NODEAU_API_KEY"])

result = client.embeddings.create(
    model="qwen3-embedding-0.6b-q8_0",
    input=["first text", "second text"],
)
vectors = [item.embedding for item in result.data]

Got a whole corpus to embed? Batch inference takes a file of embedding requests and works through it on your GPUs.

Reranking#

Give it a query and a handful of documents, and it scores how well each one answers the query. It's the step that makes retrieval sharp.

Terminal
nodeau run bge-reranker-v2-m3-q8_0 --port 8082
Terminal
curl http://127.0.0.1:8082/v1/rerank \
  -H "Authorization: Bearer $NODEAU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query": "how do I add a machine",
       "documents": ["nodeau fleet invite prints a code for the new machine",
                     "GGUF is a model file format",
                     "Sourdough needs a long, cold fermentation"]}'

The reply has a results list, one entry per document, each with the index of the document you sent and a relevance_score. Higher means more relevant, and a score can be negative, so compare them with each other rather than against zero.

A common retrieval pattern puts the two together: embed your documents once, find the nearest few for each question, rerank those, and hand the best ones to a chat model as context.

Tool calling#

Ordinary OpenAI tool calling. You describe your tools, the model may answer with tool_calls, and your code runs the tool and sends the result back in a message with the role tool. Nodeau hands you the call, and running the tool is up to you.

Terminal
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $NODEAU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role":"user","content":"What is the weather in Oslo?"}],
    "max_tokens": 2048,
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Current weather for a city",
        "parameters": {
          "type": "object",
          "properties": {"city": {"type": "string"}},
          "required": ["city"]
        }
      }
    }]
  }'

The curated models with tool calling are the Qwen3.5 family and the flagship. A model you import claims a capability only once it has proved it.

Structured output#

Ask for JSON that follows your schema:

Terminal
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $NODEAU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role":"user","content":"Describe a GPU as JSON."}],
    "max_tokens": 2048,
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "gpu",
        "schema": {
          "type": "object",
          "properties": {"name": {"type":"string"}, "vram_gb": {"type":"integer"}},
          "required": ["name", "vram_gb"]
        }
      }
    }
  }'

Give structured output a budget of at least 1024 tokens. A reasoning model thinks first, and the JSON can't start until the thinking is done.

Image input#

Give a model an image and ask about what it sees. The two Gemma 4 models in the catalog read images: gemma-4-e4b-qat-q4-0 on the 8 GB rung and gemma-4-12b-qat-q4-0 on the 12 GB rung.

Terminal
nodeau run gemma-4-e4b-qat-q4-0

Images go in the ordinary OpenAI content-parts shape, as a base64 data URL:

Python
import base64, os
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8080/v1",
                api_key=os.environ["NODEAU_API_KEY"])

with open("photo.png", "rb") as f:
    image = base64.b64encode(f.read()).decode()

reply = client.chat.completions.create(
    model="gemma-4-e4b-qat-q4-0",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this picture?"},
            {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image}"}},
        ],
    }],
    max_tokens=2048,
)
print(reply.choices[0].message.content)

These docs describe the current release, the beta channel. Spotted something wrong or missing? Tell us. A report from a machine we have never seen is the most useful thing we get.