Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

LLMhop

One port, many models: A tiny, stateless HTTP router for OpenAI-compatible LLM inference backends.

LLMhop peeks at the model field of an incoming OpenAI-compatible request and reverse-proxies it to the matching backend. It is primarily designed for single-model inference servers like vLLM and sglang that serve one model per process and need a thin model-aware gateway in front of them, but it works with any OpenAI-compatible backend (including multi-model servers and hosted providers) whenever you want to consolidate several upstreams behind a single endpoint.

Features

  • OpenAI-compatible reverse proxy, model router and request dispatcher for self-hosted LLM inference.
  • Path-based routing under /route/{model}/ for clients and services that cannot put a model field into the body.
  • Native GET /v1/models and GET /v1/models/{model} endpoints served directly from the config, so clients can discover every backend behind the single endpoint.
  • Unauthenticated GET /health for load balancers and probes, plus sd_notify readiness so systemd reports the service as started only once the port answers.
  • Stateless single-binary HTTP service: no database, no cache, no background workers, safe behind any load balancer.
  • Zero external dependencies: pure Go, no third-party packages, no CGO.
  • Works with any OpenAI API-compatible backend, self-hosted or remote: vLLM, sglang, TabbyAPI, Aphrodite, Ollama, LocalAI, OpenRouter, together.ai, DeepInfra, etc.
  • Ships as a static binary, a minimal Docker image and a hardened NixOS module that can optionally spin up llama.cpp, sglang or vLLM workers alongside the router.

How it works

  1. Client sends a request with a JSON body containing {"model": "..."}.
  2. LLMhop reads the model field and looks it up in its config.
  3. The request is forwarded verbatim to the configured backend URL.
  4. Unknown models return 404.

GET /v1/models and GET /v1/models/{model} are answered by LLMhop itself from the configured models, never proxied, so the catalog reflects exactly what clients may ask for. Everything else is dispatched by its model field as above, unless it is path routed. Invalid JSON request bodies return 400 Bad Request. Missing or non-string model values also return 400 Bad Request. When authTokens is set, all routes (the models API included) require a valid bearer token.

Path routing

A request under /route/{model}/ selects its backend by path instead of body. LLMhop strips the /route/{model} prefix and forwards the rest unchanged (method, query, headers and body), so POST /route/production/detect reaches the production backend as POST /detect. The body is never parsed or rewritten, so it need not be JSON or carry a model field. Unknown models return 404, and authentication, request limits, header injection and unix socket upstreams apply exactly as for body routing. A model name containing / is written as %2F, as in /route/Qwen%2FQwen3-8B/v1/chat/completions.

curl $LLMHOP/route/production/detect -d '{"text": "..."}'

A backend marked "unlisted": true is routed like any other but left out of both model endpoints. That is for services that are not inference models and should not look like one, such as the watermark detector below: they still reach clients through LLMhop’s listener, bearer tokens and header injection, but never show up as something to send a completion to.

Health

GET /health is served by LLMhop itself and is the one route that never requires a token, so probes and load balancers do not need a credential:

{ "status": "ok", "models": 3 }

The model count lets a downstream check assert that the proxy came up with the catalog it expects, not merely that the process is listening. It counts exactly what GET /v1/models advertises, so unlisted backends are excluded. Under systemd the same guarantee comes for free: LLMhop sends READY=1 only after the listener is bound, so a Type=notify unit stays in activating until requests are actually served.

Authentication

LLMhop can optionally gate incoming requests with a list of bearer tokens and inject per-model Authorization (or any other) headers when forwarding to the backend. Both sides are opt-in: leave authTokens and models.*.headers unset and headers are forwarded verbatim.

When authTokens is set, the router validates the incoming Authorization: Bearer <token> header (constant-time compare) and then strips it before forwarding, so the client-facing token never leaks upstream. Per-model headers are applied last, so a configured Authorization always wins over whatever the client sent.

Configuration

Create a config.json:

{
  "host": "127.0.0.1",
  "port": 8080,
  "authTokens": ["${cred:llmhop.client-token}"],
  "models": {
    "llama-3-8b": {
      "url": "http://localhost:30000"
    },
    "qwen3-8b": {
      "url": "unix:///run/llmhop/vllm-qwen3-8b/http.sock"
    },
    "openai-gpt-4o": {
      "url": "https://api.openai.com",
      "headers": {
        "Authorization": "Bearer ${cred:openai-key}"
      }
    }
  }
}

host defaults to every interface and port to 8080. IPv6 literals are written plain ("host": "::1") and bracketed internally.

A model url is either an absolute http(s) URL or unix:///<socket path>. A socket URL carries no path prefix, so requests go to the root of the server listening on it.

Each model additionally takes "unlisted": true, which keeps the backend routable by name, in the body or the path, while hiding it from GET /v1/models and GET /v1/models/{model}.

Secret references

String values inside authTokens and models.*.headers are expanded at startup, so no plaintext secret ever has to live in the config file:

  • ${cred:name}: read the systemd credential name from $CREDENTIALS_DIRECTORY, where systemd puts it. This is the same reference the NixOS module rewrites for the model backends, so one spelling covers every service: those servers receive the credential’s path because they open the file themselves, while llmhop reads its own config and so receives its contents.
  • ${env:NAME}: read from the NAME environment variable.
  • ${file:/absolute/path}: read from a file llmhop is pointed at directly. The path must be absolute, since credentials are addressed by name with ${cred:name}.
  • $$: a literal $, so $${env:NAME} stays as written. A bare $NAME is rejected rather than read from the environment.

A single trailing newline is trimmed from a file’s contents.

Unresolved references are a hard startup error.

Validation

Unknown keys are rejected rather than ignored, so a misspelled maxBodyBytes fails loudly instead of silently falling back to its default, and every model url must be an absolute http(s) or unix URL.

--check runs the full startup path (parsing, validation, router construction) and exits without binding a port:

llmhop --check --config config.json

Secret references are left unexpanded in this mode, so a config can be validated where the referenced files and environment variables do not exist, such as a CI job or a Nix build. The NixOS module uses exactly this to validate the generated config at build time.

Request limits

LLMhop buffers each request body in memory so it can peek at the model field before forwarding. Path-routed bodies are streamed instead, since the backend is known from the path. To keep a single request from exhausting memory, the body is capped at 100 MiB by default, whether buffered or streamed. Bodies beyond the cap are rejected with 413 Request Entity Too Large. A declared Content-Length above the cap is rejected before reading the body. At most 8 proxied requests are active at once by default. Additional requests receive 503 Service Unavailable immediately. Set either limit to 0 to disable it, or adjust both for larger multimodal payloads and expected concurrency:

{ "maxBodyBytes": 524288000, "maxConcurrentRequests": 4 }

Running

# native
llmhop --config config.json

# nix
nix run github:mirkolenz/llmhop -- --config config.json

# docker
docker run --rm -p 8080:8080 -v ./config.json:/config.json ghcr.io/mirkolenz/llmhop --config /config.json

NixOS module

A hardened systemd service is provided out of the box. Add LLMhop to your flake inputs and import the module into your system configuration:

{
  inputs = {
    nixpkgs.url = "github:nixos/nixpkgs/nixos-unstable";
    llmhop = {
      url = "github:mirkolenz/llmhop";
      inputs.nixpkgs.follows = "nixpkgs";
    };
  };
  outputs =
    { nixpkgs, llmhop, ... }:
    {
      nixosConfigurations.myhost = nixpkgs.lib.nixosSystem {
        system = "x86_64-linux";
        modules = [
          llmhop.nixosModules.default
          {
            services.llmhop = {
              enable = true;
              port = 8080;
              openFirewall = true;
              settings.models = {
                "llama-3-8b".url = "http://localhost:30000";
                "qwen-2.5-7b".url = "http://localhost:30001";
              };
            };
          }
        ];
      };
    };
}

The unit runs as the llmhop system user with aggressive sandboxing (ProtectSystem, PrivateTmp, restricted syscalls and address families, no new privileges, …) and restarts on failure.

The module and the binary are deliberately coupled in these places:

  • Every listener is a llmhop-<name>.socket unit, handed to the one service through socket activation, so llmhop binds nothing itself and runs with SocketBindDeny=any. The top-level port, host, socket, socketUser, socketGroup and socketMode options define the default listener, and each entry of listen adds another one with the same options.
  • A listener with a port binds host:port, with host an IP literal. The port joins the same global port registry the inference backends use, so a backend model reusing it fails evaluation instead of leaving one of the two services unable to bind, and openFirewall opens it. The default listener keeps port = 8080 unless changed.
  • A listener without a port is a unix socket at socket (default <socketDirectory>/<name>.sock), for a reverse proxy on the same host. systemd applies its socketUser, socketGroup and socketMode and removes the file on stop, so listen.caddy.socketGroup = "caddy" grants Caddy access. TCP and socket listeners can be combined freely.
  • The generated config is validated at build time by the binary itself (llmhop -check), so a typo or a malformed model URL fails nixos-rebuild rather than the service. The schema therefore lives in exactly one place, the Go Config struct, instead of being mirrored in Nix. Validation is skipped when the target platform cannot be executed by the build machine (cross-compiled deployments).
  • The unit is Type=notify, matching the binary’s readiness signal, so anything ordered after llmhop.service can assume it serves.
  • The same package ships llmhop-notify, which the native backends prefix to every model server’s command line. None of them speak sd_notify, so it polls /health and reports readiness on their behalf, letting the worker units be Type=notify too. It stays the unit’s main process and exits with the server’s status, so a model that dies while loading fails its unit immediately instead of being waited out until TimeoutStartSec.

The NixOS module is split into two exports. nixosModules.default ships the reverse proxy and the native systemd backends (llama.cpp, and vLLM and SGLang from prebuilt wheels), with no dependency on quadlet-nix, so it stays compatible with non-NixOS deployers such as system-manager. nixosModules.quadlet includes all of that and additionally provides the container variants of llama.cpp, vLLM, and SGLang, pulling in the quadlet-nix dependency they require. Import the latter only if you need llama-cpp-quadlet, vllm-quadlet, or sglang-quadlet.

Inference backends

The module can also run the inference servers themselves, so you don’t have to wire up llama.cpp, sglang or vLLM by hand. Each backend exposes a models attrset under services.llmhop.<backend> and every entry becomes one isolated worker bound to a unix socket or a loopback port, with the matching route registered automatically with llmhop. All three backends can be enabled side by side and mixed freely in the same configuration.

Every native worker stays in activating until its server answers /health, so systemctl start <backend>-<model> returns only once the model is actually servable rather than merely spawned. Cold starts download weights and profile the GPU, so that wait can be long: TimeoutStartSec allows an hour. The container variants get the same guarantee from their Notify=healthy health check.

llama.cpp runs as a native, hardened systemd system unit under DynamicUser, and the default vllm and sglang backends run the same way from prebuilt wheels, except under a dedicated system user (see below). All three engines can instead run as Podman containers through quadlet-nix, via the suffixed llama-cpp-quadlet, vllm-quadlet, and sglang-quadlet options. They are rootful system units by default, matching Quadlet itself and requiring no host UID configuration. Set quadlet.user to run them as rootless systemd user units instead. The module can create a dedicated lingering account, or target an account managed elsewhere.

The dedicated user mode remains useful for NVIDIA systems affected by NVIDIA/nvidia-container-toolkit#648. nvidia-cdi-hook runs as an OCI createContainer hook inside the container’s user namespace and can fail to read the OCI bundle’s config.json with some UID-mapped namespaces. Running the Quadlet under a real user’s systemd manager avoids that system-manager launch path while retaining rootless Podman.

Native workers (llama-cpp, vllm, sglang) are plain system units, so they are managed with the usual systemctl status <backend>-<model> and journalctl -u <backend>-<model>. The container variants are managed with the quadletctl command of quadlet-nix, which finds the manager of each unit on its own, whether rootful or rootless, for example quadletctl systemctl status <backend>-<model>, quadletctl journalctl <backend>-<model> -f, or quadletctl podman <backend>-<model> ps. quadletctl list shows all units with their owner and state, and quadletctl shell <backend>-<model> opens a shell as the owner of a unit.

services.llmhop = {
  enable = true;
  llama-cpp = {
    enable = true;
    models."qwen3-8b".settings.hf-repo = "unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL";
  };
  sglang = {
    enable = true;
    package = inputs.llmhop.legacyPackages.${pkgs.system}.mkUvEnv { workspaceRoot = ./sglang-env; };
    models."qwen3-coder" = {
      port = 19001;
      model = "Qwen/Qwen3-8B";
      settings.reasoning-parser = "qwen3";
    };
  };
  vllm = {
    enable = true;
    package = inputs.llmhop.legacyPackages.${pkgs.system}.mkUvEnv { workspaceRoot = ./vllm-env; };
    models."llama-3-8b".model = "meta-llama/Meta-Llama-3-8B-Instruct";
  };
};

See the options reference for the full list of per-backend options.

Listeners

llama.cpp and vLLM workers, and native watermark detectors, default to port = null, which binds the unix socket <socketDirectory>/<unit>/http.sock instead of a TCP port. services.llmhop.socketDirectory defaults to /run/llmhop, shared with llmhop’s own socket listeners, and must stay below /run, since each backend directory is a RuntimeDirectory= of its unit. Sockets claim nothing from the host’s port space, and only llmhop and the backend’s own workers can connect to them, instead of every local process. Access is granted to llmhop through the backend’s group, which it joins via services.llmhop.supplementaryGroups, or through a default ACL for services.llmhop.user on Quadlet sockets. The read-only socket option of each model holds the resulting path. Set port to bind 127.0.0.1:<port> instead, for example to reach a worker without llmhop. SGLang and Quadlet detectors have no socket support upstream, so they always take a port.

Containers see their socket directory at /run/llmhop/socket, mounted with Podman’s U option, so sockets work with any User= and any UserNS=, auto included. quadlet.mountOptions.socket replaces those options, for example with [ "U" "z" ] on SELinux hosts, beside quadlet.mountOptions.credentials for the credential mount.

Settings rendering

modelSettings and settings are rendered into the model server’s own CLI flags. Values keep their Nix type: strings, paths and store paths are passed verbatim, and anything else is serialised to JSON, so 0.6 stays 0.6 and an attribute set becomes the JSON object that options like vLLM’s --speculative-config parse.

  • true collapses to --<key>.
  • null and empty lists are dropped.
  • llama.cpp and vLLM render false as --no-<key>, because their parsers register a negated twin for every boolean. A flag with no such twin (an on-only one, or a tri-state one taking on|off|auto) has to be omitted or given its value explicitly rather than set to false.
  • SGLang drops false instead, since its CLI pairs --enable-X with --disable-X rather than auto-negating. Write the negated key explicitly, for example disable-radix-cache = true;.
  • vLLM and SGLang hand every element of a list to one flag (--<key> a b), which is what most of their multi-value options take. The few that expect a repeated flag have to be written out one value at a time.
  • llama.cpp repeats the flag once per element (--<key> a --<key> b), all its hand-rolled parser understands.

Quadlet execution and user namespaces

The Quadlet backends separate the host account that invokes Podman from the identity used inside each container. With no quadlet.user, quadlet-nix installs system units and Podman runs rootfully:

services.llmhop.vllm-quadlet = {
  enable = true;
  tag = "latest";
  models."qwen3-8b".model = "Qwen/Qwen3-8B";
};

llama.cpp uses the same interface, but selects one of the upstream server image variants and configures the model through llama-server flags:

services.llmhop.llama-cpp-quadlet = {
  enable = true;
  tag = "server-cuda";
  models."qwen3-8b".settings.hf-repo = "unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL";
};

To keep the previous dedicated-user layout, set a positive UID. llmhop creates the matching system user and group, enables linger, creates its home, and asks NixOS to allocate subordinate IDs:

services.llmhop.vllm-quadlet.quadlet.user.uid = 503;

Every account detail remains configurable:

services.llmhop.vllm-quadlet.quadlet.user = {
  name = "inference";
  uid = 503;
  group = "inference";
  gid = 503;
  home = "/var/lib/inference";
};

The subordinate ID ranges are NixOS-allocated by default. Pick them yourself through the native user options, which merge with what the module sets:

users.users.inference = {
  autoSubUidGidRange = false;
  subUidRanges = [
    {
      startUid = 300000;
      count = 65536;
    }
  ];
  subGidRanges = [
    {
      startGid = 300000;
      count = 65536;
    }
  ];
};

Set manage = false to select an existing account without changing it:

services.llmhop.vllm-quadlet.quadlet.user = {
  manage = false;
  name = "inference";
  uid = 1000;
  group = "inference";
};

Native Quadlet sections are exposed at backend level and on each model. Backend settings apply to every generated container, then model settings override individual keys:

services.llmhop.vllm-quadlet = {
  quadlet.containerConfig = {
    User = "1000";
    UserNS = "auto:size=65536";
  };

  models."qwen3-8b".quadlet.containerConfig.GroupAdd = [ "keep-groups" ];
};

containerConfig accepts every upstream [Container] key, including User, Group, GroupAdd, UserNS, UIDMap, GIDMap, SubUIDMap, SubGIDMap, PodmanArgs, and GlobalArgs. serviceConfig, unitConfig, quadletConfig, and extraConfig provide the corresponding systemd and Quadlet escape hatches. The SGLang gateway has the same options under gateway.quadlet.

The Hugging Face cache mount is configured separately because host ownership depends on the selected mapping:

services.llmhop.vllm-quadlet.cache = {
  directory = "/var/cache/vllm";
  containerDirectory = "/cache/huggingface";
  user = "100000";
  group = "100000";
  mountOptions = [ "idmap" ];
};

Set cache.manage = false when another module owns the host directory. Podman validates incompatible namespace combinations during the build, using the same generator that consumes the final unit.

Native vLLM and SGLang from prebuilt wheels

The default vLLM and SGLang backends run as native systemd units built from upstream’s prebuilt wheels: no Podman, and the same sandboxing as the llama.cpp backend. They need a named system user rather than DynamicUser: the /var/lib/private layout DynamicUser implies hands the state and cache directories to the unit as noexec ID-mapped mounts, and these runtimes dlopen kernels they compiled into that cache. Every native backend, like llmhop itself, runs as its user and group, which the module declares while they keep their default names. vLLM, SGLang and llmhop default to an account named after them, optionally pinned with uid and gid. llama.cpp defaults to user = null, a DynamicUser per worker, and runs as the named user instead once user is set. vLLM and SGLang lean heavily on dev snapshots and architecture-specific builds, so there is no one-derivation-fits-all version, and you pin yours in a tiny uv workspace and build the package with the flake’s mkUvEnv helper.

# vllm-env/pyproject.toml — your single version knob; edit and run `uv lock` to follow upstream.
#   [project]
#   name = "vllm-env"
#   requires-python = "==3.12.*"
#   dependencies = [ "vllm==0.16.2" ]   # or a nightly via [tool.uv.sources] / [[tool.uv.index]]

services.llmhop.vllm = {
  enable = true;
  package = inputs.llmhop.legacyPackages.${pkgs.system}.mkUvEnv {
    workspaceRoot = ./vllm-env; # directory holding pyproject.toml + uv.lock
  };
  models."llama-3-8b".model = "meta-llama/Meta-Llama-3-8B-Instruct";
};

The interpreter follows from requires-python: mkUvEnv builds against the lowest one the lock admits, which is what uv resolved against, so nothing names a Python version twice. Pass python to override that, and read the choice back from passthru.python.

mkUvEnv installs the wheels, so no GPU or C++ toolchain runs at build time, and patches them for NixOS by baking the GPU driver runpath into the closure. The driver itself is host state, so enable hardware.graphics and your vendor configuration (hardware.nvidia, the amdgpu kernel driver, …) as usual. services.llmhop.sglang works identically, launched via python -m sglang.launch_server.

Runtime JIT compilers

vLLM and SGLang compile kernels while they serve, through flashinfer, DeepGEMM, tilelang or torch.compile. Those compilers invoke nvcc themselves and look for it on PATH or below /usr/local/cuda unless CUDA_HOME names a prefix, and the CUDA wheels cannot be that prefix: they carry a runtime toolkit, with no libcudart.so namelink and no driver stub to link against. The flake exposes mkCudaHome for the prefix, and the unit takes it through environment:

services.llmhop.vllm.environment.CUDA_HOME = "${
  inputs.llmhop.legacyPackages.${pkgs.system}.mkCudaHome {
    packages = with pkgs.${config.services.llmhop.vllm.package.cudaPackagesAttr}; [
      cuda_nvcc
      cuda_cudart
      cuda_crt
      cccl
      cuda_nvrtc
      cuda_cuobjdump
      libcublas
      libcurand
    ];
  }
}";

The CUDA line is not named a second time: the environment exposes passthru.cudaPackagesAttr, the cudaPackages_* attribute matching the line its own wheels were built for, taken from the nvidia-cuda-runtime the lock resolved. It names an attribute rather than holding a package set, the way python.pythonAttr does in nixpkgs, so the packages come from your pkgs, with your configuration and overlays, rather than from this flake’s. One uv lock therefore moves the wheels and the prefix together, and a CUDA line nixpkgs has not packaged fails evaluation by name. Components drift within a line, libcublas running ahead of cuda_cudart, so the runtime is the anchor rather than any single wheel.

The helper owns only the layout: it merges every output but static, since symlinkJoin links outputs and drops the propagation that would otherwise carry the headers, and it adds the lib64 and lib64/stubs paths that some compilers link against and that only the retired runfile installer ever produced. The package list is yours, the same way buildInputs is: it follows from which JIT paths your models take, and it has to stay on the CUDA line the wheels were built for.

GPUs other than NVIDIA

Nothing in the module is CUDA-specific. Every GPU worker joins the render and video groups inside a PrivateUsers = "identity" namespace that keeps those group IDs intact, which is what opening /dev/kfd and /dev/dri/renderD* takes. Runtime kernel caches are redirected into the unit’s cache root for every stack at once, and the NCCL_* defaults cover AMD too, since RCCL reads the same variables.

Only the wheels differ per vendor:

StackNative (uv) backendsWheels
NVIDIA CUDAvLLM, SGLangPublished on PyPI, so a plain vllm==<version> pin resolves them.
AMD ROCmvLLMPublished per ROCm release at https://wheels.vllm.ai/rocm/<version>/<rocm>. SGLang’s are still landing upstream.
Intel XPUvLLMPublished per release at https://wheels.vllm.ai/<version>/xpu, and they need torch’s own XPU index alongside.

The index URLs move with every release, so take them from vLLM’s installation docs rather than from here. A non-CUDA workspace differs from a CUDA one only in where the wheel comes from:

# vllm-env/pyproject.toml — the ROCm version is part of both the pin and the index URL.
[project]
name = "vllm-env"
requires-python = "==3.12.*"
dependencies = [ "vllm==0.30.0" ]

[[tool.uv.index]]
name = "vllm-rocm"
url = "https://wheels.vllm.ai/rocm/0.30.0/rocm723"
explicit = true

[tool.uv.sources]
vllm = { index = "vllm-rocm" }

An XPU workspace takes the same shape with the XPU index, plus https://download.pytorch.org/whl/xpu for torch. Upstream reaches that second index with --index-strategy unsafe-best-match, whose lockfile equivalent is index-strategy = "unsafe-best-match" under [tool.uv].

These wheels carry a matched ROCm or oneAPI build of torch, so the host contributes only the kernel driver. Expect a different set of missing native libraries than a CUDA workspace. The userspace driver of each stack is already in the venvDriverLibs default.

Missing build systems

Not every dependency ships a wheel. The few that resolve to an sdist are built from source, and pre-PEP-517 projects that assume setuptools is simply present fail the build with No module named 'setuptools' or The build backend returned an error. Declare what they need in the workspace rather than patching the Nix side, so uv and mkUvEnv read it from the same place:

# sglang-env/pyproject.toml — SGLang reaches antlr4 through omegaconf.
[tool.uv.extra-build-dependencies]
antlr4-python3-runtime = ["setuptools"]

Re-run uv lock afterwards. The key is the package name as it appears in uv.lock, and the value is whatever its build backend needs (setuptools, cython, meson-python, …).

Versions upstream leaves unconstrained

Some wheels have to move together although none of them depends on the others. vLLM’s flashinfer payloads are the recurring case: flashinfer-cubin and flashinfer-jit-cache have to match the flashinfer-python that vLLM pins by hand in requirements/cuda.txt, and flashinfer refuses to start otherwise.

Pin that package yourself, at the version the other two use, rather than leaving it to vLLM alone:

dependencies = [
  "vllm==0.30.0",
  "flashinfer-python==0.6.18",     # vLLM pins this exactly, so a bump conflicts here
  "flashinfer-cubin==0.6.18",
  "flashinfer-jit-cache==0.6.18",
]

The next uv lock after vLLM moves its pin then fails to resolve, naming both versions, instead of producing a lock that builds and dies on the GPU host.

Missing native libraries

Wheels are built for manylinux and expect a distro underneath them. Which libraries a workspace needs beyond the driver follows from what it locks, so there are no defaults: you supply them per workspace through buildInputs and runtimePaths, which are merged into every wheel but pure -any ones. nativeBuildInputs is accepted alongside them for build-time tooling an sdist needs beyond its Python build backend.

buildInputs covers libraries a wheel names in a DT_NEEDED entry. They are added to the autoPatchelf search path, so a library only lands in the runpath of a wheel that actually links it and listing one nothing needs is harmless. No wheel fails over an unresolved entry on its own, because in isolation it cannot see the siblings it will share a venv with: half of what a wheel misses at that point is another wheel. The assembled environment is checked instead, where those have resolved, and every entry still unresolved there fails the build:

mkUvEnv: unresolved library libtbb.so.12, needed by lib/python3.12/site-packages/numba/np/ufunc/tbbpool...so
mkUvEnv: supply these through `buildInputs`, list them in `venvDriverLibs` if the host provides them, or list the libraries needing them in `venvOptionalLibs` if those load on demand only.

A soname that some file in the environment carries counts as resolved, even when ldd cannot reach it from the library that needs it. Wheels rarely link their siblings through the runpath: torch preloads the CUDA wheels on import, and torchcodec expects torch to be loaded already, so the loader finds both by soname. Every other entry is either a library nixpkgs should supply, which goes into buildInputs:

buildInputs = [
  pkgs.ffmpeg_8-headless # torchcodec supports FFmpeg 4 to 8, not the default 9
  pkgs.tbb_2022          # numba's threading layer, plain `tbb` is too old for libtbb.so.12
];

or a soname only the host provides at runtime, which goes into venvDriverLibs:

venvDriverLibs = [ "libcuda.so*" "libnvidia-*.so*" ];   # NVIDIA's userspace driver

venvDriverLibs defaults to the userspace driver of every stack vLLM publishes wheels for, NVIDIA, ROCm and XPU alike, since no build can resolve those anywhere. Setting it replaces that default.

Or it is a dependency of a library that loads on demand only, which goes into venvOptionalLibs as a glob of that library’s path relative to the environment root:

venvOptionalLibs = [
  "*/torchcodec/libtorchcodec_*[!8].so"    # variants for the FFmpeg majors not supplied
  "*/nvshmem_bootstrap_mpi.so.3"           # nvshmem plugins for an MPI launcher
];

The first build of a workspace names what it found, as does every later one the moment a wheel starts wanting something new, or nixpkgs moves a library to a soname the wheels were not built against.

Wheels tagged for any platform skip the ELF fixups, since patching the thousands of cubins in flashinfer-cubin would take longer than the rest of the environment. Should one ship host code all the same, the environment fails to build and names it, and hostWheels gives it the fixups back:

hostWheels = [ "some-wheel" ];

runtimePaths covers the other kind, reached by a bare dlopen("libfoo.so") from Python via cffi or ctypes. Nothing announces those in the ELF, so no build ever fails over one and no runpath resolves it; the environment builds cleanly and the import dies:

OSError: cannot load library 'libsndfile.so': cannot open shared object file

Only importing finds them, so run the modules you care about once after a version bump. Entries here are appended to the runpath of every wheel rather than a chosen one, because the object issuing the dlopen is generally not the package that appears in the traceback — soundfile fails, but the call comes from cffi’s _cffi_backend:

runtimePaths = [ "${pkgs.lib.getLib pkgs.libsndfile}/lib" ];   # soundfile, reached through cffi

NCCL reaches InfiniBand the same way: it dlopens libibverbs by bare name, and when that fails it falls back to TCP sockets without an error. A multi-node workspace therefore needs rdma-core here, besides buildInputs for the nvshmem and cuFile transports that link it:

buildInputs = [ pkgs.rdma-core ];                               # nvshmem, cuFile
runtimePaths = [ "${pkgs.lib.getLib pkgs.rdma-core}/lib" ];     # NCCL

Each model defaults to the backend’s package but can pin its own with models.<name>.package, so a single model can follow a nightly build for a freshly-released architecture while the rest stay on the stable pin.

Because a unit only goes active once it is healthy, startupOrdering (on by default) is effective here: workers boot one at a time in ascending name order, each finishing its GPU-memory profiling before the next begins, which is what keeps two models sharing a device from racing into an OOM.

The container variants live under services.llmhop.llama-cpp-quadlet, vllm-quadlet, and sglang-quadlet. A backend’s native and container variants emit the same <backend>-<model> units and are therefore mutually exclusive, so enable at most one variant per backend.

Inference server credentials

Credentials are granted to individual model services, never inherited from a backend. This keeps one compromised worker from reading another model’s keys. Assign the same Nix value to multiple models when sharing is intentional.

services.llmhop.llama-cpp.models."qwen3-8b" = {
  credentials.api-keys = "/run/secrets/qwen-api-keys";
  settings = {
    hf-repo = "unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL";
    api-key-file = "\${cred:api-keys}";
  };
};

credentials.<name> takes one of three sources:

credentials = {
  # A file, loaded with `LoadCredential=`.
  api-keys = "/run/secrets/qwen-api-keys";
  # An encrypted file, loaded with `LoadCredentialEncrypted=`.
  tls-key = {
    source = "/run/secrets/qwen-tls-key.cred";
    encrypted = true;
  };
  # The credential of the same name in the system credential store, imported with `ImportCredential=`.
  "llmhop.hf-token" = { };
};

Sources must lie outside the Nix store, which every local user can read, so a store path fails evaluation. An imported credential is looked up in /etc/credstore, /etc/credstore.encrypted and the other credential store directories, and among the credentials passed to the system, and decrypted as needed. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token or vllm.watermark-key. It needs no path in Nix at all, so it is the simplest way to provision a secret, as a plain file in the root-only /etc/credstore:

(umask 077; systemd-ask-password -n > /etc/credstore/llmhop.hf-token)

To keep it encrypted at rest with the host key, and the TPM if present, write it to /etc/credstore.encrypted instead:

systemd-ask-password -n | systemd-creds encrypt --name=llmhop.hf-token - /etc/credstore.encrypted/llmhop.hf-token

systemd-creds embeds the name into an encrypted credential, and systemd refuses to load it under any other, so encrypt with --name=<name>. For rootless Quadlets, the selected user’s systemd manager must be able to read the source, and its encrypted credentials must be created with systemd-creds encrypt --user.

${cred:name} expands to the read-only credential path, not its contents. Unknown references fail evaluation. Native workers use their systemd credential directory through the %d specifier. Quadlet workers mount only that unit’s credential directory at /run/llmhop/credentials.

vLLM and SGLang also accept a per-model YAML file through their --config flag. That is an ordinary file-taking setting, so it needs no dedicated option:

services.llmhop.vllm-quadlet.models."qwen3-8b" = {
  model = "Qwen/Qwen3-8B";
  credentials."config.yaml" = "/run/secrets/qwen-vllm.yaml";
  settings.config = "\${cred:config.yaml}";
};

vLLM requires the .yaml or .yml extension, which the credential name carries into its path.

llama.cpp has no general server config file, so use its file-taking settings such as api-key-file, ssl-key-file, and ssl-cert-file instead.

The default root user in rootful containers and the mapped root user in rootless containers can read the credential mount. If [Container] User= selects another identity, its user namespace mapping must map the systemd unit owner to that container UID. The module does not weaken credential modes to make an incompatible mapping work. Use an idmapped credential mount when the selected namespace does not already provide that mapping:

models."qwen3-8b".quadlet.mountOptions.credentials = [
  "idmap=uids=0-1000-1;gids=0-1000-1"
];

Literal values remain available for non-secret development configuration:

services.llmhop.sglang.models.test.settings.api-key = "development-only";

Such values enter the world-readable Nix store and should not be used for production secrets.

vLLM watermarking

Recent vLLM revisions include watermark generation and detector primitives, but the OpenAI server does not expose the detector as an endpoint. Both vLLM modules can watermark a model’s output and run a small detector server as a named service next to their model workers. Both sides take the same watermark option, an attribute set of vLLM’s WatermarkConfig fields without the key:

services.llmhop.vllm = let
  watermark.algorithm = "dual_key_gumbel";
in {
  models.qwen = {
    model = "Qwen/Qwen3-8B";
    inherit watermark;
  };
  detectors.production = {
    tokenizer = "Qwen/Qwen3-8B";
    inherit watermark;
    settings.p-value-threshold = 0.01;
  };
};

The key is the credential vllm.watermark-key, an unsigned 64-bit integer that every service with a watermark imports from the system credential store:

(umask 077; od -An -N8 -tu8 /dev/urandom | tr -d ' \n' > /etc/credstore/vllm.watermark-key)

An encrypted credential works the same, under the same name.

To load it from elsewhere, give the services the credential through any of the credential sources, e.g. credentials."vllm.watermark-key" = "/run/secrets/watermark-key";. That is also how two models use different keys.

The key never enters the Nix store, the command line, or the journal. nix/pkgs/mkVllmWatermark reads it from the credential: its serve command merges the key into the --watermark-config of vllm serve in-process, and its detect command runs the detector. vLLM redacts the configuration in its own logs, and validation errors omit their input. Setting watermark-config in settings or key in watermark fails evaluation.

Both commands validate the configuration with vLLM’s own WatermarkConfig, so every field and algorithm the installed release accepts is accepted here too, and a typo fails at startup rather than silently changing the watermark. Speculative decoding needs dual_key_gumbel, or allow_target_only_watermarking = true for the others. vLLM offers no detector factory, so the detector picks its class and arguments by vLLM’s naming conventions. An algorithm added upstream works without changes when it follows them, and fails loudly rather than detecting with mismatched parameters otherwise. p-value-threshold is the only detection-side setting.

tokenizer and watermark must match the generation side. tokenizer and the listener are options of their own, since llmhop owns the latter. Each detector is a separate vllm-detector-<name> service serving POST /detect, with readiness taken from GET /health. They do not receive a GPU device in Quadlet mode.

The script lives in nix/pkgs/mkVllmWatermark rather than coming from vLLM’s example, which has no unix socket support, no health endpoint, no file-based key, and covers only gumbel. mkVllmWatermark, exposed under legacyPackages, checks it against a vLLM environment at build time: the detector must import and detect with every algorithm the release declares, so a uv lock that breaks the vLLM modules it relies on fails the build rather than the unit. Native services run that checked script with their package. Quadlet containers mount the unchecked script into the selected image and run it with the image’s interpreter, so there the same failure surfaces at startup. A detector’s script replaces it with a server of your own, which receives the same flags after detect. Both require a vLLM revision containing vllm.v1.watermarking, which no release before 0.30.0 has.

Routing detectors through llmhop

A detector binds only to a unix socket or host loopback and is registered with llmhop, so clients reach it at llmhop’s own address under llmhop’s bearer tokens rather than on a second, unauthenticated port. The attribute name is the routing key, so it shares one namespace with every backend’s model names and a collision fails evaluation.

Detectors are registered unlisted, so they never appear in GET /v1/models. Select one by path routing, so the body carries only the candidate text:

curl https://llmhop.example.com/route/production/detect \
  -H "Authorization: Bearer $LLMHOP_TOKEN" \
  -d '{"text":"candidate text"}'

The response contains score, p_value, num_scored_tokens, and is_watermarked.

This puts llmhop’s authentication in front of a server that has none of its own. A detector given a port is still unauthenticated there, like every model worker given one, so anything else on the host can reach it directly. On its default socket, only llmhop can.

SGLang and llama.cpp workers can coexist with the detector, but they only produce detectable text if they implement the same watermark generation algorithm and parameters.

llmhop credentials

The generated config file lives in the world-readable Nix store, so secrets should never be placed in services.llmhop.settings directly. Instead, reference them via ${cred:...} and hand them to llmhop through its credentials option, which takes the same three sources as the model backends: a path for systemd’s LoadCredential=, an entry with encrypted = true for LoadCredentialEncrypted=, and { } for ImportCredential= from the system credential store. A path is any file outside the Nix store, so anything that produces one works: agenix or sops-nix outputs, a manually-managed file, or a path emitted by your own secret-provisioning tool.

services.llmhop = {
  credentials = {
    "llmhop.client-token" = { };
    openai-key = "/run/secrets/openai-key";
  };
  settings = {
    authTokens = [ "\${cred:llmhop.client-token}" ];
    models."openai-gpt-4o" = {
      url = "https://api.openai.com";
      headers.Authorization = "Bearer \${cred:openai-key}";
    };
  };
};

${cred:...} references are resolved against $CREDENTIALS_DIRECTORY, which systemd exposes as a per-unit tmpfs accessible only to this service, compatible with DynamicUser and the rest of the sandbox. ${env:...} picks up variables the unit inherits, for example through systemd.services.llmhop.serviceConfig.EnvironmentFile, for secret tooling that only produces environment files.

Module internals

Background for the NixOS module implementation in nix/modules/. None of it is needed to use the module, but it records decisions that are hard to re-derive from the code alone.

CLI dialects

Two settings shapes have no portable rendering, so each backend declares how its parser reads them. The table is keyed by the unsuffixed service name, so a native backend and its Quadlet twin share one entry.

negateBools says the parser registers a --no-<key> twin for every boolean, which argparse’s BooleanOptionalAction and llama.cpp’s paired flags both do. SGLang instead pairs --enable-X with --disable-X and rejects --no-X.

listStyle picks between handing every element to one flag (--key a b), what argparse nargs and clap multi-value options take, and emitting the flag once per element (--key a --key b), all that llama.cpp’s hand-rolled parser understands. llama.cpp also takes only --key value, never --key=value. Both argparse backends register a few options in the other style, so this is the dominant form for a backend rather than a guarantee for every flag.

Neither axis can be delegated to nixpkgs. lib.cli.toCommandLine renders neither: its optionFormat never sees the value, and list handling is hardcoded to repeat-style. The mkBool and mkList hooks of lib.cli.toGNUCommandLine could, but it is deprecated as of nixpkgs 25.11 and warns on every evaluation.

Values are rendered by Nix type: strings, paths and derivations verbatim, everything else through JSON. That keeps 0.6 from becoming 0.600000 (what toString makes of a float) and turns an attribute set into the JSON object that options like vLLM’s --speculative-config parse.

Worker hardening

Native workers start from a universal systemd-exec(5) baseline shared with the llmhop reverse proxy. Quadlet workers skip it, because Podman handles isolation at the container level. SocketBind* is not part of the baseline: it pairs with a per-unit SocketBindAllow that only worker units with a port declare. SocketBind* only governs IP sockets, so a worker on a unix socket keeps SocketBindDeny = "any".

GPU relaxations

  • PrivateDevices = false: GPU acceleration needs raw device access, /dev/nvidia* for CUDA and /dev/kfd plus /dev/dri/renderD* for ROCm and Level Zero. The upstream NVIDIA NixOS modules disable it for the same reason.
  • SupplementaryGroups = [ "render" "video" ]: /dev/nvidia* is world-readable, but systemd’s default udev rules leave the AMD and Intel nodes group-owned, so a worker running as a real user cannot open them without joining those groups.
  • PrivateUsers = "identity": those groups only mean anything if their GIDs survive into the worker’s user namespace, and the baseline’s PrivateUsers = true maps everything but the unit’s own identity to nobody. identity keeps the namespace but maps the first 65536 IDs one-to-one.
  • MemoryDenyWriteExecute = false: runtime kernel compilation mmaps PROT_WRITE|PROT_EXEC pages. torch-inductor and triton, the CUDA driver’s PTX to SASS pass, and the SPIR-V JIT behind SYCL and Level Zero all do it.
  • LimitMEMLOCK = "infinity": page-locked memory draws from RLIMIT_MEMLOCK, which systemd otherwise caps at its 8 MiB default. Every stack pins the host side of its device buffers, and llama.cpp’s --mlock pins the weights outright. Too low a limit reports OOM despite free VRAM.
  • ProcSubset = "all": the baseline’s pid hides everything in /proc that is not a process directory, but psutil, torch and NUMA discovery all read /proc/meminfo and /proc/cpuinfo, so the engine dies before it reaches the GPU. ProtectProc still keeps other users’ process directories invisible.

NCCL relaxations

getifaddrs() opens an AF_NETLINK socket during its interface scan, so that family is re-added to RestrictAddressFamilies. The bootstrap, proxy and RAS listeners bind ephemeral (port 0) TCP sockets, which SocketBindDeny = "any" refuses. The bind hook only ever sees port 0, never the assigned port, so allow-all-TCP is the tightest workable rule, and it already covers the worker’s own listener. UDP stays denied.

NCCL is additionally kept on loopback and off any InfiniBand fabric, since there is none on a single node. RCCL is an API clone of NCCL and reads the same variables, so this covers AMD as well; oneCCL (Intel) uses CCL_* and ignores them.

Caches

Every accelerator stack compiles kernels on first use and caches them next to $HOME, which ProtectSystem = "strict" makes read-only, so each one is redirected into the unit’s own cache root, the single place systemctl clean can reach. All of the variables are set unconditionally: one belonging to a stack that is not installed is never read, which is cheaper than tracking which host has which vendor. MIOpen (ROCm) needs two of them, or its kernel database and its compiled-kernel cache land under separate $HOME roots. SYCL and the Intel compute runtime only cache their SPIR-V to ISA compilation when asked to, which is what turns a multi-minute JIT into a one-time cost across restarts.

Unix sockets

A workload without a port binds http.sock in a directory of its own below services.llmhop.socketDirectory, and llmhop reaches it through a unix:// model URL. There is one directory per unit rather than one flat directory of <unit>.sock files. A flat directory would have to be writable by every worker, letting one worker delete or squat another’s socket and intercept its traffic.

connect() needs search permission on each directory and write permission on the socket. The root is root:root with mode 0711, so anyone may pass through it but not list it. Each directory has mode 0710, search only, since llmhop knows the name. vLLM and llama.cpp bind with the umask and never chmod, so workers run with UMask = "0007" (and containers with Umask=0007), which keeps the socket group-writable. llmhop’s own socket listeners live in the same root, but systemd creates those with the ownership and mode of their listener options.

A native worker’s directory and socket belong to the backend’s group, and llmhop joins every such group through services.llmhop.supplementaryGroups. This is the usual systemd pattern for a client of a socket. An ACL cannot work here, because systemd drops every ACL of an exec directory when it chowns it to a non-root unit. llama.cpp defaults to a DynamicUser with the static group, so its state and cache stay below the 0700 /var/{lib,cache}/private despite the umask. Workers of one backend can therefore reach each other’s sockets, but not those of another backend.

A container user maps to an unpredictable host UID and GID under UserNS=, so no group can be named in advance. Quadlet sockets rely on a default ACL user:<services.llmhop.user>:-wx on the root instead, which every directory and socket below it inherits. llmhop therefore runs as a named system user rather than a DynamicUser, whose UID no ACL could name in advance. The ACL survives, because systemd never chowns these directories: rootful units run as root, and rootless ones use tmpfiles. The kernel masks a new socket’s mode with the umask even below a default ACL, and the ACL mask follows the group bits, so Umask=0007 keeps the ACL’s write. The owning group’s own entry comes from the root’s 0711 and grants no write.

Each directory is a RuntimeDirectory= of its unit, created on start and removed on stop, crashes included. So no stale socket can make a restarted server fail with EADDRINUSE, although neither server removes one itself. A rootful container mounts its directory with U (quadlet.mountOptions.socket), so Podman hands it to whatever host UID the container user maps to under UserNS=, and ACLs survive the chown. The directory mode is a default rather than enforced, so serviceConfig.RuntimeDirectoryMode can still override it.

A user manager keeps RuntimeDirectory= below its own $XDG_RUNTIME_DIR, which llmhop cannot enter. Rootless containers therefore use a tmpfiles directory owned by the Podman account, and an ExecStartPre= clears it. It runs through podman unshare, since after the U chown only the account’s user namespace may delete there.

Python backends built from wheels

The vLLM and SGLang backends run from prebuilt wheels and need more than the hardening baseline.

They get a toolchain on PATH, because these runtimes compile at runtime and look one up the FHS way. Triton builds its CUDA driver shim on the first kernel launch and searches $CC, then gcc/clang on PATH. torch’s cpp_extension and flashinfer’s JIT drive their builds through ninja, which nixpkgs patches to posix_spawnp("sh"), so it needs a shell on PATH rather than at /bin/sh. ctypes falls back to invoking gcc and ld once nixpkgs’ patched ldconfig lookup returns nothing. A unit otherwise has none of them.

HOME is pointed at the cache root because the service user has none, so $HOME would be / and every library that reaches for ~ (flashinfer’s JIT workspace, among others) would hit the read-only root.

LD_LIBRARY_PATH and TRITON_LIBCUDA_PATH are about prebuilt wheels finding host driver libraries rather than about GPUs: a nixpkgs-built worker resolves the same libraries from the runpath autoAddDriverRunpath gave it at build time. mkUvEnv bakes that runpath into the wheels too, so libcuda.so.1 and its ROCm and Level Zero counterparts resolve via RPATH. LD_LIBRARY_PATH additionally covers the host driver libraries the framework dlopens by name from Python during GPU-memory profiling, such as libnvidia-ml.so.1, which RPATH does not reach. Triton locates libcuda.so.1 by shelling out to /sbin/ldconfig -p, which does not exist on NixOS, so its JIT backend dies with a FileNotFoundError the moment a kernel is compiled; TRITON_LIBCUDA_PATH is the upstream escape hatch and short-circuits the lookup entirely.

These backends run as a real system user rather than under DynamicUser. The /var/lib/private layout DynamicUser implies makes systemd hand StateDirectory and CacheDirectory over as ID-mapped mounts, which are unconditionally noexec and beyond the reach of ExecPaths=, and these runtimes compile kernels into that cache and dlopen them back. llama.cpp compiles nothing at runtime and defaults to DynamicUser.

Lifecycle

Model workers restart with Restart = "always" rather than on-failure. vLLM and SGLang catch an EngineCore death, shut the API server down gracefully, and exit 0, so on-failure would leave a crashed worker dead. A shared StartLimitBurst of 3 errors per hour still breaks crash loops, so journald surfaces the underlying error instead of an endless restart. TimeoutStartSec allows an hour, which covers cold-start model downloads plus GPU memory profiling.

Workers are chained by ascending name during startup. GPU-memory profiling races otherwise: two workers booting on the same device each see it as fully free and race to claim their share, leading to OOM. The chain is only meaningful because a worker is held in activating until its server reports itself ready, which llmhop-notify does for native units and Notify=healthy for containers.

llmhop-notify is resolved from this flake rather than from services.llmhop.package: it is a module implementation detail, and a deployer-supplied llmhop build need not ship it at all, in which case getExe' would not catch that and every worker would sit in activating until TimeoutStartSec.

Containers

The container rootfs stays writable. ML runtimes scatter JIT and compile caches across version-dependent HOME paths, so an immutable rootfs would need an ever-growing tmpfs allow-list. /tmp is a tmpfs for fast scratch, which is where torch inductor puts /tmp/torchinductor_root.

Core

services.llmhop.enable

Whether to enable llmhop reverse proxy.

Type: boolean

Default:

false

Example:

true

services.llmhop.package

The llmhop package to use.

Type: package

Default:

pkgs.callPackage ./package.nix { }

services.llmhop.credentials

Credentials granted to llmhop through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.client-token.

Reference them from settings as ${cred:<name>}, the same spelling the model backends use: llmhop reads its own config, so the reference expands to the credential’s contents rather than to its path.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.gid

GID of the declared group. null lets NixOS allocate one.

Type: null or (unsigned integer, meaning >=0)

Default:

config.services.llmhop.uid

services.llmhop.group

Primary group of user. The module declares it while it keeps its default name, any other group is the deployer’s to declare.

Type: string

Default:

config.services.llmhop.user

services.llmhop.host

IP address to bind port to. The default binds every interface, leaving access control to the firewall. IPv6 literals are written plain (e.g. ::1).

Type: string

Default:

""

Example:

"127.0.0.1"

services.llmhop.listen

Addresses llmhop serves besides the default listener, which the top-level port, host, socket, socketUser, socketGroup and socketMode options define. Each is a llmhop-<name>.socket unit handing its socket to llmhop through socket activation. systemd applies the ownership and mode of a unix socket and removes it on stop.

Type: attribute set of (submodule)

Default:

{ }

Example:

{
  caddy = {
    socketGroup = "caddy";
  };
}

services.llmhop.listen.<name>.host

IP address to bind port to. The default binds every interface, leaving access control to the firewall. IPv6 literals are written plain (e.g. ::1).

Type: string

Default:

""

Example:

"127.0.0.1"

services.llmhop.listen.<name>.port

TCP port to listen on, registered in the global port registry so a backend reusing it fails evaluation. null listens on the unix socket socket instead.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

null

services.llmhop.listen.<name>.socket

Unix socket to listen on while port is null.

Type: string

Default:

"${config.services.llmhop.socketDirectory}/‹name›.sock"

services.llmhop.listen.<name>.socketGroup

Group of socket, typically the one of a reverse proxy in front of llmhop.

Type: string

Default:

"root"

Example:

"caddy"

services.llmhop.listen.<name>.socketMode

Mode of socket. Connecting needs write permission.

Type: string

Default:

"0660"

services.llmhop.listen.<name>.socketUser

Owner of socket.

Type: string

Default:

"root"

services.llmhop.openFirewall

Whether to open the port of every TCP listener in the host firewall.

Type: boolean

Default:

false

services.llmhop.port

TCP port to listen on, registered in the global port registry so a backend reusing it fails evaluation. null listens on the unix socket socket instead.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

8080

services.llmhop.settings

Configuration written to the JSON config file passed to llmhop. See the upstream Config struct for available fields. host and port only apply outside socket activation, so the module’s listeners come from the options of the same name and listen instead.

The generated file is validated at build time by the binary itself, so unknown keys and malformed model URLs fail nixos-rebuild rather than the service.

Type: JSON value

Default:

{ }

Example:

{
  models = {
    gpt-4 = {
      url = "https://api.openai.com";
    };
  };
}

services.llmhop.socket

Unix socket to listen on while port is null.

Type: string

Default:

"${config.services.llmhop.socketDirectory}/default.sock"

services.llmhop.socketDirectory

Directory of every unix socket llmhop serves or connects to: the default path of each socket listener, and one directory per workload without a port. Those are RuntimeDirectory=s except under a rootless Quadlet user, hence the /run prefix. Each path component must start with a letter, digit, or underscore and contain only those characters, dots, and hyphens.

Type: string matching the pattern /run(/[[:alnum:]][[:alnum:].-]*)+

Default:

"/run/llmhop"

services.llmhop.socketGroup

Group of socket, typically the one of a reverse proxy in front of llmhop.

Type: string

Default:

"root"

Example:

"caddy"

services.llmhop.socketMode

Mode of socket. Connecting needs write permission.

Type: string

Default:

"0660"

services.llmhop.socketUser

Owner of socket.

Type: string

Default:

"root"

services.llmhop.supplementaryGroups

Groups llmhop joins through SupplementaryGroups=. Every native backend adds its group, which owns the sockets of its workers.

Type: list of string

Default:

[ ]

Example:

[
  "inference"
]

services.llmhop.uid

UID of the declared user. null lets NixOS allocate one.

Type: null or (unsigned integer, meaning >=0)

Default:

null

Example:

503

services.llmhop.user

System user the units run as. The module declares it while it keeps its default name, any other user is the deployer’s to declare.

Type: string

Default:

"llmhop"

llama-cpp

services.llmhop.llama-cpp.enable

Whether to enable llama.cpp model serving via systemd, fronted by llmhop.

Type: boolean

Default:

false

Example:

true

services.llmhop.llama-cpp.package

The llama-cpp package to use.

Type: package

Default:

pkgs.llama-cpp

services.llmhop.llama-cpp.environment

Environment variables set on every model service. Merged with services.llmhop.llama-cpp.models.<name>.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.llama-cpp.environmentFile

File in KEY=VALUE format forwarded to every service. Use only for upstream features that require environment variables, such as HF_TOKEN for gated Hugging Face repositories. Environment variables are not systemd credentials and are visible to every model, so prefer credentials for any secret a server can read from a file. Loaded before services.llmhop.llama-cpp.models.<name>.environmentFile, so per-model files override these entries.

Type: null or absolute path

Default:

null

Example:

"/etc/llama-cpp/.env"

services.llmhop.llama-cpp.group

Static primary group of the units, also beside a DynamicUser. The module declares it while it keeps its default name, any other group is the deployer’s to declare.

Type: string

Default:

"llama-cpp"

services.llmhop.llama-cpp.modelSettings

CLI flags forwarded to the model server for every model. Rendered as --<key> <value>, with false as --no-<key> and a list repeating the flag. See settings rendering for the full rules. Merged with services.llmhop.llama-cpp.models.<name>.settings; per-model entries take precedence.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp.models

Models to serve. Each entry produces one systemd service running llama-server; the attribute name is the routing key surfaced through llmhop and the OpenAI model field.

GPU selection is done via build-specific environment variables on environment (top-level or per-model), since llama.cpp runs as a host process — no CDI involved. Common variables: CUDA_VISIBLE_DEVICES (CUDA), HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES (ROCm), GGML_VK_VISIBLE_DEVICES (Vulkan), ZE_AFFINITY_MASK (SYCL).

Type: attribute set of (submodule)

Default:

{ }

Example:

{
  "qwen3-8b" = {
    settings = {
      hf-repo = "unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL";
      temperature = 1.0;
      top-k = 20;
    };
    # Pin this model to a specific GPU. The right variable depends on
    # the llama.cpp build: CUDA_VISIBLE_DEVICES for CUDA,
    # HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES for ROCm,
    # GGML_VK_VISIBLE_DEVICES for Vulkan, ZE_AFFINITY_MASK for SYCL.
    environment.CUDA_VISIBLE_DEVICES = "0";
  };
}

services.llmhop.llama-cpp.models.<name>.enable

Whether to enable model ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.llama-cpp.models.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.llama-cpp.models.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.llama-cpp.models.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.llama-cpp.models.<name>.environment

Additional environment variables set on this model’s service. Merged with services.llmhop.llama-cpp.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.llama-cpp.models.<name>.environmentFile

File in KEY=VALUE format forwarded to this model’s service. Loaded after services.llmhop.llama-cpp.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.llama-cpp.models.<name>.name

Canonical identifier for this model. Used for the unit name (llama-cpp-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.llama-cpp.models.<name>.port

Loopback host port that llama-server binds to. Must be unique per enabled model; the gateway (llmhop) reaches each backend at http://127.0.0.1:<port>.

null binds the unix socket socket instead, which claims no port and only llmhop can connect to.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

null

services.llmhop.llama-cpp.models.<name>.serviceConfig

Extra [Service] settings merged into this workload’s llama-cpp-<name> unit after the hardened baseline and backend-specific relaxations. The module retains ownership of ExecStart, KillMode, and Type because they implement readiness supervision as one lifecycle contract.

Type: attribute set of anything

Default:

{ }

Example:

{
  MemoryHigh = "64G";
}

services.llmhop.llama-cpp.models.<name>.settings

CLI flags forwarded to the model server for this model. Rendered as --<key> <value>, with false as --no-<key> and a list repeating the flag. See settings rendering for the full rules. Merged with services.llmhop.llama-cpp.modelSettings; per-model entries take precedence. The flags llmhop derives from the model options (its served name and listener) always win over both.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp.models.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/llama-cpp-<name>/http.sock while port is null, else null

services.llmhop.llama-cpp.models.<name>.unitConfig

Extra [Unit] settings merged into this workload’s llama-cpp-<name> unit after the shared baseline.

Ordering and dependency directives (After=, Requires=, Wants=) do not belong here: NixOS renders those from the after, requires and wants options, so a definition of the same key in unitConfig conflicts with it instead of merging. Declare them on systemd.services."llama-cpp-<name>" from your own module, where the module system concatenates them with what this one sets.

Type: attribute set of anything

Default:

{ }

Example:

{
  StartLimitBurst = 10;
}

services.llmhop.llama-cpp.openFilesLimit

File descriptor limit (LimitNOFILE) applied to every llama-cpp systemd unit. Increase if the server logs accept: Too many open files under concurrent load.

Type: positive integer, meaning >0

Default:

1048576

services.llmhop.llama-cpp.user

System user the units run as, which is the deployer’s to declare. null allocates one per unit through DynamicUser=.

Type: null or string

Default:

null

llama-cpp-quadlet

services.llmhop.llama-cpp-quadlet.enable

Whether to enable llama.cpp model serving via Quadlet, fronted by llmhop.

Type: boolean

Default:

false

Example:

true

services.llmhop.llama-cpp-quadlet.cache.containerDirectory

Path at which the cache is mounted inside every model container.

Type: string

Default:

"/root/.cache/llama.cpp"

services.llmhop.llama-cpp-quadlet.cache.directory

Host directory bind-mounted as the Hugging Face cache.

Type: absolute path

Default:

"/var/cache/llama-cpp"

services.llmhop.llama-cpp-quadlet.cache.environmentVariable

Environment variable set on every container to point its runtime at containerDirectory.

Type: string

Default:

"LLAMA_CACHE"

services.llmhop.llama-cpp-quadlet.cache.group

Host group used when cache.manage is enabled.

Type: string

Default:

if config.services.llmhop.llama-cpp-quadlet.quadlet.user == null then
  "root"
else
  config.services.llmhop.llama-cpp-quadlet.quadlet.user.group

services.llmhop.llama-cpp-quadlet.cache.manage

Whether llmhop creates the host cache directory with systemd-tmpfiles, owned by user/group and mode 0700.

Type: boolean

Default:

true

services.llmhop.llama-cpp-quadlet.cache.mountOptions

Options appended to the cache’s Quadlet Volume= entry. This can be used for Podman ownership mechanisms such as U, idmap, or SELinux relabeling.

Type: list of string

Default:

[ ]

Example:

[
  "idmap"
]

services.llmhop.llama-cpp-quadlet.cache.user

Host owner used when cache.manage is enabled. Override it when a UIDMap/idmap mapping makes the container see a different owner than the host account running Podman.

Type: string

Default:

if config.services.llmhop.llama-cpp-quadlet.quadlet.user == null then
  "root"
else
  config.services.llmhop.llama-cpp-quadlet.quadlet.user.name

services.llmhop.llama-cpp-quadlet.devices

Devices exposed to every model container — passed verbatim as Quadlet AddDevice= lines. Accepts both CDI references (recommended: nvidia.com/gpu=…, amd.com/gpu=…, intel.com/gpu=…, …) and raw host device paths (e.g. /dev/dri/renderD128). For CDI, the corresponding spec must be generated on the host (e.g. nvidia-ctk cdi generate). Defaults to [ "nvidia.com/gpu=all" ] when hardware.nvidia-container-toolkit.enable is set, otherwise empty (CPU-only). Per-model devices overrides this.

Type: list of string

Default:

if config.hardware.nvidia-container-toolkit.enable then
  [ "nvidia.com/gpu=all" ]
else
  [ ]

Example:

[
  "amd.com/gpu=all"
]

services.llmhop.llama-cpp-quadlet.environment

Environment variables set on every model service. Merged with services.llmhop.llama-cpp-quadlet.models.<name>.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.llama-cpp-quadlet.environmentFile

File in KEY=VALUE format forwarded to every service. Use only for upstream features that require environment variables, such as HF_TOKEN for gated Hugging Face repositories. Environment variables are not systemd credentials and are visible to every model, so prefer credentials for any secret a server can read from a file. Loaded before services.llmhop.llama-cpp-quadlet.models.<name>.environmentFile, so per-model files override these entries.

Type: null or absolute path

Default:

null

Example:

"/etc/llama-cpp-quadlet/.env"

services.llmhop.llama-cpp-quadlet.image

Container image used for every model worker.

Type: string

Default:

"ghcr.io/ggml-org/llama.cpp"

services.llmhop.llama-cpp-quadlet.modelSettings

CLI flags forwarded to the model server for every model. Rendered as --<key> <value>, with false as --no-<key> and a list repeating the flag. See settings rendering for the full rules. Merged with services.llmhop.llama-cpp-quadlet.models.<name>.settings; per-model entries take precedence.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models

Models served by llama-server containers. Each entry produces one llama-cpp-<name> unit and uses the attribute name as its llmhop routing key and llama.cpp --alias.

Type: attribute set of (submodule)

Default:

{ }

Example:

{
  "qwen3-8b" = {
    settings.hf-repo = "unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL";
  };
}

services.llmhop.llama-cpp-quadlet.models.<name>.enable

Whether to enable model ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.llama-cpp-quadlet.models.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.llama-cpp-quadlet.models.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.llama-cpp-quadlet.models.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.llama-cpp-quadlet.models.<name>.devices

Devices exposed to this model’s container — passed verbatim as Quadlet AddDevice= lines. Replaces (does not extend) services.llmhop.llama-cpp-quadlet.devices for this model. Use to pin a model to specific device indices (e.g. [ "nvidia.com/gpu=0" ]).

Type: list of string

Default:

config.services.llmhop.llama-cpp-quadlet.devices

Example:

[
  "nvidia.com/gpu=0"
]

services.llmhop.llama-cpp-quadlet.models.<name>.digest

Immutable digest of the container image (e.g. sha256:…). Mutually exclusive with tag.

Type: null or string

Default:

null

Example:

"sha256:a73fb0b9046fee099f7c1829d2548e6cc1740f4c2776a6855fa659ae5d0deb49"

services.llmhop.llama-cpp-quadlet.models.<name>.environment

Additional environment variables set on this model’s service. Merged with services.llmhop.llama-cpp-quadlet.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.environmentFile

File in KEY=VALUE format forwarded to this model’s service. Loaded after services.llmhop.llama-cpp-quadlet.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.llama-cpp-quadlet.models.<name>.name

Canonical identifier for this model. Used for the unit name (llama-cpp-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.llama-cpp-quadlet.models.<name>.port

Loopback host port forwarded to the container’s llama.cpp API. Must be unique per model.

null binds the unix socket socket instead, which claims no port and only llmhop can connect to.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

null

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.containerConfig

Extra [Container] settings applied to this model container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.extraConfig

Extra unit sections applied to this model container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.mountOptions.credentials

Podman volume options appended to this model container’s systemd credential mount. Use an idmap mapping when [Container] User= selects a non-root identity. The mount is always read-only.

Type: list of string

Default:

[ ]

Example:

[
  "idmap=uids=0-1000-1;gids=0-1000-1"
]

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.mountOptions.socket

Podman volume options of this model container’s socket directory mount. U hands the directory to whatever host UID the container user maps to, so the socket works under any User= and UserNS=. Add z on SELinux hosts.

Type: list of string

Default:

[
  "U"
]

Example:

[
  "U"
  "z"
]

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.quadletConfig

Extra [Quadlet] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.serviceConfig

Extra [Service] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.unitConfig

Extra [Unit] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.settings

CLI flags forwarded to the model server for this model. Rendered as --<key> <value>, with false as --no-<key> and a list repeating the flag. See settings rendering for the full rules. Merged with services.llmhop.llama-cpp-quadlet.modelSettings; per-model entries take precedence. The flags llmhop derives from the model options (its served name and listener) always win over both.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.shmSize

Size of the container’s private /dev/shm tmpfs. PyTorch and friends use shared memory for NCCL/tensor-parallel inference; upstream recommends 32g (or --ipc=host). A private tmpfs is preferred for isolation: raise the value for larger models or higher tensor-parallel sizes.

Type: string

Default:

"32g"

Example:

"64g"

services.llmhop.llama-cpp-quadlet.models.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/llama-cpp-<name>/http.sock while port is null, else null

services.llmhop.llama-cpp-quadlet.models.<name>.tag

Tag of the container image used for this model. Mutually exclusive with digest.

Type: null or string

Default:

null

services.llmhop.llama-cpp-quadlet.openFilesLimit

File descriptor limit (LimitNOFILE) applied to every llama-cpp-quadlet systemd unit. Increase if the server logs accept: Too many open files under concurrent load.

Type: positive integer, meaning >0

Default:

1048576

services.llmhop.llama-cpp-quadlet.quadlet.containerConfig

Extra [Container] settings applied to every generated container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.extraConfig

Extra unit sections applied to every generated container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.quadletConfig

Extra [Quadlet] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.serviceConfig

Extra [Service] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.unitConfig

Extra [Unit] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.user

Host account whose systemd user manager owns these Quadlets. null installs system units and runs Podman rootfully. An attribute set installs user units for its uid and runs Podman rootlessly. This is independent of [Container] User= and the container’s user namespace configuration.

A managed account gets NixOS-allocated subordinate ID ranges. Set users.users.<name>.subUidRanges and subGidRanges (with autoSubUidGidRange = false) to pick them yourself.

Type: null or (submodule)

Default:

null

services.llmhop.llama-cpp-quadlet.quadlet.user.gid

GID of the managed account’s primary group.

Type: positive integer, meaning >0

Default:

config.uid

services.llmhop.llama-cpp-quadlet.quadlet.user.group

Primary group of the managed account.

Type: string

Default:

config.name

services.llmhop.llama-cpp-quadlet.quadlet.user.home

Home directory used for rootless Podman storage. This must live on a filesystem that supports the selected storage driver.

Type: absolute path

Default:

"/var/lib/llama-cpp"

services.llmhop.llama-cpp-quadlet.quadlet.user.manage

Whether llmhop creates and configures this account. Disable this for an account managed elsewhere, including its home, linger setting, group, and subordinate ID ranges.

Type: boolean

Default:

true

services.llmhop.llama-cpp-quadlet.quadlet.user.name

Host account whose systemd user manager owns the Quadlets.

Type: string

Default:

"llama-cpp"

services.llmhop.llama-cpp-quadlet.quadlet.user.uid

UID of the systemd user manager that owns the Quadlets.

Type: positive integer, meaning >0

Example:

503

services.llmhop.llama-cpp-quadlet.startupOrdering

Whether to chain enabled model services by ascending name during startup. GPU-memory profiling races otherwise: two workers booting on the same device each see it as fully free and race to claim their share, leading to OOM. Disable only when each model pins itself to a dedicated device via its own devices.

Type: boolean

Default:

true

services.llmhop.llama-cpp-quadlet.tag

Default tag of the container image used for models that do not set their own tag or digest.

Type: string

Example:

"server-cuda"

sglang

services.llmhop.sglang.enable

Whether to enable SGLang model serving via systemd (native host process), fronted by llmhop.

Type: boolean

Default:

false

Example:

true

services.llmhop.sglang.package

Package providing the SGLang Python environment, launched as bin/python -m sglang.launch_server.

No default on purpose: SGLang has no one-derivation-fits-all (new model architectures routinely need dev snapshots, and the wheels come in per-accelerator variants), so you build the package from a uv workspace and pin / follow upstream there. The flake exposes a helper:

inputs.llmhop.legacyPackages.${pkgs.system}.mkUvEnv {
  workspaceRoot = ./sglang-env; # your pyproject.toml + uv.lock
}

Individual models may override this with models.<name>.package.

The native module serves workers only; the SGL Model Gateway remains a sglang-quadlet feature (llmhop already routes between backends).

Type: package

Example:

inputs.llmhop.legacyPackages.${pkgs.system}.mkUvEnv {
  workspaceRoot = ./sglang-env;
}

services.llmhop.sglang.environment

Environment variables set on every model service. Merged with services.llmhop.sglang.models.<name>.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.sglang.environmentFile

File in KEY=VALUE format forwarded to every service. Use only for upstream features that require environment variables, such as HF_TOKEN for gated Hugging Face repositories. Environment variables are not systemd credentials and are visible to every model, so prefer credentials for any secret a server can read from a file. Loaded before services.llmhop.sglang.models.<name>.environmentFile, so per-model files override these entries.

Type: null or absolute path

Default:

null

Example:

"/etc/sglang/.env"

services.llmhop.sglang.gid

GID of the declared group. null lets NixOS allocate one.

Type: null or (unsigned integer, meaning >=0)

Default:

config.services.llmhop.sglang.uid

services.llmhop.sglang.group

Primary group of user. The module declares it while it keeps its default name, any other group is the deployer’s to declare.

Type: string

Default:

config.services.llmhop.sglang.user

services.llmhop.sglang.modelSettings

CLI flags forwarded to the model server for every model. Rendered as --<key> <value>, with false dropped, since this CLI pairs --enable-X with --disable-X and a list handed to a single flag. See settings rendering for the full rules. Merged with services.llmhop.sglang.models.<name>.settings; per-model entries take precedence.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang.models

Models to serve. Each enabled entry produces one systemd service named sglang-<name>; the attribute name is the routing key surfaced through llmhop as the OpenAI model field. Enabled entries are sorted by ascending name.

Type: attribute set of (submodule)

Default:

{ }

Example:

{
  "qwen3-8b" = {
    model = "Qwen/Qwen3-8B";
    port = 19001;
    settings = {
      reasoning-parser = "qwen3";
      mem-fraction-static = 0.6;
    };
  };
}

services.llmhop.sglang.models.<name>.enable

Whether to enable model ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.sglang.models.<name>.package

Package providing this model’s worker, overriding the backend-wide package. Set it for a model that needs a different sglang release than the rest — e.g. a nightly wheel for a just-released architecture — built the same way with mkUvEnv over a per-model uv workspace. Defaults to the backend-wide package.

Type: package

Default:

config.services.llmhop.sglang.package

services.llmhop.sglang.models.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.sglang.models.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.sglang.models.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.sglang.models.<name>.environment

Additional environment variables set on this model’s service. Merged with services.llmhop.sglang.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.sglang.models.<name>.environmentFile

File in KEY=VALUE format forwarded to this model’s service. Loaded after services.llmhop.sglang.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.sglang.models.<name>.model

Hugging Face repo id (or local path) passed as --model-path.

Type: string

Example:

"Qwen/Qwen3-8B"

services.llmhop.sglang.models.<name>.name

Canonical identifier for this model. Used for the unit name (sglang-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.sglang.models.<name>.port

Loopback host port sglang binds to (--host 127.0.0.1 --port <port>). Must be unique per enabled model; llmhop reaches the backend at http://127.0.0.1:<port>.

Type: 16 bit unsigned integer; between 0 and 65535 (both inclusive)

services.llmhop.sglang.models.<name>.serviceConfig

Extra [Service] settings merged into this workload’s sglang-<name> unit after the hardened baseline and backend-specific relaxations. The module retains ownership of ExecStart, KillMode, and Type because they implement readiness supervision as one lifecycle contract.

Type: attribute set of anything

Default:

{ }

Example:

{
  MemoryHigh = "64G";
}

services.llmhop.sglang.models.<name>.settings

CLI flags forwarded to the model server for this model. Rendered as --<key> <value>, with false dropped, since this CLI pairs --enable-X with --disable-X and a list handed to a single flag. See settings rendering for the full rules. Merged with services.llmhop.sglang.modelSettings; per-model entries take precedence. The flags llmhop derives from the model options (its served name and listener) always win over both.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang.models.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/sglang-<name>/http.sock while port is null, else null

services.llmhop.sglang.models.<name>.unitConfig

Extra [Unit] settings merged into this workload’s sglang-<name> unit after the shared baseline.

Ordering and dependency directives (After=, Requires=, Wants=) do not belong here: NixOS renders those from the after, requires and wants options, so a definition of the same key in unitConfig conflicts with it instead of merging. Declare them on systemd.services."sglang-<name>" from your own module, where the module system concatenates them with what this one sets.

Type: attribute set of anything

Default:

{ }

Example:

{
  StartLimitBurst = 10;
}

services.llmhop.sglang.openFilesLimit

File descriptor limit (LimitNOFILE) applied to every sglang systemd unit. Increase if the server logs accept: Too many open files under concurrent load.

Type: positive integer, meaning >0

Default:

1048576

services.llmhop.sglang.startupOrdering

Whether to chain enabled model services by ascending name during startup. GPU-memory profiling races otherwise: two workers booting on the same device each see it as fully free and race to claim their share, leading to OOM. Disable only when each model pins itself to a dedicated device via environment (the variable is stack-specific: CUDA_VISIBLE_DEVICES, HIP_VISIBLE_DEVICES, ZE_AFFINITY_MASK, …).

Type: boolean

Default:

true

services.llmhop.sglang.uid

UID of the declared user. null lets NixOS allocate one.

Type: null or (unsigned integer, meaning >=0)

Default:

null

Example:

503

services.llmhop.sglang.user

System user the units run as. The module declares it while it keeps its default name, any other user is the deployer’s to declare.

Type: string

Default:

"sglang"

sglang-quadlet

services.llmhop.sglang-quadlet.enable

Whether to enable SGLang model serving via Quadlet, optionally fronted by the SGL Model Gateway.

Type: boolean

Default:

false

Example:

true

services.llmhop.sglang-quadlet.cache.containerDirectory

Path at which the cache is mounted inside every model container.

Type: string

Default:

"/root/.cache/huggingface"

services.llmhop.sglang-quadlet.cache.directory

Host directory bind-mounted as the Hugging Face cache.

Type: absolute path

Default:

"/var/cache/sglang"

services.llmhop.sglang-quadlet.cache.environmentVariable

Environment variable set on every container to point its runtime at containerDirectory.

Type: string

Default:

"HF_HOME"

services.llmhop.sglang-quadlet.cache.group

Host group used when cache.manage is enabled.

Type: string

Default:

if config.services.llmhop.sglang-quadlet.quadlet.user == null then
  "root"
else
  config.services.llmhop.sglang-quadlet.quadlet.user.group

services.llmhop.sglang-quadlet.cache.manage

Whether llmhop creates the host cache directory with systemd-tmpfiles, owned by user/group and mode 0700.

Type: boolean

Default:

true

services.llmhop.sglang-quadlet.cache.mountOptions

Options appended to the cache’s Quadlet Volume= entry. This can be used for Podman ownership mechanisms such as U, idmap, or SELinux relabeling.

Type: list of string

Default:

[ ]

Example:

[
  "idmap"
]

services.llmhop.sglang-quadlet.cache.user

Host owner used when cache.manage is enabled. Override it when a UIDMap/idmap mapping makes the container see a different owner than the host account running Podman.

Type: string

Default:

if config.services.llmhop.sglang-quadlet.quadlet.user == null then
  "root"
else
  config.services.llmhop.sglang-quadlet.quadlet.user.name

services.llmhop.sglang-quadlet.devices

Devices exposed to every model container — passed verbatim as Quadlet AddDevice= lines. Accepts both CDI references (recommended: nvidia.com/gpu=…, amd.com/gpu=…, intel.com/gpu=…, …) and raw host device paths (e.g. /dev/dri/renderD128). For CDI, the corresponding spec must be generated on the host (e.g. nvidia-ctk cdi generate). Defaults to [ "nvidia.com/gpu=all" ] when hardware.nvidia-container-toolkit.enable is set, otherwise empty (CPU-only). Per-model devices overrides this.

Type: list of string

Default:

if config.hardware.nvidia-container-toolkit.enable then
  [ "nvidia.com/gpu=all" ]
else
  [ ]

Example:

[
  "amd.com/gpu=all"
]

services.llmhop.sglang-quadlet.environment

Environment variables set on every model service. Merged with services.llmhop.sglang-quadlet.models.<name>.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.sglang-quadlet.environmentFile

File in KEY=VALUE format forwarded to every service. Use only for upstream features that require environment variables, such as HF_TOKEN for gated Hugging Face repositories. Environment variables are not systemd credentials and are visible to every model, so prefer credentials for any secret a server can read from a file. Loaded before services.llmhop.sglang-quadlet.models.<name>.environmentFile, so per-model files override these entries.

Type: null or absolute path

Default:

null

Example:

"/etc/sglang-quadlet/.env"

services.llmhop.sglang-quadlet.gateway.enable

Whether to enable the SGL Model Gateway in front of the workers. Disabled by default — llmhop already routes between every backend, and the gateway is only needed when you want SGLang’s IGW dispatch features (custom routing, prefix caching across workers, etc.) .

Type: boolean

Default:

false

Example:

true

services.llmhop.sglang-quadlet.gateway.enableMetrics

Whether to enable Prometheus metrics on the gateway.

Type: boolean

Default:

true

Example:

true

services.llmhop.sglang-quadlet.gateway.bindAddress

Host address the gateway binds its listeners to. Defaults to the loopback so external clients must go through Caddy / llmhop.

Type: string

Default:

"127.0.0.1"

services.llmhop.sglang-quadlet.gateway.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.sglang-quadlet.gateway.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.sglang-quadlet.gateway.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.sglang-quadlet.gateway.digest

Immutable digest of the gateway image. Mutually exclusive with tag.

Type: null or string

Default:

null

services.llmhop.sglang-quadlet.gateway.environment

Additional environment variables set on the gateway container.

Type: attribute set of string

Default:

{ }

services.llmhop.sglang-quadlet.gateway.environmentFile

File in KEY=VALUE format forwarded to the gateway via --env-file. Use only for upstream features that require environment variables, and prefer credentials for any secret the gateway can read from a file.

Type: null or absolute path

Default:

null

Example:

"/etc/sglang/gateway.env"

services.llmhop.sglang-quadlet.gateway.image

Container image used for the gateway.

Type: string

Default:

"docker.io/lmsysorg/sgl-model-gateway"

services.llmhop.sglang-quadlet.gateway.metricsPort

Host port the gateway exposes Prometheus metrics on. Ignored when enableMetrics is false.

Type: 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

29000

services.llmhop.sglang-quadlet.gateway.port

Host port the gateway listens on.

Type: 16 bit unsigned integer; between 0 and 65535 (both inclusive)

services.llmhop.sglang-quadlet.gateway.quadlet.containerConfig

Extra [Container] settings applied to the gateway container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.gateway.quadlet.extraConfig

Extra unit sections applied to the gateway container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.gateway.quadlet.mountOptions.credentials

Podman volume options appended to the gateway container’s systemd credential mount. Use an idmap mapping when [Container] User= selects a non-root identity. The mount is always read-only.

Type: list of string

Default:

[ ]

Example:

[
  "idmap=uids=0-1000-1;gids=0-1000-1"
]

services.llmhop.sglang-quadlet.gateway.quadlet.quadletConfig

Extra [Quadlet] settings applied to the gateway container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.gateway.quadlet.serviceConfig

Extra [Service] settings applied to the gateway container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.gateway.quadlet.unitConfig

Extra [Unit] settings applied to the gateway container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.gateway.settings

Additional CLI flags forwarded to sgl-model-gateway. Rendered as --<key> <value>, with false dropped, since this CLI pairs --enable-X with --disable-X and a list handed to a single flag. See settings rendering for the full rules. --worker-urls is rendered from the enabled models, so setting it here replaces the generated list. The listener flags (host, port, prometheus-host, prometheus-port) come from the options of the same name and always win over entries set here.

Type: attribute set of anything

Default:

{ }

Example:

{
  tls-cert-path = "/etc/sglang/tls/server.crt";
  tls-key-path = "\${cred:tls-key}";
}

services.llmhop.sglang-quadlet.gateway.tag

Default tag of the gateway image. Mutually exclusive with digest.

Type: null or string

Default:

"latest"

services.llmhop.sglang-quadlet.image

Container image used for every model worker.

Type: string

Default:

"docker.io/lmsysorg/sglang"

services.llmhop.sglang-quadlet.modelSettings

CLI flags forwarded to the model server for every model. Rendered as --<key> <value>, with false dropped, since this CLI pairs --enable-X with --disable-X and a list handed to a single flag. See settings rendering for the full rules. Merged with services.llmhop.sglang-quadlet.models.<name>.settings; per-model entries take precedence.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.models

Models to serve. Each entry produces one quadlet container; the attribute name is the routing key (advertised via --served-model-name and surfaced through both llmhop and the optional SGL Model Gateway as the OpenAI model field). Enabled entries are sorted by ascending name.

Type: attribute set of (submodule)

Default:

{ }

Example:

{
  "qwen3-8b" = {
    model = "Qwen/Qwen3-8B";
    port = 19001;
    settings = {
      reasoning-parser = "qwen3";
      tool-call-parser = "qwen3_coder";
      mem-fraction-static = 0.6;
      cuda-graph-max-bs = 4;
    };
  };
}

services.llmhop.sglang-quadlet.models.<name>.enable

Whether to enable model ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.sglang-quadlet.models.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.sglang-quadlet.models.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.sglang-quadlet.models.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.sglang-quadlet.models.<name>.devices

Devices exposed to this model’s container — passed verbatim as Quadlet AddDevice= lines. Replaces (does not extend) services.llmhop.sglang-quadlet.devices for this model. Use to pin a model to specific device indices (e.g. [ "nvidia.com/gpu=0" ]).

Type: list of string

Default:

config.services.llmhop.sglang-quadlet.devices

Example:

[
  "nvidia.com/gpu=0"
]

services.llmhop.sglang-quadlet.models.<name>.digest

Immutable digest of the container image (e.g. sha256:…). Mutually exclusive with tag.

Type: null or string

Default:

null

Example:

"sha256:a73fb0b9046fee099f7c1829d2548e6cc1740f4c2776a6855fa659ae5d0deb49"

services.llmhop.sglang-quadlet.models.<name>.environment

Additional environment variables set on this model’s service. Merged with services.llmhop.sglang-quadlet.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.sglang-quadlet.models.<name>.environmentFile

File in KEY=VALUE format forwarded to this model’s service. Loaded after services.llmhop.sglang-quadlet.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.sglang-quadlet.models.<name>.model

Hugging Face repo id (or local path) passed to the model server.

Type: string

Example:

"Qwen/Qwen2.5-7B-Instruct"

services.llmhop.sglang-quadlet.models.<name>.name

Canonical identifier for this model. Used for the unit name (sglang-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.sglang-quadlet.models.<name>.port

Loopback host port forwarded to the container’s SGLang API. Must be unique per model and must not collide with gateway.port / gateway.metricsPort when the gateway is enabled.

Type: 16 bit unsigned integer; between 0 and 65535 (both inclusive)

services.llmhop.sglang-quadlet.models.<name>.quadlet.containerConfig

Extra [Container] settings applied to this model container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.models.<name>.quadlet.extraConfig

Extra unit sections applied to this model container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.models.<name>.quadlet.mountOptions.credentials

Podman volume options appended to this model container’s systemd credential mount. Use an idmap mapping when [Container] User= selects a non-root identity. The mount is always read-only.

Type: list of string

Default:

[ ]

Example:

[
  "idmap=uids=0-1000-1;gids=0-1000-1"
]

services.llmhop.sglang-quadlet.models.<name>.quadlet.quadletConfig

Extra [Quadlet] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.models.<name>.quadlet.serviceConfig

Extra [Service] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.models.<name>.quadlet.unitConfig

Extra [Unit] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.models.<name>.settings

CLI flags forwarded to the model server for this model. Rendered as --<key> <value>, with false dropped, since this CLI pairs --enable-X with --disable-X and a list handed to a single flag. See settings rendering for the full rules. Merged with services.llmhop.sglang-quadlet.modelSettings; per-model entries take precedence. The flags llmhop derives from the model options (its served name and listener) always win over both.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.models.<name>.shmSize

Size of the container’s private /dev/shm tmpfs. PyTorch and friends use shared memory for NCCL/tensor-parallel inference; upstream recommends 32g (or --ipc=host). A private tmpfs is preferred for isolation: raise the value for larger models or higher tensor-parallel sizes.

Type: string

Default:

"32g"

Example:

"64g"

services.llmhop.sglang-quadlet.models.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/sglang-<name>/http.sock while port is null, else null

services.llmhop.sglang-quadlet.models.<name>.tag

Tag of the container image used for this model. Mutually exclusive with digest.

Type: null or string

Default:

null

services.llmhop.sglang-quadlet.openFilesLimit

File descriptor limit (LimitNOFILE) applied to every sglang-quadlet systemd unit. Increase if the server logs accept: Too many open files under concurrent load.

Type: positive integer, meaning >0

Default:

1048576

services.llmhop.sglang-quadlet.quadlet.containerConfig

Extra [Container] settings applied to every generated container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.quadlet.extraConfig

Extra unit sections applied to every generated container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.quadlet.quadletConfig

Extra [Quadlet] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.quadlet.serviceConfig

Extra [Service] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.quadlet.unitConfig

Extra [Unit] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.sglang-quadlet.quadlet.user

Host account whose systemd user manager owns these Quadlets. null installs system units and runs Podman rootfully. An attribute set installs user units for its uid and runs Podman rootlessly. This is independent of [Container] User= and the container’s user namespace configuration.

A managed account gets NixOS-allocated subordinate ID ranges. Set users.users.<name>.subUidRanges and subGidRanges (with autoSubUidGidRange = false) to pick them yourself.

Type: null or (submodule)

Default:

null

services.llmhop.sglang-quadlet.quadlet.user.gid

GID of the managed account’s primary group.

Type: positive integer, meaning >0

Default:

config.uid

services.llmhop.sglang-quadlet.quadlet.user.group

Primary group of the managed account.

Type: string

Default:

config.name

services.llmhop.sglang-quadlet.quadlet.user.home

Home directory used for rootless Podman storage. This must live on a filesystem that supports the selected storage driver.

Type: absolute path

Default:

"/var/lib/sglang"

services.llmhop.sglang-quadlet.quadlet.user.manage

Whether llmhop creates and configures this account. Disable this for an account managed elsewhere, including its home, linger setting, group, and subordinate ID ranges.

Type: boolean

Default:

true

services.llmhop.sglang-quadlet.quadlet.user.name

Host account whose systemd user manager owns the Quadlets.

Type: string

Default:

"sglang"

services.llmhop.sglang-quadlet.quadlet.user.uid

UID of the systemd user manager that owns the Quadlets.

Type: positive integer, meaning >0

Example:

503

services.llmhop.sglang-quadlet.startupOrdering

Whether to chain enabled model services by ascending name during startup. GPU-memory profiling races otherwise: two workers booting on the same device each see it as fully free and race to claim their share, leading to OOM. Disable only when each model pins itself to a dedicated device via its own devices.

Type: boolean

Default:

true

services.llmhop.sglang-quadlet.tag

Default tag of the container image used for models that do not set their own tag or digest.

Type: string

Example:

"latest"

vllm

services.llmhop.vllm.enable

Whether to enable vLLM model serving via systemd (native host process), fronted by llmhop.

Type: boolean

Default:

false

Example:

true

services.llmhop.vllm.package

Package providing the vllm CLI at bin/vllm.

No default on purpose: vLLM has no one-derivation-fits-all (new model architectures routinely need dev snapshots, and the wheels come in per-accelerator variants), so you build the package from a uv workspace and pin / follow upstream there. The flake exposes a helper:

inputs.llmhop.legacyPackages.${pkgs.system}.mkUvEnv {
  workspaceRoot = ./vllm-env; # your pyproject.toml + uv.lock
}

Individual models may override this with models.<name>.package.

Type: package

Example:

inputs.llmhop.legacyPackages.${pkgs.system}.mkUvEnv {
  workspaceRoot = ./vllm-env;
}

services.llmhop.vllm.detectors

Standalone watermark detection services.

Type: attribute set of (submodule)

Default:

{ }

services.llmhop.vllm.detectors.<name>.enable

Whether to enable watermark detector ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.vllm.detectors.<name>.package

Environment this detector runs in.

Type: package

Default:

config.services.llmhop.vllm.package

services.llmhop.vllm.detectors.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.vllm.detectors.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.vllm.detectors.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.vllm.detectors.<name>.environment

Additional environment variables set on this watermark detector’s service. Merged with services.llmhop.vllm.environment; per-watermark detector entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.vllm.detectors.<name>.environmentFile

File in KEY=VALUE format forwarded to this watermark detector’s service. Loaded after services.llmhop.vllm.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.vllm.detectors.<name>.name

Canonical identifier for this watermark detector. Used for the unit name (vllm-detector-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.vllm.detectors.<name>.port

Loopback port on which the detector serves /detect. null binds the unix socket socket instead, which claims no port and only llmhop can connect to.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

null

services.llmhop.vllm.detectors.<name>.script

Detector server run by the detector’s Python interpreter. It receives detect, --tokenizer, --watermark-config, --watermark-key-file, the listener flags (--host and --port, or --uds) and settings, and must serve GET /health and POST /detect. The native default is checked against package at build time, the Quadlet one only at startup, since the image is opaque to the build.

Type: absolute path

Default: the script built by mkVllmWatermark

services.llmhop.vllm.detectors.<name>.serviceConfig

Extra [Service] settings merged into this workload’s vllm-detector-<name> unit after the hardened baseline and backend-specific relaxations. The module retains ownership of ExecStart, KillMode, and Type because they implement readiness supervision as one lifecycle contract.

Type: attribute set of anything

Default:

{ }

Example:

{
  MemoryHigh = "64G";
}

services.llmhop.vllm.detectors.<name>.settings

Arguments passed to the detector script. p-value-threshold is the only detection-side setting. tokenizer, the watermark flags and the listener (host, port, uds) are derived from the options and always win over entries set here. Rendered as --<key> <value>, with false as --no-<key> and a list handed to a single flag. See settings rendering for the full rules.

Type: attribute set of anything

Default:

{ }

Example:

{
  p-value-threshold = 0.01;
}

services.llmhop.vllm.detectors.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/vllm-detector-<name>/http.sock while port is null, else null

services.llmhop.vllm.detectors.<name>.tokenizer

Tokenizer used to encode candidate text. It must exactly match the tokenizer used for watermarked generation.

Type: string

Example:

"Qwen/Qwen3-8B"

services.llmhop.vllm.detectors.<name>.unitConfig

Extra [Unit] settings merged into this workload’s vllm-detector-<name> unit after the shared baseline.

Ordering and dependency directives (After=, Requires=, Wants=) do not belong here: NixOS renders those from the after, requires and wants options, so a definition of the same key in unitConfig conflicts with it instead of merging. Declare them on systemd.services."vllm-detector-<name>" from your own module, where the module system concatenates them with what this one sets.

Type: attribute set of anything

Default:

{ }

Example:

{
  StartLimitBurst = 10;
}

services.llmhop.vllm.detectors.<name>.watermark

vLLM’s WatermarkConfig without key, validated at startup by the installed release. A generating worker and its detector must share it.

The key, an unsigned 64-bit integer, is the credential vllm.watermark-key. It is imported from the system credential store, for example the root-only file /etc/credstore/vllm.watermark-key, unless credentials."vllm.watermark-key" names another source. An encrypted credential must be created under that name, e.g. with systemd-creds encrypt --name=vllm.watermark-key.

Type: attribute set of anything

Default:

{ }

Example:

{
  algorithm = "dual_key_gumbel";
  context_width = 4;
}

services.llmhop.vllm.environment

Environment variables set on every model service. Merged with services.llmhop.vllm.models.<name>.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.vllm.environmentFile

File in KEY=VALUE format forwarded to every service. Use only for upstream features that require environment variables, such as HF_TOKEN for gated Hugging Face repositories. Environment variables are not systemd credentials and are visible to every model, so prefer credentials for any secret a server can read from a file. Loaded before services.llmhop.vllm.models.<name>.environmentFile, so per-model files override these entries.

Type: null or absolute path

Default:

null

Example:

"/etc/vllm/.env"

services.llmhop.vllm.gid

GID of the declared group. null lets NixOS allocate one.

Type: null or (unsigned integer, meaning >=0)

Default:

config.services.llmhop.vllm.uid

services.llmhop.vllm.group

Primary group of user. The module declares it while it keeps its default name, any other group is the deployer’s to declare.

Type: string

Default:

config.services.llmhop.vllm.user

services.llmhop.vllm.modelSettings

CLI flags forwarded to the model server for every model. Rendered as --<key> <value>, with false as --no-<key> and a list handed to a single flag. See settings rendering for the full rules. Merged with services.llmhop.vllm.models.<name>.settings; per-model entries take precedence.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm.models

Models to serve. Each enabled entry produces one systemd service named vllm-<name>; the attribute name is the routing key surfaced through llmhop as the OpenAI model field. Enabled entries are sorted by ascending name.

Type: attribute set of (submodule)

Default:

{ }

Example:

{
  "qwen2-5-7b" = {
    model = "Qwen/Qwen2.5-7B-Instruct";
  };
  "llama-3-8b" = {
    model = "meta-llama/Meta-Llama-3-8B-Instruct";
    settings.max-model-len = 8192;
  };
}

services.llmhop.vllm.models.<name>.enable

Whether to enable model ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.vllm.models.<name>.package

Package providing this model’s worker, overriding the backend-wide package. Set it for a model that needs a different vllm release than the rest — e.g. a nightly wheel for a just-released architecture — built the same way with mkUvEnv over a per-model uv workspace. Defaults to the backend-wide package.

Type: package

Default:

config.services.llmhop.vllm.package

services.llmhop.vllm.models.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.vllm.models.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.vllm.models.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.vllm.models.<name>.environment

Additional environment variables set on this model’s service. Merged with services.llmhop.vllm.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.vllm.models.<name>.environmentFile

File in KEY=VALUE format forwarded to this model’s service. Loaded after services.llmhop.vllm.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.vllm.models.<name>.model

Hugging Face repo id (or local path) passed as the vllm serve positional argument.

Type: string

Example:

"Qwen/Qwen2.5-7B-Instruct"

services.llmhop.vllm.models.<name>.name

Canonical identifier for this model. Used for the unit name (vllm-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.vllm.models.<name>.port

Loopback host port vllm binds to (--host 127.0.0.1 --port <port>). Must be unique per enabled model; llmhop reaches the backend at http://127.0.0.1:<port>.

null binds the unix socket socket instead, which claims no port and only llmhop can connect to.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

null

services.llmhop.vllm.models.<name>.serviceConfig

Extra [Service] settings merged into this workload’s vllm-<name> unit after the hardened baseline and backend-specific relaxations. The module retains ownership of ExecStart, KillMode, and Type because they implement readiness supervision as one lifecycle contract.

Type: attribute set of anything

Default:

{ }

Example:

{
  MemoryHigh = "64G";
}

services.llmhop.vllm.models.<name>.settings

CLI flags forwarded to the model server for this model. Rendered as --<key> <value>, with false as --no-<key> and a list handed to a single flag. See settings rendering for the full rules. Merged with services.llmhop.vllm.modelSettings; per-model entries take precedence. The flags llmhop derives from the model options (its served name and listener) always win over both.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm.models.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/vllm-<name>/http.sock while port is null, else null

services.llmhop.vllm.models.<name>.unitConfig

Extra [Unit] settings merged into this workload’s vllm-<name> unit after the shared baseline.

Ordering and dependency directives (After=, Requires=, Wants=) do not belong here: NixOS renders those from the after, requires and wants options, so a definition of the same key in unitConfig conflicts with it instead of merging. Declare them on systemd.services."vllm-<name>" from your own module, where the module system concatenates them with what this one sets.

Type: attribute set of anything

Default:

{ }

Example:

{
  StartLimitBurst = 10;
}

services.llmhop.vllm.models.<name>.watermark

vLLM’s WatermarkConfig without key, validated at startup by the installed release. A generating worker and its detector must share it.

The key, an unsigned 64-bit integer, is the credential vllm.watermark-key. It is imported from the system credential store, for example the root-only file /etc/credstore/vllm.watermark-key, unless credentials."vllm.watermark-key" names another source. An encrypted credential must be created under that name, e.g. with systemd-creds encrypt --name=vllm.watermark-key.

Type: null or (attribute set of anything)

Default:

null

Example:

{
  algorithm = "dual_key_gumbel";
  context_width = 4;
}

services.llmhop.vllm.openFilesLimit

File descriptor limit (LimitNOFILE) applied to every vllm systemd unit. Increase if the server logs accept: Too many open files under concurrent load.

Type: positive integer, meaning >0

Default:

1048576

services.llmhop.vllm.startupOrdering

Whether to chain enabled model services by ascending name during startup. GPU-memory profiling races otherwise: two workers booting on the same device each see it as fully free and race to claim their share, leading to OOM. Disable only when each model pins itself to a dedicated device via environment (the variable is stack-specific: CUDA_VISIBLE_DEVICES, HIP_VISIBLE_DEVICES, ZE_AFFINITY_MASK, …).

Type: boolean

Default:

true

services.llmhop.vllm.uid

UID of the declared user. null lets NixOS allocate one.

Type: null or (unsigned integer, meaning >=0)

Default:

null

Example:

503

services.llmhop.vllm.user

System user the units run as. The module declares it while it keeps its default name, any other user is the deployer’s to declare.

Type: string

Default:

"vllm"

vllm-quadlet

services.llmhop.vllm-quadlet.enable

Whether to enable vLLM model serving via Quadlet, fronted by llmhop.

Type: boolean

Default:

false

Example:

true

services.llmhop.vllm-quadlet.cache.containerDirectory

Path at which the cache is mounted inside every model container.

Type: string

Default:

"/root/.cache/huggingface"

services.llmhop.vllm-quadlet.cache.directory

Host directory bind-mounted as the Hugging Face cache.

Type: absolute path

Default:

"/var/cache/vllm"

services.llmhop.vllm-quadlet.cache.environmentVariable

Environment variable set on every container to point its runtime at containerDirectory.

Type: string

Default:

"HF_HOME"

services.llmhop.vllm-quadlet.cache.group

Host group used when cache.manage is enabled.

Type: string

Default:

if config.services.llmhop.vllm-quadlet.quadlet.user == null then
  "root"
else
  config.services.llmhop.vllm-quadlet.quadlet.user.group

services.llmhop.vllm-quadlet.cache.manage

Whether llmhop creates the host cache directory with systemd-tmpfiles, owned by user/group and mode 0700.

Type: boolean

Default:

true

services.llmhop.vllm-quadlet.cache.mountOptions

Options appended to the cache’s Quadlet Volume= entry. This can be used for Podman ownership mechanisms such as U, idmap, or SELinux relabeling.

Type: list of string

Default:

[ ]

Example:

[
  "idmap"
]

services.llmhop.vllm-quadlet.cache.user

Host owner used when cache.manage is enabled. Override it when a UIDMap/idmap mapping makes the container see a different owner than the host account running Podman.

Type: string

Default:

if config.services.llmhop.vllm-quadlet.quadlet.user == null then
  "root"
else
  config.services.llmhop.vllm-quadlet.quadlet.user.name

services.llmhop.vllm-quadlet.detectors

Standalone watermark detection containers.

Type: attribute set of (submodule)

Default:

{ }

services.llmhop.vllm-quadlet.detectors.<name>.enable

Whether to enable watermark detector ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.vllm-quadlet.detectors.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.vllm-quadlet.detectors.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.vllm-quadlet.detectors.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.vllm-quadlet.detectors.<name>.digest

Immutable digest of the container image (e.g. sha256:…). Mutually exclusive with tag.

Type: null or string

Default:

null

Example:

"sha256:a73fb0b9046fee099f7c1829d2548e6cc1740f4c2776a6855fa659ae5d0deb49"

services.llmhop.vllm-quadlet.detectors.<name>.environment

Additional environment variables set on this watermark detector’s service. Merged with services.llmhop.vllm-quadlet.environment; per-watermark detector entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.vllm-quadlet.detectors.<name>.environmentFile

File in KEY=VALUE format forwarded to this watermark detector’s service. Loaded after services.llmhop.vllm-quadlet.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.vllm-quadlet.detectors.<name>.name

Canonical identifier for this watermark detector. Used for the unit name (vllm-detector-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.vllm-quadlet.detectors.<name>.port

Loopback port on which the detector serves /detect. null binds the unix socket socket instead, which claims no port and only llmhop can connect to.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

null

services.llmhop.vllm-quadlet.detectors.<name>.quadlet.containerConfig

Extra [Container] settings applied to this detector container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.detectors.<name>.quadlet.extraConfig

Extra unit sections applied to this detector container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.detectors.<name>.quadlet.mountOptions.credentials

Podman volume options appended to this detector container’s systemd credential mount. Use an idmap mapping when [Container] User= selects a non-root identity. The mount is always read-only.

Type: list of string

Default:

[ ]

Example:

[
  "idmap=uids=0-1000-1;gids=0-1000-1"
]

services.llmhop.vllm-quadlet.detectors.<name>.quadlet.mountOptions.socket

Podman volume options of this detector container’s socket directory mount. U hands the directory to whatever host UID the container user maps to, so the socket works under any User= and UserNS=. Add z on SELinux hosts.

Type: list of string

Default:

[
  "U"
]

Example:

[
  "U"
  "z"
]

services.llmhop.vllm-quadlet.detectors.<name>.quadlet.quadletConfig

Extra [Quadlet] settings applied to this detector container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.detectors.<name>.quadlet.serviceConfig

Extra [Service] settings applied to this detector container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.detectors.<name>.quadlet.unitConfig

Extra [Unit] settings applied to this detector container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.detectors.<name>.script

Detector server run by the detector’s Python interpreter. It receives detect, --tokenizer, --watermark-config, --watermark-key-file, the listener flags (--host and --port, or --uds) and settings, and must serve GET /health and POST /detect. The native default is checked against package at build time, the Quadlet one only at startup, since the image is opaque to the build.

Type: absolute path

Default: the script built by mkVllmWatermark

services.llmhop.vllm-quadlet.detectors.<name>.settings

Arguments passed to the detector script. p-value-threshold is the only detection-side setting. tokenizer, the watermark flags and the listener (host, port, uds) are derived from the options and always win over entries set here. Rendered as --<key> <value>, with false as --no-<key> and a list handed to a single flag. See settings rendering for the full rules.

Type: attribute set of anything

Default:

{ }

Example:

{
  p-value-threshold = 0.01;
}

services.llmhop.vllm-quadlet.detectors.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/vllm-detector-<name>/http.sock while port is null, else null

services.llmhop.vllm-quadlet.detectors.<name>.tag

Tag of the container image used for this detector. Mutually exclusive with digest.

Type: null or string

Default:

null

services.llmhop.vllm-quadlet.detectors.<name>.tokenizer

Tokenizer used to encode candidate text. It must exactly match the tokenizer used for watermarked generation.

Type: string

Example:

"Qwen/Qwen3-8B"

services.llmhop.vllm-quadlet.detectors.<name>.watermark

vLLM’s WatermarkConfig without key, validated at startup by the installed release. A generating worker and its detector must share it.

The key, an unsigned 64-bit integer, is the credential vllm.watermark-key. It is imported from the system credential store, for example the root-only file /etc/credstore/vllm.watermark-key, unless credentials."vllm.watermark-key" names another source. An encrypted credential must be created under that name, e.g. with systemd-creds encrypt --name=vllm.watermark-key.

Type: attribute set of anything

Default:

{ }

Example:

{
  algorithm = "dual_key_gumbel";
  context_width = 4;
}

services.llmhop.vllm-quadlet.devices

Devices exposed to every model container — passed verbatim as Quadlet AddDevice= lines. Accepts both CDI references (recommended: nvidia.com/gpu=…, amd.com/gpu=…, intel.com/gpu=…, …) and raw host device paths (e.g. /dev/dri/renderD128). For CDI, the corresponding spec must be generated on the host (e.g. nvidia-ctk cdi generate). Defaults to [ "nvidia.com/gpu=all" ] when hardware.nvidia-container-toolkit.enable is set, otherwise empty (CPU-only). Per-model devices overrides this.

Type: list of string

Default:

if config.hardware.nvidia-container-toolkit.enable then
  [ "nvidia.com/gpu=all" ]
else
  [ ]

Example:

[
  "amd.com/gpu=all"
]

services.llmhop.vllm-quadlet.environment

Environment variables set on every model service. Merged with services.llmhop.vllm-quadlet.models.<name>.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.vllm-quadlet.environmentFile

File in KEY=VALUE format forwarded to every service. Use only for upstream features that require environment variables, such as HF_TOKEN for gated Hugging Face repositories. Environment variables are not systemd credentials and are visible to every model, so prefer credentials for any secret a server can read from a file. Loaded before services.llmhop.vllm-quadlet.models.<name>.environmentFile, so per-model files override these entries.

Type: null or absolute path

Default:

null

Example:

"/etc/vllm-quadlet/.env"

services.llmhop.vllm-quadlet.image

Container image used for every model worker.

Type: string

Default:

"docker.io/vllm/vllm-openai"

services.llmhop.vllm-quadlet.modelSettings

CLI flags forwarded to the model server for every model. Rendered as --<key> <value>, with false as --no-<key> and a list handed to a single flag. See settings rendering for the full rules. Merged with services.llmhop.vllm-quadlet.models.<name>.settings; per-model entries take precedence.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.models

Models to serve. Each entry produces one quadlet container; the attribute name is the routing key. Enabled entries are sorted by ascending name.

Type: attribute set of (submodule)

Default:

{ }

Example:

{
  "qwen2-5-7b" = {
    model = "Qwen/Qwen2.5-7B-Instruct";
  };
  "llama-3-8b" = {
    model = "meta-llama/Meta-Llama-3-8B-Instruct";
    settings.max-model-len = 8192;
  };
}

services.llmhop.vllm-quadlet.models.<name>.enable

Whether to enable model ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.vllm-quadlet.models.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.vllm-quadlet.models.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.vllm-quadlet.models.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.vllm-quadlet.models.<name>.devices

Devices exposed to this model’s container — passed verbatim as Quadlet AddDevice= lines. Replaces (does not extend) services.llmhop.vllm-quadlet.devices for this model. Use to pin a model to specific device indices (e.g. [ "nvidia.com/gpu=0" ]).

Type: list of string

Default:

config.services.llmhop.vllm-quadlet.devices

Example:

[
  "nvidia.com/gpu=0"
]

services.llmhop.vllm-quadlet.models.<name>.digest

Immutable digest of the container image (e.g. sha256:…). Mutually exclusive with tag.

Type: null or string

Default:

null

Example:

"sha256:a73fb0b9046fee099f7c1829d2548e6cc1740f4c2776a6855fa659ae5d0deb49"

services.llmhop.vllm-quadlet.models.<name>.environment

Additional environment variables set on this model’s service. Merged with services.llmhop.vllm-quadlet.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.vllm-quadlet.models.<name>.environmentFile

File in KEY=VALUE format forwarded to this model’s service. Loaded after services.llmhop.vllm-quadlet.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.vllm-quadlet.models.<name>.model

Hugging Face repo id (or local path) passed to the model server.

Type: string

Example:

"Qwen/Qwen2.5-7B-Instruct"

services.llmhop.vllm-quadlet.models.<name>.name

Canonical identifier for this model. Used for the unit name (vllm-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.vllm-quadlet.models.<name>.port

Loopback host port forwarded to the container’s vLLM API. Must be unique per model.

null binds the unix socket socket instead, which claims no port and only llmhop can connect to.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

null

services.llmhop.vllm-quadlet.models.<name>.quadlet.containerConfig

Extra [Container] settings applied to this model container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.models.<name>.quadlet.extraConfig

Extra unit sections applied to this model container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.models.<name>.quadlet.mountOptions.credentials

Podman volume options appended to this model container’s systemd credential mount. Use an idmap mapping when [Container] User= selects a non-root identity. The mount is always read-only.

Type: list of string

Default:

[ ]

Example:

[
  "idmap=uids=0-1000-1;gids=0-1000-1"
]

services.llmhop.vllm-quadlet.models.<name>.quadlet.mountOptions.socket

Podman volume options of this model container’s socket directory mount. U hands the directory to whatever host UID the container user maps to, so the socket works under any User= and UserNS=. Add z on SELinux hosts.

Type: list of string

Default:

[
  "U"
]

Example:

[
  "U"
  "z"
]

services.llmhop.vllm-quadlet.models.<name>.quadlet.quadletConfig

Extra [Quadlet] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.models.<name>.quadlet.serviceConfig

Extra [Service] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.models.<name>.quadlet.unitConfig

Extra [Unit] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.models.<name>.settings

CLI flags forwarded to the model server for this model. Rendered as --<key> <value>, with false as --no-<key> and a list handed to a single flag. See settings rendering for the full rules. Merged with services.llmhop.vllm-quadlet.modelSettings; per-model entries take precedence. The flags llmhop derives from the model options (its served name and listener) always win over both.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.models.<name>.shmSize

Size of the container’s private /dev/shm tmpfs. PyTorch and friends use shared memory for NCCL/tensor-parallel inference; upstream recommends 32g (or --ipc=host). A private tmpfs is preferred for isolation: raise the value for larger models or higher tensor-parallel sizes.

Type: string

Default:

"32g"

Example:

"64g"

services.llmhop.vllm-quadlet.models.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/vllm-<name>/http.sock while port is null, else null

services.llmhop.vllm-quadlet.models.<name>.tag

Tag of the container image used for this model. Mutually exclusive with digest.

Type: null or string

Default:

null

services.llmhop.vllm-quadlet.models.<name>.watermark

vLLM’s WatermarkConfig without key, validated at startup by the installed release. A generating worker and its detector must share it.

The key, an unsigned 64-bit integer, is the credential vllm.watermark-key. It is imported from the system credential store, for example the root-only file /etc/credstore/vllm.watermark-key, unless credentials."vllm.watermark-key" names another source. An encrypted credential must be created under that name, e.g. with systemd-creds encrypt --name=vllm.watermark-key.

Type: null or (attribute set of anything)

Default:

null

Example:

{
  algorithm = "dual_key_gumbel";
  context_width = 4;
}

services.llmhop.vllm-quadlet.openFilesLimit

File descriptor limit (LimitNOFILE) applied to every vllm-quadlet systemd unit. Increase if the server logs accept: Too many open files under concurrent load.

Type: positive integer, meaning >0

Default:

1048576

services.llmhop.vllm-quadlet.quadlet.containerConfig

Extra [Container] settings applied to every generated container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.quadlet.extraConfig

Extra unit sections applied to every generated container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.quadlet.quadletConfig

Extra [Quadlet] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.quadlet.serviceConfig

Extra [Service] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.quadlet.unitConfig

Extra [Unit] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.vllm-quadlet.quadlet.user

Host account whose systemd user manager owns these Quadlets. null installs system units and runs Podman rootfully. An attribute set installs user units for its uid and runs Podman rootlessly. This is independent of [Container] User= and the container’s user namespace configuration.

A managed account gets NixOS-allocated subordinate ID ranges. Set users.users.<name>.subUidRanges and subGidRanges (with autoSubUidGidRange = false) to pick them yourself.

Type: null or (submodule)

Default:

null

services.llmhop.vllm-quadlet.quadlet.user.gid

GID of the managed account’s primary group.

Type: positive integer, meaning >0

Default:

config.uid

services.llmhop.vllm-quadlet.quadlet.user.group

Primary group of the managed account.

Type: string

Default:

config.name

services.llmhop.vllm-quadlet.quadlet.user.home

Home directory used for rootless Podman storage. This must live on a filesystem that supports the selected storage driver.

Type: absolute path

Default:

"/var/lib/vllm"

services.llmhop.vllm-quadlet.quadlet.user.manage

Whether llmhop creates and configures this account. Disable this for an account managed elsewhere, including its home, linger setting, group, and subordinate ID ranges.

Type: boolean

Default:

true

services.llmhop.vllm-quadlet.quadlet.user.name

Host account whose systemd user manager owns the Quadlets.

Type: string

Default:

"vllm"

services.llmhop.vllm-quadlet.quadlet.user.uid

UID of the systemd user manager that owns the Quadlets.

Type: positive integer, meaning >0

Example:

503

services.llmhop.vllm-quadlet.startupOrdering

Whether to chain enabled model services by ascending name during startup. GPU-memory profiling races otherwise: two workers booting on the same device each see it as fully free and race to claim their share, leading to OOM. Disable only when each model pins itself to a dedicated device via its own devices.

Type: boolean

Default:

true

services.llmhop.vllm-quadlet.tag

Default tag of the container image used for models that do not set their own tag or digest.

Type: string

Example:

"latest"