Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

services.llmhop.llama-cpp-quadlet.enable

Whether to enable llama.cpp model serving via Quadlet, fronted by llmhop.

Type: boolean

Default:

false

Example:

true

services.llmhop.llama-cpp-quadlet.cache.containerDirectory

Path at which the cache is mounted inside every model container.

Type: string

Default:

"/root/.cache/llama.cpp"

services.llmhop.llama-cpp-quadlet.cache.directory

Host directory bind-mounted as the Hugging Face cache.

Type: absolute path

Default:

"/var/cache/llama-cpp"

services.llmhop.llama-cpp-quadlet.cache.environmentVariable

Environment variable set on every container to point its runtime at containerDirectory.

Type: string

Default:

"LLAMA_CACHE"

services.llmhop.llama-cpp-quadlet.cache.group

Host group used when cache.manage is enabled.

Type: string

Default:

if config.services.llmhop.llama-cpp-quadlet.quadlet.user == null then
  "root"
else
  config.services.llmhop.llama-cpp-quadlet.quadlet.user.group

services.llmhop.llama-cpp-quadlet.cache.manage

Whether llmhop creates the host cache directory with systemd-tmpfiles, owned by user/group and mode 0700.

Type: boolean

Default:

true

services.llmhop.llama-cpp-quadlet.cache.mountOptions

Options appended to the cache’s Quadlet Volume= entry. This can be used for Podman ownership mechanisms such as U, idmap, or SELinux relabeling.

Type: list of string

Default:

[ ]

Example:

[
  "idmap"
]

services.llmhop.llama-cpp-quadlet.cache.user

Host owner used when cache.manage is enabled. Override it when a UIDMap/idmap mapping makes the container see a different owner than the host account running Podman.

Type: string

Default:

if config.services.llmhop.llama-cpp-quadlet.quadlet.user == null then
  "root"
else
  config.services.llmhop.llama-cpp-quadlet.quadlet.user.name

services.llmhop.llama-cpp-quadlet.devices

Devices exposed to every model container — passed verbatim as Quadlet AddDevice= lines. Accepts both CDI references (recommended: nvidia.com/gpu=…, amd.com/gpu=…, intel.com/gpu=…, …) and raw host device paths (e.g. /dev/dri/renderD128). For CDI, the corresponding spec must be generated on the host (e.g. nvidia-ctk cdi generate). Defaults to [ "nvidia.com/gpu=all" ] when hardware.nvidia-container-toolkit.enable is set, otherwise empty (CPU-only). Per-model devices overrides this.

Type: list of string

Default:

if config.hardware.nvidia-container-toolkit.enable then
  [ "nvidia.com/gpu=all" ]
else
  [ ]

Example:

[
  "amd.com/gpu=all"
]

services.llmhop.llama-cpp-quadlet.environment

Environment variables set on every model service. Merged with services.llmhop.llama-cpp-quadlet.models.<name>.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.llama-cpp-quadlet.environmentFile

File in KEY=VALUE format forwarded to every service. Use only for upstream features that require environment variables, such as HF_TOKEN for gated Hugging Face repositories. Environment variables are not systemd credentials and are visible to every model, so prefer credentials for any secret a server can read from a file. Loaded before services.llmhop.llama-cpp-quadlet.models.<name>.environmentFile, so per-model files override these entries.

Type: null or absolute path

Default:

null

Example:

"/etc/llama-cpp-quadlet/.env"

services.llmhop.llama-cpp-quadlet.image

Container image used for every model worker.

Type: string

Default:

"ghcr.io/ggml-org/llama.cpp"

services.llmhop.llama-cpp-quadlet.modelSettings

CLI flags forwarded to the model server for every model. Rendered as --<key> <value>, with false as --no-<key> and a list repeating the flag. See settings rendering for the full rules. Merged with services.llmhop.llama-cpp-quadlet.models.<name>.settings; per-model entries take precedence.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models

Models served by llama-server containers. Each entry produces one llama-cpp-<name> unit and uses the attribute name as its llmhop routing key and llama.cpp --alias.

Type: attribute set of (submodule)

Default:

{ }

Example:

{
  "qwen3-8b" = {
    settings.hf-repo = "unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL";
  };
}

services.llmhop.llama-cpp-quadlet.models.<name>.enable

Whether to enable model ‹name›.

Type: boolean

Default:

true

Example:

true

services.llmhop.llama-cpp-quadlet.models.<name>.credentials

Credentials granted exclusively to this service through systemd. A path outside the Nix store uses LoadCredential=. The attribute form can select LoadCredentialEncrypted= for a systemd-creds encrypted source, which must be encrypted under the same name, or omit source to import the credential of that name from the system credential store. That store is shared by every service, so prefix an imported name with its service, as in llmhop.hf-token.

Reference the resulting read-only file from settings as ${cred:<name>}. The module resolves the reference to the native or container credential path without copying its contents to the Nix store or command line.

Type: attribute set of ((submodule) or absolute path convertible to it)

Default:

{ }

Example:

{
  api-keys = "/run/secrets/api-keys";
  tls-key = {
    source = "/run/secrets/tls-key.cred";
    encrypted = true;
  };
  "llmhop.hf-token" = { };
}

services.llmhop.llama-cpp-quadlet.models.<name>.credentials.<name>.encrypted

Whether to load and decrypt source with LoadCredentialEncrypted=. Imported credentials are decrypted as needed.

Type: boolean

Default:

false

services.llmhop.llama-cpp-quadlet.models.<name>.credentials.<name>.source

File or socket from which systemd loads the credential. null imports the credential of the same name with ImportCredential= from the system credential store, such as /etc/credstore and /etc/credstore.encrypted, and from the credentials passed to the system.

Type: null or absolute path not in the Nix store

Default:

null

services.llmhop.llama-cpp-quadlet.models.<name>.devices

Devices exposed to this model’s container — passed verbatim as Quadlet AddDevice= lines. Replaces (does not extend) services.llmhop.llama-cpp-quadlet.devices for this model. Use to pin a model to specific device indices (e.g. [ "nvidia.com/gpu=0" ]).

Type: list of string

Default:

config.services.llmhop.llama-cpp-quadlet.devices

Example:

[
  "nvidia.com/gpu=0"
]

services.llmhop.llama-cpp-quadlet.models.<name>.digest

Immutable digest of the container image (e.g. sha256:…). Mutually exclusive with tag.

Type: null or string

Default:

null

Example:

"sha256:a73fb0b9046fee099f7c1829d2548e6cc1740f4c2776a6855fa659ae5d0deb49"

services.llmhop.llama-cpp-quadlet.models.<name>.environment

Additional environment variables set on this model’s service. Merged with services.llmhop.llama-cpp-quadlet.environment; per-model entries take precedence.

Type: attribute set of string

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.environmentFile

File in KEY=VALUE format forwarded to this model’s service. Loaded after services.llmhop.llama-cpp-quadlet.environmentFile, so its entries override global ones. Prefer credentials for any secret the server can read from a file. Must be readable by the user systemd reads it as.

Type: null or absolute path

Default:

null

services.llmhop.llama-cpp-quadlet.models.<name>.name

Canonical identifier for this model. Used for the unit name (llama-cpp-<name>) and as the routing key registered with llmhop, which clients send in the model field. Shares one namespace with every other routing key, so a collision fails evaluation.

Defaults to the attribute key, so the key itself must match the required label format.

Type: string matching the pattern [[:alnum:]][[:alnum:].-]*

Default:

"‹name›"

services.llmhop.llama-cpp-quadlet.models.<name>.port

Loopback host port forwarded to the container’s llama.cpp API. Must be unique per model.

null binds the unix socket socket instead, which claims no port and only llmhop can connect to.

Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)

Default:

null

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.containerConfig

Extra [Container] settings applied to this model container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.extraConfig

Extra unit sections applied to this model container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.mountOptions.credentials

Podman volume options appended to this model container’s systemd credential mount. Use an idmap mapping when [Container] User= selects a non-root identity. The mount is always read-only.

Type: list of string

Default:

[ ]

Example:

[
  "idmap=uids=0-1000-1;gids=0-1000-1"
]

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.mountOptions.socket

Podman volume options of this model container’s socket directory mount. U hands the directory to whatever host UID the container user maps to, so the socket works under any User= and UserNS=. Add z on SELinux hosts.

Type: list of string

Default:

[
  "U"
]

Example:

[
  "U"
  "z"
]

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.quadletConfig

Extra [Quadlet] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.serviceConfig

Extra [Service] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.quadlet.unitConfig

Extra [Unit] settings applied to this model container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.settings

CLI flags forwarded to the model server for this model. Rendered as --<key> <value>, with false as --no-<key> and a list repeating the flag. See settings rendering for the full rules. Merged with services.llmhop.llama-cpp-quadlet.modelSettings; per-model entries take precedence. The flags llmhop derives from the model options (its served name and listener) always win over both.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.models.<name>.shmSize

Size of the container’s private /dev/shm tmpfs. PyTorch and friends use shared memory for NCCL/tensor-parallel inference; upstream recommends 32g (or --ipc=host). A private tmpfs is preferred for isolation: raise the value for larger models or higher tensor-parallel sizes.

Type: string

Default:

"32g"

Example:

"64g"

services.llmhop.llama-cpp-quadlet.models.<name>.socket

Unix socket the server binds, derived from port.

Type: null or string (read only)

Default: <services.llmhop.socketDirectory>/llama-cpp-<name>/http.sock while port is null, else null

services.llmhop.llama-cpp-quadlet.models.<name>.tag

Tag of the container image used for this model. Mutually exclusive with digest.

Type: null or string

Default:

null

services.llmhop.llama-cpp-quadlet.openFilesLimit

File descriptor limit (LimitNOFILE) applied to every llama-cpp-quadlet systemd unit. Increase if the server logs accept: Too many open files under concurrent load.

Type: positive integer, meaning >0

Default:

1048576

services.llmhop.llama-cpp-quadlet.quadlet.containerConfig

Extra [Container] settings applied to every generated container. Keys use Quadlet’s native PascalCase names, including User, UserNS, UIDMap, GIDMap, SubUIDMap, and SubGIDMap.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.extraConfig

Extra unit sections applied to every generated container after all generated sections. This is the final escape hatch for settings that do not fit one of the dedicated *Config options.

Type: attribute set of attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.quadletConfig

Extra [Quadlet] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.serviceConfig

Extra [Service] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.unitConfig

Extra [Unit] settings applied to every generated container.

Type: attribute set of anything

Default:

{ }

services.llmhop.llama-cpp-quadlet.quadlet.user

Host account whose systemd user manager owns these Quadlets. null installs system units and runs Podman rootfully. An attribute set installs user units for its uid and runs Podman rootlessly. This is independent of [Container] User= and the container’s user namespace configuration.

A managed account gets NixOS-allocated subordinate ID ranges. Set users.users.<name>.subUidRanges and subGidRanges (with autoSubUidGidRange = false) to pick them yourself.

Type: null or (submodule)

Default:

null

services.llmhop.llama-cpp-quadlet.quadlet.user.gid

GID of the managed account’s primary group.

Type: positive integer, meaning >0

Default:

config.uid

services.llmhop.llama-cpp-quadlet.quadlet.user.group

Primary group of the managed account.

Type: string

Default:

config.name

services.llmhop.llama-cpp-quadlet.quadlet.user.home

Home directory used for rootless Podman storage. This must live on a filesystem that supports the selected storage driver.

Type: absolute path

Default:

"/var/lib/llama-cpp"

services.llmhop.llama-cpp-quadlet.quadlet.user.manage

Whether llmhop creates and configures this account. Disable this for an account managed elsewhere, including its home, linger setting, group, and subordinate ID ranges.

Type: boolean

Default:

true

services.llmhop.llama-cpp-quadlet.quadlet.user.name

Host account whose systemd user manager owns the Quadlets.

Type: string

Default:

"llama-cpp"

services.llmhop.llama-cpp-quadlet.quadlet.user.uid

UID of the systemd user manager that owns the Quadlets.

Type: positive integer, meaning >0

Example:

503

services.llmhop.llama-cpp-quadlet.startupOrdering

Whether to chain enabled model services by ascending name during startup. GPU-memory profiling races otherwise: two workers booting on the same device each see it as fully free and race to claim their share, leading to OOM. Disable only when each model pins itself to a dedicated device via its own devices.

Type: boolean

Default:

true

services.llmhop.llama-cpp-quadlet.tag

Default tag of the container image used for models that do not set their own tag or digest.

Type: string

Example:

"server-cuda"