services.llmhop.llama-cpp.enable
Whether to enable llama.cpp model serving via systemd, fronted by llmhop.
Type: boolean
Default:
false
Example:
true
services.llmhop.llama-cpp.package
The llama-cpp package to use.
Type: package
Default:
pkgs.llama-cpp
services.llmhop.llama-cpp.environment
Environment variables set on every model service.
Merged with services.llmhop.llama-cpp.models.<name>.environment; per-model
entries take precedence.
Type: attribute set of string
Default:
{ }
services.llmhop.llama-cpp.environmentFile
File in KEY=VALUE format forwarded to every service.
Use only for upstream features that require environment variables,
such as HF_TOKEN for gated Hugging Face repositories. Environment
variables are not systemd credentials and are visible to every model,
so prefer credentials for any secret a server can read from a file.
Loaded before services.llmhop.llama-cpp.models.<name>.environmentFile, so
per-model files override these entries.
Type: null or absolute path
Default:
null
Example:
"/etc/llama-cpp/.env"
services.llmhop.llama-cpp.group
Static primary group of the units, also beside a DynamicUser. The
module declares it while it keeps its default name, any other group
is the deployer’s to declare.
Type: string
Default:
"llama-cpp"
services.llmhop.llama-cpp.modelSettings
CLI flags forwarded to the model server for every model.
Rendered as --<key> <value>, with false as --no-<key> and a list repeating the flag. See settings rendering for the full rules.
Merged with services.llmhop.llama-cpp.models.<name>.settings; per-model
entries take precedence.
Type: attribute set of anything
Default:
{ }
services.llmhop.llama-cpp.models
Models to serve.
Each entry produces one systemd service running llama-server; the
attribute name is the routing key surfaced through llmhop and the OpenAI
model field.
GPU selection is done via build-specific environment variables on
environment (top-level or per-model), since llama.cpp runs as a host
process — no CDI involved. Common variables: CUDA_VISIBLE_DEVICES
(CUDA), HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES (ROCm),
GGML_VK_VISIBLE_DEVICES (Vulkan), ZE_AFFINITY_MASK (SYCL).
Type: attribute set of (submodule)
Default:
{ }
Example:
{
"qwen3-8b" = {
settings = {
hf-repo = "unsloth/Qwen3-8B-GGUF:UD-Q4_K_XL";
temperature = 1.0;
top-k = 20;
};
# Pin this model to a specific GPU. The right variable depends on
# the llama.cpp build: CUDA_VISIBLE_DEVICES for CUDA,
# HIP_VISIBLE_DEVICES / ROCR_VISIBLE_DEVICES for ROCm,
# GGML_VK_VISIBLE_DEVICES for Vulkan, ZE_AFFINITY_MASK for SYCL.
environment.CUDA_VISIBLE_DEVICES = "0";
};
}
services.llmhop.llama-cpp.models.<name>.enable
Whether to enable model ‹name›.
Type: boolean
Default:
true
Example:
true
services.llmhop.llama-cpp.models.<name>.credentials
Credentials granted exclusively to this service through systemd.
A path outside the Nix store uses LoadCredential=. The attribute form
can select LoadCredentialEncrypted= for a systemd-creds encrypted
source, which must be encrypted under the same name, or omit source
to import the credential of that name from the system credential store.
That store is shared by every service, so prefix an imported name with
its service, as in llmhop.hf-token.
Reference the resulting read-only file from settings as
${cred:<name>}. The module resolves the reference to the native or
container credential path without copying its contents to the Nix store
or command line.
Type: attribute set of ((submodule) or absolute path convertible to it)
Default:
{ }
Example:
{
api-keys = "/run/secrets/api-keys";
tls-key = {
source = "/run/secrets/tls-key.cred";
encrypted = true;
};
"llmhop.hf-token" = { };
}
services.llmhop.llama-cpp.models.<name>.credentials.<name>.encrypted
Whether to load and decrypt source with LoadCredentialEncrypted=.
Imported credentials are decrypted as needed.
Type: boolean
Default:
false
services.llmhop.llama-cpp.models.<name>.credentials.<name>.source
File or socket from which systemd loads the credential. null
imports the credential of the same name with ImportCredential=
from the system credential store, such as /etc/credstore and
/etc/credstore.encrypted, and from the credentials passed to the
system.
Type: null or absolute path not in the Nix store
Default:
null
services.llmhop.llama-cpp.models.<name>.environment
Additional environment variables set on this model’s service.
Merged with services.llmhop.llama-cpp.environment; per-model entries
take precedence.
Type: attribute set of string
Default:
{ }
services.llmhop.llama-cpp.models.<name>.environmentFile
File in KEY=VALUE format forwarded to this model’s service.
Loaded after services.llmhop.llama-cpp.environmentFile, so its entries
override global ones. Prefer credentials for any secret the server
can read from a file. Must be readable by the user systemd reads it as.
Type: null or absolute path
Default:
null
services.llmhop.llama-cpp.models.<name>.name
Canonical identifier for this model. Used for the unit name
(llama-cpp-<name>) and as the routing key registered with llmhop,
which clients send in the model field. Shares one namespace with
every other routing key, so a collision fails evaluation.
Defaults to the attribute key, so the key itself must match the required label format.
Type: string matching the pattern [[:alnum:]][[:alnum:].-]*
Default:
"‹name›"
services.llmhop.llama-cpp.models.<name>.port
Loopback host port that llama-server binds to. Must be unique per
enabled model; the gateway (llmhop) reaches each backend at
http://127.0.0.1:<port>.
null binds the unix socket socket instead, which claims no port
and only llmhop can connect to.
Type: null or 16 bit unsigned integer; between 0 and 65535 (both inclusive)
Default:
null
services.llmhop.llama-cpp.models.<name>.serviceConfig
Extra [Service] settings merged into this workload’s
llama-cpp-<name> unit after the hardened baseline and
backend-specific relaxations. The module retains ownership of
ExecStart, KillMode, and Type because they implement readiness
supervision as one lifecycle contract.
Type: attribute set of anything
Default:
{ }
Example:
{
MemoryHigh = "64G";
}
services.llmhop.llama-cpp.models.<name>.settings
CLI flags forwarded to the model server for this model.
Rendered as --<key> <value>, with false as --no-<key> and a list repeating the flag. See settings rendering for the full rules.
Merged with services.llmhop.llama-cpp.modelSettings; per-model entries
take precedence. The flags llmhop derives from the model options
(its served name and listener) always win over both.
Type: attribute set of anything
Default:
{ }
services.llmhop.llama-cpp.models.<name>.socket
Unix socket the server binds, derived from port.
Type: null or string (read only)
Default:
<services.llmhop.socketDirectory>/llama-cpp-<name>/http.sock while port is null, else null
services.llmhop.llama-cpp.models.<name>.unitConfig
Extra [Unit] settings merged into this workload’s
llama-cpp-<name> unit after the shared baseline.
Ordering and dependency directives (After=, Requires=, Wants=)
do not belong here: NixOS renders those from the after, requires
and wants options, so a definition of the same key in unitConfig
conflicts with it instead of merging. Declare them on
systemd.services."llama-cpp-<name>" from your own module,
where the module system concatenates them with what this one sets.
Type: attribute set of anything
Default:
{ }
Example:
{
StartLimitBurst = 10;
}
services.llmhop.llama-cpp.openFilesLimit
File descriptor limit (LimitNOFILE) applied to every llama-cpp systemd unit.
Increase if the server logs accept: Too many open files under concurrent load.
Type: positive integer, meaning >0
Default:
1048576
services.llmhop.llama-cpp.user
System user the units run as, which is the deployer’s to declare.
null allocates one per unit through DynamicUser=.
Type: null or string
Default:
null