Environment variables
This page lists every environment variable SSEBench reads or sets. The tables are generated from the registry in docs/reference/env.yaml, and just docs-check and the test suite fail when the code uses a variable the registry lacks.
On the host
Provider keys
Put the keys of your model providers in .env in the repository root. The LiteLLM proxy reads provider keys only from this file, when it starts.
ANTHROPIC_API_KEY=sk-ant-...
OPENAI_API_KEY=sk-...
GOOGLE_API_KEY=...| Variable | Default | Used by | Description |
|---|---|---|---|
ANTHROPIC_API_KEY | unset | LiteLLM proxy, CLI doctor, sse.ai | Key of the Anthropic models in models/anthropic-claude.yaml. sse.ai also uses it when it runs OpenCode outside a task container. |
OPENAI_API_KEY | unset | LiteLLM proxy, CLI doctor | Key of the OpenAI models in models/openai-gpt.yaml. |
GOOGLE_API_KEY | unset | LiteLLM proxy, CLI doctor | Key of the Gemini models in models/google-gemini.yaml. |
Set the key of each provider you use and leave out the others: the proxy still lists every model, but a model fails when it is called without its key. A model you add can use any other variable name, for example:
AZURE_API_KEY=...
AZURE_API_BASE=https://your-resource.openai.azure.com/
HF_TOKEN=hf_...In model files, refer to a variable with the os.environ/ prefix instead of writing the key itself:
- model_name: gpt-5.1
litellm_params:
model: openai/gpt-5.1
api_key: os.environ/OPENAI_API_KEYWARNING
Never write API keys into files that are committed. .env is listed in .gitignore.
After you change .env, restart the proxy with just launch. After you change models/, the next just launch or ssebench run rebuilds the proxy image.
SSEBench settings
just setup writes .env from .env.example; see Configuration files. The CLI reads these settings from the environment or from .env in the SSEBench home, with the environment taking precedence, and the just recipes and Compose load .env too.
| Variable | Default | Used by | Description |
|---|---|---|---|
LITELLM_MASTER_KEY | generated by just setup | CLI, LiteLLM proxy, web UI | Admin key of the local LiteLLM proxy; the CLI uses it to create a key for each run, and the web UI to make the key of a run's AI assistant. Required. |
POSTGRES_PASSWORD | generated by just setup | LiteLLM proxy | Password of the proxy's Postgres database; letters and digits only, as it is part of a URL. Postgres keeps the password it was created with, so after changing it remove the stack's volume. Required. |
POSTGRES_USER | litellm | LiteLLM proxy | User of the proxy's Postgres database. |
LITELLM_PORT | 4000 | CLI, LiteLLM proxy, web UI | Host port of the LiteLLM proxy. |
LITELLM_BIND | 127.0.0.1 | Compose | Host address that the LiteLLM proxy's port is published on. The proxy holds the master key and the provider keys, so it listens on loopback only; set 0.0.0.0 to reach it from other machines, and only on a network you trust. |
COMPOSE_PROJECT_NAME | ssebench | CLI, Compose | Compose project of the proxy stack. Its containers, networks (<project>_default, and the internal <project>_agents that run containers join by default) and database volume (<project>_postgres_data) carry this name, so stacks with different names and ports run side by side. |
SSEBENCH_REGISTRY | ghcr.io/42-b3yond-6ug/ssebench | CLI, just, Compose, catalog service | Registry prefix of every image SSEBench builds or pulls, and of the image names that the catalog service returns. |
SSEBENCH_CATALOG | the bundled pilot manifest | CLI, just, web UI | Task catalog that ssebench run, ssebench tasks list and the web UI get tasks from: the path or URL of a manifest.json, a dataset directory, or the URL of a catalog service. --catalog overrides it. |
SSEBENCH_IMAGES_LOCK | images.lock.json next to the manifest, if there is one | CLI, catalog service | Path of the images lock that pins case images by digest, instead of the one next to the manifest: ssebench run and ssebench-catalog serve (--lock) use it, and ssebench run also takes a URL; see Case images. A lock that is missing or invalid is an error. |
SSEBENCH_HOME | see CLI | CLI | Directory that holds agents/, images/, runtime/, models/, deploy/ and datasets/, normally the repository root. Unset, a package install without a clone uses the copy of these directories in the package. |
SSEBENCH_ENV_FILE | .env in the repository root | CLI, Compose | The .env file that the LiteLLM container reads provider keys from. The CLI sets it to the .env in the workspace; set it yourself only when you run docker compose on deploy/compose/docker-compose.yaml directly. |
SSEBENCH_DEMO_PROJECT | ssebench-demo | CLI, Compose | Compose project of the demo. It must differ from COMPOSE_PROJECT_NAME: just demo-down deletes the project's database volume, and the demo refuses to touch a project whose containers it did not create. |
SSEBENCH_DEMO_WEBUI_PORT | 3001 | CLI, Compose | Host port of the demo's web UI, on 127.0.0.1. |
SSEBENCH_DEMO_CATALOG_PORT | 8090 | CLI, Compose | Host port of the demo's catalog service, on 127.0.0.1. |
SSEBENCH_BACKEND | docker | CLI, web UI | Runner backend that ssebench run and ssebench runs use when --backend is not given: where the run's containers execute. The web UI runs those commands, so it works with the backend named here. Installed extensions can add backends; see Runner backends. |
SSEBENCH_PREBUILT | unset | CLI | 1, true, yes or on makes ssebench run use the published agent images of the task, under SSEBENCH_REGISTRY, instead of building the case, tool and agent layers; the same as --prebuilt. See Prebuilt images. |
SSEBENCH_EGRESS | restricted | CLI, web UI | Network egress of a run when --egress is not given: restricted reaches the LiteLLM proxy and nothing else, open also has internet access. The web UI's launches use it unless the launch names a policy. The Helm chart sets it from runs.egress. See Integrity and egress. |
SSEBENCH_RUNTIME_IMAGE | <registry>/runtime:<version> without a checkout, else unset | CLI | Runtime image that the tool layers copy the daemon and the entrypoint from, as <registry>/runtime:<version> does. Unset, a checkout compiles both from its sources, and an installation without a checkout uses the published image of its version. Set it to skip the compilation in a checkout. The image must exist for the platform of the run, which is linux/amd64 for a pilot task. |
SSEBENCH_VERSION | set by the CLI | CLI, Compose | SSEBench version that the Compose files pass as the VERSION build argument of the images they build. The CLI sets it from VERSION; it only matters when you run docker compose build yourself. |
Kubernetes backend
ssebench run --backend kubernetes reads these settings the same way. See Kubernetes for what they configure.
| Variable | Default | Used by | Description |
|---|---|---|---|
SSEBENCH_K8S_NAMESPACE | the namespace of the kubeconfig context, or of the pod ssebench runs in, else default | CLI | Namespace where the runs' Jobs, Secrets and NetworkPolicies are created. |
SSEBENCH_K8S_CONTEXT | the kubeconfig's current context | CLI | Context of the kubeconfig to use. The kubeconfig itself is found the way kubectl finds it (KUBECONFIG, else ~/.kube/config); inside a pod the backend uses the pod's service account. |
SSEBENCH_K8S_PROXY_URL | http://litellm.<proxy namespace>.svc:4000 | CLI | URL of the LiteLLM proxy as a run's pod reaches it: the run's SSE_BASE_URL. Its port must be the port the proxy's pods listen on, since the run's network policy opens that port. |
SSEBENCH_K8S_PROXY_HOST_URL | the proxy's Service URL inside a cluster, else http://localhost:$LITELLM_PORT | CLI | URL of the LiteLLM proxy as the ssebench process reaches it, to create and read each run's key. From outside the cluster, forward the proxy's port (kubectl port-forward svc/litellm 4000) and set LITELLM_PORT or this. |
SSEBENCH_K8S_PROXY_NAMESPACE | the runs' namespace | CLI | Namespace of the LiteLLM proxy, for its default URL and the run's network policy. |
SSEBENCH_K8S_PROXY_SELECTOR | app.kubernetes.io/name=litellm | CLI | Labels of the LiteLLM proxy's pods, as key=value,key=value. A restricted run's network policy lets its pod reach the pods that match, and no others. |
SSEBENCH_K8S_RUNTIME_CLASS | unset | CLI | runtimeClassName of the run's pod, for a sandboxed runtime such as gVisor or Kata Containers. |
SSEBENCH_K8S_IMAGE_PULL_POLICY | IfNotPresent | CLI | imagePullPolicy of the run's containers. IfNotPresent uses an image that is already on the node, such as one loaded into a kind cluster. |
SSEBENCH_K8S_IMAGE_PULL_SECRETS | unset | CLI | Names of image pull Secrets in the runs' namespace, separated by commas. |
SSEBENCH_K8S_RESOURCES | requests of 1 CPU, 2Gi memory and 2Gi ephemeral storage; limits of 4 CPUs, 8Gi and 20Gi | CLI | Resource requests and limits of the task container, as JSON or YAML in the shape of a container's resources, for example {"requests": {"cpu": "2"}, "limits": {"memory": "16Gi"}}. It replaces the defaults as a whole. |
SSEBENCH_K8S_TTL_SECONDS | 3600 | CLI | How long the cluster keeps a run's finished Job, its pod and its logs (ttlSecondsAfterFinished). A run kept with --keep-container has none. |
SSEBENCH_K8S_DEADLINE_SLACK | 1800 | CLI | Seconds added to the run's --timeout to make the Job's activeDeadlineSeconds: the time for pulling images, starting and grading. When it runs out the cluster deletes the pod and the run's results with it. |
KUBERNETES_SERVICE_HOST | set by Kubernetes in every pod | CLI | The backend takes its presence to mean that ssebench runs inside a cluster, where the proxy is reached through its Service. |
Inside the task container
ssebench run and the container's entrypoint set these variables in every task container. Agents and plugins read them.
| Variable | Default | Used by | Description |
|---|---|---|---|
SSE_API_KEY | set by ssebench run | agents, plugins, sse.ai | The run's LiteLLM key, which can use only the selected model; empty in a reference run. |
SSE_BASE_URL | set by ssebench run | agents, plugins, sse.ai | URL of the LiteLLM proxy, http://litellm:4000; empty in a reference run. |
SSE_MODEL_NAME | set by ssebench run | agents, plugins, sse.ai | The selected model, as named in models/*.yaml; none in a reference run. |
SSE_ARCHIVE | /tmp/sse-archive | entrypoint, daemon, evaluator, agents | The agent's archive directory, /tmp/sse-archive, writable by the agent: dialog.jsonl and whatever the agent side writes go there. The entrypoint requires it. The graded outputs go to SSE_RESULTS instead, root-only. |
SSE_RESULTS | /var/lib/ssebench/results | entrypoint, daemon, evaluator | The run's results directory, root-only: the grade (result.json), the graded patch (final.patch), the commit log and the run's logs. Its parent is made 0700 root, so neither the agent nor the task runner can reach it. The CLI mounts the run directory here and its archive/ subdirectory at SSE_ARCHIVE. Falls back to SSE_ARCHIVE when unset. |
SSE_DIFFICULTY | 2 | daemon, MCP server | The difficulty level, from 0 to 4. The MCP server decides from it which checks test_patch runs, and the daemon refuses the withheld bencher actions on its agent-facing listeners. |
TIMEOUT | 14400 in the entrypoint, 1800 in the evaluator | entrypoint, evaluator, agents | How long the agent may run, in seconds (--timeout). The evaluator uses the same limit for grading. |
SSE_KEEP_ALIVE | 0 | entrypoint | 1 keeps the container running after grading (--keep-container), for the web UI. |
SSE_PLUGINS | the plugins plugins.yaml enables | entrypoint | Comma-separated plugins to run, set by ssebench run --plugin; when it is set, it replaces the enabled field of plugins.yaml, and an empty value runs none. See Plugins and hooks. |
SSE_DAEMON_SOCKET | /tmp/sse.sock | entrypoint, daemon, SDK | The daemon's agent-facing Unix socket, mode 0666. The entrypoint sets it for every process it starts; in sidecar mode ssebench run sets it to /run/ssebench/sse.sock, on a root-owned volume the two containers share. Without it, the daemon serves HTTP only. |
SSE_ADMIN_SOCKET | /run/ssebench/admin.sock | entrypoint, daemon | The daemon's privileged Unix socket, mode 0600, root only: grading, the reference patch and phase changes. In sidecar mode it is on the volume the two containers share, and the task container's entrypoint sets it. The daemon binds it only when this is set; the entrypoint sets it, and points the evaluator's SSE_DAEMON_SOCKET at it. The SDK's sse.reference.get_reference_patch reads it to reach the admin socket. |
The components in the container also read these settings, which have working defaults:
| Variable | Default | Used by | Description |
|---|---|---|---|
SSE_AGENT_USER | model | daemon | The agent's user, whose processes the daemon kills when the agent phase ends. |
SSE_RUNNER_USER | sse-runner | daemon | The unprivileged user the daemon runs the task's build, PoC and test scripts as, when it runs as root. A dedicated uid with no groups, neither the agent nor root; the tool layer creates it. The daemon refuses to start as root without it. |
SSE_RUNNER_DIR | /var/lib/ssebench-runner | daemon | The root of the task runner's scratch copies, reachable only by the runner and root (0710). Each check runs in a private copy of the project here. |
SSE_DAEMON_TIMEOUT | 300 | entrypoint | Seconds to wait for the daemon's socket. |
SSE_MCP_TIMEOUT | 300 | entrypoint | Seconds to wait for the MCP server. |
SSE_DEBUG | unset | entrypoint | Any non-empty value turns on debug logs. |
SSE_PLUGIN_NAME | unset | plugins | Set by the entrypoint for a plugin it runs; the plugin's name. |
SSE_PLUGIN_HOOK | unset | plugins | Set by the entrypoint for a plugin it runs; the hook it runs at, such as after-grading. |
SSE_ORACLE_FUZZ | unset | oracle plugin | 1 makes the oracle plugin fuzz the patched project after its review; it installs AFL++, so the run needs --egress open. |
SSE_BENCH_PATH | /ssebench | daemon | Directory with the task's config.yaml, scripts and files. |
SSE_HTTP_PORT | 4263 | daemon | The daemon's agent-facing HTTP port. |
SSE_DAEMON_WORKERS | 4 | daemon | Worker threads for each of the daemon's listeners. A tool call occupies its worker until the command ends. Without a limit, actix starts one worker per host CPU for every listener. |
SSE_REPO_PATH | /ssebench-repo | daemon | Clean clone of the project that grading applies the agent's diff to and builds. |
SSE_AGENT_DOCKER | unset | SDK | host:port of the daemon's HTTP listener, used when SSE_DAEMON_SOCKET is not set. |
MCP_LOG_DIR | /tmp/mcp/logs | MCP server | Where the MCP server writes the full logs of long check results; see Long logs. |
AGENT_DURATION | 0 | evaluator | The agent's run time in seconds; the entrypoint sets it for the evaluator. |
SSE_METRIC_AGENT_TIMEOUT | unset | evaluator | Set to true by the entrypoint for the evaluator when the agent hit its time limit. |
CLAUDE | unset | claude-code agent | Path of the Claude Code executable; the claude-code agent image sets it. |
OPENCODE_CONFIG_CONTENT | unset | OpenCode | OpenCode's configuration as JSON; sse.ai and the opencode agent set it for the OpenCode server they start. |
Difficulty levels
SSE_DIFFICULTY decides which checks the agent's test_patch tool runs. Final grading always runs every check the task has.
| Level | Name |
|---|---|
| 0 | FULL_ASSISTANCE |
| 1 | NO_INTENT_TEST |
| 2 | NO_FUTURE_TEST (default) |
| 3 | BUILD_ONLY |
| 4 | NO_BUILD |
See Difficulty levels for what each level allows and how it is enforced, and Daemon HTTP API for the daemon's gate.
Web UI
The web UI server reads these; see WebUI.
| Variable | Default | Used by | Description |
|---|---|---|---|
SSEBENCH_CATALOG | the bundled pilot manifest | CLI, just, web UI | Task catalog that ssebench run, ssebench tasks list and the web UI get tasks from: the path or URL of a manifest.json, a dataset directory, or the URL of a catalog service. --catalog overrides it. |
SSEBENCH_WEBUI_HOST | 127.0.0.1 | web UI | Bind address of the API server and of the Vite dev and preview servers. |
PORT | 3001 | web UI | API server port; the Vite servers forward /api there. |
SSEBENCH_WEBUI_TOKEN | unset | web UI | Access token, at least 16 characters. Required when the bind address is not loopback. |
SSEBENCH_WEBUI_CORS_ORIGINS | unset | web UI | Comma-separated origins, besides the server's own, allowed to call the API, such as https://bench.example.org. |
SSEBENCH_WEBUI_TERMINAL | 1 | web UI | 0 turns off the container terminal. The demo's web UI has it off (0) unless you set this to 1. Hosted mode (SSEBENCH_WEBUI_HOSTED) turns it off whatever this says. |
SSEBENCH_WEBUI_HOSTED | 0 | web UI | 1 runs the web UI as a read-only viewer, for a showcase or a shared server. It refuses to launch, stop or remove runs, opens no terminal, and lets the AI assistant run no commands in a container; runs and finished runs can still be read. See Security model. |
SSEBENCH_CLI | uv run ssebench | web UI | The command that runs the ssebench CLI, split on white space and started without a shell in SSEBENCH_PATH, such as ssebench or uv run --project /opt/ssebench ssebench. The web UI finds, stops, reaches and launches runs through it. The web UI image sets it to ssebench. |
SSEBENCH_PATH | the repository root | web UI | SSEBench checkout: the CLI runs there, so results/ is read and written there, and models/ and agents/ are read from there. |
SSEBENCH_LOCAL_TASKS | $SSEBENCH_PATH/datasets/pilot | web UI | Local dataset offered in the launcher. |
SSEBENCH_CATALOG_URL | unset | web UI | Former name of SSEBENCH_CATALOG; deprecated, and read only when that is unset. |
PTY_ENABLE_CLEANUP | true | web UI | false turns off the cleanup of terminal processes whose browser went away. |
PTY_CLEANUP_INTERVAL | 30000 | web UI | Milliseconds between two cleanups of terminal processes. |
PTY_ORPHAN_THRESHOLD | 300000 | web UI | Milliseconds without activity after which a terminal process is cleaned up. |
Catalog service
The catalog service reads these; each has a command-line option too. See Dataset manifest.
| Variable | Default | Used by | Description |
|---|---|---|---|
SSEBENCH_REGISTRY | ghcr.io/42-b3yond-6ug/ssebench | CLI, just, Compose, catalog service | Registry prefix of every image SSEBench builds or pulls, and of the image names that the catalog service returns. |
SSEBENCH_IMAGES_LOCK | images.lock.json next to the manifest, if there is one | CLI, catalog service | Path of the images lock that pins case images by digest, instead of the one next to the manifest: ssebench run and ssebench-catalog serve (--lock) use it, and ssebench run also takes a URL; see Case images. A lock that is missing or invalid is an error. |
SSEBENCH_MANIFEST | datasets/pilot/manifest.json | catalog service | The manifest.json to serve (--manifest); /srv/manifest.json in the catalog image. |
SSEBENCH_CATALOG_PORT | 8080 | catalog service | TCP port to listen on (--port). |
Development
| Variable | Default | Used by | Description |
|---|---|---|---|
SSEBENCH_INTEGRITY_IMAGE | the tool image that ssebench run builds for gjson-196-bf4efcb | integrity tests | Sandbox tool image that tests/integrity attacks. |
SSEBENCH_INTEGRITY_SIDECAR_ENV_IMAGE | the task image that ssebench run --mode sidecar builds for gjson-196-bf4efcb | integrity tests | Sidecar task image, with the daemon, that tests/integrity attacks. |
SSEBENCH_INTEGRITY_SIDECAR_AGENT_IMAGE | the sidecar runtime image of this version | integrity tests | Sidecar agent runtime image that tests/integrity runs the fake agent in. |
SSEBENCH_INTEGRITY_PLUGINS | unset | integrity tests | Plugins the sandbox container of the integrity tests enables, passed as SSE_PLUGINS; the image must have them installed. |
SSEBENCH_INTEGRITY_SOURCE | /src/gjson | integrity tests | Source directory inside those images. |
SSE_DAEMON_HTTP | http://localhost:4263 | integrity tests | The daemon's HTTP listener that tests/integrity/fake_agent.sh probes; in sidecar mode, the task container's. |
SSEBENCH_SMOKE_TASK | gjson-196-bf4efcb | tests/e2e/smoke.sh | Task of the end-to-end smoke run. |
SSEBENCH_SMOKE_MODEL | claude-sonnet-4-6 | tests/e2e/smoke.sh | Model the smoke run names; the dummy agent never calls it. |
SSEBENCH_UPDATE_SNAPSHOTS | unset | SDK tests | 1 rewrites the snapshot of the task prompt in sdk/python/tests/snapshots/ instead of comparing with it; see Task prompt. |
Using the variables in an agent
Python
import os
# LLM access
api_key = os.environ["SSE_API_KEY"]
base_url = os.environ["SSE_BASE_URL"]
model = os.environ["SSE_MODEL_NAME"]
# Run settings
difficulty = int(os.environ.get("SSE_DIFFICULTY", "2"))
timeout = int(os.environ.get("TIMEOUT", "3600"))
archive_path = os.environ.get("SSE_ARCHIVE", "/tmp/sse-archive")Shell
#!/bin/bash
echo "Model: $SSE_MODEL_NAME"
echo "Difficulty: $SSE_DIFFICULTY"
# List the models this run's key can use
curl -H "Authorization: Bearer $SSE_API_KEY" "$SSE_BASE_URL/models"Troubleshooting
- A model fails with an authentication error. Check that its key is set in
.envunder the name its model file uses, then runjust launchso the proxy picks it up. - The proxy does not start. Run
just doctor.LITELLM_MASTER_KEYandPOSTGRES_PASSWORDmust be set (just setupwrites them), andLITELLM_PORTmust be free. - You want to see a container's variables. Start the run with
--keep-container, then rundocker exec <container> env.
Next steps
- Configuration files:
.env, the model, agent, plugin and task files - Add a model: configure model providers
- MCP server: the
test_patchtool - CLI: the options of
ssebench run