Skip to content

Environment variables ​

This page lists every environment variable SSEBench reads or sets. The tables are generated from the registry in docs/reference/env.yaml, and just docs-check and the test suite fail when the code uses a variable the registry lacks.

On the host ​

Provider keys ​

Put the keys of your model providers in .env in the repository root. The LiteLLM proxy reads provider keys only from this file, when it starts.

sh
ANTHROPIC_API_KEY=sk-ant-...
OPENAI_API_KEY=sk-...
GOOGLE_API_KEY=...
VariableDefaultUsed byDescription
ANTHROPIC_API_KEYunsetLiteLLM proxy, CLI doctor, sse.aiKey of the Anthropic models in models/anthropic-claude.yaml. sse.ai also uses it when it runs OpenCode outside a task container.
OPENAI_API_KEYunsetLiteLLM proxy, CLI doctorKey of the OpenAI models in models/openai-gpt.yaml.
GOOGLE_API_KEYunsetLiteLLM proxy, CLI doctorKey of the Gemini models in models/google-gemini.yaml.

Set the key of each provider you use and leave out the others: the proxy still lists every model, but a model fails when it is called without its key. A model you add can use any other variable name, for example:

sh
AZURE_API_KEY=...
AZURE_API_BASE=https://your-resource.openai.azure.com/
HF_TOKEN=hf_...

In model files, refer to a variable with the os.environ/ prefix instead of writing the key itself:

yaml
- model_name: gpt-5.1
  litellm_params:
    model: openai/gpt-5.1
    api_key: os.environ/OPENAI_API_KEY

WARNING

Never write API keys into files that are committed. .env is listed in .gitignore.

After you change .env, restart the proxy with just launch. After you change models/, the next just launch or ssebench run rebuilds the proxy image.

SSEBench settings ​

just setup writes .env from .env.example; see Configuration files. The CLI reads these settings from the environment or from .env in the SSEBench home, with the environment taking precedence, and the just recipes and Compose load .env too.

VariableDefaultUsed byDescription
LITELLM_MASTER_KEYgenerated by just setupCLI, LiteLLM proxy, web UIAdmin key of the local LiteLLM proxy; the CLI uses it to create a key for each run, and the web UI to make the key of a run's AI assistant. Required.
POSTGRES_PASSWORDgenerated by just setupLiteLLM proxyPassword of the proxy's Postgres database; letters and digits only, as it is part of a URL. Postgres keeps the password it was created with, so after changing it remove the stack's volume. Required.
POSTGRES_USERlitellmLiteLLM proxyUser of the proxy's Postgres database.
LITELLM_PORT4000CLI, LiteLLM proxy, web UIHost port of the LiteLLM proxy.
LITELLM_BIND127.0.0.1ComposeHost address that the LiteLLM proxy's port is published on. The proxy holds the master key and the provider keys, so it listens on loopback only; set 0.0.0.0 to reach it from other machines, and only on a network you trust.
COMPOSE_PROJECT_NAMEssebenchCLI, ComposeCompose project of the proxy stack. Its containers, networks (<project>_default, and the internal <project>_agents that run containers join by default) and database volume (<project>_postgres_data) carry this name, so stacks with different names and ports run side by side.
SSEBENCH_REGISTRYghcr.io/42-b3yond-6ug/ssebenchCLI, just, Compose, catalog serviceRegistry prefix of every image SSEBench builds or pulls, and of the image names that the catalog service returns.
SSEBENCH_CATALOGthe bundled pilot manifestCLI, just, web UITask catalog that ssebench run, ssebench tasks list and the web UI get tasks from: the path or URL of a manifest.json, a dataset directory, or the URL of a catalog service. --catalog overrides it.
SSEBENCH_IMAGES_LOCKimages.lock.json next to the manifest, if there is oneCLI, catalog servicePath of the images lock that pins case images by digest, instead of the one next to the manifest: ssebench run and ssebench-catalog serve (--lock) use it, and ssebench run also takes a URL; see Case images. A lock that is missing or invalid is an error.
SSEBENCH_HOMEsee CLICLIDirectory that holds agents/, images/, runtime/, models/, deploy/ and datasets/, normally the repository root. Unset, a package install without a clone uses the copy of these directories in the package.
SSEBENCH_ENV_FILE.env in the repository rootCLI, ComposeThe .env file that the LiteLLM container reads provider keys from. The CLI sets it to the .env in the workspace; set it yourself only when you run docker compose on deploy/compose/docker-compose.yaml directly.
SSEBENCH_DEMO_PROJECTssebench-demoCLI, ComposeCompose project of the demo. It must differ from COMPOSE_PROJECT_NAME: just demo-down deletes the project's database volume, and the demo refuses to touch a project whose containers it did not create.
SSEBENCH_DEMO_WEBUI_PORT3001CLI, ComposeHost port of the demo's web UI, on 127.0.0.1.
SSEBENCH_DEMO_CATALOG_PORT8090CLI, ComposeHost port of the demo's catalog service, on 127.0.0.1.
SSEBENCH_BACKENDdockerCLI, web UIRunner backend that ssebench run and ssebench runs use when --backend is not given: where the run's containers execute. The web UI runs those commands, so it works with the backend named here. Installed extensions can add backends; see Runner backends.
SSEBENCH_PREBUILTunsetCLI1, true, yes or on makes ssebench run use the published agent images of the task, under SSEBENCH_REGISTRY, instead of building the case, tool and agent layers; the same as --prebuilt. See Prebuilt images.
SSEBENCH_EGRESSrestrictedCLI, web UINetwork egress of a run when --egress is not given: restricted reaches the LiteLLM proxy and nothing else, open also has internet access. The web UI's launches use it unless the launch names a policy. The Helm chart sets it from runs.egress. See Integrity and egress.
SSEBENCH_RUNTIME_IMAGE<registry>/runtime:<version> without a checkout, else unsetCLIRuntime image that the tool layers copy the daemon and the entrypoint from, as <registry>/runtime:<version> does. Unset, a checkout compiles both from its sources, and an installation without a checkout uses the published image of its version. Set it to skip the compilation in a checkout. The image must exist for the platform of the run, which is linux/amd64 for a pilot task.
SSEBENCH_VERSIONset by the CLICLI, ComposeSSEBench version that the Compose files pass as the VERSION build argument of the images they build. The CLI sets it from VERSION; it only matters when you run docker compose build yourself.

Kubernetes backend ​

ssebench run --backend kubernetes reads these settings the same way. See Kubernetes for what they configure.

VariableDefaultUsed byDescription
SSEBENCH_K8S_NAMESPACEthe namespace of the kubeconfig context, or of the pod ssebench runs in, else defaultCLINamespace where the runs' Jobs, Secrets and NetworkPolicies are created.
SSEBENCH_K8S_CONTEXTthe kubeconfig's current contextCLIContext of the kubeconfig to use. The kubeconfig itself is found the way kubectl finds it (KUBECONFIG, else ~/.kube/config); inside a pod the backend uses the pod's service account.
SSEBENCH_K8S_PROXY_URLhttp://litellm.<proxy namespace>.svc:4000CLIURL of the LiteLLM proxy as a run's pod reaches it: the run's SSE_BASE_URL. Its port must be the port the proxy's pods listen on, since the run's network policy opens that port.
SSEBENCH_K8S_PROXY_HOST_URLthe proxy's Service URL inside a cluster, else http://localhost:$LITELLM_PORTCLIURL of the LiteLLM proxy as the ssebench process reaches it, to create and read each run's key. From outside the cluster, forward the proxy's port (kubectl port-forward svc/litellm 4000) and set LITELLM_PORT or this.
SSEBENCH_K8S_PROXY_NAMESPACEthe runs' namespaceCLINamespace of the LiteLLM proxy, for its default URL and the run's network policy.
SSEBENCH_K8S_PROXY_SELECTORapp.kubernetes.io/name=litellmCLILabels of the LiteLLM proxy's pods, as key=value,key=value. A restricted run's network policy lets its pod reach the pods that match, and no others.
SSEBENCH_K8S_RUNTIME_CLASSunsetCLIruntimeClassName of the run's pod, for a sandboxed runtime such as gVisor or Kata Containers.
SSEBENCH_K8S_IMAGE_PULL_POLICYIfNotPresentCLIimagePullPolicy of the run's containers. IfNotPresent uses an image that is already on the node, such as one loaded into a kind cluster.
SSEBENCH_K8S_IMAGE_PULL_SECRETSunsetCLINames of image pull Secrets in the runs' namespace, separated by commas.
SSEBENCH_K8S_RESOURCESrequests of 1 CPU, 2Gi memory and 2Gi ephemeral storage; limits of 4 CPUs, 8Gi and 20GiCLIResource requests and limits of the task container, as JSON or YAML in the shape of a container's resources, for example {"requests": {"cpu": "2"}, "limits": {"memory": "16Gi"}}. It replaces the defaults as a whole.
SSEBENCH_K8S_TTL_SECONDS3600CLIHow long the cluster keeps a run's finished Job, its pod and its logs (ttlSecondsAfterFinished). A run kept with --keep-container has none.
SSEBENCH_K8S_DEADLINE_SLACK1800CLISeconds added to the run's --timeout to make the Job's activeDeadlineSeconds: the time for pulling images, starting and grading. When it runs out the cluster deletes the pod and the run's results with it.
KUBERNETES_SERVICE_HOSTset by Kubernetes in every podCLIThe backend takes its presence to mean that ssebench runs inside a cluster, where the proxy is reached through its Service.

Inside the task container ​

ssebench run and the container's entrypoint set these variables in every task container. Agents and plugins read them.

VariableDefaultUsed byDescription
SSE_API_KEYset by ssebench runagents, plugins, sse.aiThe run's LiteLLM key, which can use only the selected model; empty in a reference run.
SSE_BASE_URLset by ssebench runagents, plugins, sse.aiURL of the LiteLLM proxy, http://litellm:4000; empty in a reference run.
SSE_MODEL_NAMEset by ssebench runagents, plugins, sse.aiThe selected model, as named in models/*.yaml; none in a reference run.
SSE_ARCHIVE/tmp/sse-archiveentrypoint, daemon, evaluator, agentsThe agent's archive directory, /tmp/sse-archive, writable by the agent: dialog.jsonl and whatever the agent side writes go there. The entrypoint requires it. The graded outputs go to SSE_RESULTS instead, root-only.
SSE_RESULTS/var/lib/ssebench/resultsentrypoint, daemon, evaluatorThe run's results directory, root-only: the grade (result.json), the graded patch (final.patch), the commit log and the run's logs. Its parent is made 0700 root, so neither the agent nor the task runner can reach it. The CLI mounts the run directory here and its archive/ subdirectory at SSE_ARCHIVE. Falls back to SSE_ARCHIVE when unset.
SSE_DIFFICULTY2daemon, MCP serverThe difficulty level, from 0 to 4. The MCP server decides from it which checks test_patch runs, and the daemon refuses the withheld bencher actions on its agent-facing listeners.
TIMEOUT14400 in the entrypoint, 1800 in the evaluatorentrypoint, evaluator, agentsHow long the agent may run, in seconds (--timeout). The evaluator uses the same limit for grading.
SSE_KEEP_ALIVE0entrypoint1 keeps the container running after grading (--keep-container), for the web UI.
SSE_PLUGINSthe plugins plugins.yaml enablesentrypointComma-separated plugins to run, set by ssebench run --plugin; when it is set, it replaces the enabled field of plugins.yaml, and an empty value runs none. See Plugins and hooks.
SSE_DAEMON_SOCKET/tmp/sse.sockentrypoint, daemon, SDKThe daemon's agent-facing Unix socket, mode 0666. The entrypoint sets it for every process it starts; in sidecar mode ssebench run sets it to /run/ssebench/sse.sock, on a root-owned volume the two containers share. Without it, the daemon serves HTTP only.
SSE_ADMIN_SOCKET/run/ssebench/admin.sockentrypoint, daemonThe daemon's privileged Unix socket, mode 0600, root only: grading, the reference patch and phase changes. In sidecar mode it is on the volume the two containers share, and the task container's entrypoint sets it. The daemon binds it only when this is set; the entrypoint sets it, and points the evaluator's SSE_DAEMON_SOCKET at it. The SDK's sse.reference.get_reference_patch reads it to reach the admin socket.

The components in the container also read these settings, which have working defaults:

VariableDefaultUsed byDescription
SSE_AGENT_USERmodeldaemonThe agent's user, whose processes the daemon kills when the agent phase ends.
SSE_RUNNER_USERsse-runnerdaemonThe unprivileged user the daemon runs the task's build, PoC and test scripts as, when it runs as root. A dedicated uid with no groups, neither the agent nor root; the tool layer creates it. The daemon refuses to start as root without it.
SSE_RUNNER_DIR/var/lib/ssebench-runnerdaemonThe root of the task runner's scratch copies, reachable only by the runner and root (0710). Each check runs in a private copy of the project here.
SSE_DAEMON_TIMEOUT300entrypointSeconds to wait for the daemon's socket.
SSE_MCP_TIMEOUT300entrypointSeconds to wait for the MCP server.
SSE_DEBUGunsetentrypointAny non-empty value turns on debug logs.
SSE_PLUGIN_NAMEunsetpluginsSet by the entrypoint for a plugin it runs; the plugin's name.
SSE_PLUGIN_HOOKunsetpluginsSet by the entrypoint for a plugin it runs; the hook it runs at, such as after-grading.
SSE_ORACLE_FUZZunsetoracle plugin1 makes the oracle plugin fuzz the patched project after its review; it installs AFL++, so the run needs --egress open.
SSE_BENCH_PATH/ssebenchdaemonDirectory with the task's config.yaml, scripts and files.
SSE_HTTP_PORT4263daemonThe daemon's agent-facing HTTP port.
SSE_DAEMON_WORKERS4daemonWorker threads for each of the daemon's listeners. A tool call occupies its worker until the command ends. Without a limit, actix starts one worker per host CPU for every listener.
SSE_REPO_PATH/ssebench-repodaemonClean clone of the project that grading applies the agent's diff to and builds.
SSE_AGENT_DOCKERunsetSDKhost:port of the daemon's HTTP listener, used when SSE_DAEMON_SOCKET is not set.
MCP_LOG_DIR/tmp/mcp/logsMCP serverWhere the MCP server writes the full logs of long check results; see Long logs.
AGENT_DURATION0evaluatorThe agent's run time in seconds; the entrypoint sets it for the evaluator.
SSE_METRIC_AGENT_TIMEOUTunsetevaluatorSet to true by the entrypoint for the evaluator when the agent hit its time limit.
CLAUDEunsetclaude-code agentPath of the Claude Code executable; the claude-code agent image sets it.
OPENCODE_CONFIG_CONTENTunsetOpenCodeOpenCode's configuration as JSON; sse.ai and the opencode agent set it for the OpenCode server they start.

Difficulty levels ​

SSE_DIFFICULTY decides which checks the agent's test_patch tool runs. Final grading always runs every check the task has.

LevelName
0FULL_ASSISTANCE
1NO_INTENT_TEST
2NO_FUTURE_TEST (default)
3BUILD_ONLY
4NO_BUILD

See Difficulty levels for what each level allows and how it is enforced, and Daemon HTTP API for the daemon's gate.

Web UI ​

The web UI server reads these; see WebUI.

VariableDefaultUsed byDescription
SSEBENCH_CATALOGthe bundled pilot manifestCLI, just, web UITask catalog that ssebench run, ssebench tasks list and the web UI get tasks from: the path or URL of a manifest.json, a dataset directory, or the URL of a catalog service. --catalog overrides it.
SSEBENCH_WEBUI_HOST127.0.0.1web UIBind address of the API server and of the Vite dev and preview servers.
PORT3001web UIAPI server port; the Vite servers forward /api there.
SSEBENCH_WEBUI_TOKENunsetweb UIAccess token, at least 16 characters. Required when the bind address is not loopback.
SSEBENCH_WEBUI_CORS_ORIGINSunsetweb UIComma-separated origins, besides the server's own, allowed to call the API, such as https://bench.example.org.
SSEBENCH_WEBUI_TERMINAL1web UI0 turns off the container terminal. The demo's web UI has it off (0) unless you set this to 1. Hosted mode (SSEBENCH_WEBUI_HOSTED) turns it off whatever this says.
SSEBENCH_WEBUI_HOSTED0web UI1 runs the web UI as a read-only viewer, for a showcase or a shared server. It refuses to launch, stop or remove runs, opens no terminal, and lets the AI assistant run no commands in a container; runs and finished runs can still be read. See Security model.
SSEBENCH_CLIuv run ssebenchweb UIThe command that runs the ssebench CLI, split on white space and started without a shell in SSEBENCH_PATH, such as ssebench or uv run --project /opt/ssebench ssebench. The web UI finds, stops, reaches and launches runs through it. The web UI image sets it to ssebench.
SSEBENCH_PATHthe repository rootweb UISSEBench checkout: the CLI runs there, so results/ is read and written there, and models/ and agents/ are read from there.
SSEBENCH_LOCAL_TASKS$SSEBENCH_PATH/datasets/pilotweb UILocal dataset offered in the launcher.
SSEBENCH_CATALOG_URLunsetweb UIFormer name of SSEBENCH_CATALOG; deprecated, and read only when that is unset.
PTY_ENABLE_CLEANUPtrueweb UIfalse turns off the cleanup of terminal processes whose browser went away.
PTY_CLEANUP_INTERVAL30000web UIMilliseconds between two cleanups of terminal processes.
PTY_ORPHAN_THRESHOLD300000web UIMilliseconds without activity after which a terminal process is cleaned up.

Catalog service ​

The catalog service reads these; each has a command-line option too. See Dataset manifest.

VariableDefaultUsed byDescription
SSEBENCH_REGISTRYghcr.io/42-b3yond-6ug/ssebenchCLI, just, Compose, catalog serviceRegistry prefix of every image SSEBench builds or pulls, and of the image names that the catalog service returns.
SSEBENCH_IMAGES_LOCKimages.lock.json next to the manifest, if there is oneCLI, catalog servicePath of the images lock that pins case images by digest, instead of the one next to the manifest: ssebench run and ssebench-catalog serve (--lock) use it, and ssebench run also takes a URL; see Case images. A lock that is missing or invalid is an error.
SSEBENCH_MANIFESTdatasets/pilot/manifest.jsoncatalog serviceThe manifest.json to serve (--manifest); /srv/manifest.json in the catalog image.
SSEBENCH_CATALOG_PORT8080catalog serviceTCP port to listen on (--port).

Development ​

VariableDefaultUsed byDescription
SSEBENCH_INTEGRITY_IMAGEthe tool image that ssebench run builds for gjson-196-bf4efcbintegrity testsSandbox tool image that tests/integrity attacks.
SSEBENCH_INTEGRITY_SIDECAR_ENV_IMAGEthe task image that ssebench run --mode sidecar builds for gjson-196-bf4efcbintegrity testsSidecar task image, with the daemon, that tests/integrity attacks.
SSEBENCH_INTEGRITY_SIDECAR_AGENT_IMAGEthe sidecar runtime image of this versionintegrity testsSidecar agent runtime image that tests/integrity runs the fake agent in.
SSEBENCH_INTEGRITY_PLUGINSunsetintegrity testsPlugins the sandbox container of the integrity tests enables, passed as SSE_PLUGINS; the image must have them installed.
SSEBENCH_INTEGRITY_SOURCE/src/gjsonintegrity testsSource directory inside those images.
SSE_DAEMON_HTTPhttp://localhost:4263integrity testsThe daemon's HTTP listener that tests/integrity/fake_agent.sh probes; in sidecar mode, the task container's.
SSEBENCH_SMOKE_TASKgjson-196-bf4efcbtests/e2e/smoke.shTask of the end-to-end smoke run.
SSEBENCH_SMOKE_MODELclaude-sonnet-4-6tests/e2e/smoke.shModel the smoke run names; the dummy agent never calls it.
SSEBENCH_UPDATE_SNAPSHOTSunsetSDK tests1 rewrites the snapshot of the task prompt in sdk/python/tests/snapshots/ instead of comparing with it; see Task prompt.

Using the variables in an agent ​

Python ​

python
import os

# LLM access
api_key = os.environ["SSE_API_KEY"]
base_url = os.environ["SSE_BASE_URL"]
model = os.environ["SSE_MODEL_NAME"]

# Run settings
difficulty = int(os.environ.get("SSE_DIFFICULTY", "2"))
timeout = int(os.environ.get("TIMEOUT", "3600"))
archive_path = os.environ.get("SSE_ARCHIVE", "/tmp/sse-archive")

Shell ​

sh
#!/bin/bash
echo "Model: $SSE_MODEL_NAME"
echo "Difficulty: $SSE_DIFFICULTY"

# List the models this run's key can use
curl -H "Authorization: Bearer $SSE_API_KEY" "$SSE_BASE_URL/models"

Troubleshooting ​

  • A model fails with an authentication error. Check that its key is set in .env under the name its model file uses, then run just launch so the proxy picks it up.
  • The proxy does not start. Run just doctor. LITELLM_MASTER_KEY and POSTGRES_PASSWORD must be set (just setup writes them), and LITELLM_PORT must be free.
  • You want to see a container's variables. Start the run with --keep-container, then run docker exec <container> env.

Next steps ​

Released under the Apache License 2.0.