Skip to content

CLI ​

The ssebench command builds the images for a task and runs one agent on one task with one model. From a clone of the repository, run it as uv run ssebench in the repository root or any directory below it; see Working directory. Without a clone, run uvx ssebench or install the package from PyPI; see Without a clone.

ssebench --version prints the SSEBench version, which is also the tag of the tool layer and agent images the CLI builds; see Releasing and versioning. ssebench <command> --help prints the options of a command.

CommandDescription
ssebench runRun a benchmark
ssebench build-caseBuild case images
ssebench dataset validateCheck every task folder of a dataset against the task schema
ssebench dataset manifestValidate a dataset and write its manifest.json
ssebench dataset schemaWrite the JSON Schemas of the task config, dataset.yaml, the manifest and the images lock
ssebench dataset verifyGrade tasks with the reference and dummy agents and check that they grade as sound tasks do
ssebench dataset publishPush the case images of verified tasks to a registry
ssebench dataset lockWrite the images lock, which pins the published case images by digest
ssebench tasks listList the tasks of a catalog or a local dataset
ssebench runs listList the runs that exist on the backend, running or not
ssebench runs inspectShow one run on the backend
ssebench runs logsPrint the output of a run's container
ssebench runs stopStop a run's container, leaving it in place
ssebench runs removeRemove a run's containers and volumes, running or not
ssebench runs endpointPrint the URL at which a port of a running run can be reached from here
ssebench runs execRun a command in a run's container
ssebench runs resultsList the finished runs in results/, whether or not they still have containers
ssebench proxyStart, rebuild or stop the local LiteLLM proxy
ssebench initWrite .env with generated secrets, models/ and results/ to the current directory
ssebench demo upStart the demo stack, run one agent on one task and show the run in the web UI
ssebench demo downRemove everything the demo created
ssebench doctorCheck that this host can build and run benchmarks

ssebench run ​

Runs one agent × model × task combination and writes the results to a directory of its own, results/<task>/<model>/<agent>/<run-id>/.

sh
uv run ssebench run --local datasets/pilot --task <task-id> --agent <agent> --model <model>
sh
ssebench run [-h] [--model NAME] --agent NAME --task ID [--local DIR]
             [--catalog PATH|URL] [--build] [--mode MODE] [--backend NAME] [--prebuilt]
             [--tool-layer NAME] [--plugin NAME] [--run-id ID] [--timeout SECONDS]
             [--difficulty LEVEL] [--keep-container] [--egress POLICY]
OptionDefaultDescription
--model NAMEModel name, as defined in models/*.yaml; required for every agent but reference, which uses none
--agent NAME(required)Agent name, a directory under agents/
--task ID(required)Task ID, the name of the task's folder
--local DIRDataset directory that contains the task folder, for example datasets/pilot; the case image is built from the folder
--catalog PATH|URL$SSEBENCH_CATALOG, else the bundled pilot manifestTask catalog: a manifest.json path or URL, a dataset directory, or the URL of a catalog service; used when --local is not given
--buildoffBuild the task's case image from its folder instead of pulling the published one; needs a checkout or --catalog with the task folder next to it. With --local the image is always built
--mode MODEsandboxExecution mode: sandbox, or sidecar (experimental)
--backend NAMEdocker, or $SSEBENCH_BACKENDWhere the run's containers execute: docker or kubernetes; installed extensions can add more. kubernetes needs --prebuilt
--prebuiltoffUse the published agent images of the task under $SSEBENCH_REGISTRY, pulling them, instead of building the case, tool and agent layers; excludes --tool-layer, --plugin and --build. Also enabled by SSEBENCH_PREBUILT=1
--tool-layer NAMEsandboxTool layer to build in sandbox mode; installed extensions can add more
--plugin NAME[]Run a plugin in this run (repeatable), instead of those plugins.yaml enables; sandbox mode only
--run-id IDName the run: its directory is results/TASK/MODEL/AGENT/ID, and its containers get the label ssebench.run-id=ID, so a tool that starts the run can find them. 1 to 64 letters, digits, '.', '_' or '-', and not latest. The run is refused if that directory exists. Default: the UTC time the command started and six random hex digits, such as 20260929-153012-a1b2c3
--timeout SECONDS3600How long the agent may run
--difficulty LEVEL2 = NO_FUTURE_TESTWhich checks the agent's test_patch tool may run, from 0 (all) to 4 (none). One of 0, 1, 2, 3, 4
--keep-containeroffKeep the container after the run, for example to inspect it from the web UI
--egress POLICY$SSEBENCH_EGRESS, else restrictedNetwork egress of the run container: restricted reaches the LiteLLM proxy but not the internet; open also has internet access, for tasks that need network at test time

--model is optional with --agent reference; see Reference runs. --mode is explained in Sandbox and sidecar, --difficulty in Difficulty levels, --tool-layer in Extension points, --backend in Runner backends, --prebuilt in Prebuilt images, and --egress in Integrity and egress; the egress policy is recorded as config.egress in the run summary. --keep-container leaves the container for the web UI to inspect; in sidecar mode it leaves both containers and their volumes. --run-id names the run directory; without it the CLI makes an ID, so running a combination again adds a run and replaces none; see Run IDs and repeated runs.

Before it starts the proxy or builds an image, ssebench run checks the arguments: the agent must exist and have a valid agent.yaml, --model must be defined in models/*.yaml, and .env must hold the provider key that the model needs, except for the dummy and reference agents, which make no model calls. Each failure is one line on standard error, with the valid choices where there are some, and exit status 1. A failed image build ends the same way, after Docker's own output.

Without --local, the task comes from the task catalog, and its case image is pulled from $SSEBENCH_REGISTRY. When the pull fails, the CLI says so and builds the image from the task's folder instead, if a copy of the folder whose files match the catalog is available locally: next to a local manifest, or in datasets/<dataset>/ of the SSEBench home, as for the bundled pilot dataset. --build skips the pull and builds the image from that folder, and fails before starting anything when there is no such folder. With --local the image is always built, so --build changes nothing.

In order, run:

  1. starts the LiteLLM proxy if it is not running, rebuilding its image first when models/ changed, and waits for it to become healthy;
  2. checks that the proxy knows the model, and creates a key for this run that can use only that model;
  3. builds the case, tool and agent image layers;
  4. runs the task container, which runs the agent and then the evaluator. In sidecar mode, it starts the task container with the daemon, then runs the agent container, which runs the agent and then the evaluator, and removes the task container afterwards;
  5. writes the grade, the run settings and the model spend to summary.json in the run directory, and adds the run settings to the run's result.json as config.

Every image and container of a run uses the platform of the task's case image: the host's architecture if the task's manifest arch lists it, otherwise the first architecture listed. When that is not the host's, as for a pilot task on an arm64 host, run warns once, before it starts anything, that the task runs under emulation. See Architectures.

Reference runs ​

--agent reference runs the reference agent, which applies the task's reference patch, the known upstream fix, instead of asking a model for one. It checks the task and the grader rather than a model: on a sound task, every check passes. It makes no model calls, so it needs no provider key.

sh
uv run ssebench run --local datasets/pilot --task gjson-196-bf4efcb --agent reference

A reference run differs from other runs in these ways:

  • --model is optional. The run's model is none: no proxy key is created, the spend is 0, and the results go to results/<task>/none/reference/<run-id>/. Any other --model is ignored, with a warning. The LiteLLM proxy still starts, as the container joins its network.
  • run copies the file that files.patch names out of the case image and mounts it read-only at /reference/patch.diff in the container (in sidecar mode, the agent container). No other agent gets this mount. A task without files.patch cannot be run this way.
  • result.json and the summary record config.reference_run: true, the container carries the label ssebench.reference-run=true, the web UI marks the result as a reference run, and just report leaves reference runs out of the scores.

ssebench tasks list ​

Lists the tasks of a catalog or of a local dataset. It reads nothing over the network unless the catalog is a URL, so it works offline with the bundled pilot manifest.

sh
uv run ssebench tasks list
uv run ssebench tasks list --json --catalog https://catalog.example.org
sh
ssebench tasks list [-h] [--catalog PATH|URL | --local DIR] [--json]
OptionDefaultDescription
--catalog PATH|URL$SSEBENCH_CATALOG, else the bundled pilot manifestTask catalog: a manifest.json path or URL, a dataset directory, or the URL of a catalog service
--local DIRDataset directory: list its task folders instead
--jsonoffPrint a JSON array of the tasks without files and metadata, with image names prefixed by the registry and the case image tagged with the dataset version, or named by its digest when the images lock pins it

Without --json, it prints a table of task IDs, languages and projects. With --json, it prints the tasks as the catalog service returns them from GET /tasks.

Task catalog ​

A catalog is a dataset manifest. --catalog, or the SSEBENCH_CATALOG setting when the option is not given, names it as one of:

ValueManifest read
path of a file, such as datasets/pilot/manifest.jsonthe file
path of a directory, such as datasets/pilotmanifest.json in it
http(s) URL whose path ends in .jsonthe URL
any other http(s) URL, such as a catalog service at http://localhost:8080manifest.json under it

With neither, the catalog is datasets/pilot/manifest.json in the SSEBench home, the pilot dataset's manifest that ships with SSEBench. The web UI reads the same variable and passes its catalog on to the runs it launches.

Case images ​

The case image of each task is published under the name the manifest gives it, prefixed with $SSEBENCH_REGISTRY, and tagged with the dataset's version: ghcr.io/42-b3yond-6ug/ssebench/case/pilot/gjson-196-bf4efcb:pilot-v1. Each push also has a tag that names the commit it was built and verified at, such as pilot-v1-0123abc, which no later commit reuses. Only the images that passed ssebench dataset verify are published, and they are the images that it graded. They are for linux/amd64 only.

ssebench run without --local pulls the image, so that a machine runs any task without building it. It pulls:

  • by digest, when an images.lock.json pins the task. The lock is the file of that name next to a local manifest, which is datasets/pilot/ for the bundled catalog, or the path or URL that SSEBENCH_IMAGES_LOCK names. An entry applies when the lock's dataset version is the manifest's and the entry was built from the task files that the manifest lists; otherwise the CLI logs a warning and pulls by tag. The wheel carries the pilot lock, so a release runs the images it was published with. The digest identifies the image whatever registry serves it, so a mirror that copied the images with their digests can be used with SSEBENCH_REGISTRY.
  • by tag (the dataset version) for a task the lock does not pin, for a dataset without a lock, and for a catalog service URL.

After a pull by digest, the image also gets the tag name locally, so that the layers built on it and the reference patch use exactly that image. Releasing and versioning describes how the lock is refreshed, and the pilot dataset lists the size of every image.

ssebench runs ​

Finds, watches, stops and removes runs through the runner backend, and lists finished runs. A run is named by its run ID, the value of --run-id. The web UI drives these commands with --json, so it works with every backend; they are as useful by hand:

sh
uv run ssebench runs list
uv run ssebench runs logs --follow 20260929-153012-a1b2c3
uv run ssebench runs stop 20260929-153012-a1b2c3

Every command that takes a run ID exits with status 3 if no run has it, and 4 if several runs do, which --run-id allows across different tasks, models and agents. The commands other than results use the backend that --backend, or SSEBENCH_BACKEND, names.

ssebench runs list ​

sh
ssebench runs list [-h] [--backend NAME] [--json]
OptionDefaultDescription
--backend NAMEdocker, or $SSEBENCH_BACKENDThe runner backend; installed extensions can add more
--jsonoffPrint JSON instead of text

--json prints one object: backend (its name), supports_exec (whether it can run commands in a run), and runs, newest first. Each run has run_id, name, state (created, running, exited or unknown), exit_code, image, created_at, task, model, agent, reference_run, results_dir and the run's ssebench.* labels. The container's environment, which holds the run's key, is never part of it.

ssebench runs inspect ​

sh
ssebench runs inspect [-h] [--backend NAME] [--json] RUN_ID
OptionDefaultDescription
--backend NAMEdocker, or $SSEBENCH_BACKENDThe runner backend; installed extensions can add more
--jsonoffPrint JSON instead of text
RUN_ID(required)The run's ID

ssebench runs logs ​

sh
ssebench runs logs [-h] [--backend NAME] [--follow] RUN_ID
OptionDefaultDescription
--backend NAMEdocker, or $SSEBENCH_BACKENDThe runner backend; installed extensions can add more
RUN_ID(required)The run's ID
--followoffKeep printing until the container exits

ssebench runs stop ​

sh
ssebench runs stop [-h] [--backend NAME] [--grace SECONDS] RUN_ID
OptionDefaultDescription
--backend NAMEdocker, or $SSEBENCH_BACKENDThe runner backend; installed extensions can add more
RUN_ID(required)The run's ID
--grace SECONDS20How long the entrypoint may take to clean up before the container is killed

ssebench runs remove ​

sh
ssebench runs remove [-h] [--backend NAME] RUN_ID
OptionDefaultDescription
--backend NAMEdocker, or $SSEBENCH_BACKENDThe runner backend; installed extensions can add more
RUN_ID(required)The run's ID

It removes the containers and volumes and leaves the run's results/ directory alone. It fails if the run is still there afterwards.

ssebench runs endpoint ​

sh
ssebench runs endpoint [-h] [--backend NAME] [--json] RUN_ID PORT
OptionDefaultDescription
--backend NAMEdocker, or $SSEBENCH_BACKENDThe runner backend; installed extensions can add more
--jsonoffPrint JSON instead of text
RUN_ID(required)The run's ID
PORT(required)A TCP port of the run's container

With --json it prints {"url": "http://host:port"}. The URL is one that this machine can use, whatever the backend: on Docker the container's address on its network.

ssebench runs exec ​

sh
ssebench runs exec [-h] [--backend NAME] [--user USER] [--workdir DIR] [--tty] [--stdin]
                   [--env NAME]
                   RUN_ID COMMAND [COMMAND ...]
OptionDefaultDescription
--backend NAMEdocker, or $SSEBENCH_BACKENDThe runner backend; installed extensions can add more
RUN_ID(required)The run's ID
--user USERRun as this user of the container
--workdir DIRWorking directory in the container
--ttyoffAllocate a terminal
--stdinoffKeep standard input open
--env NAME[]Pass the variable NAME with the value it has here (repeatable); the value never appears in a command line
COMMAND(required)The command and its arguments
sh
uv run ssebench runs exec 20260929-153012-a1b2c3 --tty -- bash

A backend that cannot run commands in a run exits with status 1.

ssebench runs results ​

sh
ssebench runs results [-h] [--json] [--dir DIR]
OptionDefaultDescription
--jsonoffPrint JSON instead of text
--dir DIRresults in the working directoryThe results directory

It reads the summary.json of every run directory under the results directory, at any depth, and needs no backend and no container. --json prints {"runs": [...]}, newest first, each with run_id, task, model, agent, mode, reference_run, status (the grade's), started_at and dir, the run directory. Directories without a summary.json, such as a run in progress, and summaries that do not parse are left out.

ssebench build-case ​

Builds the case images of a dataset without running anything. Each image is built for the platform that a run of its task uses, and the command warns once when that is emulated; see Architectures.

sh
uv run ssebench build-case --benchmarks datasets/pilot --tasks gjson-196-bf4efcb
sh
ssebench build-case [-h] [--benchmarks DIR] [--tasks IDS] [--force]
OptionDefaultDescription
--benchmarks DIRdatasets/pilot in the SSEBench homeDataset directory
--tasks IDSevery taskComma-separated task IDs
--forceoffRebuild images even when they are up to date; without it, an image is rebuilt only when the task's files changed since it was built

ssebench dataset ​

Checks a dataset and writes the files generated from it. See Dataset manifest for the layout, the task config and the manifest these commands work with. Each command exits 1 when it fails.

sh
uv run ssebench dataset validate
uv run ssebench dataset manifest --check

ssebench dataset validate ​

Checks every task folder: the config against the schema, the ID against the folder name, the Dockerfile's base image, and that every path in the config is a file the Dockerfile copies into the image. Lists every problem, per task.

sh
ssebench dataset validate [-h] [DIR]
OptionDefaultDescription
DIRdatasets/pilot in the SSEBench homeDataset directory

ssebench dataset manifest ​

Validates the dataset and writes its manifest.json. --check keeps the commit the file records.

sh
ssebench dataset manifest [-h] [-o FILE] [--check] [--generated-from COMMIT] [DIR]
OptionDefaultDescription
DIRdatasets/pilot in the SSEBench homeDataset directory
-o, --output FILEmanifest.json in the datasetOutput file, or - for stdout
--checkoffFail if the file is out of date; write nothing
--generated-from COMMITnone; with --check, the commit recorded in the fileRecord the repository commit the manifest is generated from

ssebench dataset schema ​

Writes the JSON Schemas of the task config, dataset.yaml, the manifest and the images lock, which Configuration files describes.

sh
ssebench dataset schema [-h] [-o DIR] [--check]
OptionDefaultDescription
-o, --output DIRdatasets/schema in the SSEBench homeOutput directory
--checkoffFail if the files are out of date; write nothing

ssebench dataset lock ​

Writes the dataset's images.lock.json, which pins the published case image of each task by digest, so that ssebench run pulls exactly that image. The output keeps the images of the current lock whose task files are unchanged, adds the ones in --records, and depends on nothing else. A dataset with no published images has a lock with no images. --check fails when the file is missing, is for another dataset version or pins a task the dataset does not have; a lock that lags behind a changed task passes.

sh
ssebench dataset lock [-h] [--records DIR] [-o FILE] [--check] [DIR]
OptionDefaultDescription
DIRdatasets/pilot in the SSEBench homeDataset directory
--records DIRnone; keep the current imagesDirectory of publication records to add
-o, --output FILEthe images lock in the datasetOutput file
--checkoffFail if the dataset's lock is missing, is for another version or names other tasks; write nothing

ssebench dataset publish ​

Pushes the case images of verified tasks to a registry and records the digests. It publishes the image that ssebench dataset verify graded and no other: a task with no result, a task that did not pass, and an image that was rebuilt after the verification are refused. Log in to the registry first. The Dataset workflow runs it for every task that passes, on a release tag and on a manual run that asks for it, and never from a private repository.

sh
uv run ssebench dataset verify gjson-196-bf4efcb
uv run ssebench dataset publish --registry ghcr.io/owner/repo gjson-196-bf4efcb
sh
ssebench dataset publish [-h] [--dir DIR] [-o OUTPUT] --registry PREFIX
                         [--revision COMMIT]
                         [TASK ...]
OptionDefaultDescription
TASKevery task with a resultTasks to publish
--dir DIRdatasets/pilot in the SSEBench homeDataset directory
-o, --output OUTPUTresults/dataset-verifyVerification output
--registry PREFIX(required)Registry prefix, such as ghcr.io/owner/repo
--revision COMMITHEAD of the repository that holds the datasetFull commit of the SSEBench repository that the dataset is at

ssebench dataset verify ​

Grades tasks end to end, each with two ssebench runs: the reference agent, whose upstream fix must pass every check the task has, and the dummy agent, whose unmodified project must build and pass its functional tests while every proof of concept still triggers the bug. The dummy's intent tests should fail too; when they pass, the summary warns that they do not check the fix. It starts the LiteLLM proxy once, like ssebench run, and needs the base images of the tasks (make -C images/base-images <name>).

sh
uv run ssebench dataset verify gjson-196-bf4efcb
uv run ssebench dataset verify --changed-since origin/main --jobs 4

It writes, under the output directory, one JSON file per task in tasks/, the log of every run in logs/, the run directories, one for each attempt, in runs/results/, and summary.md and summary.json, a table of every task result in the directory by check. It exits 1 when a task does not grade as expected. The Dataset workflow runs it for the tasks a pull request changes, and weekly for all of them.

sh
ssebench dataset verify [-h] [--dir DIR] [--changed-since REV] [--list] [-j JOBS]
                        [-o OUTPUT] [--model MODEL] [--difficulty DIFFICULTY]
                        [--timeout TIMEOUT] [--retries RETRIES] [--summarize]
                        [TASK ...]
OptionDefaultDescription
TASKevery taskTasks to verify
--dir DIRdatasets/pilot in the SSEBench homeDataset directory
--changed-since REVVerify the tasks that changed since the merge base of REV and HEAD, including those whose base image changed under images/base-images
--listoffPrint the selected tasks and their base images as JSON; run nothing
-j, --jobs JOBS2Tasks to verify in parallel
-o, --output OUTPUTresults/dataset-verifyOutput directory
--model MODELclaude-sonnet-4-6Model of the dummy runs, which make no model calls
--difficulty DIFFICULTY2Difficulty level of the runs
--timeout TIMEOUT3600Time limit of each run's agent and grading
--retries RETRIES0Times to repeat a run that ends without a result, as when a build loses the network
--summarizeoffWrite the summary of the task results already in OUTPUT/tasks/; run nothing

ssebench proxy ​

Starts, rebuilds or stops the local LiteLLM proxy, the Compose stack in deploy/compose/. just launch and just stop call it.

sh
uv run ssebench proxy up
sh
ssebench proxy [-h] [--volumes] [--rebuild] ACTION
OptionDefaultDescription
ACTION(required)up: build the proxy image if it is missing or older than models/, start the stack and wait until the proxy is healthy; build: only build the image, if it is missing or older than models/; down: stop the stack and keep its database volume
--volumesoffWith down, also remove the database volume: the next start creates a new database with the POSTGRES_PASSWORD of .env, and the proxy's stored keys and spend records are gone
--rebuildoffWith up or build, rebuild the proxy image even if it is current

The image is tagged with the SSEBench version and carries a hash of models/*.yaml and images/litellm/ in its ssebench.litellm-config label, which is how up, build and ssebench run tell that it is out of date. The stack's Compose project is COMPOSE_PROJECT_NAME (default ssebench), and the proxy listens on LITELLM_PORT (default 4000); see SSEBench settings.

ssebench init ​

Sets up the current directory as a workspace, the way just setup does in a checkout. Run it once, in an empty directory, when you run ssebench from a package install.

sh
ssebench init
sh
ssebench init [-h]

It takes no options.

It writes three things, and leaves each one alone if it already exists, so it is safe to run again:

  • .env, from the packaged .env.example, with a generated master key for the LiteLLM proxy and a generated password for its database. Only you can read the file. Add the keys of your model providers to it.
  • models/, a copy of the model definitions in the package. ssebench run and ssebench proxy up build the proxy image from this directory, and rebuild it when you change a file; see Add a model.
  • results/, where runs write their results.

In a checkout, the workspace is the checkout, which already has models/, so the command only writes .env and results/.

ssebench demo ​

Starts, and removes, the local demo: the LiteLLM proxy and its database, the task catalog and the web UI in a Compose project of their own (SSEBENCH_DEMO_PROJECT, default ssebench-demo), and one run of an agent that the web UI shows. just demo and just demo-down call it.

sh
uv run ssebench demo up
sh
ssebench demo up [-h] [--task ID] [--agent NAME] [--model NAME] [--timeout SECONDS]
                 [--build]
OptionDefaultDescription
--task IDgjson-196-bf4efcbTask ID, from the bundled pilot dataset; pick a fast one
--agent NAMEreferenceAgent name, a directory under agents/; reference applies the task's known fix
--model NAMEModel name, as defined in models/*.yaml; required for every agent but reference
--timeout SECONDS3600How long the agent may run
--buildoffBuild every image from this checkout instead of pulling the published ones

up checks the host as ssebench doctor does, then pulls the images of the catalog, the web UI and the task at this checkout's version, and builds those that the registry does not have. The LiteLLM proxy image is always built from models/, as ssebench proxy up does. It starts the stack, waits until the proxy, the catalog and the web UI answer, and runs ssebench run --keep-container in the background, because a kept container keeps that command running. When the evaluator has graded the run, up checks that the web UI shows it, prints the result and the address of the web UI, and returns. The agent's --model needs its provider key in .env, and up stops before it builds anything when the key is missing. Running up again replaces the previous run.

The web UI listens on 127.0.0.1:3001 (SSEBENCH_DEMO_WEBUI_PORT), the catalog on 127.0.0.1:8090 (SSEBENCH_DEMO_CATALOG_PORT) and the proxy on LITELLM_PORT. The web UI container shares the host's network and mounts the Docker socket, so the demo needs a Linux Docker engine, and its terminal is off unless you set SSEBENCH_WEBUI_TERMINAL=1.

sh
uv run ssebench demo down
sh
ssebench demo down [-h]

It takes no options.

down removes the containers of runs on the demo's networks, then the Compose project with its database volume. It touches nothing else: not the stack you run with just launch, not other containers, not the images and not results/. It refuses to work on a Compose project that has containers the demo did not create.

ssebench doctor ​

Checks that the host can build images and run benchmarks, and prints a fix for each problem it finds.

sh
uv run ssebench doctor
sh
ssebench doctor [-h] [--json]
OptionDefaultDescription
--jsonoffPrint the checks as JSON, with the provider keys that each model in models/ needs and .env lacks

It checks Docker, buildx and Compose, the free disk space where Docker keeps its images, the CPU architecture (a warning on a host that is not amd64, where every pilot task runs under emulation), .env and the proxy secrets in it, whether the LiteLLM proxy answers on its port, and which provider keys that models/*.yaml refers to are set. It exits with status 1 when a required check fails; warnings do not change the exit status.

Commands from extensions ​

Installed packages can add subcommands; ssebench --help lists them. See Extension points.

Working directory ​

The CLI finds agents/, images/, runtime/, datasets/ and the Compose file in the SSEBench home, which is the first of:

  1. the directory in SSEBENCH_HOME;
  2. the nearest directory at or above the working directory whose pyproject.toml has a [tool.ssebench] table, which is the repository root;
  3. the repository the CLI was installed from, for an editable install;
  4. the copy of these directories that the ssebench package carries; see Without a clone.

From outside the repository, run uv run --project /path/to/ssebench ssebench ... or set SSEBENCH_HOME. The CLI reads .env in the workspace; variables set in the environment take precedence over it. Paths you pass, such as --local, --catalog and --benchmarks, are relative to the working directory, and so is results/.

The workspace holds what you edit and what runs produce: .env, models/ and results/. In a checkout it is the SSEBench home. Without a clone it is the working directory, so run ssebench from the directory that ssebench init set up.

Without a clone ​

ssebench runs from the wheel alone, as uvx ssebench or after pip install ssebench. A checkout, or SSEBENCH_HOME, still takes precedence when there is one.

The wheel carries, under ssebench/_data/ and laid out like the repository:

  • the agent definitions in agents/, with their Dockerfiles;
  • models/ and the Compose file with the LiteLLM image's sources;
  • the tool layer Dockerfiles, the sources of the SDK, the evaluator and the MCP server, uv.lock, and runtime/plugins/;
  • the pilot manifest, datasets/pilot/manifest.json, and the digests of its published case images, datasets/pilot/images.lock.json, but not the task folders.

What a run needs beyond that comes from the registry and the network:

  • The task's case image is pulled from $SSEBENCH_REGISTRY/case/<dataset>/<task>, by the digest in the packaged images.lock.json when it pins the task. Without a task folder to build it from, a failed pull ends the run.
  • The tool layer and the agent image are built on top of the case image from the packaged files, as in a checkout. The daemon and the entrypoint, which need Rust and Go to build, come from the published runtime image of the same version, passed as --build-context runtime=docker-image://$SSEBENCH_REGISTRY/runtime:<version>. The Python environments of the evaluator and the MCP server are built from the packaged uv.lock, which downloads packages.
  • The proxy image is built from the packaged Dockerfile with the models in the workspace's models/, and rebuilt when they change.

Commands that read task folders, such as ssebench build-case and ssebench dataset validate without a directory, need a checkout or an explicit directory.

Released under the Apache License 2.0.