Configuration files
SSEBench reads these files. Paths are relative to the repository root, the SSEBench home.
| File | Read by | What it configures |
|---|---|---|
.env | the CLI, just, Compose | Local secrets, provider keys and settings |
models/*.yaml | the LiteLLM proxy image | The models --model can name |
agents/<agent>/agent.yaml | the CLI | An agent's image |
runtime/plugins/plugins.yaml | the CLI, the entrypoint | Plugins and their hooks |
datasets/<dataset>/<task>/sse/config.yaml | the CLI, the daemon | A task |
datasets/<dataset>/dataset.yaml | the CLI | A dataset's version |
.env
.env in the repository root holds local secrets and settings. just setup writes it from .env.example, with a generated LiteLLM master key and Postgres password. It is ignored by git; never commit it. The CLI reads it from the SSEBench home, and variables set in the environment take precedence over it.
| Variable | Default | Description |
|---|---|---|
LITELLM_MASTER_KEY | generated by just setup | Admin key of the local LiteLLM proxy; the CLI uses it to create a key for each run, and the web UI to make the key of a run's AI assistant. Required. |
POSTGRES_PASSWORD | generated by just setup | Password of the proxy's Postgres database; letters and digits only, as it is part of a URL. Postgres keeps the password it was created with, so after changing it remove the stack's volume. Required. |
ANTHROPIC_API_KEY | unset | Key of the Anthropic models in models/anthropic-claude.yaml. sse.ai also uses it when it runs OpenCode outside a task container. |
OPENAI_API_KEY | unset | Key of the OpenAI models in models/openai-gpt.yaml. |
GOOGLE_API_KEY | unset | Key of the Gemini models in models/google-gemini.yaml. |
LITELLM_PORT | 4000 | Host port of the LiteLLM proxy. |
LITELLM_BIND | 127.0.0.1 | Host address that the LiteLLM proxy's port is published on. The proxy holds the master key and the provider keys, so it listens on loopback only; set 0.0.0.0 to reach it from other machines, and only on a network you trust. |
COMPOSE_PROJECT_NAME | ssebench | Compose project of the proxy stack. Its containers, networks (<project>_default, and the internal <project>_agents that run containers join by default) and database volume (<project>_postgres_data) carry this name, so stacks with different names and ports run side by side. |
SSEBENCH_REGISTRY | ghcr.io/42-b3yond-6ug/ssebench | Registry prefix of every image SSEBench builds or pulls, and of the image names that the catalog service returns. |
SSEBENCH_CATALOG | the bundled pilot manifest | Task catalog that ssebench run, ssebench tasks list and the web UI get tasks from: the path or URL of a manifest.json, a dataset directory, or the URL of a catalog service. --catalog overrides it. |
It may set any other variable too, such as the key of a provider you add or the settings in Environment variables. After changing it, restart the proxy with just launch.
models/*.yaml
Each file is a list of LiteLLM model definitions. The LiteLLM proxy image combines every .yaml and .yml file in models/ into its configuration, and the proxy offers each model under its model_name, which is what ssebench run --model takes. just launch and ssebench run rebuild the image when models/ changes.
- model_name: claude-sonnet-4-6
litellm_params:
model: anthropic/claude-sonnet-4-6
api_key: os.environ/ANTHROPIC_API_KEY
model_info:
input_cost_per_token: 3.0e-06
output_cost_per_token: 1.5e-05Add a model describes the keys, and the LiteLLM documentation all of them. The models defined now:
--model | LiteLLM model | Key | File |
|---|---|---|---|
claude-opus-5-5 | anthropic/claude-opus-5-5 | ANTHROPIC_API_KEY | models/anthropic-claude.yaml |
claude-sonnet-5-5 | anthropic/claude-sonnet-5-5 | ANTHROPIC_API_KEY | models/anthropic-claude.yaml |
claude-opus-4-6 | anthropic/claude-opus-4-6 | ANTHROPIC_API_KEY | models/anthropic-claude.yaml |
claude-sonnet-4-6 | anthropic/claude-sonnet-4-6 | ANTHROPIC_API_KEY | models/anthropic-claude.yaml |
claude-opus-4-5 | anthropic/claude-opus-4-5 | ANTHROPIC_API_KEY | models/anthropic-claude.yaml |
claude-sonnet-4-5 | anthropic/claude-sonnet-4-5 | ANTHROPIC_API_KEY | models/anthropic-claude.yaml |
claude-haiku-4-5 | anthropic/claude-haiku-4-5 | ANTHROPIC_API_KEY | models/anthropic-claude.yaml |
gemini-3.1-pro | gemini/gemini-3.1-pro-preview | GOOGLE_API_KEY | models/google-gemini.yaml |
gemini-3-flash | gemini/gemini-3-flash-preview | GOOGLE_API_KEY | models/google-gemini.yaml |
gpt-5.2 | openai/gpt-5.2 | OPENAI_API_KEY | models/openai-gpt.yaml |
gpt-5.1 | openai/gpt-5.1 | OPENAI_API_KEY | models/openai-gpt.yaml |
gpt-5.1-codex-max | openai/gpt-5.1-codex-max | OPENAI_API_KEY | models/openai-gpt.yaml |
gpt-5.1-codex | openai/gpt-5.1-codex | OPENAI_API_KEY | models/openai-gpt.yaml |
agent.yaml
Every directory in agents/ is an agent: ssebench run --agent <directory> builds its Dockerfile on top of the tool layer and runs it. agent.yaml next to the Dockerfile names the agent's image. Unknown keys are rejected. Add an agent builds one from scratch.
name: claude-code| Key | Type | Required | Description |
|---|---|---|---|
name | string | yes | Name of the agent, used in the names of its images. |
version | string | Tag of the agent image; the SSEBench version when it is not set. |
plugins.yaml
runtime/plugins/plugins.yaml lists the plugins in runtime/plugins/, each a folder with an executable run.sh, and when they run. runtime/plugins/schema.json is its JSON Schema; the CLI validates the file against it when it builds the tool layer, and the entrypoint again when the container starts. See Plugins and hooks.
- name: artifact
enabled: false
hook: after-grading
llm: false
timeout: 5Each entry has:
| Key | Type | Required | Description |
|---|---|---|---|
name | string | yes | Name of the plugin; it must equal the name of its folder. |
enabled | boolean | yes | Whether the plugin runs in every run; the --plugin option of ssebench run selects plugins for one run instead. |
hook | before-agent | on-agent | after-agent | before-grading | on-grading | after-grading | yes | When the plugin runs: before (blocking), on (in parallel with) or after (blocking) the agent or the grading. |
llm | boolean | yes | Whether the plugin gets the run's model: SSE_BASE_URL, SSE_API_KEY and SSE_MODEL_NAME. |
timeout | integer | yes | Time limit of the plugin, in minutes; it is stopped when it runs longer. |
sse/config.yaml
The task config, in sse/ of each task folder, describes the task to the CLI, the daemon and the grader. Paths in it are relative to /ssebench in the case image. Unknown keys are rejected. Dataset manifest has an example and explains the layout of a task folder and the checks the grader derives from the config; Add a task shows how to write one. datasets/schema/task.schema.json is its JSON Schema, which ssebench dataset schema exports.
| Key | Type | Required | Description |
|---|---|---|---|
id | string | yes | Task ID, equal to the task's folder name. |
project | string | yes | Name of the upstream project. |
repository | string | yes | URL of the upstream repository. |
language | c | go | rust | yes | Language of the project. The case image builds on the base image of that language. |
source | string | yes | Absolute path of the project's source tree in the case image. |
task_description | object | yes | What the agent is told about the vulnerability. At least one field must be set. |
task_description.issue | string | Issue text given to the agent. | |
task_description.crash_report | list of string | Report files given to the agent, such as the upstream issue or a sanitizer log. | |
task_description.bug_description | string | Short description of the bug given to the agent. | |
scripts | object | yes | Scripts that the grader and the agent's test_patch tool run. |
scripts.build | string | Builds the project in the source directory. | |
scripts.run | string | Runs the proof of concept given as its argument. | |
scripts.test | string | Runs the project's tests. | |
files | object | yes | Reference material. None of it is shown to the agent. |
files.patch | string | The upstream fix, as a diff against the source directory. | |
files.future_test | string | Hidden tests of the fix, usually those added upstream, as a diff. The intent test applies it and runs scripts.test. | |
files.security_test | string | Regression test for the vulnerability, as a diff. Kept for reference: the grader does not apply it, and its security check runs the proofs of concept instead. | |
files.intent_test | string | Tests of the intended behaviour, as a diff. Kept for reference: the intent test uses future_test. | |
files.poc | list of string | Proof-of-concept inputs, each run with scripts.run. | |
sanitizer | string | Sanitizer the build enables, such as address. | |
type | string | Class of the bug, such as Heap Buffer Overflow or a CWE. | |
binary | string | Program or harness that the proofs of concept exercise. | |
trigger_commit | string | Upstream revision that the case image builds. | |
patch_commit | string | Upstream commits that fix the bug, comma-separated. | |
reference | list of string | Links to the advisory, the report and the fix. | |
originality | public | crafted | Origin of the proofs of concept: public when they come from the public report or advisory, possibly adapted; crafted when they were written for the task because the report has no reproducer. |
task_description needs at least one of its keys.
dataset.yaml
dataset.yaml, next to the task folders of a dataset, holds the dataset's version, which the manifest records. datasets/schema/dataset.schema.json is its JSON Schema.
version: pilot-v1| Key | Type | Required | Description |
|---|---|---|---|
version | string | yes | Dataset version, such as pilot-v1. |