Skip to content

Configuration files ​

SSEBench reads these files. Paths are relative to the repository root, the SSEBench home.

FileRead byWhat it configures
.envthe CLI, just, ComposeLocal secrets, provider keys and settings
models/*.yamlthe LiteLLM proxy imageThe models --model can name
agents/<agent>/agent.yamlthe CLIAn agent's image
runtime/plugins/plugins.yamlthe CLI, the entrypointPlugins and their hooks
datasets/<dataset>/<task>/sse/config.yamlthe CLI, the daemonA task
datasets/<dataset>/dataset.yamlthe CLIA dataset's version

.env ​

.env in the repository root holds local secrets and settings. just setup writes it from .env.example, with a generated LiteLLM master key and Postgres password. It is ignored by git; never commit it. The CLI reads it from the SSEBench home, and variables set in the environment take precedence over it.

VariableDefaultDescription
LITELLM_MASTER_KEYgenerated by just setupAdmin key of the local LiteLLM proxy; the CLI uses it to create a key for each run, and the web UI to make the key of a run's AI assistant. Required.
POSTGRES_PASSWORDgenerated by just setupPassword of the proxy's Postgres database; letters and digits only, as it is part of a URL. Postgres keeps the password it was created with, so after changing it remove the stack's volume. Required.
ANTHROPIC_API_KEYunsetKey of the Anthropic models in models/anthropic-claude.yaml. sse.ai also uses it when it runs OpenCode outside a task container.
OPENAI_API_KEYunsetKey of the OpenAI models in models/openai-gpt.yaml.
GOOGLE_API_KEYunsetKey of the Gemini models in models/google-gemini.yaml.
LITELLM_PORT4000Host port of the LiteLLM proxy.
LITELLM_BIND127.0.0.1Host address that the LiteLLM proxy's port is published on. The proxy holds the master key and the provider keys, so it listens on loopback only; set 0.0.0.0 to reach it from other machines, and only on a network you trust.
COMPOSE_PROJECT_NAMEssebenchCompose project of the proxy stack. Its containers, networks (<project>_default, and the internal <project>_agents that run containers join by default) and database volume (<project>_postgres_data) carry this name, so stacks with different names and ports run side by side.
SSEBENCH_REGISTRYghcr.io/42-b3yond-6ug/ssebenchRegistry prefix of every image SSEBench builds or pulls, and of the image names that the catalog service returns.
SSEBENCH_CATALOGthe bundled pilot manifestTask catalog that ssebench run, ssebench tasks list and the web UI get tasks from: the path or URL of a manifest.json, a dataset directory, or the URL of a catalog service. --catalog overrides it.

It may set any other variable too, such as the key of a provider you add or the settings in Environment variables. After changing it, restart the proxy with just launch.

models/*.yaml ​

Each file is a list of LiteLLM model definitions. The LiteLLM proxy image combines every .yaml and .yml file in models/ into its configuration, and the proxy offers each model under its model_name, which is what ssebench run --model takes. just launch and ssebench run rebuild the image when models/ changes.

yaml
- model_name: claude-sonnet-4-6
  litellm_params:
    model: anthropic/claude-sonnet-4-6
    api_key: os.environ/ANTHROPIC_API_KEY
  model_info:
    input_cost_per_token: 3.0e-06
    output_cost_per_token: 1.5e-05

Add a model describes the keys, and the LiteLLM documentation all of them. The models defined now:

--modelLiteLLM modelKeyFile
claude-opus-5-5anthropic/claude-opus-5-5ANTHROPIC_API_KEYmodels/anthropic-claude.yaml
claude-sonnet-5-5anthropic/claude-sonnet-5-5ANTHROPIC_API_KEYmodels/anthropic-claude.yaml
claude-opus-4-6anthropic/claude-opus-4-6ANTHROPIC_API_KEYmodels/anthropic-claude.yaml
claude-sonnet-4-6anthropic/claude-sonnet-4-6ANTHROPIC_API_KEYmodels/anthropic-claude.yaml
claude-opus-4-5anthropic/claude-opus-4-5ANTHROPIC_API_KEYmodels/anthropic-claude.yaml
claude-sonnet-4-5anthropic/claude-sonnet-4-5ANTHROPIC_API_KEYmodels/anthropic-claude.yaml
claude-haiku-4-5anthropic/claude-haiku-4-5ANTHROPIC_API_KEYmodels/anthropic-claude.yaml
gemini-3.1-progemini/gemini-3.1-pro-previewGOOGLE_API_KEYmodels/google-gemini.yaml
gemini-3-flashgemini/gemini-3-flash-previewGOOGLE_API_KEYmodels/google-gemini.yaml
gpt-5.2openai/gpt-5.2OPENAI_API_KEYmodels/openai-gpt.yaml
gpt-5.1openai/gpt-5.1OPENAI_API_KEYmodels/openai-gpt.yaml
gpt-5.1-codex-maxopenai/gpt-5.1-codex-maxOPENAI_API_KEYmodels/openai-gpt.yaml
gpt-5.1-codexopenai/gpt-5.1-codexOPENAI_API_KEYmodels/openai-gpt.yaml

agent.yaml ​

Every directory in agents/ is an agent: ssebench run --agent <directory> builds its Dockerfile on top of the tool layer and runs it. agent.yaml next to the Dockerfile names the agent's image. Unknown keys are rejected. Add an agent builds one from scratch.

yaml
name: claude-code
KeyTypeRequiredDescription
namestringyesName of the agent, used in the names of its images.
versionstringTag of the agent image; the SSEBench version when it is not set.

plugins.yaml ​

runtime/plugins/plugins.yaml lists the plugins in runtime/plugins/, each a folder with an executable run.sh, and when they run. runtime/plugins/schema.json is its JSON Schema; the CLI validates the file against it when it builds the tool layer, and the entrypoint again when the container starts. See Plugins and hooks.

yaml
- name: artifact
  enabled: false
  hook: after-grading
  llm: false
  timeout: 5

Each entry has:

KeyTypeRequiredDescription
namestringyesName of the plugin; it must equal the name of its folder.
enabledbooleanyesWhether the plugin runs in every run; the --plugin option of ssebench run selects plugins for one run instead.
hookbefore-agent | on-agent | after-agent | before-grading | on-grading | after-gradingyesWhen the plugin runs: before (blocking), on (in parallel with) or after (blocking) the agent or the grading.
llmbooleanyesWhether the plugin gets the run's model: SSE_BASE_URL, SSE_API_KEY and SSE_MODEL_NAME.
timeoutintegeryesTime limit of the plugin, in minutes; it is stopped when it runs longer.

sse/config.yaml ​

The task config, in sse/ of each task folder, describes the task to the CLI, the daemon and the grader. Paths in it are relative to /ssebench in the case image. Unknown keys are rejected. Dataset manifest has an example and explains the layout of a task folder and the checks the grader derives from the config; Add a task shows how to write one. datasets/schema/task.schema.json is its JSON Schema, which ssebench dataset schema exports.

KeyTypeRequiredDescription
idstringyesTask ID, equal to the task's folder name.
projectstringyesName of the upstream project.
repositorystringyesURL of the upstream repository.
languagec | go | rustyesLanguage of the project. The case image builds on the base image of that language.
sourcestringyesAbsolute path of the project's source tree in the case image.
task_descriptionobjectyesWhat the agent is told about the vulnerability. At least one field must be set.
task_description.issuestringIssue text given to the agent.
task_description.crash_reportlist of stringReport files given to the agent, such as the upstream issue or a sanitizer log.
task_description.bug_descriptionstringShort description of the bug given to the agent.
scriptsobjectyesScripts that the grader and the agent's test_patch tool run.
scripts.buildstringBuilds the project in the source directory.
scripts.runstringRuns the proof of concept given as its argument.
scripts.teststringRuns the project's tests.
filesobjectyesReference material. None of it is shown to the agent.
files.patchstringThe upstream fix, as a diff against the source directory.
files.future_teststringHidden tests of the fix, usually those added upstream, as a diff. The intent test applies it and runs scripts.test.
files.security_teststringRegression test for the vulnerability, as a diff. Kept for reference: the grader does not apply it, and its security check runs the proofs of concept instead.
files.intent_teststringTests of the intended behaviour, as a diff. Kept for reference: the intent test uses future_test.
files.poclist of stringProof-of-concept inputs, each run with scripts.run.
sanitizerstringSanitizer the build enables, such as address.
typestringClass of the bug, such as Heap Buffer Overflow or a CWE.
binarystringProgram or harness that the proofs of concept exercise.
trigger_commitstringUpstream revision that the case image builds.
patch_commitstringUpstream commits that fix the bug, comma-separated.
referencelist of stringLinks to the advisory, the report and the fix.
originalitypublic | craftedOrigin of the proofs of concept: public when they come from the public report or advisory, possibly adapted; crafted when they were written for the task because the report has no reproducer.

task_description needs at least one of its keys.

dataset.yaml ​

dataset.yaml, next to the task folders of a dataset, holds the dataset's version, which the manifest records. datasets/schema/dataset.schema.json is its JSON Schema.

yaml
version: pilot-v1
KeyTypeRequiredDescription
versionstringyesDataset version, such as pilot-v1.

Released under the Apache License 2.0.