Skip to content

Daemon HTTP API ​

ssebench-daemon runs in every task container. It knows the task: the project's source tree, the build, run and test scripts, and the reference material. The Python SDK, the MCP server, the evaluator and the web UI use it to read the task, to build and test the project, and to grade the patch. Agents normally use it through the SDK or the MCP server's test_patch tool.

Every request and response body is JSON. The API is described in OpenAPI 3.1 in sdk/daemon/openapi.yaml, which a test in the daemon's crate checks against the routes it serves, their access rules and their responses; the tables on this page are generated from it. Errors have the body {"error": "<message>"}.

Listeners ​

The daemon serves the same routes on up to three listeners. Only the admin socket is privileged.

ListenerAddressUsed by
Agent-facing Unix socketSSE_DAEMON_SOCKET, /tmp/sse.sock in sandbox mode; mode 0666the agent, the SDK in the agent's process, the MCP server
Agent-facing HTTPport SSE_HTTP_PORT (4263) on all interfacesthe web UI, from the host
Admin Unix socketSSE_ADMIN_SOCKET, /run/ssebench/admin.sock in sandbox mode; mode 0600, root onlythe entrypoint and the evaluator

The daemon binds the admin socket only when SSE_ADMIN_SOCKET is set, and the Unix socket only when SSE_DAEMON_SOCKET is set; without the latter it serves HTTP alone. The entrypoint sets both; see Environment variables.

The agent runs as the unprivileged user model, so it reaches the agent-facing listeners but not the admin socket. The daemon enforces the rules that keep the answer from the agent on the agent-facing listeners, whatever client calls it; the integrity model explains why:

  • Difficulty gate. The daemon reads SSE_DIFFICULTY when it starts and answers 403 to the bencher actions that the level withholds; see Difficulty gate. The MCP server applies the same levels to test_patch.
  • Agent phase. Once the agent phase ends — the entrypoint calls POST /admin/agent_exited on the admin socket when the agent process exits — the agent-facing listeners answer 403 to every tool, and the daemon kills the agent user's processes.
  • Admin-only endpoints. POST /prepare_grading, POST /admin/agent_exited and GET /reference/patch answer 403 on the agent-facing listeners at all times; only the admin socket serves them. Post-run tooling on the host reads the reference patch from the task folder or the run's results instead.

The admin socket is never gated, so grading runs every check the task has. The evaluator's SDK talks to it because the entrypoint points the evaluator's SSE_DAEMON_SOCKET at the admin socket. Post-agent tooling reads the reference patch with sse.reference.get_reference_patch().

Endpoints ​

EndpointServed onSummary
GET /versionevery listenerThe daemon's version.
GET /projectevery listenerThe task's metadata, without the reference material.
GET /capabilitiesevery listenerWhich checks the task supports.
POST /tool/{name}every listener, until the agent phase ends; then admin socket onlyRun an action of a tool.
GET /diffevery listenerThe agent's changes so far, as a unified diff.
GET /filesevery listenerThe files the agent changed, with line counts.
GET /agent/dialogevery listenerThe agent's dialog, from $SSE_ARCHIVE/dialog.jsonl.
GET /resultevery listenerThe grade, once the evaluator has written it.
POST /prepare_gradingadmin socket onlySave the agent's diff and apply it to the copy of the repository that is graded.
GET /final_diffevery listenerThe agent's diff that POST /prepare_grading saved.
POST /admin/agent_exitedadmin socket onlyEnd the agent phase.
GET /reference/patchadmin socket onlyThe reference patch, the answer to the task.

Difficulty gate ​

The bencher actions of POST /tool/{name} on the agent-facing listeners, by difficulty level:

Levelbuildfunction_testrun_pocintent_test
0 FULL_ASSISTANCErunsrunsrunsruns
1 NO_INTENT_TESTrunsrunsruns403
2 NO_FUTURE_TESTrunsruns403403
3 BUILD_ONLYruns403403403
4 NO_BUILD403403403403

The admin socket runs every action at every level. bash is never gated. See Difficulty levels for what the levels mean to the agent.

Agent-facing endpoints ​

These endpoints answer on every listener.

GET /version ​

The daemon's version.

The SDK compares it with its own version and warns when they differ.

Served on: every listener.

StatusBodyDescription
200VersionThe version.

GET /project ​

The task's metadata, without the reference material.

The task config, with the contents of the report files and of the build and test scripts instead of their paths. The reference patch and the hidden tests are left out.

Served on: every listener.

StatusBodyDescription
200MetadataThe task's metadata.

GET /capabilities ​

Which checks the task supports.

Derived from the scripts and files the task config names. The difficulty level can still withhold a supported check from the agent.

Served on: every listener.

StatusBodyDescription
200CapabilitiesThe task's capabilities.

POST /tool/{name} ​

Run an action of a tool.

bencher builds and tests a private copy of the source tree, running the task's scripts as the unprivileged runner user; bash runs a command in one interactive shell in the source tree, as the unprivileged user model. The daemon runs one action at a time. A script that runs and fails is a 200 with a non-zero code. The agent-facing listeners serve tools during the agent phase only, and never with grading: true.

Served on: every listener, until the agent phase ends; then admin socket only.

ParameterInTypeRequiredDescription
namepathbencher | bashyesThe tool.
actionquerystringyesThe action: build, run_poc, function_test or intent_test for bencher. bash has one action and ignores the value; the SDK sends execute.

Request body (application/json): GradingArgument | PocArgument | BashArgument. The action's argument, sent with Content-Type: application/json.

StatusBodyDescription
200ScriptResultThe action ran; its script's exit code and output.
400The action parameter is missing, or the body is not JSON.
403ErrorOn an agent-facing listener: a bencher action that the difficulty level withholds, an argument with grading: true, or any tool after the agent phase.
500ErrorAn unknown tool or action, an invalid argument, a script the task does not have, or hidden tests that do not apply (git apply failed).

GET /diff ​

The agent's changes so far, as a unified diff.

The diff of the source tree against the task's single initial commit, as grading captures it: every change, committed or not, and every new file that is not ignored, whatever the agent did with git.

Served on: every listener.

StatusBodyDescription
200DiffThe diff; empty when nothing changed.
500Errorgit failed.

GET /files ​

The files the agent changed, with line counts.

Served on: every listener.

StatusBodyDescription
200FilesThe files in GET /diff, with their status and line counts.
500Errorgit failed.

GET /agent/dialog ​

The agent's dialog, from $SSE_ARCHIVE/dialog.jsonl.

The entries the agent wrote in the dialog protocol. Lines that are not JSON, or have no seq, are skipped.

Served on: every listener.

ParameterInTypeRequiredDescription
sincequeryintegernoReturn only the entries whose seq is greater, for polling.
StatusBodyDescription
200DialogThe entries in file order; none when the file does not exist.
500ErrorThe file exists but cannot be read.

GET /result ​

The grade, once the evaluator has written it.

Served on: every listener.

StatusBodyDescription
200EvaluationResultThe grade, or available: false before the evaluator has written it or when it cannot be parsed.

GET /final_diff ​

The agent's diff that POST /prepare_grading saved.

Served on: every listener.

StatusBodyDescription
200DiffThe saved diff; empty before grading.
500ErrorThe saved diff cannot be read.

Admin endpoints ​

These endpoints answer on the admin socket. The agent-facing listeners refuse them with 403 at all times.

POST /prepare_grading ​

Save the agent's diff and apply it to the copy of the repository that is graded.

Starts a new session of the runner user, which kills its processes and discards its caches and workspaces, writes the agent's diff to final.patch and its commit messages to commits.log in the results directory ($SSE_RESULTS), then applies the diff to the clean clone of the repository in $SSE_REPO_PATH, which the bencher actions use with grading: true. The evaluator calls it once, after the agent has finished.

Served on: admin socket only.

StatusBodyDescription
200SuccessThe diff is saved and applied.
403ErrorThe request came from an agent-facing listener.
500Errorgit failed, or the diff does not apply (git apply failed).

POST /admin/agent_exited ​

End the agent phase.

The entrypoint calls it once the agent process has exited. From then on the agent-facing listeners refuse tools, and every process of the agent's user in the daemon's container is killed. Calling it again has no other effect.

Served on: admin socket only.

StatusBodyDescription
200SuccessThe agent phase has ended.
403ErrorThe request came from an agent-facing listener.
500ErrorA process of the agent's user could not be killed.

GET /reference/patch ​

The reference patch, the answer to the task.

The upstream fix that the task config names as files.patch. Only the admin socket serves it: the agent-facing listeners are reachable from other containers on the run network, also after the run. Tools on the host read it from the task folder or the run's results directory.

Served on: admin socket only.

StatusBodyDescription
200DiffThe patch; empty when the task has none or it cannot be read.
403ErrorThe request came from an agent-facing listener.

Schemas ​

Version ​

FieldTypeRequiredDescription
versionstringyesThe daemon's version, which is the SSEBench version.

Metadata ​

FieldTypeRequiredDescription
idstringyesTask ID.
projectstringyesName of the upstream project.
languagestringyesLanguage of the project, such as c, go or rust.
sourcestringyesAbsolute path of the source tree that the agent edits.
task_descriptionTaskDescriptionyes
pocarray of stringyesAbsolute paths of the proof-of-concept inputs, to pass to run_poc.
build_scriptstring | nullyesContents of the build script.
test_scriptstring | nullyesContents of the test script.

TaskDescription ​

FieldTypeRequiredDescription
issuestring | nullyesIssue text.
crash_reportarray of string | nullyesContents of the report files.
bug_descriptionstring | nullyesShort description of the bug.

Capabilities ​

FieldTypeRequiredDescription
can_buildbooleanyesThe task has a build script.
can_run_pocbooleanyesThe task has a run script and at least one proof of concept.
poc_countintegeryesNumber of proofs of concept.
has_function_testbooleanyesThe task has a test script.
has_intent_testbooleanyesThe task has a test script and hidden tests (files.future_test).

GradingArgument ​

The argument of build, function_test and intent_test.

FieldTypeRequiredDescription
gradingbooleannoUse the graded copy of the repository in $SSE_REPO_PATH instead of the source tree that the agent edits. Default: false.

PocArgument ​

The argument of run_poc. It runs in the output of the last build.

FieldTypeRequiredDescription
pocstringyesPath of the proof of concept, one of poc in the metadata.

BashArgument ​

The argument of bash.

FieldTypeRequiredDescription
commandstringyesShell command. It runs in the persistent shell, so cd and export carry over to the next command.

ScriptResult ​

FieldTypeRequiredDescription
codeintegeryesExit code; -1 for run_poc before any build.
stdoutstringyesStandard output.
stderrstringyesStandard error.

Diff ​

FieldTypeRequiredDescription
diffstringyesUnified diff, with binary changes as git binary patches.

Files ​

FieldTypeRequiredDescription
filesarray of ChangedFileyes

ChangedFile ​

FieldTypeRequiredDescription
pathstringyesPath relative to the source tree.
statusmodified | added | deletedyesA new file, staged or not, is added; a renamed file is a deleted file and an added one.
additionsintegeryesLines added; 0 for a binary file.
deletionsintegeryesLines deleted; 0 for a binary file.

Dialog ​

FieldTypeRequiredDescription
entriesarray of objectyesDialog entries, as the dialog protocol defines them.

EvaluationResult ​

FieldTypeRequiredDescription
availablebooleanyesWhether the evaluator has written the grade.
patch_resultPatchResultno
runtime_resultRuntimeResultno

PatchResult ​

The grade; a check the task does not have is null.

FieldTypeRequiredDescription
statuspassed | failed | errornopassed when every check that ran passed, failed when a check failed or the patch did not apply, error when no check ran and the patch was not graded. Absent from the results of older evaluators.
build_successboolean | nullnoWhether the patched project builds.
pov_passedinteger | nullnoNumber of proofs of concept that no longer trigger the vulnerability.
pov_totalinteger | nullnoNumber of proofs of concept run.
func_test_successboolean | nullnoWhether the project's tests pass.
intent_test_successboolean | nullnoWhether the tests pass with the hidden tests applied.
error_msgstring | nullnoThe first failure, or with status error, why the patch was not graded.
error_logstring | nullnoOutput of the step that failed first.

RuntimeResult ​

FieldTypeRequiredDescription
agent_durationintegeryesHow long the agent ran, in seconds.
agent_timeoutbooleanyesWhether the agent hit its time limit.
evaluator_timeoutbooleanyesWhether grading hit its time limit.

Success ​

FieldTypeRequiredDescription
success`true
`yes

Error ​

FieldTypeRequiredDescription
errorstringyesWhat went wrong.

Released under the Apache License 2.0.