Dataset manifest
A dataset is a folder under datasets/ with one folder per task. Each task has a config, sse/config.yaml, and each dataset has a generated manifest.json that lists its tasks for the CLI, the catalog service and CI.
Tasks and datasets explains the concepts; this page is the reference, and Add a task walks through writing and checking a task.
The Pydantic models in bench/src/ssebench/tasks/ (metadata.py for the task config, manifest.py for the manifest and dataset.yaml) are the definition of these files. The JSON Schemas in datasets/schema/ are exported from them, for use by other tools.
Layout
datasets/
├── schema/ # exported by `ssebench dataset schema`
│ ├── task.schema.json # sse/config.yaml
│ ├── dataset.schema.json # dataset.yaml
│ └── manifest.schema.json # manifest.json
└── pilot/
├── dataset.yaml # the dataset version
├── manifest.json # generated by `ssebench dataset manifest`
└── <task-id>/
├── Dockerfile # case image: the project at the vulnerable commit
└── sse/
├── config.yaml # the task config → /ssebench/config.yaml
├── build.sh # builds the project → /ssebench/scripts/build.sh
├── run.sh # runs one PoC → /ssebench/scripts/run.sh
├── test.sh # runs the tests → /ssebench/scripts/test.sh
├── pocs/ # PoC inputs → /ssebench/pocs/
├── reports/ # what the agent is told → /ssebench/reports/
└── diffs/ # reference patch and hidden tests → /ssebench/diffs/Every subdirectory of a dataset is a task. A task folder may hold more build inputs next to sse/, such as the vendored crate sources of the Rust tasks.
The rules for a task folder:
- The folder name is the task ID. It is used for
--task, image names and result paths, and the config'sidmust be equal to it. IDs are letters and digits separated by.,_or-, up to 128 characters. The case image iscase/<dataset>/<id>in lowercase, so two IDs may not differ only in case. - The Dockerfile builds on a pinned SSEBench base image. It declares
ARG SSEBENCH_REGISTRYbefore the firstFROM, and its lastFROMis${SSEBENCH_REGISTRY}/<base image>:<version>, for example${SSEBENCH_REGISTRY}/base-generic-go:1.0.0. The base image needs a version tag or a digest;latestis rejected. See Base images. - The Dockerfile copies
sse/to/ssebenchas shown above. Paths in the config are relative to/ssebenchin the case image, and each one must be a file that the Dockerfile copies from the task folder.
Base images
The base images hold the toolchain of each language and are defined in images/base-images/: generic-c, generic-go and generic-rust. Their inputs are pinned: the upstream images by digest, and the ccache release that generic-c downloads by checksum.
A task builds on the base images of one SSEBench release, named by the release's version tag; every pilot task uses 1.0.0. The manifest records each task's base image, so jq -r '.tasks[].base' datasets/pilot/manifest.json | sort -u lists the images a dataset needs. Docker pulls them when it builds a case image. make -C images/base-images (or just base-images) builds them from your checkout instead: besides the current version and latest, it tags them with every version that a dataset manifest names, so that case builds use them.
Moving a task to the base images of a later release changes its Dockerfile, and so the dataset version.
Network access
Building a case image uses the network: the Dockerfile clones the upstream project at a pinned commit and installs its distribution packages, Go modules, crates and toolchains. Pin versions where the upstream project does not: the Rust tasks whose upstream project commits no Cargo.lock carry one next to their Dockerfile.
Grading does not use the network. Under the default egress policy (see --egress), a run container reaches only the LiteLLM proxy, so everything that build.sh, run.sh and test.sh need, with or without the reference patch and the hidden tests applied, must already be in the case image: vendored Go modules or the module cache, fetched crates, toolchains and test data. To check a task, run its scripts in the case image without a network, for example:
docker run --rm --network none <case image> \
bash -c 'cd <source> && /ssebench/scripts/build.sh && /ssebench/scripts/test.sh'Task config
sse/config.yaml describes the task to the CLI, the daemon and the grader. Keys that are not listed here are rejected. For example:
id: gjson-196-bf4efcb
project: gjson
repository: https://github.com/tidwall/gjson
language: go
source: /src/gjson
task_description:
crash_report:
- reports/crash_report_1.txt
- reports/issue.md
bug_description:
scripts:
build: scripts/build.sh
run: scripts/run.sh
test: scripts/test.sh
files:
patch: diffs/patch.diff
poc:
- pocs/poc.go
future_test: diffs/test.diff
intent_test:
sanitizer:
type: Slice bounds out of range
binary:
trigger_commit: 9f58baa7a613f89dfdc764c39e47fd3a15606153
patch_commit: bf4efcb3c18d1825b2988603dea5909140a5302b
reference:
- https://github.com/tidwall/gjson
- https://github.com/tidwall/gjson/commit/bf4efcb3c18d1825b2988603dea5909140a5302b
- https://github.com/tidwall/gjson/issues/196
originality: public| Key | Required | Meaning |
|---|---|---|
id | yes | Task ID, equal to the folder name. |
project | yes | Name of the upstream project. |
repository | yes | URL of the upstream repository, https://…. |
language | yes | Language of the project, lowercase: c, go or rust. |
source | yes | Absolute path of the project's source tree in the case image. The agent edits it, and the scripts build and test it. |
task_description | yes | What the agent is told; at least one of the keys below. The task prompt includes every one that is set. |
task_description.issue | Issue text. | |
task_description.crash_report | Report files, such as the upstream issue or a sanitizer log. | |
task_description.bug_description | Short description of the bug. | |
scripts | yes | Scripts that the grader and the agent's test_patch tool run. |
scripts.build | Builds the project. | |
scripts.run | Runs the proof of concept given as its argument. It exits 0 when the vulnerability no longer triggers. | |
scripts.test | Runs the project's tests. | |
files | yes | Reference material; none of it is shown to the agent. |
files.patch | The upstream fix, as a diff against the source tree. | |
files.poc | Proof-of-concept inputs, each run with scripts.run. | |
files.future_test | Hidden tests of the fix, usually those added upstream, as a diff. The intent test applies it and runs scripts.test. | |
files.security_test | A regression test for the vulnerability, as a diff. Kept for reference; the grader does not apply it. Leave it unset when it would repeat future_test, as in every pilot task. | |
files.intent_test | Tests of the intended behaviour, as a diff. Kept for reference; the intent test uses future_test. Leave it unset when it would repeat future_test, as in every pilot task. | |
sanitizer | Sanitizer the build enables, such as address. | |
type | Class of the bug, such as Heap Buffer Overflow or a CWE. | |
binary | Program or harness that the proofs of concept exercise. | |
trigger_commit | Upstream revision that the case image builds. | |
patch_commit | Upstream commits that fix the bug, comma-separated. | |
reference | Links to the advisory, the report and the fix. | |
originality | Where the proofs of concept come from; see below. |
originality is public when the proofs of concept come from the public report or advisory, possibly adapted to the task's harness, and crafted when they were written for the task because the public report has no reproducer. The Go tasks of pilot record it; the others leave it out.
Checks
The grader runs every check that the config makes possible, in this order; see Grading pipeline. The difficulty level only limits which of them the agent's test_patch tool may run.
| Check | Needs | What it does |
|---|---|---|
build | scripts.build | Builds the patched project. |
poc | scripts.run and files.poc | Runs every proof of concept. This is the security check. |
function_test | scripts.test | Runs the project's existing tests. |
intent_test | scripts.test and files.future_test | Applies the hidden tests and runs the tests again. |
Dataset version
dataset.yaml holds the dataset's version, which the manifest records:
version: pilot-v1Change it whenever tasks are added, removed or changed in a way that can change their results.
Manifest
manifest.json lists the tasks of one dataset version. It is generated from the task folders and committed next to them:
{
"dataset": "pilot",
"version": "pilot-v1",
"tasks": [
{
"id": "gjson-196-bf4efcb",
"language": "go",
"project": "gjson",
"repository": "https://github.com/tidwall/gjson",
"base": "base-generic-go:1.0.0",
"image": "case/pilot/gjson-196-bf4efcb",
"arch": ["amd64"],
"checks": ["build", "poc", "function_test", "intent_test"],
"files": {
"Dockerfile": "94e4db62…",
"sse/config.yaml": "3c51886a…"
},
"metadata": { "id": "gjson-196-bf4efcb", "project": "gjson", "…": "…" }
}
]
}| Key | Meaning |
|---|---|
dataset | Dataset name, the name of its folder. |
version | Dataset version, from dataset.yaml. |
generated_from | Commit of the SSEBench repository the manifest was generated from. Optional; the committed manifest leaves it out, and release builds can record it. |
tasks | One entry per task, sorted by id. |
tasks[].id | Task ID. |
tasks[].language, project, repository | From the task config. |
tasks[].base | Base image of the case image, relative to the registry, with its version: the last FROM of the task's Dockerfile without ${SSEBENCH_REGISTRY}/, such as base-generic-go:1.0.0. |
tasks[].image | Case image, relative to the registry: case/<dataset>/<id>, lowercase. |
tasks[].arch | Platforms the task runs on, in order of preference. A run uses the host's architecture if it is listed, else the first; see Architectures. amd64 for every task today, since the case images are built and verified for amd64 only. |
tasks[].checks | The checks the grader can run for the task. |
tasks[].files | SHA-256 of every file in the task folder, by path relative to the folder. |
tasks[].metadata | The task config, validated, with every key present. |
Image names carry no registry and no tag. To pull a task's case image, prefix it with $SSEBENCH_REGISTRY and tag it with the dataset's version, for example ghcr.io/42-b3yond-6ug/ssebench/case/pilot/gjson-196-bf4efcb:pilot-v1. The pilot dataset page describes the published images.
Images lock
images.lock.json, next to manifest.json, pins the published case image of each task by digest. ssebench dataset lock writes it, and datasets/schema/images-lock.schema.json describes it:
{
"dataset": "pilot",
"version": "pilot-v1",
"images": {
"gjson-196-bf4efcb": {
"digest": "sha256:9c4f…",
"files_sha256": "5d1e…",
"revision": "0123456789abcdef0123456789abcdef01234567"
}
}
}| Key | Meaning |
|---|---|
dataset, version | The dataset and version of the manifest that the lock goes with. A lock for another version is ignored. |
images | One entry per published task, sorted by ID; a task without an entry is pulled by tag. |
images[].digest | Digest of the image in the registry. docker pull <registry>/<image>@<digest> pulls exactly that image, from any registry that holds a copy with its digests. |
images[].files_sha256 | SHA-256 of the task's files in the manifest that the image was built from. The entry applies only while the manifest lists the same files. |
images[].revision | Commit of the repository that the image was built and verified at. |
A publishing run of the Dataset workflow produces the lock; see Releasing and versioning for how it gets into the repository and the release. A lock that lags behind a task that changed is harmless: ssebench run skips the entries whose files changed, and ssebench dataset lock drops them.
Commands
uv run ssebench dataset validate [DIR]
uv run ssebench dataset manifest [DIR] [-o FILE] [--check] [--generated-from COMMIT]
uv run ssebench dataset schema [-o DIR] [--check]
uv run ssebench dataset lock [DIR] [--records DIR] [-o FILE] [--check]DIR defaults to datasets/pilot; see the CLI reference.
validatechecks every task folder against the rules above and lists every problem it finds, per task. It exits 1 if any task is invalid.manifestvalidates the dataset and writesmanifest.jsoninto it. The output depends only on the task folders. Since the manifest records a checksum of every task file, regenerate it after changing anything in a task folder;--checkfails, with a diff, when the committed file is out of date.schemawrites the JSON Schemas;--checkfails when they are out of date.lockwrites the images lock: the images of the current lock whose task files are unchanged, and those of the publication records thatssebench dataset publishwrote.--checkfails when the lock is missing, is for another dataset version or pins a task the dataset does not have.
The JSON Schemas also work with other validators, for example:
check-jsonschema --schemafile datasets/schema/task.schema.json datasets/pilot/*/sse/config.yamlThey cannot express the rules that involve other files, such as the id matching the folder name or the config's paths existing; ssebench dataset validate checks those.
Catalog service
The catalog service (catalog/, Go) serves a manifest over HTTP. In each task it returns, base and image are prefixed with the registry, and image is tagged with the dataset's version, or named by its digest (<registry>/case/pilot/<id>@sha256:…) when the images lock next to the manifest pins the task to an image built from the files the manifest lists. A client can pull the image it gets as it is.
go run ./catalog/cmd/ssebench-catalog serve --manifest datasets/pilot/manifest.json| Option | Environment | Default | Description |
|---|---|---|---|
--manifest FILE | SSEBENCH_MANIFEST | datasets/pilot/manifest.json | Manifest to serve |
--lock FILE | SSEBENCH_IMAGES_LOCK | images.lock.json next to the manifest, if there is one | Images lock that pins the case images by digest; a lock that was asked for must exist |
--registry PREFIX | SSEBENCH_REGISTRY | ghcr.io/42-b3yond-6ug/ssebench | Registry prefix of the images |
--port PORT | SSEBENCH_CATALOG_PORT | 8080 | TCP port |
| Endpoint | Returns |
|---|---|
GET /tasks | Every task, sorted by id, without files and metadata |
GET /tasks/{id} | One task |
GET /tasks/{id}/metadata | The task config |
GET /manifest.json | The manifest as loaded, with image names relative to the registry and without a tag |
docker build -f catalog/Dockerfile ., from the repository root, builds an image that serves the committed datasets/pilot/manifest.json and images.lock.json.
The service is optional. The CLI and the web UI read a manifest directly from a file or URL, and default to the bundled datasets/pilot/manifest.json, so they list tasks without any server or network. To use a service, set SSEBENCH_CATALOG (or pass ssebench run --catalog) to its base URL, such as http://localhost:8080; they read its GET /manifest.json and prefix image names with their own $SSEBENCH_REGISTRY, and tag them with the dataset version. See Task catalog.