Skip to content

Dataset manifest ​

A dataset is a folder under datasets/ with one folder per task. Each task has a config, sse/config.yaml, and each dataset has a generated manifest.json that lists its tasks for the CLI, the catalog service and CI.

Tasks and datasets explains the concepts; this page is the reference, and Add a task walks through writing and checking a task.

The Pydantic models in bench/src/ssebench/tasks/ (metadata.py for the task config, manifest.py for the manifest and dataset.yaml) are the definition of these files. The JSON Schemas in datasets/schema/ are exported from them, for use by other tools.

Layout ​

datasets/
├── schema/                    # exported by `ssebench dataset schema`
│   ├── task.schema.json       # sse/config.yaml
│   ├── dataset.schema.json    # dataset.yaml
│   └── manifest.schema.json   # manifest.json
└── pilot/
    ├── dataset.yaml           # the dataset version
    ├── manifest.json          # generated by `ssebench dataset manifest`
    └── <task-id>/
        ├── Dockerfile         # case image: the project at the vulnerable commit
        └── sse/
            ├── config.yaml    # the task config  → /ssebench/config.yaml
            ├── build.sh       # builds the project → /ssebench/scripts/build.sh
            ├── run.sh         # runs one PoC     → /ssebench/scripts/run.sh
            ├── test.sh        # runs the tests   → /ssebench/scripts/test.sh
            ├── pocs/          # PoC inputs       → /ssebench/pocs/
            ├── reports/       # what the agent is told → /ssebench/reports/
            └── diffs/         # reference patch and hidden tests → /ssebench/diffs/

Every subdirectory of a dataset is a task. A task folder may hold more build inputs next to sse/, such as the vendored crate sources of the Rust tasks.

The rules for a task folder:

  • The folder name is the task ID. It is used for --task, image names and result paths, and the config's id must be equal to it. IDs are letters and digits separated by ., _ or -, up to 128 characters. The case image is case/<dataset>/<id> in lowercase, so two IDs may not differ only in case.
  • The Dockerfile builds on a pinned SSEBench base image. It declares ARG SSEBENCH_REGISTRY before the first FROM, and its last FROM is ${SSEBENCH_REGISTRY}/<base image>:<version>, for example ${SSEBENCH_REGISTRY}/base-generic-go:1.0.0. The base image needs a version tag or a digest; latest is rejected. See Base images.
  • The Dockerfile copies sse/ to /ssebench as shown above. Paths in the config are relative to /ssebench in the case image, and each one must be a file that the Dockerfile copies from the task folder.

Base images ​

The base images hold the toolchain of each language and are defined in images/base-images/: generic-c, generic-go and generic-rust. Their inputs are pinned: the upstream images by digest, and the ccache release that generic-c downloads by checksum.

A task builds on the base images of one SSEBench release, named by the release's version tag; every pilot task uses 1.0.0. The manifest records each task's base image, so jq -r '.tasks[].base' datasets/pilot/manifest.json | sort -u lists the images a dataset needs. Docker pulls them when it builds a case image. make -C images/base-images (or just base-images) builds them from your checkout instead: besides the current version and latest, it tags them with every version that a dataset manifest names, so that case builds use them.

Moving a task to the base images of a later release changes its Dockerfile, and so the dataset version.

Network access ​

Building a case image uses the network: the Dockerfile clones the upstream project at a pinned commit and installs its distribution packages, Go modules, crates and toolchains. Pin versions where the upstream project does not: the Rust tasks whose upstream project commits no Cargo.lock carry one next to their Dockerfile.

Grading does not use the network. Under the default egress policy (see --egress), a run container reaches only the LiteLLM proxy, so everything that build.sh, run.sh and test.sh need, with or without the reference patch and the hidden tests applied, must already be in the case image: vendored Go modules or the module cache, fetched crates, toolchains and test data. To check a task, run its scripts in the case image without a network, for example:

sh
docker run --rm --network none <case image> \
    bash -c 'cd <source> && /ssebench/scripts/build.sh && /ssebench/scripts/test.sh'

Task config ​

sse/config.yaml describes the task to the CLI, the daemon and the grader. Keys that are not listed here are rejected. For example:

yaml
id: gjson-196-bf4efcb
project: gjson
repository: https://github.com/tidwall/gjson
language: go
source: /src/gjson
task_description:
  crash_report:
  - reports/crash_report_1.txt
  - reports/issue.md
  bug_description:
scripts:
  build: scripts/build.sh
  run: scripts/run.sh
  test: scripts/test.sh
files:
  patch: diffs/patch.diff
  poc:
  - pocs/poc.go
  future_test: diffs/test.diff
  intent_test:
sanitizer:
type: Slice bounds out of range
binary:
trigger_commit: 9f58baa7a613f89dfdc764c39e47fd3a15606153
patch_commit: bf4efcb3c18d1825b2988603dea5909140a5302b
reference:
- https://github.com/tidwall/gjson
- https://github.com/tidwall/gjson/commit/bf4efcb3c18d1825b2988603dea5909140a5302b
- https://github.com/tidwall/gjson/issues/196
originality: public
KeyRequiredMeaning
idyesTask ID, equal to the folder name.
projectyesName of the upstream project.
repositoryyesURL of the upstream repository, https://….
languageyesLanguage of the project, lowercase: c, go or rust.
sourceyesAbsolute path of the project's source tree in the case image. The agent edits it, and the scripts build and test it.
task_descriptionyesWhat the agent is told; at least one of the keys below. The task prompt includes every one that is set.
task_description.issueIssue text.
task_description.crash_reportReport files, such as the upstream issue or a sanitizer log.
task_description.bug_descriptionShort description of the bug.
scriptsyesScripts that the grader and the agent's test_patch tool run.
scripts.buildBuilds the project.
scripts.runRuns the proof of concept given as its argument. It exits 0 when the vulnerability no longer triggers.
scripts.testRuns the project's tests.
filesyesReference material; none of it is shown to the agent.
files.patchThe upstream fix, as a diff against the source tree.
files.pocProof-of-concept inputs, each run with scripts.run.
files.future_testHidden tests of the fix, usually those added upstream, as a diff. The intent test applies it and runs scripts.test.
files.security_testA regression test for the vulnerability, as a diff. Kept for reference; the grader does not apply it. Leave it unset when it would repeat future_test, as in every pilot task.
files.intent_testTests of the intended behaviour, as a diff. Kept for reference; the intent test uses future_test. Leave it unset when it would repeat future_test, as in every pilot task.
sanitizerSanitizer the build enables, such as address.
typeClass of the bug, such as Heap Buffer Overflow or a CWE.
binaryProgram or harness that the proofs of concept exercise.
trigger_commitUpstream revision that the case image builds.
patch_commitUpstream commits that fix the bug, comma-separated.
referenceLinks to the advisory, the report and the fix.
originalityWhere the proofs of concept come from; see below.

originality is public when the proofs of concept come from the public report or advisory, possibly adapted to the task's harness, and crafted when they were written for the task because the public report has no reproducer. The Go tasks of pilot record it; the others leave it out.

Checks ​

The grader runs every check that the config makes possible, in this order; see Grading pipeline. The difficulty level only limits which of them the agent's test_patch tool may run.

CheckNeedsWhat it does
buildscripts.buildBuilds the patched project.
pocscripts.run and files.pocRuns every proof of concept. This is the security check.
function_testscripts.testRuns the project's existing tests.
intent_testscripts.test and files.future_testApplies the hidden tests and runs the tests again.

Dataset version ​

dataset.yaml holds the dataset's version, which the manifest records:

yaml
version: pilot-v1

Change it whenever tasks are added, removed or changed in a way that can change their results.

Manifest ​

manifest.json lists the tasks of one dataset version. It is generated from the task folders and committed next to them:

json
{
  "dataset": "pilot",
  "version": "pilot-v1",
  "tasks": [
    {
      "id": "gjson-196-bf4efcb",
      "language": "go",
      "project": "gjson",
      "repository": "https://github.com/tidwall/gjson",
      "base": "base-generic-go:1.0.0",
      "image": "case/pilot/gjson-196-bf4efcb",
      "arch": ["amd64"],
      "checks": ["build", "poc", "function_test", "intent_test"],
      "files": {
        "Dockerfile": "94e4db62…",
        "sse/config.yaml": "3c51886a…"
      },
      "metadata": { "id": "gjson-196-bf4efcb", "project": "gjson", "…": "…" }
    }
  ]
}
KeyMeaning
datasetDataset name, the name of its folder.
versionDataset version, from dataset.yaml.
generated_fromCommit of the SSEBench repository the manifest was generated from. Optional; the committed manifest leaves it out, and release builds can record it.
tasksOne entry per task, sorted by id.
tasks[].idTask ID.
tasks[].language, project, repositoryFrom the task config.
tasks[].baseBase image of the case image, relative to the registry, with its version: the last FROM of the task's Dockerfile without ${SSEBENCH_REGISTRY}/, such as base-generic-go:1.0.0.
tasks[].imageCase image, relative to the registry: case/<dataset>/<id>, lowercase.
tasks[].archPlatforms the task runs on, in order of preference. A run uses the host's architecture if it is listed, else the first; see Architectures. amd64 for every task today, since the case images are built and verified for amd64 only.
tasks[].checksThe checks the grader can run for the task.
tasks[].filesSHA-256 of every file in the task folder, by path relative to the folder.
tasks[].metadataThe task config, validated, with every key present.

Image names carry no registry and no tag. To pull a task's case image, prefix it with $SSEBENCH_REGISTRY and tag it with the dataset's version, for example ghcr.io/42-b3yond-6ug/ssebench/case/pilot/gjson-196-bf4efcb:pilot-v1. The pilot dataset page describes the published images.

Images lock ​

images.lock.json, next to manifest.json, pins the published case image of each task by digest. ssebench dataset lock writes it, and datasets/schema/images-lock.schema.json describes it:

json
{
  "dataset": "pilot",
  "version": "pilot-v1",
  "images": {
    "gjson-196-bf4efcb": {
      "digest": "sha256:9c4f…",
      "files_sha256": "5d1e…",
      "revision": "0123456789abcdef0123456789abcdef01234567"
    }
  }
}
KeyMeaning
dataset, versionThe dataset and version of the manifest that the lock goes with. A lock for another version is ignored.
imagesOne entry per published task, sorted by ID; a task without an entry is pulled by tag.
images[].digestDigest of the image in the registry. docker pull <registry>/<image>@<digest> pulls exactly that image, from any registry that holds a copy with its digests.
images[].files_sha256SHA-256 of the task's files in the manifest that the image was built from. The entry applies only while the manifest lists the same files.
images[].revisionCommit of the repository that the image was built and verified at.

A publishing run of the Dataset workflow produces the lock; see Releasing and versioning for how it gets into the repository and the release. A lock that lags behind a task that changed is harmless: ssebench run skips the entries whose files changed, and ssebench dataset lock drops them.

Commands ​

sh
uv run ssebench dataset validate [DIR]
uv run ssebench dataset manifest [DIR] [-o FILE] [--check] [--generated-from COMMIT]
uv run ssebench dataset schema [-o DIR] [--check]
uv run ssebench dataset lock [DIR] [--records DIR] [-o FILE] [--check]

DIR defaults to datasets/pilot; see the CLI reference.

  • validate checks every task folder against the rules above and lists every problem it finds, per task. It exits 1 if any task is invalid.
  • manifest validates the dataset and writes manifest.json into it. The output depends only on the task folders. Since the manifest records a checksum of every task file, regenerate it after changing anything in a task folder; --check fails, with a diff, when the committed file is out of date.
  • schema writes the JSON Schemas; --check fails when they are out of date.
  • lock writes the images lock: the images of the current lock whose task files are unchanged, and those of the publication records that ssebench dataset publish wrote. --check fails when the lock is missing, is for another dataset version or pins a task the dataset does not have.

The JSON Schemas also work with other validators, for example:

sh
check-jsonschema --schemafile datasets/schema/task.schema.json datasets/pilot/*/sse/config.yaml

They cannot express the rules that involve other files, such as the id matching the folder name or the config's paths existing; ssebench dataset validate checks those.

Catalog service ​

The catalog service (catalog/, Go) serves a manifest over HTTP. In each task it returns, base and image are prefixed with the registry, and image is tagged with the dataset's version, or named by its digest (<registry>/case/pilot/<id>@sha256:…) when the images lock next to the manifest pins the task to an image built from the files the manifest lists. A client can pull the image it gets as it is.

sh
go run ./catalog/cmd/ssebench-catalog serve --manifest datasets/pilot/manifest.json
OptionEnvironmentDefaultDescription
--manifest FILESSEBENCH_MANIFESTdatasets/pilot/manifest.jsonManifest to serve
--lock FILESSEBENCH_IMAGES_LOCKimages.lock.json next to the manifest, if there is oneImages lock that pins the case images by digest; a lock that was asked for must exist
--registry PREFIXSSEBENCH_REGISTRYghcr.io/42-b3yond-6ug/ssebenchRegistry prefix of the images
--port PORTSSEBENCH_CATALOG_PORT8080TCP port
EndpointReturns
GET /tasksEvery task, sorted by id, without files and metadata
GET /tasks/{id}One task
GET /tasks/{id}/metadataThe task config
GET /manifest.jsonThe manifest as loaded, with image names relative to the registry and without a tag

docker build -f catalog/Dockerfile ., from the repository root, builds an image that serves the committed datasets/pilot/manifest.json and images.lock.json.

The service is optional. The CLI and the web UI read a manifest directly from a file or URL, and default to the bundled datasets/pilot/manifest.json, so they list tasks without any server or network. To use a service, set SSEBENCH_CATALOG (or pass ssebench run --catalog) to its base URL, such as http://localhost:8080; they read its GET /manifest.json and prefix image names with their own $SSEBENCH_REGISTRY, and tag them with the dataset version. See Task catalog.

Released under the Apache License 2.0.