Evaluation methods

oci-sealed-execution

OCI sealed execution (reference)

Sealed container run next to the data. Every result on the Open Code Infrastructure (OCI) names the evaluation method and version that produced it, and a method version is reviewed before its results are published. The declarations below are what that review examines.

Sealed containerReference implementation

About this method

Who provides it, how a submission runs, and who the four parties named in the declarations are.

Slug
oci-sealed-execution
How it runs
Sealed container run next to the data
Provider
Reference implementation
Versions
1 — shown below.
Registered
Parties
  • data host — the institution that holds the dataset and its reference labels
  • model developer — the participant whose model is being evaluated
  • platform operator — the team that runs the Open Code Infrastructure
  • method provider — the party that built and operates the evaluation method

Version 1.0.0

provisional

Provisional: the evaluation method that produced this result has not yet passed review, so it is excluded from published reporting. Declared ; not yet reviewed.

Threat model

Who might try to learn something they should not, what this method does about each of them, what it takes for granted, and what it makes no promise about.

Adversaries

PartyCapabilityDefended
model developersubmits arbitrary code as a container image that executes against host datayes
platform operatorreads worker logs and can inspect the hostno
data hostholds the data and the ground truth in the clearno

A party marked no is named so nobody mistakes it for a guarantee: this method does not defend against it.

Assumptions

  • Container isolation as configured holds: --network none, cap-drop ALL, no-new-privileges, read-only root
  • The registry digest pinned at dispatch is the image that runs
  • The host kernel is not compromised

What has to hold for the defences above to work. If an assumption fails, so does the guarantee that rests on it.

Out of scope

A threat model with nothing out of scope is rejected on entry; naming the boundaries is the point.

  • A malicious or compromised data host — it already holds the data in the clear
  • A malicious platform operator — operator log access is deliberately not defended against
  • Side channels measurable from within the container, such as timing or resource observation
  • Statistical inference about the dataset from legitimately returned metrics

Disclosure profile

Who gets to see what while an evaluation runs, what that guarantee ultimately rests on, and whether a result can be reproduced.

Trust anchor
The guarantees rest on a contract — an operating agreement between the parties, not a technical mechanism.
Key governance
No encryption keys. Isolation is enforced by runtime configuration and digest pinning; the trust anchor is the operating agreement with the data host rather than an attestation root.
Reproducible
Yes. Re-run the pinned image digest against the same task; scoring is deterministic given identical predictions.

Who observes what

Everything each party can see during an evaluation, including what it already holds.

model developer
only the classified failure code and its own metrics; never stdout, stderr, or any host data
data host
everything — it owns the data and the ground truth
platform operator
worker logs including container stdout/stderr, which are never returned to the participant
method provider
nothing beyond the operator view; the reference route has no separate provider

Operational envelope

The limits a submission must be designed to run within. The runtime and memory caps are enforced by the sandbox, not merely documented. Memory is given in mebibytes (MiB) and gibibytes (GiB).

Permitted operations
  • container inference over the mounted /input
  • write predictions to /output
Arithmetic precision
unconstrained — plaintext execution, participant chooses
Maximum runtime
1 h (3,600 seconds)
Maximum memory
16 GiB (16,384 MiB)
Model constraints
Any architecture that runs offline in the sandbox: no network, read-only root, /output the only writable path and size-capped, non-root user, all capabilities dropped.
Fidelity gap
Not yet measured. The fidelity gap is the difference between a score produced through this method and the same model scored in the clear; it is reported once measured.

1 version declared. Back to evaluation methods.