oci-predictions-scoring
OCI predictions scoring (reference)
Predictions file scored on the platform. Every result on the Open Code Infrastructure (OCI) names the evaluation method and version that produced it, and a method version is reviewed before its results are published. The declarations below are what that review examines.
About this method
Who provides it, how a submission runs, and who the four parties named in the declarations are.
- Slug
- oci-predictions-scoring
- How it runs
- Predictions file scored on the platform
- Provider
- Reference implementation
- Versions
- 1 — shown below.
- Registered
- Parties
- data host — the institution that holds the dataset and its reference labels
- model developer — the participant whose model is being evaluated
- platform operator — the team that runs the Open Code Infrastructure
- method provider — the party that built and operates the evaluation method
Declarations are frozen once review begins
Version 1.0.0
provisionalProvisional: the evaluation method that produced this result has not yet passed review, so it is excluded from published reporting. Declared ; not yet reviewed.
Threat model
Who might try to learn something they should not, what this method does about each of them, what it takes for granted, and what it makes no promise about.
Adversaries
| Party | Capability | Defended |
|---|---|---|
| model developer | submits predictions and may probe the ground truth by repeated scored submissions | yes |
| platform operator | holds the ground truth and reads all logs | no |
A party marked no is named so nobody mistakes it for a guarantee: this method does not defend against it.
Assumptions
- No participant code executes: a predictions file is data, not a program
- Ground truth never leaves the OCI and is never returned in any response
- Scored-submission quotas bound how much a participant can learn by repetition
What has to hold for the defences above to work. If an assumption fails, so does the guarantee that rests on it.
Out of scope
A threat model with nothing out of scope is rejected on entry; naming the boundaries is the point.
- A malicious platform operator, who holds the ground truth by construction
- Statistical inference about the answer key from the participant's own returned metrics
- Anything about how the participant produced the predictions — the model itself is never seen
Disclosure profile
Who gets to see what while an evaluation runs, what that guarantee ultimately rests on, and whether a result can be reproduced.
- Trust anchor
- The guarantees rest on a contract — an operating agreement between the parties, not a technical mechanism.
- Key governance
- No keys. The ground truth is held server-side and the guarantee is that it is never returned — enforced by the read boundary, not by encryption.
- Reproducible
- Yes. Scoring is a pure function of the predictions and the pinned ground truth; re-scoring the same file yields identical metrics.
Who observes what
Everything each party can see during an evaluation, including what it already holds.
- model developer
- its own metrics only; never the ground truth, never another participant's predictions
- data host
- everything it already owns
- platform operator
- predictions, ground truth and scores
- method provider
- nothing; the reference route has no separate provider
Operational envelope
The limits a submission must be designed to run within. The runtime and memory caps are enforced by the sandbox, not merely documented. Memory is given in mebibytes (MiB) and gibibytes (GiB).
- Permitted operations
- submit a predictions map keyed on the task's published item identifiers
- Arithmetic precision
- exact — integer labels
- Maximum runtime
- 1 min (60 seconds)
- Maximum memory
- 512 MiB
- Model constraints
- None. No participant code runs, so no architecture constraint applies; the constraint is the quota on scored submissions.
- Fidelity gap
- Not yet measured. The fidelity gap is the difference between a score produced through this method and the same model scored in the clear; it is reported once measured.
1 version declared. Back to evaluation methods.