Evaluation

idrid-dr-grading

IDRiD — diabetic retinopathy severity grading (demo)

Predictions submitted for this task are scored against reference labels held by the platform. The labels are never published and are not part of any response — only the metrics below come back.

Grading1 submission

About this task

What is being scored, on which dataset, and where the referable cut-off sits.

Slug
idrid-dr-grading
Task kind
Grading
Scored dataset
idrid-grading-demo
Grades
5 classes (grade 0 to 4)
Referable threshold
2 — referable means grade ≥ 2; anything below that counts as non-referable.
Items
30 — listed in full below; predictions are keyed on these identifiers.
Created

Item identifiers

Every item this task scores against. A predictions file is a map keyed on these identifiers, and a sealed container receives the same set at run time as index.json on its /input mount. They are identifiers only — the reference labels behind them are held by the platform and are not part of any response.

  • IDRiD_001
  • IDRiD_002
  • IDRiD_003
  • IDRiD_004
  • IDRiD_005
  • IDRiD_006
  • IDRiD_007
  • IDRiD_008
  • IDRiD_009
  • IDRiD_010
  • IDRiD_011
  • IDRiD_012
  • IDRiD_013
  • IDRiD_014
  • IDRiD_015
  • IDRiD_016
  • IDRiD_018
  • IDRiD_029
  • IDRiD_030
  • IDRiD_032
  • IDRiD_037
  • IDRiD_038
  • IDRiD_039
  • IDRiD_041
  • IDRiD_043
  • IDRiD_063
  • IDRiD_073
  • IDRiD_074
  • IDRiD_085
  • IDRiD_101

Read this list rather than generating it — the identifiers are not guaranteed to be contiguous or densely numbered. Items you omit are permitted and reported as reduced coverage; identifiers this task does not recognise are a validation failure, so a mismatched naming convention fails loudly instead of scoring zero.

Results

Ordered best first by QWK — quadratic-weighted kappa, the headline metric: agreement with the reference grades on a −1 to 1 scale where 1 is perfect and 0 is chance. The referable rates are measured against the grade ≥ 2 cut-off. Coverage is the share of the reference set a submission actually predicted. Results are provisional until the evaluation method that produced them is approved.

  1. demo-baseline-v1

    Evaluation method: legacy

    scored

    Scored before the evaluation-route registry existed (ADR-0018 / WP5). This result carries no route declaration and is excluded from published reporting.

    QWK
    0.908
    Accuracy
    70%
    Referable sensitivity
    88.9%
    Referable specificity
    83.3%
    Coverage
    100%

1 of 1 submissions scored · 0 published. Back to evaluation tasks.