idrid-dr-grading
IDRiD — diabetic retinopathy severity grading (demo)
Predictions submitted for this task are scored against reference labels held by the platform. The labels are never published and are not part of any response — only the metrics below come back.
About this task
What is being scored, on which dataset, and where the referable cut-off sits.
- Slug
- idrid-dr-grading
- Task kind
- Grading
- Scored dataset
- idrid-grading-demo
- Grades
- 5 classes (grade 0 to 4)
- Referable threshold
- 2 — referable means grade ≥ 2; anything below that counts as non-referable.
- Items
- 30 — listed in full below; predictions are keyed on these identifiers.
- Created
Item identifiers
Every item this task scores against. A predictions file is a map keyed on these identifiers, and a sealed container receives the same set at run time as index.json on its /input mount. They are identifiers only — the reference labels behind them are held by the platform and are not part of any response.
- IDRiD_001
- IDRiD_002
- IDRiD_003
- IDRiD_004
- IDRiD_005
- IDRiD_006
- IDRiD_007
- IDRiD_008
- IDRiD_009
- IDRiD_010
- IDRiD_011
- IDRiD_012
- IDRiD_013
- IDRiD_014
- IDRiD_015
- IDRiD_016
- IDRiD_018
- IDRiD_029
- IDRiD_030
- IDRiD_032
- IDRiD_037
- IDRiD_038
- IDRiD_039
- IDRiD_041
- IDRiD_043
- IDRiD_063
- IDRiD_073
- IDRiD_074
- IDRiD_085
- IDRiD_101
Read this list rather than generating it — the identifiers are not guaranteed to be contiguous or densely numbered. Items you omit are permitted and reported as reduced coverage; identifiers this task does not recognise are a validation failure, so a mismatched naming convention fails loudly instead of scoring zero.
Metrics are per task, not a global ranking
Results
Ordered best first by QWK — quadratic-weighted kappa, the headline metric: agreement with the reference grades on a −1 to 1 scale where 1 is perfect and 0 is chance. The referable rates are measured against the grade ≥ 2 cut-off. Coverage is the share of the reference set a submission actually predicted. Results are provisional until the evaluation method that produced them is approved.
demo-baseline-v1
Evaluation method: legacy
scoredScored before the evaluation-route registry existed (ADR-0018 / WP5). This result carries no route declaration and is excluded from published reporting.
- QWK
- 0.908
- Accuracy
- 70%
- Referable sensitivity
- 88.9%
- Referable specificity
- 83.3%
- Coverage
- 100%
1 of 1 submissions scored · 0 published. Back to evaluation tasks.