Evaluation

Tasks

1 task

These are benchmarking tasks. A model’s predictions are scored against ground truth held by the platform — the reference labels are never published, and no submission can read them. Only the resulting metrics come back.