Fetching latest headlines…

Dev

Same score, 23 places apart: the normalization maths behind a hackathon judging platform

Dev.toUnited States · NORTH AMERICA

Two projects in the official DOGFOOD data set both had a raw average score of 3.44, from three judges each. After I normalized the scores, one finished #37. The other finished #14. Same number. 23 pla...

0 views0 likes0 comments

Two projects in the official DOGFOOD data set both had a raw average score of 3.44, from three judges each.

After I normalized the scores, one finished #37. The other finished #14.

Same number. 23 places apart. This is the story of why, and of the other technical problems I hit while building Verdict, a platform that runs a hackathon end to end: events, teams, deadline-enforced submissions, rubric judging, public voting, certificates, webhooks and an embeddable gallery. It was built for the DOGFOOD hackathon by @partnerships_raptors.

Stack: FastAPI, SQLAlchemy 2.0 async, PostgreSQL, Redis, Next.js. Repo: https://github.com/Lakshya-Varshney/Verdict

1. The problem: a raw average cannot tell who was asked

The obvious way to rank projects is to average judge scores. It silently assumes every judge uses the scale the same way.

They do not. On the official data (40 projects, 30 judges, 126 review rows), judges with at least 5 reviews had personal averages from 3.11 to 4.22, with the same spread (standard deviation 0.81 for both). That is a full point of pure leniency.

Now look at the two projects:

Project Raw mean (rank) Its reviewers' own average Normalized rank
Flat Meadow 3.44 (#24) 3.97 (generous panel: 4.22, 3.61, 4.08) #37
Glass Signal 3.44 (#26) 3.48 (close to the pooled mean 3.57) #14

Flat Meadow got a below-par verdict from each of its reviewers. Glass Signal got a normal one from each of its reviewers. The raw average sees two identical numbers.

2. The fix: per-judge z-scores

For each judge, compute their mean and standard deviation over everything they scored, then convert each score into "how far above or below this judge's own usual":

z(judge, project) = (score - judge_mean) / judge_stddev
normalized(project) = mean of z over the judges who reviewed it

Why z-scores and not min-max or rank averaging? They correct both leniency (the mean) and spread (the standard deviation), and they stay well-defined when projects have different numbers of reviews. Two reviews on one project and five on another is fine, because each project is just the mean over whoever reviewed it.

Result on the real data: 38 of 40 projects change rank, Spearman correlation 0.864 against the raw ranking. I also recomputed everything independently with plain statistics and no application code, and the two agree to floating-point precision.

3. The edge case that decides whether the maths is correct

A z-score divides by the judge's standard deviation. What if a judge gave 4/4/4 on every project? Or reviewed only one project? Then the standard deviation is zero and there is nothing to divide by.

The tempting shortcut is to use that judge's raw score. I rejected it, because it mixes two scales. Raw scores sit around 3.0. Z-scores sit around 0. If you average them together, every project a flat judge touched jumps by roughly that judge's raw score, and the judge with no opinion ends up as the loudest voice in the ranking.

The design that works: standardize the flat judge's raw score against the pooled mean and standard deviation of every score in the event.

if not flat:
    z = (score - mu) / sigma                 # judge's own scale
elif pooled_sd > 1e-12:
    z = (score - pooled_mu) / pooled_sd      # same scale as everyone else
else:
    z = 0.0

The flat judge still contributes a real signal (higher raw score means higher z), it lands on the same scale as everyone else, and it is never silently trusted: they appear in a zero_variance_judges list and in the audit log. The results payload also separates "identical scores everywhere" from "only one review", so the UI never mislabels a judge who simply had too little data. There is a regression test for exactly this: test_zero_variance_judge_does_not_swamp_z_scores.

On the official data, four judges take this path.

4. Real data fights back: the duplicate project row

The official fixtures.json has 41 project rows for 40 teams. prj_41 repeats prj_07: same team, same title ("Dry Harbour"), same repo, different id. Deduplicating by id would never catch it, so I merge by team plus title.

The interesting part is the consequence. Three judges had scored both ids, so their criterion rows collided after the merge: 9 rows, taking 378 score rows down to 369. One of those judges had scored the original at 2.33 and the duplicate at 3.67. Under the documented rule (the later review wins), all three of her remaining projects score 3.67, and she correctly lands in the zero-variance path from section 3.

Data cleaning is not neutral. It changes the maths downstream. So I wrote the rule down (later review wins, averaging is the alternative) and surfaced every flagged judge in the results payload.

5. Role isolation, tested by attacking the running system

Reading the code tells you what you intended. Probing the live API tells you what you built. I ran adversarial probes against the running system and found 12 real problems, each now covered by a regression test. The ones that taught me the most:

  • An organizer could grant themselves admin. Every permission check accepts admin from any event, so one self-grant meant site-wide control. Only an admin can grant admin now.
  • Any organizer could read every other event's audit log. 98 of the first 100 rows returned belonged to other events. Audit access is scoped to the caller's own events.
  • Judges could change scores after results were published. The judging window was enforced in the UI, not the API. A UI button is not an access control.
  • The vote count was hidden in one field and leaked in another. While voting was open count was null, but vote_count still carried the real number.
  • [email protected] counted as a new voter. Email voting now uses a canonical mailbox: lowercase, drop the +tag, ignore Gmail dots.

Roles are read from the database on every request and scoped per event, so cross-event isolation is structural rather than a convention.

6. A 13 second gallery, and lazy="raise"

With the real 40-project data the gallery took 13 s and /roles took 18 s, because SQLAlchemy was eager-loading every relationship.

I set lazy="raise" on every relationship. Any accidental lazy load now throws instead of quietly running extra queries, so every query has to be written explicitly. It forces more typing and gives predictable performance. test_real_scale_endpoints_stay_fast keeps it that way.

7. An audit log you can verify, not just trust

An append-only audit log is normally append-only by convention. Anyone with database access can still edit history.

In Verdict, every audit row contains the hash of the previous row. GET /admin/audit/verify and an offline script both detect tampering. To keep the chain linear, appends take a Postgres advisory lock (pg_advisory_xact_lock), so two concurrent transactions can never read the same tip hash and fork the history.

The same principle runs through the certificates: Ed25519 signatures instead of HMAC, so anyone can verify a certificate offline with only a public key, with no shared secret. Webhooks use a transactional outbox, so events cannot be lost or invented, and no message broker is needed to run offline.

8. The scope decision: normalized absolute scores

Judging can be built two ways: judges score each project against a rubric, or judges compare two projects at a time (pairwise, Bradley-Terry style). I shipped normalized absolute scoring.

It works directly with the weighted rubric, the result is recomputed from raw scores on demand (never trusted from storage), and I could check it independently against the real data. That made the correctness work possible: the normalization proof, the role-isolation tests and the verifiable audit trail.

What I would carry into the next build

  1. Test the maths on messy real data, not a clean example. My clean example proved the idea. The duplicate row and the flat judges decided the actual design.
  2. Attack the running system early. Role-isolation bugs hide in code that reads correctly.
  3. Make failures loud. lazy="raise" and server-side checks turn "slow and silently wrong" into "fails immediately".

The full normalization proof, the threat model and all the tests are in the repo: https://github.com/Lakshya-Varshney/Verdict

Built for DOGFOOD. Tagging @partnerships_raptors for the Write Up Quest.

Comments (0)

Sign in to join the discussion

Be the first to comment!