←  work / case-studies
case study · academic submission

The screen where a student agrees to lose marks.

An AI evaluation platform for professional education. My part was the student's side of it — everything between opening an assignment and the moment a submission becomes a grade, including the one screen where they decide to submit anyway.

open the live product ↗ the real build — signing in is one click, and the panel at the bottom-right runs the blocked-link and unreadable-scan paths on demand
four-week client engagement designed and built in code team of six · I owned the build

01The problem

A student thinks the job is upload → submit → done. The system needs complete → readable → evaluable before grading can start at all. Nearly every platform reports that gap after submission, as a failure message.

By then it is not information, it is a verdict. The deadline has passed, the file is with the instructor, and the marks for whatever was missing are already gone. The student's first accurate picture of their own submission arrives at the exact moment they can no longer act on it.

Four separate gaps sit inside that, and they fail differently:

The client's own priority was not student experience — it was instructor load. Every escalation, every "I didn't know I needed that", every grade dispute lands on faculty. That reframed the work: nothing here earns its place unless it also reduces what reaches the instructor.

02The constraint that shaped everything

Most clients hide infrastructure cost from designers. This one put it on the table in week one: reading a document with a model is expensive, and at institutional scale the difference between checking something instantly and checking it in a queue is the difference between a business and a science project.

That is a design constraint wearing an engineering costume. It decides what a student can be told and when — so it decides the shape of every screen after upload. A product that promises real-time everything either lies or bankrupts the person paying for it.

So validation runs in three tiers, and the interface is honest about which one you are in:

The split is not a compromise forced by cost. It is the correct division anyway: the fast checks are the ones you can act on standing at your laptop, and the slow one is the one you want to walk away from.

the two non-negotiables

Stated as architecture, not preference

The client fixed two rules on day one: no evidence, no credit — the system only awards for what it can actually find — and the instructor is the final authority. Stating them as constraints rather than opinions removed an entire class of design argument before it could start.

The consequence runs through every screen: the AI never scores anything. It reads, and it reports what it found and where. Every post-submission screen carries the same sentence — your instructor decides your final grade — because in a market where the adoption risk is "a machine is marking my work", trust is built by repetition, not by a tooltip somebody dismissed once.

03The thread, screen by screen

All of this is the running build, not a mockup. The countdowns are live, the data is seeded, and you can open any of it in the demo.

The student dashboard. One assignment card carries the state, the deadline countdown, the separate early-feedback countdown, and three remaining feedback rounds shown as segments.
01 · dashboardOne card, one assignment, one action. Two countdowns sit side by side because they are different deadlines and only one of them can still help you — the assignment is due in 1d 8h, the feedback window shuts in 1d 2h. The three segments are the rounds still available. Nothing here needs reading; it needs glancing at.

The nudge under the card is the whole incentive design in one line: submit early to unlock feedback. Early submission is not framed as diligence, it is framed as the only way to get something the late submitter cannot have.

Assignment detail. The early-feedback deadline sits in its own card beside the brief, and the content is split across Assignment Brief, What to Submit, and Grading Rubric tabs.
02 · the briefThe rubric is a tab, at the top, before the upload zone — because the single most common way to lose marks is not knowing what was being marked. The early-feedback card is deliberately louder than the deadline: it is the one the student can still do something about.
Upload screen after a successful check. Ticks for upload, format, size, page count, readable text, detected handwriting, and each of the two links found in the document verified separately.
03 · the fast checksT1 and T2, itemised as they land. It reads the document rather than the file: twelve pages, text readable throughout, handwriting on pages 8–9, and the two links found inside it opened and checked one at a time. A student who has been told "upload failed" all their life can see exactly what passed.

Showing the process here is right, and showing it during the rubric analysis would be wrong — the same mechanism, opposite call. On this screen every line is something the student can act on. During analysis they can only wait, and a list of checks ticking past is just a countdown to a verdict.

The same screen with one link failing. Google Drive report is marked permission blocked, with an inline instruction giving the exact menu path to fix it and a Check again action.
04 · the blocked linkThe failure that costs the most marks and is the least visible to the person who caused it. Every error on this screen names the thing, states the cause, and gives the actual fix — Open Google Drive → Share → set to "Anyone with the link can view" — then offers to re-check. An error that only says what went wrong is half an error.
Analysis report. Problems are listed first and expanded, matches are collapsed below, and a rubric coverage ring reads sixty-eight per cent in amber.
05 · the reportProblems first, expanded; the things that passed are collapsed underneath. The missing criterion carries a fix first flag, what to add, and where the gap is — "Section 3, pages 6–8". The counter at the foot says two attempts left, which is the difference between a report and a verdict.
The consent dialog. Three criteria are listed by name with their exact weight and the consequence of proceeding, tagged high and medium risk, above a warning-coloured Submit anyway button.
06 · the consent momentThe screen this case study is about. Each criterion is named, carries its exact weight, and states what proceeding costs. The reminder that there is still time sits between the list and the buttons. "Go back & improve" is the quiet one; "Submit anyway" is the loud one — deliberately, because the weight should sit on the path that cannot be undone.

04The decision I'd defend

Every platform designs the success state and the failure state. Almost nobody designs the one in between, which is the only one that is actually hard.

the problem

The warning state is the dangerous one

Failures are easy: the system blocks, the student fixes, the system unblocks. Nobody has to decide anything.

A warning asks a nineteen-year-old to make an irreversible decision, under deadline pressure, on incomplete information. Read it as a blocker and they abandon a submission that was fine. Read it as noise and they lose marks they never understood they were agreeing to lose. Both readings are reasonable responses to “some issues were found.”

And it is the state that generates the dispute. Not the failure — the failure is unambiguous. The dispute comes from the student who proceeded without understanding what proceeding meant.

the reframe

Consent has to be specific enough to be indefensible to argue with

A generic “I understand” is worth nothing to either side. It does not tell the student what they are giving up, and it does not give the institution anything to point at when the grade is challenged.

So the dialog does not summarise. It names each criterion, prints its exact weight, and states the consequence in the student's own currencywill not receive evidence credit for this criterion — 10% of your grade. Three named risks, not "3 issues".

Then it does the thing that makes it fair rather than merely defensible: it says you still have 1d 2h, and resubmitting cannot lower your grade. Informed consent is only real when refusing is genuinely available.

Submission sent screen with a five-minute countdown reading four minutes fifty-eight, a draining progress bar, and the line: your instructor hasn't seen this yet.
07 · the five minutes afterThe counterpart to consent. One sentence does most of the work — your instructor hasn't seen this yet — because the panic in that moment is social, not technical. Five minutes, visibly draining, then it locks. A safety net that stayed open indefinitely would just move the deadline.

The argument for spending this much design on one dialog only holds if it pays off later, so here is where it lands. Two weeks on, the same student is looking at a grade.

The grade result. A rubric table gives a score, weight and weighted contribution for each of six criteria, with the instructor's written feedback split into strengths and areas for improvement.
08 · the gradeNot a number — an arithmetic the student can follow. Every criterion shows its score, its weight, and what it actually contributed, so 7.1 is a sum rather than a judgement. The instructor's own words sit underneath, split into what worked and what did not. This is the screen that decides whether the student trusts the platform on the next assignment.
The re-evaluation request. Criteria are listed lowest score first with checkboxes, beside a panel that shows the instructor's feedback for whichever criteria are selected.
09 · disagreeing wellIt is called request re-evaluation, never dispute. Criteria are ordered lowest first, because that is what the student came for. Selecting one shows the instructor's reasoning next to the box you are typing in — so the request is a response to an argument rather than a complaint about a number. One request per assignment, seven-day window.

That is the return on the consent screen. A student who agreed to a specific, named, weighted risk is a student who can be shown that record, and an instructor facing a challenge has something better than memory. The dialog is not there to protect the institution from the student — it is there so the conversation two weeks later is about the work.

05Two comparisons the build makes better than I can

Both pairs are the same screen in two real states of the shipped build — nothing here is reconstructed or mocked up. Drag anywhere on the image to compare.

Before: the analysis report on the first attempt, sixty-eight per cent coverage, one criterion missing.
After: the same report on the third attempt, one hundred per cent coverage, two criteria resolved.
attempt 1 attempt 3 drag anywhere
The entire argument for the product, in one screen twice. Left: 68% coverage, a criterion scored zero, similarity above the limit, "1 rubric area missing" in red. Right: the same student after two rounds — 100%, and a since V2 card that names what changed and does not make them hunt for it. The loop is the value; everything else is plumbing that makes the loop possible before the deadline.
Before: the dashboard in the building tier, with a coaching prompt and momentum framing.
After: the identical dashboard in the established tier, with a stretch prompt and confident framing.
building established drag anywhere
The same dashboard for two students. Coaching prompt becomes stretch prompt; "keep the momentum going" becomes "your edge is showing"; the nudge shifts from unlocking feedback to unlocking instructor review of stretch work. Note what does not change: the layout is identical, and the lower tier is never labelled or shown a comparison. Tiering that the weaker student can see is just ranking with a friendlier name.

The tiers are driven by behaviour, not grades — whether you submit early and use your rounds, not what you scored. A 65 who iterates is a more useful signal, and a more actionable one, than an 85 who submits at 11:58pm. Three tiers, not five: more granularity produces more anxiety and no more product.

06Three times the work had to be redone

The demo where the client agrees with everything is the one that taught you nothing.

scope

A third of week one did not survive the first client demo

We presented the prototype as a walk through one student's journey rather than a tour of features — lead with the anxious first-timer, then break the story with the students who behave differently. It worked, in the sense that it produced redirection instead of polite agreement.

He came back with three changes: support a single-document submission alongside the multi-file model, because faculty load makes consolidation matter more than structure; split validation into fast and queued rather than promising real-time; and add incentive layers to pull submissions earlier. About a third of what we had built stopped being relevant that afternoon.

What held was the four principles underneath it. We changed the surfaces without relitigating the philosophy, which is the only reason a week-two scope expansion did not become a spiral.

over-promised

I designed three features I had no authority to promise

A pre-submission validation preview, server-side storage of the consent record, and preserved validation state on resubmission so a student replacing one element does not re-run everything. All three make the product meaningfully better. All three depend on architecture decisions that were not mine to make.

The preview was the worst of them — I had it in the prototype before I understood what running the expensive tier speculatively would cost. Designing around infrastructure you have not asked about is how you produce a beautiful thing that cannot ship.

They ended up documented as dependencies rather than designed around, which is the honest outcome but not a satisfying one. The right move was a feasibility check in week one, not a discovery in week three.

reverted

I nearly let the design system get talked out of me

I built this alongside an AI coding agent, and it is very good at producing something plausible and very willing to drift. Mid-build it proposed flattening the card system to hairline dividers — cleaner, more current, genuinely defensible in isolation.

I took it, looked at it, and put it back. The cards are not decoration: they are what makes a validation result read as one object with a status, on a screen where six of them stack. Flat dividers turn six statuses into one long list.

The general version, which I now write down at the start of a build: a suggestion that is defensible in isolation can still be a regression, because the thing it breaks is somewhere else on the screen. Two more went the same way — a carousel flattened into a static list, and a status pill quietly swapped for a component-library default.

07What this is not

What I think is genuinely good here: the consent moment, which I would defend in any room; the tier split falling out of a cost constraint rather than being bolted onto it; and the decision to spend the design effort on the warning state instead of the success state.

What honestly limits it:

One housekeeping note, since the screenshots are the evidence: the institution branding in them is a neutral placeholder. The real build carries a named university as demo data, and putting that on my own site would imply a customer relationship that does not exist.

08What this rests on

The competitive audit behind the design covered eleven public platforms, read for three things — how grading is made transparent, how scanned and handwritten input is handled, and the gap between a system validating something and a student understanding it:

  1. 1Gradescope · Turnitin · Canvas LMS · Google Classroom · Blackboard — grading, submission and feedback surfaces
  2. 2Google Cloud Vision · Azure AI Vision · Amazon Textract — OCR behaviour, confidence reporting and failure modes on handwriting
  3. 3ETS · FICO · IBM ODM — how automated scoring and rules engines explain a decision to the person it lands on

The build carries its own reasoning: a decision log recording every design and code call with its rationale, a defence document mapping each screen to the behaviour it is trying to produce, and a design system written down so the thing stays coherent across a four-week sprint. Those live in the repository, not in a deck.

The platform is named because it is publicly deployed under that name; the founder is not named here. Screens are the running build, captured at 2× — click any of them to read one at full size.

←  back to selected work talk about this one →