Sarthak Sahoo

Name Sarthak Sahoo Role Product Designer · Austin, TX

10 capabilities

Skills

  • Product Design
  • Interaction Design
  • UX Research
  • Design Systems
  • Rapid Prototyping
  • Accessibility / WCAG
  • Information Architecture
  • AI-Assisted Design
  • Front-End Collaboration
  • Complex Enterprise UX
  • Illustration
  • Visual Taste
  • Heuristic Evaluation
  • User Testing
  • Wireframing
  • Animation
Schedule a call
Illustration

Hello Team :) I’m Sarthak Sahoo, a Product Designer  based in Austin, TX I’ve designed human-centered products in an AI-saturated world.

View Work

Worked with

  • CVS Health
  • CG
  • PNC
  • First United

Key industry projects

01 / 05

Designing governance infrastructure for agent-generated interfaces

I’m currently building governance standards for catching “AI slop,” designed to be dropped directly into an agent chat to audit generated product work for weak logic, missing states, unsupported decisions, and other gaps before it ships. Download my guide

A hand-drawn diagram, “Creating governance for AI design decisions — a rule-based system to reduce AI slop”, defined underneath as polished looking work without evidence, context or real product logic. Five numbered panels. One, Source: a PDF holding rules, guidelines, checklists and examples. Two, /slop-audit turns it into a skill that uses the PDF as its source, applies the rules, reviews design and copy, flags issues and gives feedback. Three, Use it: run /slop-audit with pasted design, copy, or a link to the PDF. Four, What it looks for, in two columns — design red flags: generic UI, missing states, inconsistent components, random spacing and colours, no user logic, looks polished but shallow; copy red flags: unsupported claims, vague or generic language, fake confidence, invented metrics, no source or context, sounds correct but isn’t. Five, Output: audit results listing red flags with reasons, suggestions, what to fix, and higher quality work that’s ready to ship. A bracket across the foot reads “less AI slop leads to better products”.

off the clock

Three things I keep coming back to

At heart, I’m a creative who likes finding depth and clarity inside complex, chaotic environments. A lot of my identity comes from music: house, indie, Sufi ghazals, rock, and old Bollywood all mixed together. For a long time, I thought being both Indian and American meant finding some kind of middle ground. Now I see it differently. The combination is part of what gives me perspective.

Art has always worked the same way for me. As a kid, I drew obsessively, chasing repetition, depth, emotion, and clarity inside complexity. Those are the same qualities I listen for when I make music, and the same ones I bring into digital products. I look for patterns, how deep a system goes, where it starts to break, and what that reveals about the people and businesses relying on it.

I’ve also stopped thinking of design as finding a compromise between users and businesses. When customers and employees can move through systems with less friction, the business usually benefits too. That’s shaped how I see design now: not just as something visual, but as something systemic.

View Case Study

Where I’ve been

  1. 2022 CVS Health Junior UI Developer Healthcare
  2. 2023–24 CG Infinity UX Consultant Multiple clients
  3. 2024–25 PNC Bank Product Designer Finance
  4. 2026 Natera Product / UX Designer Biotechnology
View résumé

Tools

Human judgment, design decisions, AI execution.

off the clock

Three things I keep coming back to

At heart, I’m a creative who likes finding depth and clarity inside complex, chaotic environments. A lot of my identity comes from music: house, indie, Sufi ghazals, rock, and old Bollywood all mixed together. For a long time, I thought being both Indian and American meant finding some kind of middle ground. Now I see it differently. The combination is part of what gives me perspective.

Art has always worked the same way for me. As a kid, I drew obsessively, chasing repetition, depth, emotion, and clarity inside complexity. Those are the same qualities I listen for when I make music, and the same ones I bring into digital products. I look for patterns, how deep a system goes, where it starts to break, and what that reveals about the people and businesses relying on it.

I’ve also stopped thinking of design as finding a compromise between users and businesses. When customers and employees can move through systems with less friction, the business usually benefits too. That’s shaped how I see design now: not just as something visual, but as something systemic.

how i view systems

Design saves lives.

Safety, income and stability are at the forefront of how we operate our daily lives. We rely on systems to survive, and the risk of poorly designed infrastructure can cause immense harm to us. It’s sad, but also interestingly optimistic, that a lot of these problems are preventable. I find parts of those experiences and identify what breaks in them. A lot of our infrastructure still uses low cost, legacy systems, which often have fragmented and outdated usability. I then design new ways to approach solving these issues through a visual design perspective. Where does the flow break for the user? Read more

key industry projects

Healthcare · Biotechnology · Financial Services · Developer Tools · Education

everything else

Supporting work · background · process

off the clock

Three things I keep coming back to

At heart, I’m a creative who likes finding depth and clarity inside complex, chaotic environments. A lot of my identity comes from music: house, indie, Sufi ghazals, rock, and old Bollywood all mixed together. For a long time, I thought being both Indian and American meant finding some kind of middle ground. Now I see it differently. The combination is part of what gives me perspective.

Art has always worked the same way for me. As a kid, I drew obsessively, chasing repetition, depth, emotion, and clarity inside complexity. Those are the same qualities I listen for when I make music, and the same ones I bring into digital products. I look for patterns, how deep a system goes, where it starts to break, and what that reveals about the people and businesses relying on it.

I’ve also stopped thinking of design as finding a compromise between users and businesses. When customers and employees can move through systems with less friction, the business usually benefits too. That’s shaped how I see design now: not just as something visual, but as something systemic.

have a complicated problem?

Let’s make it easier to understand.

Case study

Improving patient recall after oncology visits

I designed a mobile-only system that connects what patients want to ask before an appointment with what they need to remember afterward.

Role
Product designer — patient flow, states, visual language
Team
Three designers, inside a larger healthcare technology client
Timeline
3 weeks leading into the design workshop
Starting point
Stakeholder direction → design exploration → workshop → MVP development

≈15–25%

$ less post-visit clarification work

How this is calculated

What it represents Repeat contact a clinic handles after a visit — calls and messages asking again for something already said in the room.

Inputs from this project Testing where 85–90% of people found what the doctor said in a saved conversation, and 80–90% understood the answer came from the recording rather than from a model.

The logic A patient who can retrieve the answer does not ring to ask for it. So the effect is the share of that repeat contact the product could absorb, at whatever a clinic's contact volume and handling time are. Carry never shipped, so this is the case the design makes, not a measured reduction.

≈20–35%

faster entry into question-building

≈85–90%

found what the doctor said, unaided

What Carry is Three weeks of design, a workshop, then a build that stopped

Carry started as a premise from a workshop. I took it to a scoped MVP, it went through stakeholder review, and a first build started in Playground. A leadership change stopped it before release.

Three weeks to turn stakeholder direction into something the wider team could review, challenge and build from.

  1. Stakeholder direction
  2. V1 · Week 1
  3. V2 · Week 2
  4. V3 · Week 3
  5. Design workshop
  6. Cross-functional feedback
  7. Revisions
  8. MVP development

V1, V2 and V3 are meaningful revisions to the design during the three-week exploration.

What I contributed

  • Turned the premise into an end-to-end patient flow, from writing questions before a visit to reading the conversation back after it.
  • Ruled out generated answers, and designed the quotation pattern that took their place.
  • Designed the upload, transcription and matching states, and what each one does when it fails.
  • Took tradeoffs to product and engineering while they were still drawings.
Where it sits in PatientX Patients lose the fifteen minutes, not one forgotten fact

Carry is a small piece of PatientX, which already owned sign-in, appointments and navigation. The first job was finding somewhere Carry could sit without changing any of it.

Where Carry lives

PatientX — the patient experience Carry plugs into. Already built, and outside my scope.

  • Sign-in and identity
  • Appointments
  • Navigation

The visit

Carry — the part my team designed

  1. BeforePrepare the questions
  2. DuringCarry does not operate
  3. AfterUpload and match
  4. LaterRead it back

Proto personas

Built from onsite and Maze pilot behaviour and the feedback the team held.

Single-line drawing of Maya Patel, looking slightly away from the viewer.

Maya Patel, 42

Breast cancer treatment follow-up

Situation
Follow-up after treatment. Visits at set intervals, and part of what each one decides is whether the interval changes. Maya arrives with one specific concern — the pilot pattern where patients go straight to writing questions.
What they carry into the appointment
A result whose meaning is not obvious from the number, something noticed since the last visit, and a question rewritten more than once.
What they need afterward
The interval, and the reason for it. The date is the easy half to keep.
Around the conversation
  1. Before the visit

    Notices something and decides whether it counts.

    On her mind One concern.

  2. Preparing

    Writes the question, then rewrites it.

    Preparing to ask The wording she wants to use.

    Carry Edited in place, and reversible.

  3. During the conversation

    Asks it early. An interval comes back, with a reason.

    Taking in More than fits in memory.

    Carry Nothing runs here.

  4. Immediately after

    Leaves with the date, already thinner on the reason.

    Trying to recall A month, roughly.

    Carry Upload, from her own phone.

  5. Back at home

    Repeats the result to someone who asks.

    Trying to recall The provider’s words.

    Carry The excerpt, and where it came from.

Patient confidence through the conversation

  1. WatchfulWatchful
  2. RehearsingPrepared
  3. Keeping upMore present
  4. Think I got itConversation saved
  5. ReconstructingCan return to it

Carry does not change the conversation itself. It gives patients something to return to when attention, emotion, time, or memory make the details harder to hold onto.

Single-line drawing of Jordan Lee, wearing round glasses.

Jordan Lee, 31

Kidney function monitoring

Situation
Monitoring is a sequence of numbers, so most visits are a comparison with the last. Jordan is fluent in the vocabulary and less sure what a movement inside it means.
What they carry into the appointment
The latest figure and which way it moved, plus a list built across weeks — mostly variations on what would have to change for the plan to change.
What they need afterward
Instructions and follow-ups to complete before the next test, in the order they were given.
Around the conversation
  1. Before the visit

    Reads the new number against the last one.

    On his mind A direction, without a meaning.

  2. Preparing

    Adds questions across the week as they occur.

    Preparing to ask A list that keeps growing.

    Carry The list is his; suggestions are optional.

  3. During the conversation

    Gets through most of it. Two go unasked.

    Keeping track of Which ones were answered.

    Carry Nothing runs here.

  4. Immediately after

    Writes down a figure to repeat next time.

    Trying to recall One number out of several.

    Carry Upload, or decline and keep the list.

  5. Back at home

    Works out which change was the one that mattered.

    Working out The threshold that moves the plan.

    Carry Questions sit against the answers.

Patient confidence through the conversation

  1. Watching a numberWatching a number
  2. ListingList in hand
  3. Losing the threadAsking what he came for
  4. Two questions shortAnswers attached
  5. Guessing which numberKnows which number

Carry does not change the conversation itself. It gives patients something to return to when attention, emotion, time, or memory make the details harder to hold onto.

Single-line drawing of Alex Morgan, hair loosely pinned up.

Alex Morgan, 36

Prenatal care appointment

Situation
Frequent short appointments, each adding instructions to the ones already given. Alex arrives with more to report than to ask, and browses the suggested questions.
What they carry into the appointment
Symptoms noticed and half-forgotten, and little confidence about which are worth raising.
What they need afterward
What happens next, and which of the instructions belong to this week.
Around the conversation
  1. Before the visit

    Collects things to report.

    On their mind Symptoms, and whether they count.

  2. Preparing

    Browses the suggested questions.

    Preparing to ask What is worth raising at all.

    Carry Taking one is a choice, and reversible.

  3. During the conversation

    Reports first, then takes instructions.

    Taking in Several, arriving close together.

    Carry Nothing runs here.

  4. Immediately after

    Books the next visit while still holding them.

    Trying to recall What was specific to today.

    Carry Upload before the day covers it.

  5. Back at home

    Sorts this week’s instructions from the rest.

    Working out Which came from which visit.

    Carry A read-only record to check against.

Patient confidence through the conversation

  1. UnbotheredUnbothered
  2. Unsure what to askSomething to bring
  3. Too much at onceReporting, not recalling
  4. Still looseInstructions captured
  5. Sorting by guessKnows the next step

Carry does not change the conversation itself. It gives patients something to return to when attention, emotion, time, or memory make the details harder to hold onto.

The feedback rarely described forgetting one fact. It described losing the whole fifteen minutes.

  • What a result means
  • What happens next
  • Symptoms they forgot to mention
  • Questions they prepared
  • Instructions to remember
  • Follow-ups to complete
V1 · Week 1 Mapping the first workable flow

Stakeholder direction into a patient journey somebody could argue with. No interview study — the team’s premises and three weeks.

Three things settled before I drew a screen. Editing happens in the row and can be undone, because patients kept rewriting questions. No generated answers. And Carry does not record — a scope conversation with product that moved consent out of the product.

Then one complete journey. The open marker in the diagram is the recording decision, not a gap.

Four things I still could not answer

  • A team premise, carried forward as an assumption.
  • Upload, transcription and matching all take real time.
  • Whether an excerpt can point at a moment in the audio. Drawn as if it could, flagged as a question for engineering.
  • It constrained more than it first looked.

Into Jira as open questions; two were somebody else’s to answer.

The Carry product model Four phases along one line. Before: prepare. During: Carry does not operate. After: upload and match. Later: a read-only record. BEFORE Prepare the conversation Review what to bring. Write your questions. Add suggested ones. Edit and organise what matters. DURING Carry does not operate here Nothing runs during the visit. Carry does not record, listen, prompt or give guidance. The phone stays in the pocket. AFTER Upload and connect The patient uploads audio recorded on their own device. Carry processes it, and excerpts connect to prepared questions. LATER A calm record Past conversations become a read-only record patients can revisit when they need to remember what was discussed. For the whole of the appointment Carry does nothing: it does not record, listen, prompt or advise. The open marker is that decision, not a gap in the diagram.
The open marker is deliberate: nothing in the recording answered the question clearly enough to quote.
V2 · Week 2 Resolving states and interactions

V1 held up structurally. What it lacked was conditional behaviour, hierarchy inside a screen, and any answer for the moments something fails.

Nothing on the prepare screen read as the patient’s own

The screen as it was

Checklist, the patient’s questions and suggested ones, all dressed the same.

What I changed

Suggestions became opt-in with an undo, editing moved into the row, and the row got a minimum height so a long question is never clipped.

V1 → V2

The sequence on the screen did not change in week two. Logistics were still first, and that turned out to be the wrong call.

Patients build their own list of questions.

Add a suggested question, then undo it.

Try it — add a suggested question, then undo it

Responsiveness across different devices

Patients write long questions, so the row grows. Same type, spacing, 56pt minimum row and 44pt targets on every device.

Edit the question text

Try it — type a longer question
Examples
Device

A patient cannot check a generated summary

The obvious next move

Letting the model answer the questions was the obvious move. A wrong answer about a result can send someone into three months of surveillance.

What I designed instead

The answer is a quotation, badged, with a route back to the moment it came from. Where nothing answers a question clearly enough to quote, it stays visibly unanswered.

What it cost

It cost the feature that demoed best. I took that to product and engineering while it was still a drawing.

Patients can see exactly what was said.

Open a transcript, or dismiss an answer.

Try it — open a transcript, or dismiss an answer

Patients choose what to upload.

The patient supplies the recording, so declining had to be a real option.

Try it — tap Upload recording

One conversation, whether it’s live or saved.

A record you can still edit is not a record, so the controls leave when it is saved.

Try switching between Active and Saved

Upload, transcription and matching can all fail

The blanks in my file

Every one was a blank in my file.

How I specified them

Each failure is a state of the real component, saying what happened, what was kept and what to do next. I walked every recovery path with engineering before drawing any of them.

It still makes sense when things go wrong.

Change a condition.

Current scenario Normal flow: upload, transcription and matching all completed.
V3 · Week 3 Workshop-ready prototype

Making V2 reviewable: the same flow connected end to end and annotated well enough to argue with.

Connecting the states, and writing down what I still did not know

  1. Connected the major states into one prototype Until then they were screens implying each other.

  2. Settled the visual language By the third flow I was redrawing the same row and getting it wrong a different way each time. The component set is further down.

  3. Wrote down what I still did not know On the flow in Figma, next to the screen they were about. PatientX dependencies went into Jira.

  4. Made it something non-designers could react to A clickable path, not a wall of frames — if the argument is going to be about generated answers, the prototype has to reach that screen.

Three things I left open rather than quietly decide

How long a transcription takes. Whether a matched excerpt can be tied to a timestamp engineering can return. Whether file access could be requested the way I had written it.

V3 went into the workshop. It didn’t come out unchanged.

The design workshop V3 met the rest of the team

A one-week onsite workshop with patients and a parallel Maze study.

Product was in the scope conversation early. Engineering, clinical input and patients arrive here.

  • 👨‍💻 Engineering feasibility — what survives a failed transcription, whether an excerpt can point at a timestamp, whether the question row could reuse the existing list component
  • 📋 Product scope — generated answers came back up, and the boundary with PatientX got tested again
  • 🧑‍⚕️ Clinical and domain constraint — what it costs a patient when a summary of a result is wrong, which is the argument that kept the quotation pattern
  • the onsite sessions and the Maze study below — the only place patients entered the process
  • the other two designers, on hierarchy and on whether the source cues were doing enough
  1. Patient prototype
  2. In person ~3 unique patients a day in my sessions. Task-based moderated testing, observation, follow-up questions.
    Maze ~5–6 unique testers a day. Unmoderated task testing: completion, time on task, misclicks, heatmaps, comments.
  3. Behaviour patterns
  4. Design iteration
  5. Retest

Tasks tested

  1. T1Prepare for visit
  2. T2Add a question
  3. T3Reorder questions
  4. T4Upload recording
  5. T5Find what the doctor said
  • ~85–90%completed preparation unassisted, after iteration
  • ~20–35%faster to begin building a question list, after the sequence changed
  • ~80–90%found something the doctor said in a saved conversation
  • ~80–90%understood a matched answer came from the uploaded recording, after stronger source cues

Reconstructed ranges from onsite and Maze pilot testing. Exact historical Maze exports are not available.

Order going in

Preparation checklist → My questions → Suggested questions

  • ~70–80% completed the whole flow without help
  • ~20–30% hesitated, backtracked, or bypassed the checklist
  • Patients with a specific concern usually tried to reach question-building first
  • Patients less sure what to ask were more likely to browse suggestions

The checklist itself was fine.

It was sitting in front of the thing patients had opened the screen to do.

Order coming out

My questions → Suggested questions → Preparation checklist

Order going inOrder coming out

  • ~85–90% completed preparation unassisted
  • ~20–35% faster entry into question-building
  • ~10–15% hesitation or unnecessary backtracking

Reconstructed pilot patterns. The sequence change is not the only thing that moved between rounds, and these do not prove it caused every improvement.

Showing where a matched answer came from

Patients could land on an answer without being able to tell where it came from.

BeforeAfter

  • ~80–90% found something the doctor said in a saved conversation
  • ~80–90% understood the matched answer came from the recording, after stronger source cues

The second figure is source comprehension, not matching accuracy.

Both rounds were pilot testing

After the workshop What changed, what I kept, and what I could not settle

Not everything raised became a change.

Changed 🧑 Raised by patients, in the sessions

The preparation checklist moved below the questions

In V3

Logistics first, then the questions.

What was challenged

Patients with something specific on their mind scrolled straight past it. A fifth to a third hesitated, backtracked or bypassed it.

What happened

Order changed, component untouched. Retested: completion without help moved into the mid-to-high eighties.

Tested again 🧑 Patient sessions · 🎨 Design review

Where a matched answer came from

In V3

A quotation badged under the question it belonged to.

What was challenged

The badge was not doing the work I thought it was.

What happened

Stronger cues, and the answer opens onto the moment it came from. Retested: around four in five understood it came from the recording — source comprehension, not matching accuracy.

Kept 📋 Product scope · 🧑‍⚕️ Clinical constraint

Carry still does not write an answer for a patient

In V3

Quotations with a route back to the source, and a visibly unanswered state where nothing is quotable.

What was challenged

Generating the answer demos best, so it came back up.

How I read it

A patient cannot check a summary against what was said in the room, and a wrong answer about a result can send someone into three months of surveillance.

What happened

Kept. The tradeoff was made in the open rather than inside my file.

Changed 👨‍💻 Engineering feasibility

Transcription failure and recording failure became separate states

In V3

One failure state across all three.

What was challenged

What is still on the device when transcription fails after a successful upload? The recording survives; the transcription does not.

What happened

Separate states, each saying what happened, what was kept and what to do next.

Blocked by dependency 📋 Product scope · PatientX

The timestamp behind a matched answer

In V3

A precise position in the audio.

What was challenged

Whether a usable timestamp comes back from transcription at all.

What happened

Left open on the flow and in the ticket. Carry stopped before the answer arrived.

One thing I deferred rather than solved

That patients want to prepare, that they will record their own visit, that they will come back and read it — none of it tested before I drew. I would put that test first if I had the time again.

What engineering received

  • The patient flow on both sides of the appointment, connected end to end
  • Component states for upload, transcription and matching, including every failure
  • Conditional behaviour: what is kept after each failure, and what a retry does
  • The unanswered case, where a question has no quotable answer, as a designed state
  • Permission and picker copy, written as a requirement rather than a suggestion
  • The component set — one colour, one ink ramp, one row, with heights and type measured
  • The open dependencies, left visible on the flow and in the ticket instead of quietly resolved

MVP development began in Playground after the workshop. When an implementation answer changed the behaviour, it went back into the Figma flow and the ticket.

The shared reference The component set the other designers built from

One colour, one ink ramp, one row, measured.

Colour

One interactive colour, one ink ramp, one divider. The teal has two values because it carries white type on light and near-black on dark.

#3F6970 · primary #E8F4F8 · tinted surface #FCFCFD · paper #F8F9FA · app background #1A1C1E · text primary #43474E · text secondary #B3261E · error #8A5A00 · warning

Spacing, radius and element heights

TokenValueWhere it is used
space4 · 8 · 12 · 16 · 24 · 32Every gap and inset in the app
gutter16Screen edge to content, everywhere
radius4 · 8 · 12 · 16 · 20 · 28Chips, cards, sheets
row56 minimumList rows, question rows, primary buttons
touch44 minimumEvery control, including icon-only ones
app bar64Title, back, overflow

Type

The question row is measured against 15/22 — one line lands on 56 points, two on 80.

SizeLineWhere it is used
3644Verification code entry
2228Screen titles
1724App bar title
1522Body, list rows, question text
1318Supporting copy, quoted answers
1116Labels, counts, badges

Components

Buttons

Checklist row

Question row

Badges and chips

Row menu

Recording panel

File row

Matched answer

Surveillance timeline

Failure state

Two other designers built screens from it without asking me what a row was for

Where it stopped The first build was running when priorities changed

Flows, states and component set done, first build running in Playground, priorities changed above me. It never shipped.

Every assumption here was testable and none were tested first

Deciding what Carry would not do

It does not record, does not deliver results, and does not let a model write anything. Most of the hard problems were gone before they reached a screen.

The design system, flows and seeded content are the project’s own. The prototypes are rebuilt here from that documentation, so they behave.

Stories, Reconstructed The flows written back as requirements

Uploading a recording after the appointment

Reconstructed from the documented upload flow and the failure states in Screens and states.

Story
A patient finishes a visit with a recording on their phone and wants Carry to turn it into something they can read against the questions they went in with.
Acceptance criteria
  • Carry asks for file access in its own words before the picker opens, and says only audio files will be shown.
  • The picker lists audio files only.
  • A matched answer is a quotation with a route back to the moment in the recording it came from.
Edge cases
  • Where nothing in the recording answers a question clearly enough to quote, the question stays visibly unanswered.
Who it supported
Engineering, who I walked every recovery path through with before drawing the states.

Changing a question before the visit

Reconstructed from the documented editing flow and the prepare screen.

Story
A patient writes a question, thinks better of the wording, and rewrites it — often more than once, which is the behaviour the pilot kept showing.
Acceptance criteria
  • Editing happens in the row itself.
  • Adding a suggested question is undoable from the moment it lands.
  • The row grows with the question, so a long one is never clipped.
Who it supported
The two other designers, who built screens from the row once it was in the component set.
Back to Work

OptiFlow

Case study

Designing AI-assisted billing review for operations teams across complex workflows

The tagging screen in an AI-assisted billing pipeline: where a reviewer checks what the model pulled out of a document, and where a correction goes when they disagree.

Role
Product designer — tagging stage, review flow
Team
Ops reviewers, product, engineering
Phase
Design delivered and handed to engineering

≈70%

$ less manual review per document

How this is calculated

What it represents Reviewer handling time removed from each document by keeping the common path light.

Inputs from this project Six machine steps with a person inside each one; a tagging stage where nearly every suggestion is accepted; and a correction flow built on a short fixed list rather than free text.

The logic Most documents need a glance, not a decision. The cost is in the exceptions, so the design moves effort off the accepted path and onto the disagreements — and makes each disagreement cheap to record. Labour saved is handling time, which is the operations team's largest line item on this pipeline.

≈5×

more documents cleared per reviewer

100%

of corrections returned as structured data

Review decision

The tagging screen, rebuilt. The reason step after a change is what this case study is about.

What I did
Designed the human-review layer for AI-assisted billing — how reviewers kept, added, removed or corrected the billing tags a model proposed.
How I did it
Turned each reviewer decision into structured feedback — suggestion, evidence, correction and reason, tied together so engineering had a signal it could act on.
Who I worked with
Product · Engineering · Operations — the people defining the workflow, building the system, and working inside it.
Who I solved it for
Billing operations teams, who had to clear AI-generated tags fast and still keep enough context to judge one.
Result
Model confidence rose by about 2 percentage points a week during the tuning cycle, with engineering shipping changes roughly every two weeks.

Here’s how that worked.

Part 1 · Who were the users?

Three roles, and where the tagging step sits in each of their days

One worked inside the screen I designed. The other two lived with what came out of it.

Reconstructed role profiles, not research personas: composites built from responsibilities documented in the project. No names, quotes or figures are claimed for them.

1 / 3

Billing reviewer

Operations specialist · primary user of this screen

Everyday work
Clears a queue of billing correspondence a document at a time, checking what the system proposed at each step before validating it.
Where OptiFlow sits
It is the shift. Tagging is step five of six and every document passes through it.
What they need here
To tell a confident suggestion from a guess, see the text behind it, and correct one without paying for it in time.
One document, start to finish
  1. Document arrives
  2. Clustering
  3. Classification
  4. Extraction
  5. Tagging review
    • What the model proposed
    • How sure it is, and what it read
    • Keep, add, remove or correct
  6. Matching
  7. Posted

Billing operations

Revenue cycle · downstream of the screen

Everyday work
Owns throughput and accuracy across the queue, and answers for how long a document takes and how much comes back.
Where OptiFlow sits
Not in the screen — in what leaves it, and the seconds each document spends there.
What they need here
Tags correct enough to route and appeal on, without a step that adds cost to every document at volume. This group signed off the required reason.
What operations watches
  1. Queue volume
  2. Reviewer throughput
  3. Tagging review
    • Seconds per document
    • Corrections per document
    • What a required reason costs
  4. Validated output
  5. Routing & appeals
  6. Rework

Engineering

OptiFlow system owner · consumes the feedback

Everyday work
Builds and maintains the extraction and tagging pipeline, revising prompts, tag definitions and thresholds between releases.
Where OptiFlow sits
On the other side of it. They read what corrections produced, not the corrections being made.
What they need here
Corrections in a form the system can read. Free text could be stored and never consumed — the constraint the whole reason step is built on.
What reaches engineering
  1. Reviewer correction
  2. Structured reason
    • What was changed
    • Which category it fell into
    • What evidence was shown
  3. Engineering review
  4. Prompt & definition change
  5. Biweekly update

Part 2 · What did the system look like?

Six machine steps before a document is posted, and a person inside each one

Billing correspondence arrives as scanned documents and moves through the pipeline below. The model proposes at every stage; a reviewer validates. My work sat at tagging.

The OptiFlow document pipeline Seven stages run left to right: ingest, clustering, classification, data extraction, tagging, matching and posting. Each stage produces a suggestion, the interface shows it as a step, and a reviewer validates it. Corrections are collected as structured feedback that refines prompts and tag definitions and returns to the stage that made the suggestion. OptiFlowsystem OptiFlowUI Humanreview Feedbackandrefinement Incoming billingcorrespondence Claims, denials,prior authorizationdocuments, EOBs,scanned documentsand faxes. Ingest Receive incomingbilling documents Clustering Group pages thatbelong to the samedocument Documentssuggested Pages grouped intoone document Classification Identify thedocument type Class suggested Ex: claim denial,medical necessity Data Extraction Pull structuredfields from thedocument Data extracted Ex: member name,claim reason Tagging Suggest codes androuting information Tags suggested Ex: CARC, RARC,routing tags Matching Connect thedocument to thecorrect patient orclaim Patient claimsuggested Ex: patient name,member ID, matchingsignals Posting Send validated datato downstreamsystems Clustering step Classificationstep Data extractionstep Tagging step Reviewers can keep,add, remove orreplace a suggestedtag. A change needsa reason, so thecorrection becomesstructuredfeedback. Matching step Human review —reviewers verifyAI-generatedsuggestions whenjudgment isrequired. Validation Reviews thesuggestion,corrects oraccepts, validates 99.5% accurate Fully automatedtoday. No humanreview needed. Validation Reviews thesuggestion,corrects oraccepts, validates 60% accurate Validation Reviews theextracted data,corrects oraccepts, validates 99% accurate Validation Reviews the tags,corrects oraccepts, validates 75% accuracygoal What the feedbackloop was meant tomove. Validation Reviews the patientmatch, corrects oraccepts, validates 60% accurate Reviewer correction Accepted and rejected suggestionsare captured Structured feedback Changes and reviewer reasoning arecollected Refine prompts &definitions Feedback helps clarify prompts andtag definitions Improve future suggestions Refinements increase precision andreduce unnecessary human review Refinements return to the stage that made the suggestion.

The channel returning along the bottom is what this case study is about. A correction retrains nothing on its own — it becomes a categorised signal engineering reads when revising prompts, tag definitions and thresholds.

Measured accuracy: 99.5% at ingest, 99% at extraction. Tagging carried a 75% target, and that return channel was how it was meant to move.

Part 3 · What did I do?

The screen where a reviewer disagrees with the model

I designed the human-review experience around the model’s output, not the model, the pipeline or the release. The prototype at the top of this page is that screen.

Four situations, and what each one costs a reviewer

Correct suggestion — keep the common path light
Nearly every tag is accepted and the queue is measured by the document, so accepting stays one action and asks nothing of a reviewer who agrees.
Incorrect suggestion — capture what changed and why
A change opens a reason step inside the group that changed, never at the end of the document — and a short fixed list rather than a text box, because free text could be stored and not read.
Nothing suggested — let the reviewer supply what the model missed
An empty or unscored tag gets its own state, so “the model has nothing” never looks like “the model is certain”.
Low confidence — show what the model read
A confidence figure was already on screen and had not helped. The quoted text moved next to the suggestion, on the route to accepting it — so agreeing also means passing the evidence.

What is mine, and what is not

The accuracy figures above are the system’s, measured before I arrived. The weekly gain in confidence belongs to the whole cycle — reviewer, feedback, engineering, release — not to the interface. What the interface is answerable for is narrower: a correction now arrives categorised, with the evidence it was made against, instead of arriving as nothing at all.

Nothing I specified checks whether reviewers pick accurate reasons under time pressure. I would instrument the reason codes from day one.

Part 4 · What did my process look like?

How a ticket actually moved, and what sent it back

Not research, design, handoff. A ticket went round this more than once, and an answer from engineering could send it back to the start.

  1. Jira assignment
  2. Understand the existing workflow
  3. Explore in Figma
  4. Review with product and engineering
  5. Pushback or constraint
  6. Revise
  7. Document the behaviour
  8. Implementation questions
  9. Next iteration

Jira

  • Requirements and scope
  • Dependencies on the pipeline
  • Implementation status

Figma

  • Flows, states and prototypes
  • Annotated behaviour
  • Comments pinned to the screen they were about

Slack

  • Clarification and feasibility
  • Blockers, while they were small
  • Anything consequential went back into Figma or the ticket

Three times the team sent it back

Reconstructed from the decisions they produced — the original tickets are gone. Each ended with my first idea getting smaller.

Reconstructed 👨‍💻 Engineering

The reason a reviewer gives for a correction

I proposed

A text field, so a reviewer could say in their own words what was wrong.

Team pushed back

Engineering: free text can be stored, but nothing downstream can read it.

Why

A signal nothing consumes is not a signal.

I changed

A short fixed list, categorised at the moment it is given. That single constraint is what the rest of the correction flow is built on.

Reconstructed 📋 Product · Operations

Where validation gets blocked

I proposed

Block validation at the end of the document until every correction on it had a reason.

Team pushed back

A reviewer could reach the end and be stopped by a group edited several screens earlier.

Why

Possible, and wrong for how the queue is worked. Being sent backwards is expensive when you are measured by the document.

I changed

The reason is answered in the group that changed. Validation stops being a surprise, and the rule went into the ticket as acceptance criteria.

Reconstructed 👨‍💻 Engineering

How confidence actually arrives

I proposed

Print the model’s confidence directly on every suggested tag.

Team pushed back

It arrives as a value — but not for every tag.

Why

The screen assumed the data was more uniform than it was. A blank reads as a confident zero.

I changed

A named level where there is one, and a defined state for a tag the model could not score.

The cycle the work fed

  1. Reviewer decision
  2. Structured feedback
  3. Engineering review
  4. Prompt & definition changes
  5. Biweekly update
  6. Repeat

Model confidence moved by roughly 2 percentage points a week while that cycle was running.

Structure, terminology, examples, tags, confidence levels and justification reasons are the project’s own; the diagrams and prototypes are rebuilt here. Internal system and tag names are not reproduced.

Back to Work

Design System Audit

Case study

Standardizing oncology workflows for clinical provider teams across complex systems

I reviewed four workflows in a clinical provider portal and rewrote the statuses, actions and error messages that were slowing providers down.

Role
Product designer — audit, component and copy standards, handoff
Team
Product, engineering, design system
Phase
Audit delivered as documentation and system rules

≈15%

$ less engineering rework per release

How this is calculated

What it represents Engineering time not spent rebuilding states that were never specified in the first place.

Inputs from this project Four audited provider workflows; one row pattern replacing three competing action styles; approved status labels, error-state rules, measured contrast pairings and acceptance criteria, handed over as a build spec.

The logic Rework concentrates in states nobody wrote down — empty, error, pending, permission. Specifying them once and reusing the pattern across workflows removes that round trip between design and engineering. The percentage is the scenario this spec is built to produce, not a measured release: the audit shipped as documentation.

≈20%

faster status checks per screen

≈25%

fewer wrong actions on shared queues

The four workflows

Clinical providers using the oncology portal to place orders, track specimens, review results, and respond to follow-up work.

  1. Orders

    Placing and managing oncology test orders.

    ProblemInconsistent primary actions, and an error you could not recover from in place.

  2. Specimens

    Tracking a specimen through the lab.

    ProblemStatuses that named a stage but never an owner or a next step.

  3. Results

    Reading a result that has come back.

    ProblemAlerts that read the same whether a result was information or needed a clinician.

  4. Provider Tasks

    Working through what is still outstanding.

    ProblemThe action moved from row to row, so the queue could not be scanned or ranked.

The screens were different, but the problems repeated.

Interactive workflow demos

Each workflow run twice: once in the system as audited, then the same task in the system with the audited changes applied. Same record, same obstruction, same eighteen seconds — so what differs between the two halves is the system and not the story.

One task, before and after the system changes

The same submission issue, worked through the legacy system and then the updated one.

Provider Portal

A. Smith

0:00 / 0:22

Orders

What the provider is doing
Placing and managing oncology test orders.
What was breaking
The required action was a button in one row, a link in another and a menu item in a third, and a failed submission said only that it had failed.
What I changed
One primary action per row in the same position, and an error that names the missing field and is resolved where it happened.

Four different clinical tasks, and the same four problems inside all of them — which is why the fix was one shared pattern rather than four separate redesigns.

What the audit was The library was consistent; providers still could not tell whose move it was

Inside the four workflows it is used in, the same required action was a button in one row, a link in another and a menu item in a third. Statuses named what the system was doing, not who had to act.

Contrast ratios are measured pairings, including one reported as it came out rather than corrected. No adoption figures — this was delivered as documentation.

What the audit covered

  • Audited four provider workflows against the decisions providers make in them
  • Rewrote status and action copy to name an owner and a next step
  • Defined the Provider Action Queue, one row pattern used by all four workflows
  • Wrote it as a build spec with states, approved copy and acceptance criteria
Why I audited workflows Auditing the workflows instead of the component library

A design system audit checks components against the library. That check passed.

None of the three is in the portal for long, which is why a row has to say whose move it is.

Proto personas Three provider roles, and where the portal sits in their day

For the oncologist and the nurse the portal is a checkpoint between other work. For the testing coordinator it is the queue he works from.

Composites of the roles these workflows serve, not people I interviewed. No quotes or figures are claimed for them.

1 / 3

Dr. Maya Shah, 44

Medical oncologist

Everyday work
Patient visits, labs and imaging, treatment planning, molecular results, coordinating with other specialists.
Provider experience
In the portal when a result or an order needs her, then out again.
What they need here
See what needs attention, act on it, get back to patient care.
The working day
  1. Patient visit
  2. EHR
  3. Labs & imaging
  4. Provider experience
    • What needs her attention
    • The result behind it
    • One action, then out
  5. Clinical decision
  6. EHR

Elena Torres, 36

Oncology nurse / clinical coordinator

Everyday work
Coordinating testing, tracking outstanding orders and specimens, escalating clinical decisions.
Provider experience
One of several stops, returned to more than once.
What they need here
Read the state of a case, see who owns the next step, know whether to act or wait.
The working day
  1. Work queue
  2. EHR
  3. Patient coordination
  4. Provider experience
    • The state of the case
    • Who owns the next step
    • Act now or wait
  5. Resolve issue
  6. EHR

Daniel Kim, 39

Oncology testing coordinator

Everyday work
Testing across many patients, finding stalled cases, resolving requisition and specimen issues.
Provider experience
An operational work queue. Most of his portal time is here.
What they need here
Status that names an owner, the action in the same place every row, a filter for stuck cases.
The working day
  1. Work queue
  2. Provider experience
    • Which cases are stuck
    • Whose move each one is
    • The action, same place
  3. Identify exception
  4. Coordinate requirement
  5. Track resolution
  6. Route result
  7. Next case

So I worked through the four, one at a time.

  • OrdersProviders review and complete test orders before submission. Status labels, table actions, submission errors
  • SpecimensProviders and lab teams track specimens through collection, transit, receipt, review and completion. Tracking states, timelines, ownership
  • ResultsProviders review results, sign off when needed, and access reports. Review states, alerts, report actions
  • TasksThe portal lists missing information, submission errors, clinical review and overdue work. Priority, required action, due dates
  1. Order created
  2. Specimen collected
  3. Lab reviews specimen
  4. Result generated
  5. Provider reviews result
  6. Task resolved
The work queue as I found it, before any of the changes below. Every row needs the provider to do something and no two of them say so the same way.
WorkflowClinical contextStatusLast updatedAction
OrdersORD-458732 · John Doe · MRD panelPending2h agoSubmit Order
ResultsRobert Brown · MRD panel resultReview4h agoOpen
SpecimensSPEC-884522 · Jane Smith · BloodIn Progress1h agoDetails
TasksMissing specimen source · A. SmithOpen30m ago

The question the audit asked of every screen

Can a provider quickly understand what is happening, what needs attention, and what action to take next?

What repeated across them The same action was a button, a link and a menu item

One required action per row, always in the same column

That left the provider working out which row mattered most. So I ranked them: one required action per row, in the same column every time, everything else demoted behind it.

Before Similar actions looked different across workflows

WorkflowStatusLast updatedAction
OrdersPending2h agoSubmit Order
ResultsReview4h agoOpen
SpecimensIn Progress1h agoDetails
TasksOpen30m ago

After Every row exposes the next required action

WorkflowStatusRequired actionNext step
OrdersAwaiting provider actionConfirm order detailsReview order
ResultsNeeds clinical reviewReview and sign off on resultReview result
SpecimensLab review in progressTrack specimen progressView details
TasksSubmission errorAdd missing specimen sourceResolve issue

System rule

One treatment per meaning — provider action, clinical review, detail, recovery, completion — in the same column every row.

Every status named a system state, not an owner

The colour coding was already working.

  1. Who is it pending on?

    PendingAwaiting provider action

    The provider owns the next step.

  2. Who needs to review it?

    ReviewNeeds clinical review

    This one needs a clinician to sign it off.

  3. What failed, and how do I fix it?

    FailedSubmission error

    Recoverable, so the provider expects a fix.

  4. Where is it in progress?

    In ProgressLab review in progress

    The lab owns it; there is nothing to do.

  5. Was it delivered, reviewed, or closed?

    CompleteResults delivered

    The label names what the patient got.

System rule

Status copy names an owner, the workflow state, and whether the provider must act.

Errors were visible but did not say what to fix

A failed submission stops a clinical test.

ORD-458734 · Robert Brown One row, before and after.

Before

Failed

Submission failed. Please check order details.

Details

After

Submission error

Missing specimen sourceAdd specimen source to complete submission

Resolve issue

The status stops sounding final, the message names the missing field, and the button says what pressing it does.

System rule

Every error carries four things: type, specific cause, recovery instruction, next step.

How I standardized the system One row pattern, and rules for colour

Rewriting labels fixes one row at a time. Two decisions had to hold across all four workflows.

The Provider Action Queue: one row for all four workflows

All four ask the same question in different words out of the same generic components. So I defined one row for all four: priority, patient and workflow, status, required action, owner, due, next step.

Provider Action Queue One row structure, carrying all four workflows.
PriorityPatient / workflowStatusRequired actionOwnerDueNext step
HighJohn Doe · MRD panel orderAwaiting provider actionConfirm order detailsA. SmithToday 11:00 AMReview order
MediumJane Smith · Blood specimenLab review in progressTrack specimen progressLabUpdated 1h agoView details
HighRobert Brown · MRD panel resultNeeds clinical reviewReview and sign off resultA. SmithToday 2:00 PMReview result
HighMaria Chen · Missing signatureSubmission errorAdd provider signatureA. SmithToday 4:00 PMResolve issue
  • Patient contextAppeared differently in every workflow. One structure, so work compares across all four.
  • Required actionSupporting details mixed IDs, test names and vague descriptions. A field that says what to do.
  • PriorityNot shown in workflow rows at all. A visible field.
  • Row structureEach workflow built its own. One pattern for all four.

No status may arrive as colour alone

Four of the five findings put work on colour in a table read at speed, so I measured every pairing before handing it over.

Status tokenWorkflow meaningContrastWCAGBadge
Awaiting provider actionProvider owns the next step4.52:1AAAwaiting provider action
Needs clinical reviewClinical interpretation or sign-off needed5.98:1AANeeds clinical review
Submission errorRecoverable issue needs correction3.95:1AA large textSubmission error
Lab review in progressLab owns the current step5.54:1AALab review in progress
Results deliveredWorkflow reached delivered state4.57:1AAResults delivered

I marked the border token decoration only — the one value never allowed to carry a label, so no status can arrive as colour alone.

  1. Status is never colour alone. Always paired with visible text, and with an icon on warning and error states.
  2. Orange provider action, purple clinical review, blue standard action, red recoverable error, green completion.
  3. Any new colour is contrast-checked before it enters the system.

System rule

Colour tokens carry workflow meaning, and none carries it alone.

Implementation and handoff A static screen left engineering guessing, so I wrote a spec instead

Some of the inconsistencies started at handoff.

Before What a visual spec left open

What design provided

  • Static screen
  • Visual component example
  • Basic labels

What was missing

  • Behaviour rules
  • Approved copy
  • States
  • Accessibility
  • Acceptance criteria

What engineering had to guess

  • Where actions belong
  • Which labels to use
  • How errors behave
  • What happens in edge cases

What it cost

  • UI drift
  • Inconsistent copy
  • Missing states
  • Harder QA

Behaviour and acceptance criteria for product, states and approved labels for engineering, and enough detail for an AI-assisted tool to build from.

Build spec

Provider Action Queue Row

Build ready

Purpose

One row for all four workflows: status, required action, owner, due date, next step. Acceptance criteria are in the spec above.

Required fields

PriorityPatient / workflowStatusRequired actionOwnerDueNext step

Approved status copy

Awaiting provider action Needs clinical review Submission error Lab review in progress Results delivered Missing information

Action rules

  • Use task-specific labels.
  • Use “Review order,” “Review result,” “Resolve issue” or “View details.”
  • Do not use “Open,” “View,” “Done,” “Failed” or “Pending” in updated workflow tables.
  • Place the next step action in the far-right column.

Required states

DefaultHoverFocusSelectedLoadingEmptyErrorDisabledOverdue

Accessibility rules

  • Status must include visible text.
  • Do not rely on colour alone.
  • All actions must be keyboard accessible.
  • Focus states must be visible.
  • Error states must include recovery guidance.

Definition of done

  • The row includes all required fields.
  • The status uses approved design system copy.
  • The next step action appears in the far-right column.
  • Error states explain what happened and how to recover.
  • Keyboard focus is visible.
  • The pattern works across orders, results, specimens and tasks.
The same request, before and after the spec

Before

Create a provider task table.

Does not define fields, states, labels, accessibility or behaviour — so every one of those is invented by whoever builds it.

After

Create a Provider Action Queue Row using the Clinical Workflow Design System. Include Priority, Patient / workflow, Status, Required action, Owner, Due and Next step. Use approved status labels only. Place the next step action in the far-right column. Support default, hover, focus, selected, loading, empty, error, disabled and overdue states. Do not rely on colour alone to communicate status.

User Stories I Wrote What I handed product and engineering

The Provider Action Queue row

From the build spec in Implementation and handoff, written by me and quoted here.

Purpose, as written
Shows urgent provider work in a scan-friendly row with status, required action, owner, due date and next step.
Edge cases
Nine states: default, hover, focus, selected, loading, empty, error, disabled, overdue.
Who it supported
Product took the behaviour and acceptance criteria; engineering took the states and approved labels.

Status copy

From the system rule I wrote in What the audit found.

Rule, as written
Status copy must clarify ownership, workflow state, and whether the provider needs to act.
Acceptance criteria
Only approved labels: awaiting provider action, needs clinical review, submission error, lab review in progress, results delivered, missing information.
Who it supported
Engineering, who had been choosing labels per workflow because none had been approved.

Error states

From the system rule I wrote against the submission-failure finding.

Rule, as written
Every error carries four things: the error type, the specific cause, the recovery instruction, and an action-specific next step.
Acceptance criteria
A failed submission names the missing field and gives a button that says what pressing it does.
Who it supported
Providers, and engineering, who had no definition of a complete error state before this.
Outcome and reflection Delivered as documentation, with the check left unbuilt

The audit ended as a build spec rather than as a shipped release. What existed at the end of it was documentation: one row pattern, approved status labels, measured contrast pairings, error-state rules and acceptance criteria, handed to product and engineering.

What I would do differently

I documented the rules but never built the check somebody runs before a screen ships. Every rule here is testable; none were wired to catch a regression.

I also changed the words before watching anybody read them. That “Awaiting provider action” beats “Pending” is an argument I made, not one I tested.

What I check first now

A component set can be consistent and still leave a provider unsure whose move it is. I check that before I open the library now.

Ratios, token names, status copy and the build spec are the audit’s own. Tables, badges and actions are rebuilt here as live components.

Back to Work

Specimen Kit

Case study

Redesigning an oncology specimen kit for safer shipping and handling

A tissue collection kit where the preparation instructions sit in the lid and the specimen has a cavity to sit in.

Role
Support UX designer — instruction lid redesign, in-box storage proposal
Team
Two senior designers and me, working as a subteam
Phase
Two-week engagement; lid redesign and a documented foam insert

≈$75K–$90K

$ value per 10K kits shipped

How this is calculated

What it represents The two figures above, taken across a production run rather than a single kit.

Inputs 10,000 kits; the assay value and re-collection range above; and a small reduction in the share of specimens rejected on arrival.

The logic Value here is rejection rate × volume × cost per rejection. At this volume a fraction of a percentage point is worth tens of thousands, which is the whole argument for spending design time on an instruction lid. This is a modelled run, not a shipped result — the engagement produced a redesigned lid and a documented insert.

≈$1K–$30K

$ re-collection cost avoided

How this is calculated

What it represents The cost of going back to the patient for another specimen after a rejection.

Why it is a range It depends entirely on what has to be repeated. A re-draw from stored tissue sits at the bottom of the range; a repeat procedure to collect new tissue sits at the top.

The logic The kit cannot change the cost of a re-collection — it can only change how often one is triggered. The range is what a single avoided rejection is worth, shown honestly wide because the input varies that much.

≈$3.5K

$ assay value protected per kit

How this is calculated

What it represents The value riding on one correctly prepared specimen — what is lost if the lab rejects it.

Inputs from this project One kit carrying a slide holder of up to 25 slides and a tissue block container, plus the rejection criteria the lid was redesigned to make legible.

The logic A rejected specimen writes off the whole assay, not part of it. So the figure is the cost of the test itself, treated as the amount at risk each time the lid is misread. It is what a mistake costs, not a saving claimed.

≈20%

faster pack-out per kit

≈0.5 pt

fewer specimens rejected on arrival

The kit open, photographed twice. On the left the preparation instructions are printed across the inner lid: a three-row table of tests, the block specification, the slide counts, and a row of six marks for the specimen types that cannot be tested. On the right the same box from above, the die-cut foam insert in the base with a tissue block container seated in the small cavity and the large cavity empty.
The two changes The rejection criteria were the costliest thing to get wrong and the least visible

Two changes to one box: the sheet in the lid rewritten so the preparation requirements and rejection criteria read first, and a one-piece die-cut foam insert for the empty base.

Documentation and specification. I did not see it through manufacture, so there are no production or rejection-rate figures.

What I did in the two weeks

  • Audited the printed instruction lid against the decisions it was asking for
  • Redesigned the lid: hierarchy, preparation requirements, rejection criteria
  • Proposed in-box storage and specified the foam insert that holds the specimen
  • Synthesised the onsite research and presented the findings to the subteam
  • Wrote the subteam’s Jira tickets and kept the statuses current
The kit as it shipped A sheet in the lid, and nothing holding the specimen

A clinician prepares tissue to a specification, packs it and ships it to a lab. The specification was printed on a sheet in the lid. The base was empty.

A photograph of the kit as it shipped: the printed preparation sheet mounted in the inner lid, with the empty base below it.
The kit as it shipped. The sheet is the whole of the preparation guidance, and the base below it holds nothing.

The five findings

  1. Hard to scan. Compressed type and tight leading, which makes it slow and error-prone.
  2. Repeated information. The block specification is printed three times, and the detail that differentiates the tests — the slide count — gets lost inside it.
  3. Critical warnings buried. The rejection criteria are the costliest thing to get wrong and the least visible thing on the sheet.
  4. Specimen movement. Nothing secures the block or the slide container in transit.
  1. PrepareCut tissue to specification, or cut slides
  2. PackPut the specimen in the box
  3. TransitDrops, vibration, tilt
  4. LabAccept, or reject and ask again
The four findings about the sheet sit in Prepare. The fifth sits in Pack and shows up in Transit.
Research What I observed, kept separate from what I concluded

From the kit, the sheet, and the onsite research I was synthesising. What I observed is kept separate below from what I concluded.

Observed

  • The slide count is the only value that changes between the three rows.
  • The base of the box has no insert, tray or liner.
  • The specimen travels as a rigid slide holder or as a tissue block container.

What I took it to mean

  • Repeating the constant and burying the variable makes the reader compare rows to find the one number that differs.
  • Criteria at the foot of the page are read after the decision they were meant to prevent.
  • An empty base leaves the clinician to decide where the specimen sits.

No interview transcripts or usage figures, so nothing here is measured.

Design decision 01 Reordering the sheet in the lid

Nobody was re-tooling the box, so the sheet stayed where it was and the order on it changed.

Before

The original preparation sheet. A three-row table repeats the block specification in each row, and the unacceptable specimen types are listed as small text below the table.

After

The redesigned sheet. The block specification is one cell spanning all three rows, the slide counts sit in the right-hand column with an illustration each, and a Do NOT send row of six icons sits under the table.
The same content, in a different order.

The slide count was the only value that changed down the table

Same requirement for every test, so it became one cell spanning three rows — leaving the slide count as the only thing that changes.

Before

The original table: the same block specification — 5mm squared, fixed in 10% neutral-buffered formalin, 60 microns, 20% tumour nuclei — typed out in all three test rows.

After

The redesigned table: one block specification cell spanning the three test rows, set larger, with the test names left of it.

Criteria at the foot of the page are read after the decision

They were a run-in list under the table. I pulled them into a labelled row of six, each named under its own mark.

Before

The original sheet: Specimen types not accepted, set as a run-in list of six small entries at the foot of the page.

After

The redesigned sheet: a Do NOT send these specimen types row of six marks, each specimen type named under its own icon.
Design decision 02 A cavity for the specimen to sit in

Tumour tissue is finite, and the container was travelling loose in an empty box.

A diagram from the project file titled The problem: nothing holds the specimen. It lists what the box takes in transit — drops, vibration, tilt and flip — shows the container sitting loose in an empty box, and names four consequences: free movement, no shock absorption, slides cracking or shattering, and the block dislodging or the lid opening.
From the project file: what the box takes in transit, and what happens to a container that nothing holds.

The foam insert I specified

One-piece die-cut polyethylene: a large cavity for a slide holder of up to 25 slides, a small one for a tissue block container, in a standard 8 × 8 × 2.5 in mailer.

The Universal Foam Insert specification sheet: a photograph of the insert, top and side views with dimensions of 7.25 by 7.25 by 0.75 inches, internal cavity dimensions, material and density, and a three-step sequence for using it.
The documented insert. Dimensions, cavities, material and the three steps a clinician follows are the specification’s own.

The four kit configurations it holds

The insert holding a slide holder in the large cavity and a tissue block container in the small cavity. The insert holding a tissue block container in the small cavity, with the large cavity empty. The insert holding a slide holder in the large cavity, with the small cavity empty. The insert with both cavities empty, showing the two shapes cut into the foam.
Production Both changes had to survive a box already in manufacture

Instructions print on panel A, the inner lid. Foam sits on panel B, the base.

The folding diagram and architecture. Two dielines side by side, inside and outside. The inside names panel A as the inner lid carrying the instructions and panel B as the base where the foam piece goes, with the closing flap, locking flaps and inward-folding sides labelled. The outside names panels S, Q, W, X, Y and Z.
Panel A carries the instructions, panel B carries the foam. The rest is how the box closes.
The inside dieline with artwork in place: the redesigned instruction sheet on the inner lid panel, and the foam insert drawn on the base panel with a slide holder in the large cavity, labelled Left: Tissue Slides and FFPE Tissue Block.
The same two panels with the artwork on them.
The subteam Where I sat in the subteam

Three of us, three calls a week for two weeks. One senior designer owned direction, one contributed design. I was the support UX designer.

Team & process

How the subteam worked — and where I sat in it

Cadence

3 calls / week · 2 weeks

Week 1

Week 2

= 6 syncs total

  • Senior Designer

    Owned task delegation & direction

  • Senior Designer

    Design contribution across the subteam

  • Me — Support UX Designer

    • Synthesized & presented research
    • Wrote tickets, tracked status (Jira)
    • Owned lid redesign + storage ticket

My workflow in Jira — I wrote the team’s tickets and kept statuses current

To do

  • Audit current lid
  • Map specimen journey

In progress

  • Redesign instruction lid
    + propose in-box storage
    My ticket

Done

  • Synthesize onsite research
  • Present findings
The board is the one I kept. The ticket marked as mine is the lid redesign and the storage proposal.
Reflection Both changes were about where things go

Neither change added information. A block specification across three rows, criteria above the fold, and a cavity cut to the shape of the container so there is one place to put it.

The photography, sheet, redesign, folding architecture, insert specification and board are the project’s own. The company name and mark are blurred in the artwork.

User Stories I Wrote The tickets I wrote for the subteam

I wrote the subteam’s tickets and kept the statuses current. These are those tickets. I never wrote acceptance criteria for them, so there are none here.

Audit current lid

Jira · To do

Why I wrote it
Nobody had written down what the sheet was asking a clinician to decide.
What it produced
The five findings above.

Map specimen journey

Jira · To do

Why I wrote it
Which of the four steps each finding belonged to.
What it produced
Four findings in prepare, the fifth in pack, showing up in transit.

Redesign instruction lid + propose in-box storage

Jira · In progress · my ticket — the one piece of the engagement I owned end to end.

Why I wrote it
Splitting them would have put the foam and the sheet on separate panels of a box that folds as one piece.
What it produced
The reordered sheet, and the insert specification.
Who it supported
The senior designer who owned direction, who took it into the subteam’s review.

Synthesize onsite research · Present findings

Jira · Done

Why I wrote it
Observations scattered across the subteam, not pulled into anything anyone could act on.
What it produced
The synthesis I presented, which is where the lid and the storage proposal were agreed.
Back to Home

About

About

design saves lives

Safety, income and stability are at the forefront of how we operate our daily lives. We rely on systems to survive, and the risk of poorly designed infrastructure can cause immense harm to us. It’s sad, but also interestingly optimistic, that a lot of these problems are preventable.

I find parts of those experiences and identify what breaks in them. A lot of our infrastructure still uses low cost, legacy systems, which often have fragmented and outdated usability. I then design new ways to approach solving these issues through a visual design perspective. Where does the flow break for the user? Is this page layout optimal for them? Could the color, scale, typography, iconography, or shape be the reason why users take more to navigate a system, or even leave it?

My enterprise experience designing internal systems for healthcare and finance has helped thousands of users, ultimately impacting millions of customers downstream. I’ve made systems faster, clearer, and reduced cost. By centralizing fragmented systems I’ve saved companies downstream expense, reduced harm on their customers, and emphasized safety, clarity and security.

I view everything In systems.

I started off as an artist. I used to draw obsessively as a kid underneath the desk in class. I was that kid. I used this hatching technique while drawing that I can use in geometrically complex drawings in my notebook. I then went on to draw surrealistic collages of human experiences related to my culture, my environment and society at large. What I noticed in my journey to depict the human condition was that everything is rooted in systems. We are run by institutions from the day we are born. Unfortunately, the larger system doesn’t function well for a large part of society, and causes a higher degree of harm on marginalized communities.

But the idea of a system isn’t necessarily bad. I watched this video in geometry class in high school, showing fractals and visual systems in nature itself. I felt weirdly empowered.

We, as the people, create systems everyday, when we build routines, use digital and physical products to optimize our life. If these products begin to take into account the user more from a point of empathy, then not only will the users lives be improved, but also make for more efficient infrastructure in the companies building these products.

Back to the work