← All work
Service Design Human-AI Workflows Operating Model Evaluation

Making accessibility scale — without breaking instructional quality.

A human + AI system for making instructional video accessible — designed to run reliably at institutional scale.

Role
Design Strategist
Timeline
~6 months
Team
Instructional design, AI vendor, digital learning & senior leadership
Recognition
Presented at TeachX 2026, Northwestern
The Problem

Scaling accessibility across 16,000+ learning assets

A federal mandate required universities to make digital learning content accessible, including audio descriptions for instructional video. Leadership wanted to meet it without burying staff in manual work, and saw AI as the way through.

So the project began with the obvious question: can AI detect visuals and generate descriptions well enough to replace manual work?

That question was the wrong one — and finding out why became the project.

The Reframe

Early exploration showed generation was the easy part. The same visual needed different descriptions depending on instructional context — so quality lived in context, review, and decision-making, not the AI output alone.

Can AI generate audio descriptions?
THE REAL QUESTION ↓
What operating model makes AI output trustworthy enough to scale?
My Role

Owning the system between content and technology

The instructional designers knew the content. The AI vendor knew the technology. The gap was between them — and that gap was mine to close.

I owned the stakeholder map, the end-to-end workflow, the review logic reviewers use to make calls, and the evaluation framework Phase 1 runs against. I also led the working sessions with senior leadership — the Deputy Vice Provost for Online Learning and the Chief Learning Officer — to settle governance and review ownership before any workflow was locked.

Process mapping session Working session Team collaboration Review session
The Work

Four decisions that shaped the system

Before defining process, I had to answer where AI created value, where human expertise stayed necessary, and what evidence leadership needed before investing further. Each decision moved the project from an exploratory AI capability toward a viable pilot.

1

Who actually holds decision rights?

Stakeholder mapping

What looked like a small IT and accessibility-office project turned out to involve faculty, instructional designers, platform teams, the AI vendor, and senior leadership — each with a real stake.

FindingNo one owned the review process — the single most important gap on the map.

DecisionI treated that as a blocker, not a footnote: until review ownership was assigned, any workflow would have ambiguous accountability. So I paused workflow design and led leadership sessions to settle ownership and governance first.

2

What workflow is actually possible here?

Technology & content landscape

FindingAI accuracy varied sharply by visual type — text slides, instructor footage, animation, and screencasts each needed different levels of human judgment.

TradeoffA single uniform workflow would be simpler to run but would fail on the hard content. I chose a tiered workflow — more complex to operate, but it matched effort to where judgment was actually needed.

3

How do reviewers make consistent calls?

Workflow & instructional value logic

I designed the end-to-end flow — selection to AI processing to human review to publish — specifying what AI handles, what humans handle, and where escalation sits.

FindingThe core need was the instructional value logic: a framework for which visuals warrant description and which don't. Describe too much and you bury learners in noise; describe the wrong things and you mislead them.

TradeoffCompleteness vs. usefulness. I optimized for instructional signal over exhaustive coverage — so human effort lands where it changes learning, not everywhere.

4

How will the institution know it's working?

Evaluation framework

I defined success across three lenses, each tied to who cares about it — the exact criteria Phase 1 reviewers apply to every video today:

  • Detection quality — is AI catching the right visual elements? (Vendor)
  • Description quality — accurate, clear, instructionally appropriate? (Faculty + learners)
  • Workflow efficiency — sustainable at the pace the institution needs? (Leadership)
Outcome

A repeatable system the institution can run and evaluate

Operating model diagram
A simplified view of the workflow — the teams, governance, and feedback loops behind scaling AI-assisted audio description. Click to enlarge.

The deliverable was an operating model — roles, workflow, review criteria, and evaluation metrics connected into one governable system. It replaced an ad hoc AI experiment with something the institution could actually run and evaluate.

  • Phase 1 live across 3 courses and 100 videos
  • Instructional value logic in active use — reviewers flagging visuals against my criteria right now
  • Phase 1 findings gate Phase 2 (~700 videos), building toward 16,000+ assets

My engagement ended at pilot launch — by design. The fact that it runs without me in the room is the point.

The work was presented at TeachX 2026 at Northwestern University, contributing to broader conversations on human + AI accessibility workflows in higher education.

What I'd Take Further

Track edit rate by visual category

Track the human edit rate by visual category, not just overall quality. High edit rates on specific content types feed straight back into retraining the AI — turning review into a continuous improvement loop instead of a one-time check.

Define success before the workflow

Define evaluation criteria before designing the workflow, not after. Success-first would have let me test every workflow decision against what the institution was actually optimizing for.

What This Taught Me

The technology was the easy part. The AI worked from day one. The six months went into the system of people, decisions, and governance that made it usable. The workflow was the real product.

Criteria have to work without you in the room. A reviewer on video 80 in month three needs to apply the logic cold. Designing for handoff, not just for the pilot, was the hardest part.

Want to talk through the thinking behind this?