A human + AI system for making instructional video accessible — designed to run reliably at institutional scale.
A federal mandate required universities to make digital learning content accessible, including audio descriptions for instructional video. Leadership wanted to meet it without burying staff in manual work, and saw AI as the way through.
So the project began with the obvious question: can AI detect visuals and generate descriptions well enough to replace manual work?
That question was the wrong one — and finding out why became the project.
Early exploration showed generation was the easy part. The same visual needed different descriptions depending on instructional context — so quality lived in context, review, and decision-making, not the AI output alone.
The instructional designers knew the content. The AI vendor knew the technology. The gap was between them — and that gap was mine to close.
I owned the stakeholder map, the end-to-end workflow, the review logic reviewers use to make calls, and the evaluation framework Phase 1 runs against. I also led the working sessions with senior leadership — the Deputy Vice Provost for Online Learning and the Chief Learning Officer — to settle governance and review ownership before any workflow was locked.
Before defining process, I had to answer where AI created value, where human expertise stayed necessary, and what evidence leadership needed before investing further. Each decision moved the project from an exploratory AI capability toward a viable pilot.
What looked like a small IT and accessibility-office project turned out to involve faculty, instructional designers, platform teams, the AI vendor, and senior leadership — each with a real stake.
FindingNo one owned the review process — the single most important gap on the map.
DecisionI treated that as a blocker, not a footnote: until review ownership was assigned, any workflow would have ambiguous accountability. So I paused workflow design and led leadership sessions to settle ownership and governance first.
FindingAI accuracy varied sharply by visual type — text slides, instructor footage, animation, and screencasts each needed different levels of human judgment.
TradeoffA single uniform workflow would be simpler to run but would fail on the hard content. I chose a tiered workflow — more complex to operate, but it matched effort to where judgment was actually needed.
I designed the end-to-end flow — selection to AI processing to human review to publish — specifying what AI handles, what humans handle, and where escalation sits.
FindingThe core need was the instructional value logic: a framework for which visuals warrant description and which don't. Describe too much and you bury learners in noise; describe the wrong things and you mislead them.
TradeoffCompleteness vs. usefulness. I optimized for instructional signal over exhaustive coverage — so human effort lands where it changes learning, not everywhere.
I defined success across three lenses, each tied to who cares about it — the exact criteria Phase 1 reviewers apply to every video today:

The deliverable was an operating model — roles, workflow, review criteria, and evaluation metrics connected into one governable system. It replaced an ad hoc AI experiment with something the institution could actually run and evaluate.
My engagement ended at pilot launch — by design. The fact that it runs without me in the room is the point.
The work was presented at TeachX 2026 at Northwestern University, contributing to broader conversations on human + AI accessibility workflows in higher education.
Track the human edit rate by visual category, not just overall quality. High edit rates on specific content types feed straight back into retraining the AI — turning review into a continuous improvement loop instead of a one-time check.
Define evaluation criteria before designing the workflow, not after. Success-first would have let me test every workflow decision against what the institution was actually optimizing for.
The technology was the easy part. The AI worked from day one. The six months went into the system of people, decisions, and governance that made it usable. The workflow was the real product.
Criteria have to work without you in the room. A reviewer on video 80 in month three needs to apply the logic cold. Designing for handoff, not just for the pilot, was the hardest part.