Thirteen front doors became one. A human stayed in charge.
E-Wolf Support is a confidence-gated AI routing platform for student support at FRCC. I designed it end to end as the only designer in the building. It went through two months of testing, and it's now in Phase 2 of development with $300K behind it.
The case study as a deck. Arrow through it; every slide maps to a section below.
Thirteen front doors and not one receipt.
A student needs a late add approved before Friday or she falls a semester behind. The FRCC website offers her a Contact us page. Behind it: a faculty directory that assumes she already knows a name, three campus pages listing service hours, and a wall of thirteen Formstack entries. Only four of those are actual forms. The other nine are an email address or a phone number. The AI Chat Pod can't take her request either. It answers with the same contacts, and its own logs made the case for something better: of 1,242 conversations analyzed, 789 were students actively asking for help.
Whichever door she picks, the same thing happens next: nothing. No confirmation email. No request ID. No status. If two offices are involved, she explains herself twice and hopes somebody follows up.
Staff carried the other half of the cost. Requests arrived through email, phone, walk-ins, and forms in parallel, and triage ran about fifteen minutes each. Read it, interpret it, guess the department, forward it. At 250-plus requests a day, during a 26 percent decline in enrolled headcount, the maze wasn't a quirk. It was the front door of the institution.
The before state, traced by hand in a field audit: thirteen entry points, nine of them only an email or phone listing, and not one of them sends a confirmation or creates a trackable request.
Sole designer, end to end.
One designer, thirteen intake channels, and a $300K bet that a student could stop guessing which door to knock on. I was FRCC's only designer, sitting inside Strategic Marketing and Communications. I owned the field research, the interaction and visual design, the AI system design, and the pitch that won cabinet approval. A UX researcher partnered on the 1,242-conversation analysis that tuned the thresholds, and IT partnered on the build. Constraints: four departments with four different processes, and a user base that includes prospective applicants with no login. Everything below went through two months of testing and is now in Phase 2 of development.
One door, a consent step, and a receipt in the inbox.
Now there's one door. The form takes current and prospective students, asks for two sentences in plain words, and handles the rest.
The AI shows up exactly once before submission. When it does, it asks permission first. E-Wolf proposes a cleaner version of the request: a clear subject line, the details in order, nothing invented. The student reads it and approves it or declines it. Decline it, and the original text goes through verbatim. The original wording is stored either way, so authorship never leaves the student.
The same screen declares the destination before anything is sent. I considered the quieter alternatives: route silently in the background, or reveal the split after submission. Both are faster to build, and both are worse, because a surprise after submit reads as the system doing things behind your back, and an anxious student has no appetite for surprises. So when a request spans two departments, the screen says so up front, and it splits into sibling requests, each with its own owner. Nothing falls between two queues, and nobody finds out later.
Then comes the receipt that never existed before: a request ID on screen, a confirmation email in the inbox, and a tracking link. From there the request moves through a five-step stepper, Submitted to Resolved. One notification and one email per meaningful change, never one for internal edits. That's the noise policy, and it's deliberate: the moment update emails feel like spam, students stop reading the one that matters. At the end there's a named human, not a bot.
Three roles live in the product: students who ask, staff who resolve, and administrators who watch the whole system.
Before, a student guessed between thirteen doors and waited about a week with no receipt. After: one form, an instant confirmation email, and a request ID she can track to a named human.
Students explicitly approve the AI polish before anything is submitted. The original wording is stored and recoverable, so authorship stays with the student. The multi-department split is declared up front, and each sibling request gets its own owner.
The receipt that never existed before: a request ID, a tracking link, and a confirmation email, issued the moment the student submits.
The live prototype. Click through the golden path: form, consent, confirmation, tracking.
The AI suggests. A person decides.
Think of the best front-desk person you've ever met. She knows who handles what, she reads a half-finished question and gets it to the right office, and when she isn't sure, she asks instead of guessing. That was the design brief for the routing engine.
E-Wolf works the same way, with the judgment made measurable. When a request arrives, the model reads it and scores its confidence about which of the four departments should own it. That score decides everything.
At 90 and above, the request routes itself straight into the department queue. Between 70 and 89, the model suggests a department and a person confirms or corrects it in one click. Below 70, nothing moves. The request holds in a review gate where a person decides, oldest first, with the model's weak signal shown as context, never as a decision. Oldest first is deliberate too: the requests the model understands least are usually the ones a student wrote in distress, and those shouldn't wait behind easier tickets.
The bands aren't round numbers picked in a meeting. The calibration target was blunt: at 90 and above, the model has to be right at least nine times in ten, or it doesn't get to route alone. We tuned against 1,242 real conversations from the intake channels this product replaced, and the thresholds held at 91 percent accuracy on requests the model had never seen. Two months of testing didn't move them.
The hardest call in this section wasn't a threshold. It was deciding what the student sees. Showing the confidence score felt like the transparent choice, and transparency is usually the right instinct. Here it isn't. A number like 78 percent on a support request calibrates nothing for someone without a baseline in probability. It just invites worry, and worry generates the exact follow-up calls this system exists to remove. Google reached the same conclusion with Flights price predictions and suppressed the number entirely. So staff see the score on every queue row, because for them it's information they can act on. Students see a status instead, which is the thing they actually asked for.
One rule sits under all of it: the owner of a request must belong to the department that holds it. Routing is a suggestion. Ownership is not.
Below 70 percent confidence nothing auto-routes; a person decides, oldest first. From 70 to 89 the model suggests a department and a person confirms. At 90 and above the request routes itself.
The human in the loop. Staff confirm or correct the suggested department in one click, and a routing note travels with the request into the audit trail.
Designed for the day the AI is wrong.
The harder design work was the failure states, because an AI product is defined by its worst day, not its demo. Six failure modes, six designed fallbacks, all of them exercised in testing.
Low confidence holds for a person at the gate. A request that spans departments splits into owned siblings before submit, so no team inherits work outside its remit. A declined polish submits the original verbatim, because consent that punishes refusal isn't consent. A misroute is one click to correct, and the routing note travels with the request into the audit trail, so the receiving team knows who moved it and why. An aging, unowned request surfaces on two dashboards, the team's unassigned list and the admin backlog chart, before a student ever feels the delay.
And if the model goes down entirely, intake stays up. Every new request queues as manual review and the staff tooling degrades to a plain triage list, which is exactly what existed before E-Wolf. The floor is designed, not accidental: the service can fall back to the old world, never below it. The other floor is the gate itself. Nothing auto-routes below 70. Not during a backlog, not during finals week, not if staff ask for it. That line doesn't move.
Every status change is atomic and writes an audit entry. Consent precedes every rewrite. Students never see a score. Those three sentences are the whole trust model, and every screen in the product answers to them.
Thresholds were tuned on 1,242 real conversations and hold 91 percent routing accuracy on held-out requests. The dashed lines are the 70 and 90 percent gates.
Six failure modes, six designed fallbacks. When the model is unavailable, intake stays up and everything queues for manual review. The service degrades to the before state, never below it.
Measured before and after, not projected.
Every number below is a before and after. Same operation, measured twice.
Thirteen forms, directories, and listings collapsed into one tracked entry point.
An 87 percent reduction per request, measured across the four departments.
Typical resolution fell 56 percent, and keeps being measured as volume grows.
Held-out accuracy, with thresholds tuned on 1,242 analyzed conversations.
Cabinet approved the platform and committed production funding.
These are operational numbers, not revenue projections. Two months of testing proved the system, and Phase 2 of development is underway on the $300K the cabinet committed.
What I'm sharpening next.
The system works. Two months of testing proved the routing, the gate, and the fallbacks, and Phase 2 of development is underway. What I'm sharpening now is the scope of the AI itself: where else it can be genuinely helpful to the student without ever crossing the lines above.
The number I'm watching is the 70 to 89 band. Every request in it costs a staff member a confirmation click, and I suspect the band is wider than it needs to be. The system is built to answer that question itself. Every confirm and every correction in that band is a labeled example, so as live volume accumulates, we re-tune the thresholds on production data. If the model keeps proving right at 85, the auto-route line moves down, the review band shrinks, and staff minutes come back without the below-70 rule ever loosening.
If this maps to problems your team is working on, I'd value a conversation.