A redesign of BlueLink, Braeburn's smart thermostat app.We rebuilt the homeowner experience around first-time setup, troubleshooting, and everyday control, rethinking both the user flows and the visual system.
Explore the case studyBluelink is a two-role system by design: HVAC contractors configure and service the equipment, while homeowners manage its everyday use. The product changes hands at a critical moment — from installation to ownership — leaving the homeowner to build a relationship with a device they did not configure themselves.
Our brief was to identify friction across this experience and uncover opportunities to create a more seamless connection between the hardware and software. Q1 focused on building the UX research foundation; in Q2, we translated those findings into the redesigned end-user experience.
How might we preserve existing end-user processes — onboarding, troubleshooting, and everyday use — while implementing an improved user experience into the Bluelink app?
In Q1 our team focused on research, building the foundation for the design work that followed. I led the contextual inquiry, user interviews, competitive analysis, and heuristic evaluation. In Q2 I focused on user flows and the high-fidelity prototype and helped build the design system, while my teammates led journey mapping and the two rounds of usability testing. Once testing gave us insights we could act on, we iterated the prototype together and carried it through to final handoff. The screens on this page are a team output; the flow and prototype work is mine.
What do users want from a value product like Bluelink?
Where does the contractor experience break down during installation and consumer handoff?
Where does the consumer experience break down during onboarding and daily use?
How effectively does the current app mirror and complement the physical hardware?
A higher user confidence score during onboarding and in error states — a clearer understanding of what went wrong when a flow fails.
Alignment between the user's reported mental model and the actual system model.
Reduced negative feedback and drop-off on core flows — adding a thermostat, creating a schedule.
“You can have the best thermostat in the world, but if it fails or you can't get it to connect and you can't talk to anybody, what good is it? What good is it?”
Interview participant · retired HVAC contractor, 15+ years — and a Braeburn homeownerThe client's material was almost entirely market-level, with no UX perspective at all. Across 20 sources we rebuilt the foundation: how people form mental models of temperature control, and how trust in smart-home products is built and broken. The number that stuck: installation splits 51% DIY / 49% professional — a near-even divide, which says there is still a lot of room to improve how each of those two paths is supported.
The client's competitive material was about the market and its trends, not the product design. We redid it — 7 brands across 15 experience criteria: install guidance, scheduling, C-wire solutions, Wi-Fi band, app onboarding, app reliability. The verdict was blunt: Bluelink is positioned as a value brand for contractors and homeowners, the price holds up, multi-thermostat monitoring is a genuine differentiator — but app reliability came out as a clear weak spot.
Four of us walked the live Bluelink app screen by screen. Result: 9 of Nielsen's 10 heuristics violated across 6 screens, 4 of them critical. The worst pair was consistency and minimalist design — every home-screen button is the same size and color, and the largest, brightest one turns out to be account information. The app also forces landscape orientation. Our shared conclusion afterwards: redesigning is cheaper than patching.
Four depth interviews spanning both sides of the handoff: a retired contractor who is also a heavy user, a manufacturer's rep who had just self-installed, and two homeowners — one a double amputee for whom smart-home automation is not a convenience but a necessity. This is where we explored what actually happens at the moment of handoff.
Fielded through QuestionPro: 103 responses, 98% completion. This round didn't ask why people buy — it asked what happens after the box is opened, which is exactly the half the client's existing market data couldn't answer.
Mainly adjust manuallyOnly 22% let the device run automatically. The app is used constantly — the automation isn't trusted with the job.
Energy reports change behaviorMore than any automation feature. Cost information follows at 43% and alerts at 44% — geofencing lands near the bottom.
Households share controlYet the app is built for a single user. Multi-user control came up repeatedly in the open-ended responses.
Friction is front-loaded into wiring and account setup — precisely the stretch where the contractor is present and the homeowner is not.
The problem isn't comprehension, it's effort — which shifted the design goal from re-explaining the interface to cutting steps out of it.
These 103 respondents are smart-thermostat owners in general, not Bluelink owners — 50% are on Nest and only 2% on Braeburn. So we used it as a category baseline: it tells us where users of premium products are still frustrated, while the interviews and heuristic evaluation carry the Bluelink-specific answers. Conflating the two would have produced the wrong conclusions, and we flagged that explicitly when reporting to the client.
The official flow does account for the handoff — the printed guide is meant to be left behind. But in every case we interviewed, the handoff was improvised in person: one contractor's method was to pull out his own phone, demo the app until the customer said they wanted it, then finish the install and leave; one homeowner wasn't even home that day. The judgments made during setup stayed in the installer's head — one homeowner, two to three years in, still didn't know geofencing existed.
A failed connection returns a code, not a diagnosis — thermostat, router, internet, phone: the interface never says which of the four links broke. So support became part of the product. Users name individual reps and cite them as the reason for brand loyalty. Support being the strong point is the bill for the product's UX debt.
Users define success as the device disappearing into the background. They aren't asking for more features — they're asking for the existing ones to be reliable. One participant abandoned geofencing outright because it was, in her words, about a fifty-fifty shot. This directly challenges the assumption that more engagement means a better experience: here, success means not needing to open it.
“I just want to turn it on, off, and change the temp. I don't need anything else.”
Uses scheduling, remote access, Alexa integration, geofencing and energy-saving modes daily.
“Once you get it, it works great.” / “It's pretty straightforward.”
Watched YouTube videos, called a coworker, phoned Braeburn support, or had a professional install it. “Straightforward” only in retrospect.
“I don't actually open the app.” / “I connected them and that was pretty much it.”
90% of control happens on the phone. Checks temperature remotely, pre-cools before returning home, adjusts from the couch instead of walking to the device.
“It's a stupid gripe. It's ridiculous. Not a big deal.” — about re-authenticating on every login.
Mentions it unprompted and describes it in detail. Friction memorable enough to surface in a research interview is friction that has accumulated.
This comparison was the most useful instrument in the whole study. Users almost never state their pain points directly — they rationalize them away as first-world problems. So we stopped listening to how they rated the product and listened to what they described themselves doing.
Heavy app users with variable schedules. Manual adjusters for events, travel and arrivals. Core need: fast, reliable remote control.
Consistent routines, deep users of scheduling. Retirees and shift workers with fixed hours. Rarely open the app once configured.
Low app engagement, variable schedule. Distrustful of complexity, prefer the physical device, use basic on/off only.
Deep ecosystem investment — Alexa, voice control. Predictable enough to automate everything. The segment that suffers most when connectivity fails.
The point of segmenting wasn't to categorize but to see that the same defect costs different people wildly different amounts. For a power remote user, a dropped connection is an inconvenience. For the bottom-right segment who depend on voice and automation — including the double amputee we interviewed — the same dropped connection means being stranded.
Anticipation meets confusion. Wiring terminology mismatches between brands cause real anxiety, especially when the installer isn't present. The manual is praised; cross-brand jargon is the blocker.
The lowest-sentiment moment in the journey. The thermostat → phone → router chain was described as “a little confusing” by multiple users. Feature overwhelm on first launch. Those who reach support recover; those who don't give up on advanced features entirely.
Users who persevere reach their equilibrium here. Scheduling is set once and largely forgotten; ecosystem integrations get made. But value is highly lifestyle-dependent — WFH users and retirees often skip scheduling entirely. Integrations not configured at this stage are rarely configured later.
The happiest phase — because the product disappears. Users stop noticing it, which is exactly their goal. Remote control is used constantly, geofencing delights when it works, minor annoyances are rationalized away. This is where brand loyalty forms.
When connectivity breaks, users have no idea where to start — router? ISP? thermostat? Alexa? That leads to blind troubleshooting: “I just go in and mess with it.” No system feedback on what went wrong. Users who call support recover their trust; those who self-troubleshoot lose confidence in the whole platform. For accessibility users, this phase is physically isolating.
This curve set the priorities for the entire second quarter. The product performs beautifully at stage 4 — but whether a user ever reaches stage 4 depends on whether stage 2 drove them off, and stage 5 can undo the trust built at stage 4 in a single evening. So we didn't optimize the happy stretch. We spent the design effort on the two troughs: first setup, and what happens after something breaks.
In Q2 the team split into two parallel tracks — one driving usability testing, one building the design system — merging weekly for review. This is also where scope closed around the homeowner app for good.
Two sub-teams in parallel: Sophia and Sarah mapped the full contractor-to-homeowner journey while Aaron and I broke down the app's user flows. Laid over each other, the handoff gap was impossible to miss.
Eight ideas in eight minutes, one minute each. The point wasn't to find the answer — it was to burn through the obvious solution fast so the less obvious ones had room to appear.
Once we converged we built clickable lo-fi prototypes in Figma Make, so the flows could be walked through instead of argued about in a meeting.
We presented the concepts in three tiers — conservative, moderate, aggressive — at the same time. It was the most effective communication move of the project: it turned the conversation from “is this acceptable” into “how far are you willing to go,” and got the client to name their own constraints out loud.
Round one ran on the lo-fi prototype while the design-system track started in parallel, so findings could be absorbed straight into components instead of stranded in meeting notes.
We built the design system on the 2024 Braeburn brandbook — 6 core colors, 9 semantic tokens in each of light and dark, 6 approved gradients — then built close to 70 screens of high-fidelity prototype on top of it, covering account creation, pairing, troubleshooting, scheduling, reports and settings.
Unmoderated, to check whether the Round 1 fixes actually held. With nobody there to explain, the interface had to speak for itself.
The final delivery was a package: the design system, annotated flows, and an explicit trace from each test finding to the design decision it drove — so whoever picks it up knows why each call was made, and which recommendations they're free to decline.
The largest, brightest, only-yellow button on the screen is “Service Dealer” — which is just an account information page. Visual weight runs exactly opposite to actual importance.
→ The redesign makes the current temperature the single anchor and files account info under Settings.List items on the left are rectangular while the buttons on the right are pill-shaped — two UI languages on one screen. Devices are named by model number: “7 Day Commercial 7320” asks the user to remember which unit is in which room.
→ One component language, plus custom device names — a fix traced directly to Pro1 in the competitive analysis.A bare “?” floats in the bottom-right corner with no border and no label. In testing, not one participant realized it was tappable.
→ Help became a labelled, permanent destination in the bottom navigation bar.Temperature moves one degree per tap through arrow buttons, with no way to type a number. Going from 60°F to 75°F takes fifteen taps.
→ The redesign keeps ± but adds a draggable ring with both bounds visible at once.“Following Program,” “Occupied Schedule,” “Hold” — HVAC trade language throughout. One participant, verbatim: “Permanent hold sounds scary… what's permanent?”
→ Rewritten in plain life language: Cooling → 72°F, wake up, leave home, return home, sleep.Both screens are landscape because the entire app forces landscape orientation — in a world where nearly everyone holds a phone upright. It was the fastest agreement we reached in the evaluation, and the starting point for deciding that a redesign was cheaper than a patch.
On the old home screen every button was the same size and color, with no hierarchy. Here the current temperature is the single visual anchor, with heating and cooling read off one continuous ring in the brand's semantic colors. “Feels like” and humidity were called out by name in testing as moments of clarity — they translate raw data into something a person reads instantly.
Scheduling scored 4.48 out of 5 in the survey and was simultaneously the loudest complaint in the open-ended responses — a satisfied majority hiding a badly frustrated minority. The old version buried it in a multi-layer calendar where users had to guess how weekly and monthly differ. Here the schedule is anchored to life events — wake up, leave home, return home, sleep — with the time and the temperature on the same line.
Energy reports are the strongest behavior driver in the survey (51%) and simultaneously carry the highest non-use rate (15%). The old version emailed you a CSV — you had to leave the app to read it. Here it lives inside the app with a neighborhood comparison, and the clinical name “Data Reports” became “Smart Report,” a rename participants proposed themselves.
This is the screen I care most about. The old app returned a code when connection failed — thermostat, router, internet, phone, and it never said which link broke. Here the failure is split into three symptoms a user can recognize as their own, and “Contact Support” sits at the bottom in the lightest possible treatment: still there, no longer the only way out.
Round 1 produced 9 themes, 6 pain points and 8 design opportunities. These five are the ones that actually changed the design — not the ones confirming we were right, the ones telling us where we were wrong.
The most consistent finding of the round: nearly every participant failed to recognize the word. They tried Auto first, then scheduling, then settings, and eventually found it by elimination — it carried the longest task times in the study. Once explained, everyone liked the feature. The problem was never the feature; it was the label. We renamed it “Away mode” and gave it a scenario line.
Granting Bluetooth created a firm expectation that the device would now connect on its own. When Wi-Fi configuration appeared afterwards it felt jarring and disconnected from what came before. The sequence itself communicated the wrong mental model — so we moved Wi-Fi up to be introduced alongside Bluetooth.
When a device went offline, every participant's first move was Settings. In their model, Settings is where you manage the device and Help is where you find a phone number. We had put troubleshooting under Help; this finding moved it. Separately, the help button itself was too quiet — visually identical to its neighbors, and a stressed user is the least likely to search carefully.
People's prior experience with QR codes is that scanning opens a web page, not that it searches for a nearby device. On failure they couldn't tell whether it was their error, the distance, or the app. The interesting part: the recovery flow itself was fine — once troubleshooting options were visible, nearly everyone resolved it. What was missing was expectation-setting before the scan.
Anything marked optional during setup — scheduling, away mode, energy alerts — got skipped by every participant, with the same reasoning: it isn't needed to use the thing, I'll come back to it. They don't come back. So we moved the explanation to the moment of use: the app now surfaces a short, dismissible introduction the first time someone actually opens a feature — so skipping it during setup no longer leaves them stuck later.
Remote and unmoderated on the high-fidelity prototype through Optimal Workshop — 10 tasks, numbered to match Round 1.
In Round 1, not one participant understood “geofencing.” In Round 2, before they ever saw the screen, 5 of 5 named the feature “Away” or “Away mode” unprompted, and 5 of 5 found the label immediately clear once they reached it — one describing it as better than expected, another as exactly what they thought it would be called. The cleanest before/after in the dataset.
Same feature: 3 of 5 still went hunting through Settings first and only came back to the home screen after failing; 2 of 5 missed it entirely on the first pass, one of them detouring into accessibility options. One called it clearly the hardest step of the session — it simply did not click. A rename solves a vocabulary problem. It does not solve an information-architecture problem — and I had assumed the label was the whole failure.
QR scanning failed or was unavailable for 3 of 5 — in one session no QR code appeared at all. When the device wasn't found, 4 of 5 expected the app to walk them through recovery. It didn't. 3 of 5 retried in a loop rather than looking for in-app help, and 2 of 5 showed open frustration. The troubleshooting screen I was proudest of does work — once it is found. One participant only got there because an alert pushed them.
Worth noting: we changed nothing about the Wi-Fi step. It was the smoothest flow in Round 1 and stayed smoothest in Round 2. It became the ruler — the level of clarity every other flow was measured against.
In the same report, 3 of 5 described schedule setup as clear and intuitive while 3 of 5 hesitated or backtracked at event creation — one cycling aloud between “create event” and “create schedule,” unable to choose. The honest read: improved enough that everyone got through it, not enough to call it resolved.
Four participants abandoned swiping for tapping, citing imprecision. But I won't write this up as a finding: the study ran on desktop with a mouse — judging a gesture designed for a thumb by how it feels under a cursor proves nothing. This one needs a re-test on a real device.
This was the project's starting point — Round 1's central insight was that when something fails there is no path but the phone. We asked everyone at the end of Round 2. Three of the five gave specific, attributable reasons: clear labels, being able to find features unaided, and in-app troubleshooting offering enough options to handle it themselves. One described support as a last resort for drastic hardware failures only, with the app handling everything else. No participant expressed the opposite view. The one reservation came from another participant, who named scheduling as the place he might still want help.
One note: Optimal Workshop's auto-generated summary claims all five, while the itemized evidence beneath it covers three. I report three. A tool's AI summary is a lead, not a finding — copying one straight into a report is a habit this year taught me to distrust.
The constraint is the point. The brandbook defines six core colors and no status or feature colors at all — so heating, cooling, success and error are tokens we added on top of it, each with a light and a dark value and a contrast check. What the client received was therefore not a new visual language competing with their own, but a legitimate extension of an existing brand asset — which is what determines whether it actually gets adopted.
The client had extensive market research, all of it answering why people buy. Our 103 responses asked what happens after they buy, and when they give up. Sample size was never the constraint — the framing was.
We deliberately dropped the cognitive walkthrough. With a full redesign already decided, two weeks validating a flow we were about to delete would have come straight out of the design work. Methods exist to answer questions, not to prove the process was complete.
The mild / medium / spicy format was the most effective communication move of the project. It reframed “can you accept this” into “how far do you want to go” — the bold option made the middle one feel safe, and the client volunteered their own resource constraints as a result.
Users could name Braeburn's support reps and cited them as the reason for their loyalty. It reads as a strength, but every one of those calls corresponds to something the product failed to resolve on its own. I learned to read what users praise separately from what the product owes.